Qwen3.8-27B on RX 9070 XT: one 64K-context stream, 655 tok/s prefill, 23.5 tok/s decode on 16 GB
Qwen3.8-27BA 65,100-token prompt on a single 16 GB card: 4-bit KV cache (1.1 GB for the whole context), host-staged embedding table, 99 s TTFT (was 146 s before the flash-attention fix), 23.46 tok/s decode, 14.24 GB peak. Speculative decoding is unavailable past 4K context by design.
