Qwen3.8-27B on RX 9070 XT: 4 concurrent streams at 4K context, 98.3 tok/s aggregate (custom HIP engine)
Qwen3.8-27BFour 4K-context streams on one 16 GB RDNA4 card: custom IQ4_S256 quant, 4-bit KV cache, 16K shared paged pool. Plain batched decode matches speculative decoding at this batch (98.2 vs 98.3 tok/s) with 1.65 GB more headroom.
