CMP 170HX: Qwen3.8-27B AutoRound + DFlash2 at 131k
Qwen3.8-27B-W4A16-AutoRoundvLLM 0.29 on a 64GB CMP 170HX. AutoRound W4A16 + DFlash2, production max-model-len 131072. 133 tok/s short decode, 98 tok/s with a 7.3k prefix. Not an 8k-ctx speed-run.
2 benchmarks
