Qwen3.8-27B on one Tenstorrent P150: 7.7 to 50+ tok/s, native 262K context, and lessons learned
Qwen3.8-27B-TT-Mixed-BFP4-BFP8-P150Five weeks taking Qwen3.8-27B on a single Tenstorrent Blackhole P150 from 7.7 tok/s to ~50-55 tok/s single-stream decode (DFlash2), ~1,200 tok/s prefill and native 262K context, using a custom mixed BFP4/BFP8 quant (now GPTQ). Covers the quant, every speed-test milestone, quality checks, power and thermal data, the dead ends, lessons learned and what's next.
