895 tok/s from a 27B dense model on one RTX 5080: DFlash2 + n-gram hybrid drafting on PQ2_0
Bonsai-2-27B-Ternary-CRACK-GGUFHow a 2.13-bpw ternary 27B, a self-speculative DFlash2 draft head, an n-gram drafting layer, and four small llama.cpp patches took a single RTX 5080 to 895 tok/s warm / 142 tok/s fresh - with verified runs, kernel-level analysis, and a full reproduction package.
7 benchmarks
