Qwen3.6-35B-A3B on dual RTX 3060 12GB (llama.cpp tensor + MTP, 88k)
Qwen3.6-35B-A3B-GGUFTwo RTX 3060 12GB, no NVLink. llama.cpp b10682, Unsloth UD-Q4_K_XL, 88k. Tensor split + native MTP + q4_0 draft KV: 129 tok/s code / 104 tok/s prose. 96k+MTP does not fit.
1 benchmark
