Erste SchritteLeaderboardDecode calculatorModelleReportsHardwareBenchmarksMarktplatzVermietungenProAPI-Doku
Sprache
Actual Computer — Every computer, one endpoint

Reports

Community-authored local inference reports

Published field reportsnewest

Qwen3.8-27B on one Tenstorrent P150: 7.7 to 50+ tok/s, native 262K context, and lessons learned

Qwen3.8-27B-TT-Mixed-BFP4-BFP8-P150

Five weeks taking Qwen3.8-27B on a single Tenstorrent Blackhole P150 from 7.7 tok/s to ~50-55 tok/s single-stream decode (DFlash2), ~1,200 tok/s prefill and native 262K context, using a custom mixed BFP4/BFP8 quant (now GPTQ). Covers the quant, every speed-test milestone, quality checks, power and thermal data, the dead ends, lessons learned and what's next.

BFP4 vllmTenstorrent P150 · 32 GB
3 speed tests
@Lottolabs Oct 3, 2026 0

Zeus 2× Arc Pro B70: llama.cpp SYCL FP16 tuning across Qwen3.8 and nine other model routes

Qwen3.8-27B-GGUF

On two Intel Arc Pro B70s, a b11190 SYCL FP16 build substantially improved prompt processing across 10 production routes. This report gives matched serving results, copy-heavy generation gains from n-gram drafting, rejected tuning variants, and measurement limits.

Q6_K_XL focus; Q8_0, Q4_K_M, UD-Q4_K_XL llama.cpp b11190, SYCL / Level Zero
@Captain-Tripps Sep 26, 2026 0

895 tok/s from a 27B dense model on one RTX 5080: DFlash2 + n-gram hybrid drafting on PQ2_0

Bonsai-2-27B-Ternary-CRACK-GGUF

How a 2.13-bpw ternary 27B, a self-speculative DFlash2 draft head, an n-gram drafting layer, and four small llama.cpp patches took a single RTX 5080 to 895 tok/s warm / 142 tok/s fresh - with verified runs, kernel-level analysis, and a full reproduction package.

PQ2_0 llama.cppRTX 5080 · 16 GB
7 speed tests
@zotowata Sep 21, 2026 0

FreeToken vs vLLM on one RTX 3090: BF16 MoE RAM offload compared

Qwen3.6-35B-A3B

An apples-to-apples single-RTX-3090 comparison of FreeToken expert offload and vLLM CPU offload with Qwen3.6-35B-A3B BF16. FreeToken reached 42.78 tok/s established decode versus 7.01 tok/s for tuned vLLM, while vLLM delivered lower TTFT. Includes sequential, prefill, concurrency, memory, tuning, and MTP/DFlash compatibility findings.

BF16 FreeToken,vLLM
@Lottolabs Aug 25, 2026 0