Self-reported. Single Intel Arc Pro B70 32GB. Realistic 5-user coding serving, v2 with REAL completions kept in session history (supersedes the stubbed same-day 138.3 record). Stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching ON. 5 concurrent sessions, ~8K start, 3 turns, g512, 60/60 OK. tokSOut = sum of per-stream generation rates (5 x 25.5 median per-user tok/s, all turns). Per-turn: t1 165.5, t2 127.9, t3 114.2 aggregate. TTFT 22.6-25.0s every turn: prefix cache lands 0-38% of shared tokens at C5 on this build (vs 91% at C1), so ~45-53K session tokens re-prefill per turn. MTP acceptance 43-56% under concurrency. Short-prompt C5 (203.8) is a separate record; C1 is 106.7.
Self-reported. Single Intel Arc Pro B70 32GB. Current full stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1 overlay, prefix caching ON (zero hits, unique prompts). C1 client post-first, n=5 median (92.9-105.1), calibrated real-world Pi prompt set, same prompt file as the 2026-08-18 record. p8192/g128 median 100.0; cold input 1694 tok/s at p8192/g1. Prompt-content sensitive: degenerate INDEX filler measures ~71 (44% MTP acceptance) vs 93-96% on realistic text. Cache-off record on the same patches: 112.65.
NEW identity vs cmsur82fz06svms01ga1f0z83 (BF16 draft 83.7). Same 1x B70, same image vllm/vllm-openai-xpu@sha256:f01e24f6c7ff01f1e0662234255a1372297d1dbd89d003cf13c8fad3eab1ba4f, kernels 0.1.12.3, same target SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16. Extra patches after Qwen MTP nightly+boundary: patch_gdn_mixed_split_v5.py + patch_draft_lmhead_int4.py + patch_draft_mtp_int4.py (runtime RTN INT4 of draft LM head + 5 MTP linears only; target verify stays BF16). tokSOut = client post-first at p512/g128, median n=5, C1, cache OFF (zero prefix_cache_hits_total delta): 112.65 (range 111.58-117.73). Matched same-harness BF16-draft arm was 81.20, not the older Run40 83.7 card. tokSPrefill = actual input/TTFT at p8192/g1 MTP4 n=5: 1696 (flat vs 1691). Accept 510/540=94.44% vs 95.86%. Short agentic +32.8%; 8K/16K long-ctx agentic +37%/+26%. 230 W configured cap; measured median draw 205 W at p512/g128. Speed-only: no token/KL/task-quality parity vs BF16 draft. E2 self-reported. Measured 2026-08-18.
Qwen3.8-27B (dense 27B, hybrid GDN linear+full attention, MTP head not in GGUF quant). Community quant unsloth/Qwen3.8-27B-GGUF Q4_K_M (SHA-256 7b2aec3b...cc89f1b) -- weights released 2026-08-14T15:00Z, benched same day. 230W cap (stock). llama-bench tg128 mean 22.30 t/s (+/-0.01, n=5, C1, cold cache). Cold input pp8192 mean 1121.40 t/s (+/-1.89, n=5); pp512 778.80 / pp2048 1115.99 / pp4096 1137.02. Load at -c 131072 verified, 10.2 GiB VRAM free after load. KV q8_0 K + q4_1 V, FA on. Correctness: coherent smoke 5/5, deterministic replay pass, token exactness pass; reference parity not run. Cross-cap note: at matched 150W cap this model runs tg128 mean 17.78 (n=5, separate submission); prior Qwen3.6-27B Q4_K_M ran tg128 21.3 mean (n=3, 150W, Run 9) -- comparisons across caps are directional only. Dense power scaling 150->230W measured +25% decode on this model.
Qwen3.8-27B (dense 27B, hybrid GDN linear+full attention, MTP head not in GGUF quant). Community quant unsloth/Qwen3.8-27B-GGUF Q4_K_M (SHA-256 7b2aec3b...cc89f1b) -- weights released 2026-08-14T15:00Z, benched same day. 230W cap (stock). llama-bench tg128 mean 22.30 t/s (+/-0.01, n=5, C1, cold cache). Cold input pp8192 mean 1121.40 t/s (+/-1.89, n=5); pp512 778.80 / pp2048 1115.99 / pp4096 1137.02. Load at -c 131072 verified, 10.2 GiB VRAM free after load. KV q8_0 K + q4_1 V, FA on. Correctness: coherent smoke 5/5, deterministic replay pass, token exactness pass; reference parity not run. Cross-cap note: at matched 150W cap this model runs tg128 mean 17.78 (n=5, separate submission); prior Qwen3.6-27B Q4_K_M ran tg128 21.3 mean (n=3, 150W, Run 9) -- comparisons across caps are directional only. Dense power scaling 150->230W measured +25% decode on this model.
Qwen3.8-27B (dense 27B, hybrid GDN linear+full attention, MTP head not in GGUF quant). Community quant unsloth/Qwen3.8-27B-GGUF Q4_K_M (SHA-256 7b2aec3b...cc89f1b) -- weights released 2026-08-14T15:00Z, benched same day. llama-bench tg128 mean 17.78 t/s (+/-0.17, n=5, C1, cold cache). Cold input pp8192 mean 787.45 t/s (+/-1.74, n=5). Load at -c 131072 verified, 10.2 GiB VRAM free after load. KV q8_0 K + q4_1 V, FA on. Correctness: coherent smoke 5/5, deterministic replay pass, token exactness pass; reference parity not run. At matched 150W/build/quant/flags the prior Qwen3.6-27B Q4_K_M ran tg128 21.3 mean (n=3, Run 9); 3.8 mean 17.78 (n=5) -- ~16% slower decode. Comparison directional: different checkpoints, n mismatch (3 vs 5).
Self-reported E2, not independently reproduced. C1 (max_num_seqs=1), prefix cache explicitly off (--no-enable-prefix-caching; enable_prefix_caching=False; zero prefix_cache_hits delta). Timing is client monotonic SSE: tokSOut is median client post-first decode at exact p2048/g128 n=5 = 186.61 t/s (range 174.60-201.83; families research/rag/tool/document/assistant). Additional decode cells on the same warm server: p512/g128 median 194.61 (range 140.20-220.01, acceptance 45.1% — wide family spread, not the representative scalar); p8192/g128 median 157.92 (143.50-170.25, acceptance 53.0%). tokSPrefill is the p8192/g1 n=5 median COLD INPUT RATE from client TTFT (7160 t/s, range 7117-7226), which includes scheduling and first-token work — NOT isolated engine prefill. Also p2048/g1 cold input median 6456 t/s. Spec: method=dflash n_spec=7; window acceptance 1830/3521 = 52.0%. Target is a LOCAL symmetric GPTQ INT4 G64 conversion of the named BF16 repo (not an official HF quant). Draft is a LOCAL NVFP4 E2M1→BF16 reconstruction of NVIDIA DFlash, not a published BF16 draft. Stack: vllm/vllm-openai-xpu@sha256:1da0a95485455f08588c11080b9718992fd7d434c6a965d74654903a9d999c57 plus local det image patches (native grouped-topk v2, SSU B8/W4, at::zeros grouped-GEMM); VLLM_XPU_ENABLE_XPU_GRAPH=1; PIECEWISE+FULL graphs; --async-scheduling; --quantization gptq --dtype float16 --max-model-len 16384 --gpu-memory-utilization 0.90. Configured cap 150 W; measured cell-window averages ~149-160 W; peak interval-average 179.3 W (0.5 s energy1_input on p8192/g128); pkg max 68.0 C. Deterministic raw-completion replay smoke exact_match=true (scope: smoke only, not logit/KL or task-quality). n=3 screen 20260813T080234Z is superseded and was not used. Matched-except-speculation vs prior no-spec n=5 graph campaign (p8192/g128 87.25 t/s): 1.81× on that cell only.
Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). tokSPrefill is the cold input rate at exact p8192/g128 from median client TTFT (8,368 tok/s, TTFT median 0.9789s, n=3) - a client-observed input rate including scheduling and first-token work, NOT an isolated engine prefill measurement. NO speculative decoding in this config: native MTP was tested and rejected on this stack (0% draft acceptance), N-gram rejected (temperature-0 nondeterminism), DFlash not supported for this model on this stack. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent.
Self-reported. Single Intel Arc Pro B70 32GB. Realistic 5-user coding serving, v2 with REAL completions kept in session history (supersedes the stubbed same-day 138.3 record). Stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching ON. 5 concurrent sessions, ~8K start, 3 turns, g512, 60/60 OK. tokSOut = sum of per-stream generation rates (5 x 25.5 median per-user tok/s, all turns). Per-turn: t1 165.5, t2 127.9, t3 114.2 aggregate. TTFT 22.6-25.0s every turn: prefix cache lands 0-38% of shared tokens at C5 on this build (vs 91% at C1), so ~45-53K session tokens re-prefill per turn. MTP acceptance 43-56% under concurrency. Short-prompt C5 (203.8) is a separate record; C1 is 106.7.
Self-reported. Single Intel Arc Pro B70 32GB. Current full stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1 overlay, prefix caching ON (zero hits, unique prompts). C1 client post-first, n=5 median (92.9-105.1), calibrated real-world Pi prompt set, same prompt file as the 2026-08-18 record. p8192/g128 median 100.0; cold input 1694 tok/s at p8192/g1. Prompt-content sensitive: degenerate INDEX filler measures ~71 (44% MTP acceptance) vs 93-96% on realistic text. Cache-off record on the same patches: 112.65.
NEW identity vs cmsur82fz06svms01ga1f0z83 (BF16 draft 83.7). Same 1x B70, same image vllm/vllm-openai-xpu@sha256:f01e24f6c7ff01f1e0662234255a1372297d1dbd89d003cf13c8fad3eab1ba4f, kernels 0.1.12.3, same target SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16. Extra patches after Qwen MTP nightly+boundary: patch_gdn_mixed_split_v5.py + patch_draft_lmhead_int4.py + patch_draft_mtp_int4.py (runtime RTN INT4 of draft LM head + 5 MTP linears only; target verify stays BF16). tokSOut = client post-first at p512/g128, median n=5, C1, cache OFF (zero prefix_cache_hits_total delta): 112.65 (range 111.58-117.73). Matched same-harness BF16-draft arm was 81.20, not the older Run40 83.7 card. tokSPrefill = actual input/TTFT at p8192/g1 MTP4 n=5: 1696 (flat vs 1691). Accept 510/540=94.44% vs 95.86%. Short agentic +32.8%; 8K/16K long-ctx agentic +37%/+26%. 230 W configured cap; measured median draw 205 W at p512/g128. Speed-only: no token/KL/task-quality parity vs BF16 draft. E2 self-reported. Measured 2026-08-18.
Qwen3.8-27B (dense 27B, hybrid GDN linear+full attention, MTP head not in GGUF quant). Community quant unsloth/Qwen3.8-27B-GGUF Q4_K_M (SHA-256 7b2aec3b...cc89f1b) -- weights released 2026-08-14T15:00Z, benched same day. 230W cap (stock). llama-bench tg128 mean 22.30 t/s (+/-0.01, n=5, C1, cold cache). Cold input pp8192 mean 1121.40 t/s (+/-1.89, n=5); pp512 778.80 / pp2048 1115.99 / pp4096 1137.02. Load at -c 131072 verified, 10.2 GiB VRAM free after load. KV q8_0 K + q4_1 V, FA on. Correctness: coherent smoke 5/5, deterministic replay pass, token exactness pass; reference parity not run. Cross-cap note: at matched 150W cap this model runs tg128 mean 17.78 (n=5, separate submission); prior Qwen3.6-27B Q4_K_M ran tg128 21.3 mean (n=3, 150W, Run 9) -- comparisons across caps are directional only. Dense power scaling 150->230W measured +25% decode on this model.
Qwen3.8-27B (dense 27B, hybrid GDN linear+full attention, MTP head not in GGUF quant). Community quant unsloth/Qwen3.8-27B-GGUF Q4_K_M (SHA-256 7b2aec3b...cc89f1b) -- weights released 2026-08-14T15:00Z, benched same day. 230W cap (stock). llama-bench tg128 mean 22.30 t/s (+/-0.01, n=5, C1, cold cache). Cold input pp8192 mean 1121.40 t/s (+/-1.89, n=5); pp512 778.80 / pp2048 1115.99 / pp4096 1137.02. Load at -c 131072 verified, 10.2 GiB VRAM free after load. KV q8_0 K + q4_1 V, FA on. Correctness: coherent smoke 5/5, deterministic replay pass, token exactness pass; reference parity not run. Cross-cap note: at matched 150W cap this model runs tg128 mean 17.78 (n=5, separate submission); prior Qwen3.6-27B Q4_K_M ran tg128 21.3 mean (n=3, 150W, Run 9) -- comparisons across caps are directional only. Dense power scaling 150->230W measured +25% decode on this model.
Qwen3.8-27B (dense 27B, hybrid GDN linear+full attention, MTP head not in GGUF quant). Community quant unsloth/Qwen3.8-27B-GGUF Q4_K_M (SHA-256 7b2aec3b...cc89f1b) -- weights released 2026-08-14T15:00Z, benched same day. llama-bench tg128 mean 17.78 t/s (+/-0.17, n=5, C1, cold cache). Cold input pp8192 mean 787.45 t/s (+/-1.74, n=5). Load at -c 131072 verified, 10.2 GiB VRAM free after load. KV q8_0 K + q4_1 V, FA on. Correctness: coherent smoke 5/5, deterministic replay pass, token exactness pass; reference parity not run. At matched 150W/build/quant/flags the prior Qwen3.6-27B Q4_K_M ran tg128 21.3 mean (n=3, Run 9); 3.8 mean 17.78 (n=5) -- ~16% slower decode. Comparison directional: different checkpoints, n mismatch (3 vs 5).
Self-reported E2, not independently reproduced. C1 (max_num_seqs=1), prefix cache explicitly off (--no-enable-prefix-caching; enable_prefix_caching=False; zero prefix_cache_hits delta). Timing is client monotonic SSE: tokSOut is median client post-first decode at exact p2048/g128 n=5 = 186.61 t/s (range 174.60-201.83; families research/rag/tool/document/assistant). Additional decode cells on the same warm server: p512/g128 median 194.61 (range 140.20-220.01, acceptance 45.1% — wide family spread, not the representative scalar); p8192/g128 median 157.92 (143.50-170.25, acceptance 53.0%). tokSPrefill is the p8192/g1 n=5 median COLD INPUT RATE from client TTFT (7160 t/s, range 7117-7226), which includes scheduling and first-token work — NOT isolated engine prefill. Also p2048/g1 cold input median 6456 t/s. Spec: method=dflash n_spec=7; window acceptance 1830/3521 = 52.0%. Target is a LOCAL symmetric GPTQ INT4 G64 conversion of the named BF16 repo (not an official HF quant). Draft is a LOCAL NVFP4 E2M1→BF16 reconstruction of NVIDIA DFlash, not a published BF16 draft. Stack: vllm/vllm-openai-xpu@sha256:1da0a95485455f08588c11080b9718992fd7d434c6a965d74654903a9d999c57 plus local det image patches (native grouped-topk v2, SSU B8/W4, at::zeros grouped-GEMM); VLLM_XPU_ENABLE_XPU_GRAPH=1; PIECEWISE+FULL graphs; --async-scheduling; --quantization gptq --dtype float16 --max-model-len 16384 --gpu-memory-utilization 0.90. Configured cap 150 W; measured cell-window averages ~149-160 W; peak interval-average 179.3 W (0.5 s energy1_input on p8192/g128); pkg max 68.0 C. Deterministic raw-completion replay smoke exact_match=true (scope: smoke only, not logit/KL or task-quality). n=3 screen 20260813T080234Z is superseded and was not used. Matched-except-speculation vs prior no-spec n=5 graph campaign (p8192/g128 87.25 t/s): 1.81× on that cell only.
Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). tokSPrefill is the cold input rate at exact p8192/g128 from median client TTFT (8,368 tok/s, TTFT median 0.9789s, n=3) - a client-observed input rate including scheduling and first-token work, NOT an isolated engine prefill measurement. NO speculative decoding in this config: native MTP was tested and rejected on this stack (0% draft acceptance), N-gram rejected (temperature-0 nondeterminism), DFlash not supported for this model on this stack. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent.