Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). Prefill field is cold input rate from TTFT (~4,900 tok/s at p512, ~10,400 tok/s at p8192), not isolated prefill. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent. Prefill field omitted: no isolated prefill measurement on this stack; TTFT-derived cold input rate is a different metric and is NOT reported as tokSPrefill.
Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). Prefill field is cold input rate from TTFT (~4,900 tok/s at p512, ~10,400 tok/s at p8192), not isolated prefill. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent.
Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). Prefill field is cold input rate from TTFT (~4,900 tok/s at p512, ~10,400 tok/s at p8192), not isolated prefill. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent.
Single-stream dense 27B. ThinkingCap-Qwen3.6-27B Q4_K_M GGUF. engine rate timings.predicted_per_second. q8_0 K + q4_1 V KV, FA on. 230W (dense scales with power; 79°C peak). vLLM FP8 blocked on this card (no XPU kernel, KeyError PlatformEnum.XPU). Measured 2026-08-06 Run 19.
Fresh single-B70 ThinkingCap benchmark on 2026-07-22. Independent phases with prompt cache disabled: 4,014-token prefill at temperature 0 and separate 350-token reasoning/decode task at temperature 0.2. Two warm-ups plus five measured repetitions per phase. Prefill samples=[622.0612,621.6988,621.3770,620.9613,620.5751], CV=0.094%. Decode samples=[26.6528,27.2787,27.2420,28.2763,28.0818], CV=2.421%. MTP-4 decode acceptance=1016/1154 (88.04%). Full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096; 200K context; 165 W. Build: llama.cpp b10053 0dc74e332 plus PR #25690 0bd0ec609. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw artifact SHA-256: 91eba0e2c34bf90b2d895b7d6e1819c763aadb9ed154dbb0da48b94b368d41d2.
Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 180 W. MTP-4 enabled; decode draft acceptance 225/297 (75.76%).
Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 165 W. MTP-4 enabled; decode draft acceptance 223/298 (74.83%).
Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 150 W. MTP disabled.
Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 165 W. MTP-4 enabled; decode draft acceptance 199/216 (92.13%).
Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). Prefill field is cold input rate from TTFT (~4,900 tok/s at p512, ~10,400 tok/s at p8192), not isolated prefill. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent. Prefill field omitted: no isolated prefill measurement on this stack; TTFT-derived cold input rate is a different metric and is NOT reported as tokSPrefill.
Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). Prefill field is cold input rate from TTFT (~4,900 tok/s at p512, ~10,400 tok/s at p8192), not isolated prefill. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent.
Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). Prefill field is cold input rate from TTFT (~4,900 tok/s at p512, ~10,400 tok/s at p8192), not isolated prefill. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent.
Single-stream dense 27B. ThinkingCap-Qwen3.6-27B Q4_K_M GGUF. engine rate timings.predicted_per_second. q8_0 K + q4_1 V KV, FA on. 230W (dense scales with power; 79°C peak). vLLM FP8 blocked on this card (no XPU kernel, KeyError PlatformEnum.XPU). Measured 2026-08-06 Run 19.
Fresh single-B70 ThinkingCap benchmark on 2026-07-22. Independent phases with prompt cache disabled: 4,014-token prefill at temperature 0 and separate 350-token reasoning/decode task at temperature 0.2. Two warm-ups plus five measured repetitions per phase. Prefill samples=[622.0612,621.6988,621.3770,620.9613,620.5751], CV=0.094%. Decode samples=[26.6528,27.2787,27.2420,28.2763,28.0818], CV=2.421%. MTP-4 decode acceptance=1016/1154 (88.04%). Full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096; 200K context; 165 W. Build: llama.cpp b10053 0dc74e332 plus PR #25690 0bd0ec609. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw artifact SHA-256: 91eba0e2c34bf90b2d895b7d6e1819c763aadb9ed154dbb0da48b94b368d41d2.
Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 180 W. MTP-4 enabled; decode draft acceptance 225/297 (75.76%).
Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 165 W. MTP-4 enabled; decode draft acceptance 223/298 (74.83%).
Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 150 W. MTP disabled.
Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 165 W. MTP-4 enabled; decode draft acceptance 199/216 (92.13%).