Canonical single-B70 llama-bench result at a verified 150 W power cap. Independent PP4096 and TG128 phases; five measured repetitions per phase. PP4096 samples=[1593.73, 1592.28, 1582.49, 1589.57, 1590.76]; TG128 samples=[78.1363, 78.8193, 78.8413, 78.8379, 78.8436]. PP CV=0.274%; TG CV=0.398%. K cache q8_0, V cache q4_1; full GPU offload; Flash Attention on; no CPU MoE offload; no speculative decoding. Build: llama.cpp b10053 0dc74e332 plus upstream PR #25690 0bd0ec609. Actual benchmark host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04, 150 W cap. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw artifacts checksum-verified.
Canonical single-B70 llama-bench result at a verified 150 W power cap. Independent PP4096 and TG128 phases; five measured repetitions per phase. PP4096 samples=[1605.62, 1607.22, 1594.85, 1597.68, 1601.33]; TG128 samples=[66.7979, 67.0723, 67.016, 67.0435, 67.0143]. PP CV=0.325%; TG CV=0.163%. K cache q8_0, V cache q4_1; full GPU offload; Flash Attention on; no CPU MoE offload; no speculative decoding. Build: llama.cpp b10053 0dc74e332 plus upstream PR #25690 0bd0ec609. Actual benchmark host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04, 150 W cap. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw artifacts checksum-verified.
Canonical single-B70 llama-bench result at a verified 150 W power cap. Independent PP4096 and TG128 phases; five measured repetitions per phase. PP4096 samples=[1605.32, 1603.1, 1595.12, 1604.54, 1609.38]; TG128 samples=[69.5063, 69.7579, 69.7894, 69.7738, 69.643]. PP CV=0.326%; TG CV=0.172%. K cache q8_0, V cache q4_1; full GPU offload; Flash Attention on; no CPU MoE offload; no speculative decoding. Build: llama.cpp b10053 0dc74e332 plus upstream PR #25690 0bd0ec609. Actual benchmark host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04, 150 W cap. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw artifacts checksum-verified.
PEAK SINGLE-STREAM PREFILL TEST: Measured using a contiguous 3000-token prompt block without concurrency (batch=1). Uncovers the raw 1,242 tok/s prefill capacity of the B70 without the massive context-switching overhead incurred in fleet concurrency. Hardware capped at optimal 165W/2400MHz.
POWER/FREQUENCY UNLOCKED TEST: max_freq pushed to 2800 MHz, power1_cap relaxed to 230W. Telemetry logged 196W sustained draw at 73C. Memory bandwidth bottleneck confirmed: +33% power yielded only +8% throughput over the 165W/2400MHz baseline.
q5_0-q4_1 asymmetric KV cache (92.65% tail precision, 38% less VRAM than q8_0). 150W power cap — MoE model, compute-bound. 256K context with q5_0-q4_1 KV. Agent-optimized model at massive context window. Intel Arc Pro B70 32GB. llama.cpp SYCL b9851. oneAPI 2026.0.0. GPU temp: 67C under load. Zero context scaling penalty vs 128K.
q5_0-q4_1 asymmetric KV cache (92.65% tail precision, 38% less VRAM than q8_0). 150W power cap — MoE model is compute-bound, higher power provides negligible benefit. 256K context at 150W with q5_0-q4_1 KV (would need q8_0-q8_0 for same context, costing +38% VRAM). Intel Arc Pro B70 32GB. llama.cpp SYCL b9851. oneAPI 2026.0.0. GPU temp: 67C under load. Zero context scaling penalty vs 128K.
Canonical single-B70 llama-bench result at a verified 150 W power cap. Independent PP4096 and TG128 phases; five measured repetitions per phase. PP4096 samples=[1593.73, 1592.28, 1582.49, 1589.57, 1590.76]; TG128 samples=[78.1363, 78.8193, 78.8413, 78.8379, 78.8436]. PP CV=0.274%; TG CV=0.398%. K cache q8_0, V cache q4_1; full GPU offload; Flash Attention on; no CPU MoE offload; no speculative decoding. Build: llama.cpp b10053 0dc74e332 plus upstream PR #25690 0bd0ec609. Actual benchmark host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04, 150 W cap. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw artifacts checksum-verified.
Canonical single-B70 llama-bench result at a verified 150 W power cap. Independent PP4096 and TG128 phases; five measured repetitions per phase. PP4096 samples=[1605.62, 1607.22, 1594.85, 1597.68, 1601.33]; TG128 samples=[66.7979, 67.0723, 67.016, 67.0435, 67.0143]. PP CV=0.325%; TG CV=0.163%. K cache q8_0, V cache q4_1; full GPU offload; Flash Attention on; no CPU MoE offload; no speculative decoding. Build: llama.cpp b10053 0dc74e332 plus upstream PR #25690 0bd0ec609. Actual benchmark host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04, 150 W cap. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw artifacts checksum-verified.
Canonical single-B70 llama-bench result at a verified 150 W power cap. Independent PP4096 and TG128 phases; five measured repetitions per phase. PP4096 samples=[1605.32, 1603.1, 1595.12, 1604.54, 1609.38]; TG128 samples=[69.5063, 69.7579, 69.7894, 69.7738, 69.643]. PP CV=0.326%; TG CV=0.172%. K cache q8_0, V cache q4_1; full GPU offload; Flash Attention on; no CPU MoE offload; no speculative decoding. Build: llama.cpp b10053 0dc74e332 plus upstream PR #25690 0bd0ec609. Actual benchmark host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04, 150 W cap. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw artifacts checksum-verified.
PEAK SINGLE-STREAM PREFILL TEST: Measured using a contiguous 3000-token prompt block without concurrency (batch=1). Uncovers the raw 1,242 tok/s prefill capacity of the B70 without the massive context-switching overhead incurred in fleet concurrency. Hardware capped at optimal 165W/2400MHz.
POWER/FREQUENCY UNLOCKED TEST: max_freq pushed to 2800 MHz, power1_cap relaxed to 230W. Telemetry logged 196W sustained draw at 73C. Memory bandwidth bottleneck confirmed: +33% power yielded only +8% throughput over the 165W/2400MHz baseline.
q5_0-q4_1 asymmetric KV cache (92.65% tail precision, 38% less VRAM than q8_0). 150W power cap — MoE model, compute-bound. 256K context with q5_0-q4_1 KV. Agent-optimized model at massive context window. Intel Arc Pro B70 32GB. llama.cpp SYCL b9851. oneAPI 2026.0.0. GPU temp: 67C under load. Zero context scaling penalty vs 128K.
q5_0-q4_1 asymmetric KV cache (92.65% tail precision, 38% less VRAM than q8_0). 150W power cap — MoE model is compute-bound, higher power provides negligible benefit. 256K context at 150W with q5_0-q4_1 KV (would need q8_0-q8_0 for same context, costing +38% VRAM). Intel Arc Pro B70 32GB. llama.cpp SYCL b9851. oneAPI 2026.0.0. GPU temp: 67C under load. Zero context scaling penalty vs 128K.