Get startedLeaderboardDecode calculatorModelsReportsHardwareBenchmarksMarketplaceRentalsProAPI Docs
Language
SergiioB

SergiioB

@SergiioB · Member since 2026

← Leaderboard
Mainly llama.cpp 74 approved runs

Submissions

74

74 approved public runs

Models

17

unique models

Avg output

91.8

Best 1139.8 tok/s

Avg prefill

2062

Best 10371.0 tok/s

Avg total

248.6

Best 2033.5 tok/s

Avg TTFT

4587.2ms

Best 49ms

Hardware

2

distinct setups

Reactions

2

across all runs

Personal bests

Personal bests17

Fastest approved result for each model.

3B MoE9 runs🔥
Output1139.8 tok/s
Prefill8715.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

vllm · GPTQ-Int4 · 3w ago

28B11 runs
Output224.2 tok/s
Prefill4.8 tok/s
TTFT15419ms
Depth
Intel Arc Pro B70 32GB

vllm · GPTQ-Int4 · 1w ago

3B MoE9 runs
Output186.6 tok/s
Prefill7160.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

vllm · GPTQ-INT4-G64-sym-local+DFlash-BF16-local · 2w ago

07 runs
Output164.5 tok/s
Prefill383.4 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q5_K_M · 1mo ago

01 run
Output145.9 tok/s
Prefill
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q5_K_M · 1mo ago

28B5 runs
Output112.7 tok/s
Prefill1696.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

vllm · GPTQ-Int4+draft-INT4-RTN · 1w ago

3B MoE3 runs
Output108.4 tok/s
Prefill9072.9 tok/s
TTFT320ms
Depth
Intel Arc Pro B70 32GB

vllm · GPTQ-Int4 · 1w ago

3B MoE1 run
Output87.4 tok/s
Prefill766.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q4_0 · 2w ago

35B MoE2 runs
Output73.2 tok/s
Prefill416.1 tok/s
TTFT77ms
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q4_K_M · 2mo ago

28B7 runs
Output69.3 tok/s
Prefill1754.6 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

vllm · GPTQ-Int4 · 3w ago

35B2 runs
Output62.8 tok/s
Prefill1682.4 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q4_K_XL · 1mo ago

02 runs
Output50.3 tok/s
Prefill553.9 tok/s
TTFT58ms
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q6_K · 2mo ago

28B2 runs
Output31.1 tok/s
Prefill389.6 tok/s
TTFT190ms
Depth
Intel Arc Pro B70 32GB ×2

vllm · FP8 · today

30B3 runs
Output29.2 tok/s
Prefill1301.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q4_K_XL · 2w ago

28B3 runs
Output27.9 tok/s
Prefill936.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q6_K · 3w ago

33B6 runs
Output26.6 tok/s
Prefill384.8 tok/s
TTFT
Depth0k
Intel Arc Pro B70 32GB

llama.cpp · Q4_K_M · 1mo ago

28B1 run
Output25.1 tok/s
Prefill613.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q5_K_M · 1mo ago

Hardware breakdown

Hardware breakdown2

Distinct rigs used across approved submissions.

Intel Arc Pro B70 32GB

72 runs · best 1139.8 tok/s

Intel Arc Pro B70 32GB ×2

2 runs · best 31.1 tok/s

Engines used

Engines used2

Runtime mix and top quantizations.

llama.cpp43Q4_K_M · Q4_0 · Q4_K_XL
vllm31FP8 · GPTQ-Int4 · GPTQ-Int4+draft-INT4-RTN
All runs

All runs74

Full approved benchmark history · page 2 of 3.

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-INT4-G64-sym-local

TTFT

Context

16k · 2w ago

Show all run details

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

93

Prefill tok/s

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU)

Quantization

GPTQ-INT4-G64-sym-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

default

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

0.95

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /mnt/models/nemotron-lightning-gptq-sym64-20260812 --host 0.0.0.0 --port 8001 --no-enable-prefix-caching --async-scheduling --gpu-memory-utilization 0.95 --max-model-len 16384 --enforce-eager false --trust-remote-code

Extra flags

--no-enable-prefix-caching --async-scheduling --max-model-len 16384 --enforce-eager false

Notes

Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). Prefill field is cold input rate from TTFT (~4,900 tok/s at p512, ~10,400 tok/s at p8192), not isolated prefill. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent. Prefill field omitted: no isolated prefill measurement on this stack; TTFT-derived cold input rate is a different metric and is NOT reported as tokSPrefill.

Reactions

Submitted

Aug 12, 2026, 9:06 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-INT4-G64-sym-local

TTFT

Context

16k · 2w ago

Show all run details

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

93

Prefill tok/s

4900

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU)

Quantization

GPTQ-INT4-G64-sym-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

default

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

0.95

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /mnt/models/nemotron-lightning-gptq-sym64-20260812 --host 0.0.0.0 --port 8001 --no-enable-prefix-caching --async-scheduling --gpu-memory-utilization 0.95 --max-model-len 16384 --enforce-eager false --trust-remote-code

Extra flags

--no-enable-prefix-caching --async-scheduling --max-model-len 16384 --enforce-eager false

Notes

Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). Prefill field is cold input rate from TTFT (~4,900 tok/s at p512, ~10,400 tok/s at p8192), not isolated prefill. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent.

Reactions

Submitted

Aug 12, 2026, 9:00 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-INT4-G64-sym-local

TTFT

Context

16k · 2w ago

Show all run details

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

93

Prefill tok/s

4900

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU)

Quantization

GPTQ-INT4-G64-sym-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

default

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

0.95

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /mnt/models/nemotron-lightning-gptq-sym64-20260812 --host 0.0.0.0 --port 8001 --no-enable-prefix-caching --async-scheduling --gpu-memory-utilization 0.95 --max-model-len 16384 --enforce-eager false --trust-remote-code

Extra flags

--no-enable-prefix-caching --async-scheduling --max-model-len 16384 --enforce-eager false

Notes

Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). Prefill field is cold input rate from TTFT (~4,900 tok/s at p512, ~10,400 tok/s at p8192), not isolated prefill. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent.

Reactions

Submitted

Aug 12, 2026, 8:59 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_0

TTFT

Context

16k · 2w ago

Show all run details

Model

bartowski/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF

Base model

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

87.4

Prefill tok/s

766

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

llama.cpp

Engine version

b306 SYCL (f785fc9ea) + custom Mamba2 kernel fusion + MMQ/XMX dispatch

Quantization

Q4_0

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Nemotron-3.5-Lightning-30B-A3B-Q4_0.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 16384 -b 8192 -ub 4096 -t 8 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 --dev SYCL0

Notes

Nemotron 3.5 Lightning 30B-A3B (hybrid Mamba2-Transformer MoE, 128 experts/7 active). Q4_0 quantization (17.44 GiB, requantized from Q4_K_M). Decode with native MTP speculative decoding (draft-mtp, n-max=4, 95-99% acceptance, n=5 median). No-spec decode: 68.0 t/s. Q4_0 delivers +78% over Q4_K_M (38->68 t/s no-spec) due to simpler GEMV dequantization. Measured power: 134.0W avg (cap 230W). GPU temp: ~64C. Custom kernel fusion: ssm_conv bias+SiLU, ssm_scan dt_bias+D*x, MMQ/XMX dispatch override. Build: llama.cpp SYCL f785fc9ea. KV: q8_0/q4_1, FA on.

Reactions

Submitted

Aug 12, 2026, 6:53 AM

Last edited

29.2

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_XL

TTFT

Context

131k · 2w ago

Show all run details

Model

meta-models/Muse-Glimmer-30B

Display name

Muse-Glimmer-30B

Base model

Revision

main

Family

Parameters

30B

Active params

MoE

no

Output tok/s

29.2

Prefill tok/s

1301

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

llama.cpp

Engine version

b284 SYCL (d2f83055d) GGML_SYCL_F16=ON

Quantization

Q4_K_XL

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

Spec model

dflash-draft.gguf

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Muse-Glimmer-30B-UD-Q4_K_XL.gguf -md dflash-draft.gguf --spec-type draft-dflash --spec-draft-n-max 2 -ngl 99 -ngld 99 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -b 8192 -ub 8192 -t 8 -dev SYCL0

Extra flags

--spec-type draft-dflash --spec-draft-n-max 2 --ngld 99 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0

Notes

Meta Muse-Glimmer-30B (dense 30B multimodal, vision+reasoning). Decode with DFlash speculative decoding (draft-dflash, n-max=2). p512/g128 median 29.2 t/s (max 31.8, n=5, Run 33). No-spec baseline: 25.3 t/s (Run 34). Measured at 230W cap. GGML_SYCL_F16=ON mandatory for this architecture (3.4x prefill loss without it). Build: llama.cpp SYCL d2f83055d (build 284). KV: q8_0 K + q4_1 V, FA on, -ngl 99 -ub 8192. Model: Muse-Glimmer-30B-UD-Q4_K_XL (15 GB). Multimodal verified (vision mmproj). Day-0 adoption: llama.cpp PR #26841 merged.

Reactions

Submitted

Aug 11, 2026, 9:55 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_0

TTFT

Context

16k · 2w ago

Show all run details

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

87.4

Prefill tok/s

766

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

llama.cpp

Engine version

b306 SYCL (f785fc9ea) + custom Mamba2 kernel fusion + MMQ/XMX dispatch

Quantization

Q4_0

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Nemotron-3.5-Lightning-30B-A3B-Q4_0.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 16384 -b 8192 -ub 4096 -t 8 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 --dev SYCL0

Notes

Nemotron 3.5 Lightning 30B-A3B (hybrid Mamba2-Transformer MoE, 128 experts/7 active). Q4_0 quantization (17.44 GiB, requantized from Q4_K_M). Decode with native MTP speculative decoding (draft-mtp, n-max=4, 95-99% acceptance, n=5 median). No-spec decode: 68.0 t/s. Q4_0 delivers +78% over Q4_K_M (38->68 t/s no-spec) due to simpler GEMV dequantization. Measured power: 134.0W avg (cap 230W). GPU temp: ~64C. Custom kernel fusion: ssm_conv bias+SiLU, ssm_scan dt_bias+D*x, MMQ/XMX dispatch override. Build: llama.cpp SYCL f785fc9ea. KV: q8_0/q4_1, FA on.

Reactions

Submitted

Aug 11, 2026, 9:55 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_M

TTFT

Context

16k · 2w ago

Show all run details

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

49.8

Prefill tok/s

1027

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

llama.cpp

Engine version

b306 SYCL (f785fc9ea) + custom Mamba2 kernel fusion + MMQ/XMX dispatch

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 16384 -b 8192 -ub 4096 -t 8 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 --dev SYCL0

Notes

Nemotron 3.5 Lightning 30B-A3B (hybrid Mamba2-Transformer MoE). Decode with native MTP speculative decoding (draft-mtp, n-max=4, 98.6% acceptance, n=5 median). No-spec decode: 38.2 t/s. Measured power draw: 140.3W avg (cap 230W). GPU temp: 63.9C avg. Custom kernel fusion: ssm_conv bias+SiLU fused, ssm_scan dt_bias+D*x fused, MMQ/XMX dispatch override for K-quant GEMV. Build: llama.cpp SYCL f785fc9ea. KV: q8_0 K + q4_1 V, FA on, -ngl 99 -ncmoe 0.

Reactions

Submitted

Aug 11, 2026, 8:04 PM

Last edited

26.8

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_XL

TTFT

Context

131k · 2w ago

Show all run details

Model

meta-models/Muse-Glimmer-30B

Display name

Muse-Glimmer-30B

Base model

Revision

main

Family

Parameters

30B

Active params

MoE

no

Output tok/s

26.8

Prefill tok/s

1301

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

llama.cpp

Engine version

d2f83055d (284) SYCL

Quantization

Q4_K_XL

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

Extra flags

Notes

Muse-Glimmer-30B UD-Q4_K_XL GGUF, llama.cpp SYCL d2f83055d (GGML_SYCL_F16=ON), single-stream C1. Decode = engine timings.predicted_per_second with DFlash draft n_max=2, p512/g128 median n=5, 128K ctx, 230W cap; no-spec baseline 22.5. Prefill = llama-bench pp4096. Reasoning model (decode includes reasoning tokens). E2 provisional self-report. 2026-08-10 (Run 33-35).

Reactions

Submitted

Aug 10, 2026, 7:11 PM

Last edited

26.8

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_XL

TTFT

Context

131k · 2w ago

Show all run details

Model

meta-models/Muse-Glimmer-30B

Display name

Muse-Glimmer-30B

Base model

Revision

main

Family

Parameters

30B

Active params

MoE

no

Output tok/s

26.8

Prefill tok/s

1301

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

llama.cpp

Engine version

d2f83055d (284) SYCL

Quantization

Q4_K_XL

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

Extra flags

Notes

Muse-Glimmer-30B UD-Q4_K_XL GGUF, llama.cpp SYCL d2f83055d (GGML_SYCL_F16=ON), single-stream C1. Decode = engine timings.predicted_per_second with DFlash draft n_max=2, p512/g128 median n=5, 128K ctx, 230W cap; no-spec baseline 22.5. Prefill = llama-bench pp4096. Reasoning model (decode includes reasoning tokens). E2 provisional self-report. 2026-08-10 (Run 33-35).

Reactions

Submitted

Aug 10, 2026, 7:11 PM

Last edited

Qwen3.6-27B

28B · Qwen

69.3

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

Context

131k · 3w ago

Show all run details

Model

Qwen/Qwen3.6-27B

Display name

Qwen3.6-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

69.3

Prefill tok/s

1754.6

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

0.26.1rc1.dev457+gc810e5ee9.xpu

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.88

Max running seqs

64

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 131072 --gpu-memory-utilization 0.88 --kv-cache-dtype fp8 --max-num-seqs 64 --max-num-batched-tokens 8192 --enable-prefix-caching --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":4}

Extra flags

--max-model-len 131072 --language-model-only --speculative-config {method:mtp,num_speculative_tokens:4}

Notes

Single-stream C1. vLLM XPU pinned nightly (vllm/vllm-openai-xpu@sha256:2c427ef477da092eb6f2cdbbbd24950b5fa171565b916db69d4c7bb10e68ca97), vllm-xpu-kernels 0.1.12. Model: llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GPTQ-Int4 (dense 27B, GPTQ-INT4 g128 sym, preserved BF16 MTP head). Patches: patch_mtp_nightly.py + patch_mtp_boundary.py. fp8 KV cache (REQUIRED for dense 128K; fp16 does not fit). MTP4 speculative decode, 1-layer draft. gpu-memory-utilization 0.88, scheduler 8192, max-num-seqs 64, prefix cache on with zero hit delta. tokSOut = client post-first rate at p512/g128, median n=5 (not engine-native decode); tokSPrefill = actual input tokens / client TTFT at p2048/g1, median n=5 (not llama-bench pp). 230W configured cap. E2 provisional self-reported; independent reproduction pending. Measured 2026-08-09 Run 31.

Reactions

Submitted

Aug 9, 2026, 10:57 PM

Last edited

Qwen3.6-27B

28B · Qwen

69.3

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

Context

131k · 3w ago

Show all run details

Model

Qwen/Qwen3.6-27B

Display name

Qwen3.6-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

69.3

Prefill tok/s

1754.6

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

0.26.1rc1.dev457+gc810e5ee9.xpu

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.88

Max running seqs

64

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 131072 --gpu-memory-utilization 0.88 --kv-cache-dtype fp8 --max-num-seqs 64 --max-num-batched-tokens 8192 --enable-prefix-caching --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":4}

Extra flags

--max-model-len 131072 --language-model-only --speculative-config {method:mtp,num_speculative_tokens:4}

Notes

Single-stream C1. vLLM XPU pinned nightly (vllm/vllm-openai-xpu@sha256:2c427ef477da092eb6f2cdbbbd24950b5fa171565b916db69d4c7bb10e68ca97), vllm-xpu-kernels 0.1.12. Model: llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GPTQ-Int4 (dense 27B, GPTQ-INT4 g128 sym, preserved BF16 MTP head). Patches: patch_mtp_nightly.py + patch_mtp_boundary.py. fp8 KV cache (REQUIRED for dense 128K; fp16 does not fit). MTP4 speculative decode, 1-layer draft. gpu-memory-utilization 0.88, scheduler 8192, max-num-seqs 64, prefix cache on with zero hit delta. tokSOut = client post-first rate at p512/g128, median n=5 (not engine-native decode); tokSPrefill = actual input tokens / client TTFT at p2048/g1, median n=5 (not llama-bench pp). 230W configured cap. E2 provisional self-reported; independent reproduction pending. Measured 2026-08-09 Run 31.

Reactions

Submitted

Aug 9, 2026, 10:57 PM

Last edited

Qwen3.6-27B

28B · Qwen

69.3

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

Context

131k · 3w ago

Show all run details

Model

Qwen/Qwen3.6-27B

Display name

Qwen3.6-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

69.3

Prefill tok/s

1754.6

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

0.26.1rc1.dev457+gc810e5ee9.xpu

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.88

Max running seqs

64

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 131072 --gpu-memory-utilization 0.88 --kv-cache-dtype fp8 --max-num-seqs 64 --max-num-batched-tokens 8192 --enable-prefix-caching --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":4}

Extra flags

--max-model-len 131072 --language-model-only --speculative-config {method:mtp,num_speculative_tokens:4}

Notes

Single-stream C1. vLLM XPU pinned nightly (vllm/vllm-openai-xpu@sha256:2c427ef477da092eb6f2cdbbbd24950b5fa171565b916db69d4c7bb10e68ca97), vllm-xpu-kernels 0.1.12. Model: llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GPTQ-Int4 (dense 27B, GPTQ-INT4 g128 sym, preserved BF16 MTP head). Patches: patch_mtp_nightly.py + patch_mtp_boundary.py. fp8 KV cache (REQUIRED for dense 128K; fp16 does not fit). MTP4 speculative decode, 1-layer draft. gpu-memory-utilization 0.88, scheduler 8192, max-num-seqs 64, prefix cache on with zero hit delta. tokSOut = client post-first rate at p512/g128, median n=5 (not engine-native decode); tokSPrefill = actual input tokens / client TTFT at p2048/g1, median n=5 (not llama-bench pp). 230W configured cap. E2 provisional self-reported; independent reproduction pending. Measured 2026-08-09 Run 31.

Reactions

Submitted

Aug 9, 2026, 10:57 PM

Last edited

Qwen3.6-35B-A3B

3B MoE · Qwen

1139.8

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

Context

16k · 3w ago

🔥
Show all run details

Model

Qwen/Qwen3.6-35B-A3B

Display name

Qwen3.6-35B-A3B

Base model

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

1139.8

Prefill tok/s

8715

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

64

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

vllm

Engine version

0.21.1.dev18+gec426505f XPU

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

Max running seqs

64

Scheduler delay

Num parallel

64

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 16384 --max-num-seqs 64 --max-num-batched-tokens 8192 --served-model-name Qwen3.6-35B-A3B-MTP-Preserved-GPTQ-Int4 --language-model-only

Extra flags

--max-model-len 16384 --language-model-only

Notes

CONCURRENCY MAX: 64 concurrent users. Aggregate 1967.7 t/s wall-agg, generation 1139.8 t/s, diverse 512-token prompts, median TPOT 56ms. No-MTP native int4 v4 (MTP+concurrency blocked by GDN kernel). max-num-seqs 64. 165W. Multi-user aggregate throughput, not single-stream decode. 2026-08-07.

Reactions

🔥

Submitted

Aug 7, 2026, 12:19 PM

Last edited

Qwen3.6-35B-A3B

3B MoE · Qwen

204.6

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

Context

16k · 3w ago

Show all run details

Model

Qwen/Qwen3.6-35B-A3B

Display name

Qwen3.6-35B-A3B

Base model

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

204.6

Prefill tok/s

8715

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

vllm

Engine version

0.21.1.dev18+gec426505f XPU

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

Max running seqs

64

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 16384 --max-num-seqs 64 --max-num-batched-tokens 8192 --served-model-name Qwen3.6-35B-A3B-MTP-Preserved-GPTQ-Int4 --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":4}

Extra flags

--max-model-len 16384 --language-model-only --speculative-config {method:mtp,num_speculative_tokens:4}

Notes

Single-stream MTP4 num_spec=4. vLLM 0.21 XPU int4moe + 2 patches: native int4 int8 store, BF16 MTP draft. Recurrent single-layer MTP emits 4 draft tokens/step. short/32 decode 204.6 t/s (grid, prefix-cache-off, warmup discarded). Prefill p4k 8715 t/s @230W. Decode @165W. --max-num-batched-tokens 8192. Model: Qwen3.6-35B-A3B-MTP-Preserved-GPTQ-Int4. 2026-08-07.

Reactions

Submitted

Aug 7, 2026, 12:19 PM

Last edited

Qwen3.6-27B

28B · Qwen

23.0

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_M

TTFT

Context

16k · 3w ago

Show all run details

Model

Qwen/Qwen3.6-27B

Display name

Qwen3.6-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

23

Prefill tok/s

1007

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

1935

Output tokens

128

Prefill tokens

0

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b10255+ (071327508) SYCL

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf -ngl 99 -fa on -ctk q8_0 -ctv q4_1 -c 16384 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0

Extra flags

--ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0

Notes

Single-stream dense 27B. ThinkingCap-Qwen3.6-27B Q4_K_M GGUF. engine rate timings.predicted_per_second. q8_0 K + q4_1 V KV, FA on. 230W (dense scales with power; 79°C peak). vLLM FP8 blocked on this card (no XPU kernel, KeyError PlatformEnum.XPU). Measured 2026-08-06 Run 19.

Reactions

Submitted

Aug 6, 2026, 3:05 PM

Last edited

Qwen3.6-35B-A3B

3B MoE · Qwen

69.0

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_XL

TTFT

Context

16k · 3w ago

Show all run details

Model

Qwen/Qwen3.6-35B-A3B

Display name

Qwen3.6-35B-A3B

Base model

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

69

Prefill tok/s

1498

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

1935

Output tokens

128

Prefill tokens

0

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b10255+ (071327508) SYCL

Quantization

Q4_K_XL

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 16384 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0

Notes

Single-stream SYCL llama-server. engine rate timings.predicted_per_second (per AGENTS.md §9.4). Qwen3.6-35B-A3B-UD-Q4_K_XL GGUF. q8_0 K + q4_1 V KV, FA on, -ngl 99 -ncmoe 0. 150W (MoE self-limits to ~140W). MoE baseline for vLLM MTP comparison. Measured 2026-08-06 Run 19.

Reactions

Submitted

Aug 6, 2026, 3:05 PM

Last edited

Qwen3.6-35B-A3B

3B MoE · Qwen

132.9

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

Context

16k · 3w ago

Show all run details

Model

Qwen/Qwen3.6-35B-A3B

Display name

Qwen3.6-35B-A3B

Base model

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

132.9

Prefill tok/s

7535

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

55

Output tokens

32

Prefill tokens

0

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

vllm

Engine version

0.21.1.dev18+gec426505f XPU

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

auto

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.92

Max running seqs

1

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 16384 --gpu-memory-utilization 0.92 --max-num-seqs 1 --max-num-batched-tokens 8192 --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":1} --cudagraph-capture-sizes 1 2 4 8 16 32

Extra flags

--max-model-len 16384 --language-model-only --speculative-config {method:mtp,num_speculative_tokens:1} --cudagraph-capture-sizes 1

Notes

Single-stream. vLLM 0.21 XPU + 4 in-container patches (github.com/SergiioB/intel-arc-pro-b70-inference-cookbook): native int4 dtype (uint8->int8), BF16 MTP draft, XpuFusedMoe kwarg strip, GDN spec assert->warning. MTP speculative 1 layer num_spec=1. --max-num-batched-tokens 8192 clears chunking cap. Model: llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-GPTQ-Int4. FLASH_ATTN PIECEWISE graphs. 150W. tokSOut=short/g32 best; tokSPrefill=pp2048 (1945-token prompt). Byte-identical greedy replays verified; KL audit vs eager pending. Measured 2026-08-06 Run 20.

Reactions

Submitted

Aug 6, 2026, 3:05 PM

Last edited

27.9

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q6_K

TTFT

Context

131k · 3w ago

Show all run details

Model

bottlecapai/ThinkingCap-Qwen3.6-27B

Display name

ThinkingCap-Qwen3.6-27B

Base model

Qwen3.6-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

27.9

Prefill tok/s

936

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b10255+ SYCL (master 071327508)

Quantization

Q6_K

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

draft_mtp

Spec model

Spec draft model

Spec tokens

4

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m ThinkingCap-Qwen3.6-27B-Q6_K-MTP.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75

Notes

ThinkingCap-Qwen3.6-27B Q6_K-MTP with MTP-4 spec decoding (draft-mtp n-max 4). Decode = engine rate timings.predicted_per_second, 4 diverse prompts x 3 reps steady-state (reps 2-3), 200W, 128K ctx (MTP VRAM ceiling for dense). Prefill = llama-bench pp4096 (Run 9, same build). q8_0 K + q4_1 V, FA on.

Reactions

Submitted

Aug 5, 2026, 7:59 AM

Last edited

26.3

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q5_K_M

TTFT

Context

131k · 3w ago

Show all run details

Model

bottlecapai/ThinkingCap-Qwen3.6-27B

Display name

ThinkingCap-Qwen3.6-27B

Base model

Qwen3.6-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

26.3

Prefill tok/s

936

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b10255+ SYCL (master 071327508)

Quantization

Q5_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

draft_mtp

Spec model

Spec draft model

Spec tokens

4

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m ThinkingCap-Qwen3.6-27B-Q5_K_M-MTP.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75

Notes

ThinkingCap-Qwen3.6-27B Q5_K_M-MTP with MTP-4 spec decoding (draft-mtp n-max 4). Decode = engine rate timings.predicted_per_second, 4 diverse prompts x 3 reps steady-state (reps 2-3), 200W, 128K ctx (MTP VRAM ceiling for dense). Prefill = llama-bench pp4096 (Run 9, same build). q8_0 K + q4_1 V, FA on.

Reactions

Submitted

Aug 5, 2026, 7:58 AM

Last edited

Qwen3.6-35B-A3B

3B MoE · Qwen

72.6

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_XL

TTFT

Context

33k · 3w ago

Show all run details

Model

Qwen/Qwen3.6-35B-A3B

Display name

Qwen3.6-35B-A3B

Base model

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

72.6

Prefill tok/s

2128

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

32768

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b10255+ SYCL (master 071327508)

Quantization

Q4_K_XL

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-bench -m MODEL.gguf -ngl 99 -ncmoe 0 -fa 1 -ctk q8_0 -ctv q4_1 -p 512,4096,8192,32768 -n 128 -r 3 -b 8192 -ub 4096 -t 8 -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --p 512,4096,8192,32768 --r 3 --b 8192 --dev SYCL0

Notes

llama-bench tg128 + pp4096 (Run 9). q8_0 K + q4_1 V, FA on, -ngl 99 -ncmoe 0, 150W. Master 0804 build (#25874 quantized-KV XMX FA). Ctx = bench ctx (max prompt 32K).

Reactions

Submitted

Aug 5, 2026, 6:49 AM

Last edited

27.5

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_M

TTFT

Context

205k · 1mo ago

Show all run details

Model

bottlecapai/ThinkingCap-Qwen3.6-27B

Display name

ThinkingCap-Qwen3.6-27B

Base model

Qwen3.6-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

27.5

Prefill tok/s

621.3

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

4014

Output tokens

350

Prefill tokens

0

Context length

204800

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b10053+PR25690

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

KV cache size

Prefix caching

no

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

mtp

Spec model

Spec draft model

Spec tokens

4

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

yes

MTP draft layers

Temperature

0.2

Top P

0.95

Top K

20

Min P

0

Repeat penalty

Mirostat

Command

llama-server -m ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf -dev SYCL0 --parallel 1 -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 204800 -t 8 -b 8192 -ub 4096 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75

Extra flags

SYCL/Level Zero; asymmetric KV K=q8_0,V=q4_1; prefill temperature=0; reasoning enabled

Notes

Fresh single-B70 ThinkingCap benchmark on 2026-07-22. Independent phases with prompt cache disabled: 4,014-token prefill at temperature 0 and separate 350-token reasoning/decode task at temperature 0.2. Two warm-ups plus five measured repetitions per phase. Prefill samples=[622.0612,621.6988,621.3770,620.9613,620.5751], CV=0.094%. Decode samples=[26.6528,27.2787,27.2420,28.2763,28.0818], CV=2.421%. MTP-4 decode acceptance=1016/1154 (88.04%). Full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096; 200K context; 165 W. Build: llama.cpp b10053 0dc74e332 plus PR #25690 0bd0ec609. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw artifact SHA-256: 91eba0e2c34bf90b2d895b7d6e1819c763aadb9ed154dbb0da48b94b368d41d2.

Reactions

Submitted

Jul 22, 2026, 8:24 PM

Last edited

gemma-4-31B-it

33B · Gemma

26.6

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_M

TTFT

Context

131k · 1mo ago

Show all run details

Model

google/gemma-4-31B-it

Display name

gemma-4-31B-it

Base model

gemma-4-31B

Revision

main

Family

Gemma

Parameters

33B

Active params

MoE

no

Output tok/s

26.6

Prefill tok/s

384.8

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

4021

Output tokens

350

Prefill tokens

0

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b9853

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

mtp

Spec model

Spec draft model

mtp-gemma-4-31B-it.gguf

Spec tokens

4

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

yes

MTP draft layers

Temperature

0.2

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m gemma-4-31B-it-Q4_K_M.gguf -dev SYCL0 --parallel 1 -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -t 8 -b 8192 -ub 4096 --spec-type draft-mtp --spec-draft-model mtp-gemma-4-31B-it.gguf --spec-draft-n-max 4 --spec-draft-p-min 0.75

Extra flags

SYCL/Level Zero; asymmetric KV K=q8_0,V=q4_1; prefill temperature=0

Notes

Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 180 W. MTP-4 enabled; decode draft acceptance 225/297 (75.76%).

Reactions

Submitted

Jul 22, 2026, 10:55 AM

Last edited

gemma-4-31B-it

33B · Gemma

24.8

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_M

TTFT

Context

131k · 1mo ago

Show all run details

Model

google/gemma-4-31B-it

Display name

gemma-4-31B-it

Base model

gemma-4-31B

Revision

main

Family

Gemma

Parameters

33B

Active params

MoE

no

Output tok/s

24.8

Prefill tok/s

361.2

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

4021

Output tokens

350

Prefill tokens

0

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b9853

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

mtp

Spec model

Spec draft model

mtp-gemma-4-31B-it.gguf

Spec tokens

4

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

yes

MTP draft layers

Temperature

0.2

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m gemma-4-31B-it-Q4_K_M.gguf -dev SYCL0 --parallel 1 -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -t 8 -b 8192 -ub 4096 --spec-type draft-mtp --spec-draft-model mtp-gemma-4-31B-it.gguf --spec-draft-n-max 4 --spec-draft-p-min 0.75

Extra flags

SYCL/Level Zero; asymmetric KV K=q8_0,V=q4_1; prefill temperature=0

Notes

Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 165 W. MTP-4 enabled; decode draft acceptance 223/298 (74.83%).

Reactions

Submitted

Jul 22, 2026, 10:50 AM

Last edited

gemma-4-31B-it

33B · Gemma

16.4

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_M

TTFT

Context

131k · 1mo ago

Show all run details

Model

google/gemma-4-31B-it

Display name

gemma-4-31B-it

Base model

gemma-4-31B

Revision

main

Family

Gemma

Parameters

33B

Active params

MoE

no

Output tok/s

16.4

Prefill tok/s

332.9

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

4021

Output tokens

350

Prefill tokens

0

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b9853

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

0.2

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m gemma-4-31B-it-Q4_K_M.gguf -dev SYCL0 --parallel 1 -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -t 8 -b 8192 -ub 4096

Extra flags

SYCL/Level Zero; asymmetric KV K=q8_0,V=q4_1; prefill temperature=0

Notes

Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 150 W. MTP disabled.

Reactions

Submitted

Jul 22, 2026, 10:45 AM

Last edited

Qwen3.6-27B

28B · Qwen

25.1

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q5_K_M

TTFT

Context

205k · 1mo ago

Show all run details

Model

Qwen/Qwen3.6-27B

Display name

Qwen3.6-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

25.1

Prefill tok/s

613.0

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

4014

Output tokens

350

Prefill tokens

0

Context length

204800

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b9853

Quantization

Q5_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

mtp

Spec model

Spec draft model

Spec tokens

4

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

yes

MTP draft layers

Temperature

0.2

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Qwen3.6-27B-MTP-Q5_K_M.gguf -dev SYCL0 --parallel 1 -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 204800 -t 8 -b 8192 -ub 4096 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75

Extra flags

SYCL/Level Zero; asymmetric KV K=q8_0,V=q4_1; prefill temperature=0

Notes

Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 165 W. MTP-4 enabled; decode draft acceptance 199/216 (92.13%).

Reactions

Submitted

Jul 22, 2026, 10:40 AM

Last edited

ModelHardwareEnginetok/s outprefilltok/s totalTTFTDepthShare
NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

3B MoE

Intel Arc Pro B70 32GB
vllmGPTQ-INT4-G64-sym-local93.0
Show all details for NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

93

Prefill tok/s

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU)

Quantization

GPTQ-INT4-G64-sym-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

default

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

0.95

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /mnt/models/nemotron-lightning-gptq-sym64-20260812 --host 0.0.0.0 --port 8001 --no-enable-prefix-caching --async-scheduling --gpu-memory-utilization 0.95 --max-model-len 16384 --enforce-eager false --trust-remote-code

Extra flags

--no-enable-prefix-caching --async-scheduling --max-model-len 16384 --enforce-eager false

Notes

Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). Prefill field is cold input rate from TTFT (~4,900 tok/s at p512, ~10,400 tok/s at p8192), not isolated prefill. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent. Prefill field omitted: no isolated prefill measurement on this stack; TTFT-derived cold input rate is a different metric and is NOT reported as tokSPrefill.

Reactions

Submitted

Aug 12, 2026, 9:06 PM

Last edited

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

3B MoE

Intel Arc Pro B70 32GB
vllmGPTQ-INT4-G64-sym-local93.04900.0
Show all details for NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

93

Prefill tok/s

4900

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU)

Quantization

GPTQ-INT4-G64-sym-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

default

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

0.95

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /mnt/models/nemotron-lightning-gptq-sym64-20260812 --host 0.0.0.0 --port 8001 --no-enable-prefix-caching --async-scheduling --gpu-memory-utilization 0.95 --max-model-len 16384 --enforce-eager false --trust-remote-code

Extra flags

--no-enable-prefix-caching --async-scheduling --max-model-len 16384 --enforce-eager false

Notes

Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). Prefill field is cold input rate from TTFT (~4,900 tok/s at p512, ~10,400 tok/s at p8192), not isolated prefill. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent.

Reactions

Submitted

Aug 12, 2026, 9:00 PM

Last edited

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

3B MoE

Intel Arc Pro B70 32GB
vllmGPTQ-INT4-G64-sym-local93.04900.0
Show all details for NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

93

Prefill tok/s

4900

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU)

Quantization

GPTQ-INT4-G64-sym-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

default

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

0.95

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /mnt/models/nemotron-lightning-gptq-sym64-20260812 --host 0.0.0.0 --port 8001 --no-enable-prefix-caching --async-scheduling --gpu-memory-utilization 0.95 --max-model-len 16384 --enforce-eager false --trust-remote-code

Extra flags

--no-enable-prefix-caching --async-scheduling --max-model-len 16384 --enforce-eager false

Notes

Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). Prefill field is cold input rate from TTFT (~4,900 tok/s at p512, ~10,400 tok/s at p8192), not isolated prefill. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent.

Reactions

Submitted

Aug 12, 2026, 8:59 PM

Last edited

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF

3B MoE

Intel Arc Pro B70 32GB
llama.cppQ4_087.4766.0
Show all details for NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF

Model

bartowski/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-GGUF

Base model

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

87.4

Prefill tok/s

766

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

llama.cpp

Engine version

b306 SYCL (f785fc9ea) + custom Mamba2 kernel fusion + MMQ/XMX dispatch

Quantization

Q4_0

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Nemotron-3.5-Lightning-30B-A3B-Q4_0.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 16384 -b 8192 -ub 4096 -t 8 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 --dev SYCL0

Notes

Nemotron 3.5 Lightning 30B-A3B (hybrid Mamba2-Transformer MoE, 128 experts/7 active). Q4_0 quantization (17.44 GiB, requantized from Q4_K_M). Decode with native MTP speculative decoding (draft-mtp, n-max=4, 95-99% acceptance, n=5 median). No-spec decode: 68.0 t/s. Q4_0 delivers +78% over Q4_K_M (38->68 t/s no-spec) due to simpler GEMV dequantization. Measured power: 134.0W avg (cap 230W). GPU temp: ~64C. Custom kernel fusion: ssm_conv bias+SiLU, ssm_scan dt_bias+D*x, MMQ/XMX dispatch override. Build: llama.cpp SYCL f785fc9ea. KV: q8_0/q4_1, FA on.

Reactions

Submitted

Aug 12, 2026, 6:53 AM

Last edited

Muse-Glimmer-30B

30B

Intel Arc Pro B70 32GB
llama.cppQ4_K_XL29.21301.0
Show all details for Muse-Glimmer-30B

Model

meta-models/Muse-Glimmer-30B

Display name

Muse-Glimmer-30B

Base model

Revision

main

Family

Parameters

30B

Active params

MoE

no

Output tok/s

29.2

Prefill tok/s

1301

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

llama.cpp

Engine version

b284 SYCL (d2f83055d) GGML_SYCL_F16=ON

Quantization

Q4_K_XL

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

Spec model

dflash-draft.gguf

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Muse-Glimmer-30B-UD-Q4_K_XL.gguf -md dflash-draft.gguf --spec-type draft-dflash --spec-draft-n-max 2 -ngl 99 -ngld 99 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -b 8192 -ub 8192 -t 8 -dev SYCL0

Extra flags

--spec-type draft-dflash --spec-draft-n-max 2 --ngld 99 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0

Notes

Meta Muse-Glimmer-30B (dense 30B multimodal, vision+reasoning). Decode with DFlash speculative decoding (draft-dflash, n-max=2). p512/g128 median 29.2 t/s (max 31.8, n=5, Run 33). No-spec baseline: 25.3 t/s (Run 34). Measured at 230W cap. GGML_SYCL_F16=ON mandatory for this architecture (3.4x prefill loss without it). Build: llama.cpp SYCL d2f83055d (build 284). KV: q8_0 K + q4_1 V, FA on, -ngl 99 -ub 8192. Model: Muse-Glimmer-30B-UD-Q4_K_XL (15 GB). Multimodal verified (vision mmproj). Day-0 adoption: llama.cpp PR #26841 merged.

Reactions

Submitted

Aug 11, 2026, 9:55 PM

Last edited

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

3B MoE

Intel Arc Pro B70 32GB
llama.cppQ4_087.4766.0
Show all details for NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

87.4

Prefill tok/s

766

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

llama.cpp

Engine version

b306 SYCL (f785fc9ea) + custom Mamba2 kernel fusion + MMQ/XMX dispatch

Quantization

Q4_0

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Nemotron-3.5-Lightning-30B-A3B-Q4_0.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 16384 -b 8192 -ub 4096 -t 8 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 --dev SYCL0

Notes

Nemotron 3.5 Lightning 30B-A3B (hybrid Mamba2-Transformer MoE, 128 experts/7 active). Q4_0 quantization (17.44 GiB, requantized from Q4_K_M). Decode with native MTP speculative decoding (draft-mtp, n-max=4, 95-99% acceptance, n=5 median). No-spec decode: 68.0 t/s. Q4_0 delivers +78% over Q4_K_M (38->68 t/s no-spec) due to simpler GEMV dequantization. Measured power: 134.0W avg (cap 230W). GPU temp: ~64C. Custom kernel fusion: ssm_conv bias+SiLU, ssm_scan dt_bias+D*x, MMQ/XMX dispatch override. Build: llama.cpp SYCL f785fc9ea. KV: q8_0/q4_1, FA on.

Reactions

Submitted

Aug 11, 2026, 9:55 PM

Last edited

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

3B MoE

Intel Arc Pro B70 32GB
llama.cppQ4_K_M49.81027.0
Show all details for NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

49.8

Prefill tok/s

1027

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

llama.cpp

Engine version

b306 SYCL (f785fc9ea) + custom Mamba2 kernel fusion + MMQ/XMX dispatch

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m NVIDIA-Nemotron-3.5-Lightning-30B-A3B-Q4_K_M.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 16384 -b 8192 -ub 4096 -t 8 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75 --dev SYCL0

Notes

Nemotron 3.5 Lightning 30B-A3B (hybrid Mamba2-Transformer MoE). Decode with native MTP speculative decoding (draft-mtp, n-max=4, 98.6% acceptance, n=5 median). No-spec decode: 38.2 t/s. Measured power draw: 140.3W avg (cap 230W). GPU temp: 63.9C avg. Custom kernel fusion: ssm_conv bias+SiLU fused, ssm_scan dt_bias+D*x fused, MMQ/XMX dispatch override for K-quant GEMV. Build: llama.cpp SYCL f785fc9ea. KV: q8_0 K + q4_1 V, FA on, -ngl 99 -ncmoe 0.

Reactions

Submitted

Aug 11, 2026, 8:04 PM

Last edited

Muse-Glimmer-30B

30B

Intel Arc Pro B70 32GB
llama.cppQ4_K_XL26.81301.0
Show all details for Muse-Glimmer-30B

Model

meta-models/Muse-Glimmer-30B

Display name

Muse-Glimmer-30B

Base model

Revision

main

Family

Parameters

30B

Active params

MoE

no

Output tok/s

26.8

Prefill tok/s

1301

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

llama.cpp

Engine version

d2f83055d (284) SYCL

Quantization

Q4_K_XL

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

Extra flags

Notes

Muse-Glimmer-30B UD-Q4_K_XL GGUF, llama.cpp SYCL d2f83055d (GGML_SYCL_F16=ON), single-stream C1. Decode = engine timings.predicted_per_second with DFlash draft n_max=2, p512/g128 median n=5, 128K ctx, 230W cap; no-spec baseline 22.5. Prefill = llama-bench pp4096. Reasoning model (decode includes reasoning tokens). E2 provisional self-report. 2026-08-10 (Run 33-35).

Reactions

Submitted

Aug 10, 2026, 7:11 PM

Last edited

Muse-Glimmer-30B

30B

Intel Arc Pro B70 32GB
llama.cppQ4_K_XL26.81301.0
Show all details for Muse-Glimmer-30B

Model

meta-models/Muse-Glimmer-30B

Display name

Muse-Glimmer-30B

Base model

Revision

main

Family

Parameters

30B

Active params

MoE

no

Output tok/s

26.8

Prefill tok/s

1301

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

llama.cpp

Engine version

d2f83055d (284) SYCL

Quantization

Q4_K_XL

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

Extra flags

Notes

Muse-Glimmer-30B UD-Q4_K_XL GGUF, llama.cpp SYCL d2f83055d (GGML_SYCL_F16=ON), single-stream C1. Decode = engine timings.predicted_per_second with DFlash draft n_max=2, p512/g128 median n=5, 128K ctx, 230W cap; no-spec baseline 22.5. Prefill = llama-bench pp4096. Reasoning model (decode includes reasoning tokens). E2 provisional self-report. 2026-08-10 (Run 33-35).

Reactions

Submitted

Aug 10, 2026, 7:11 PM

Last edited

Qwen3.6-27B

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int469.31754.6
Show all details for Qwen3.6-27B

Model

Qwen/Qwen3.6-27B

Display name

Qwen3.6-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

69.3

Prefill tok/s

1754.6

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

0.26.1rc1.dev457+gc810e5ee9.xpu

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.88

Max running seqs

64

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 131072 --gpu-memory-utilization 0.88 --kv-cache-dtype fp8 --max-num-seqs 64 --max-num-batched-tokens 8192 --enable-prefix-caching --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":4}

Extra flags

--max-model-len 131072 --language-model-only --speculative-config {method:mtp,num_speculative_tokens:4}

Notes

Single-stream C1. vLLM XPU pinned nightly (vllm/vllm-openai-xpu@sha256:2c427ef477da092eb6f2cdbbbd24950b5fa171565b916db69d4c7bb10e68ca97), vllm-xpu-kernels 0.1.12. Model: llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GPTQ-Int4 (dense 27B, GPTQ-INT4 g128 sym, preserved BF16 MTP head). Patches: patch_mtp_nightly.py + patch_mtp_boundary.py. fp8 KV cache (REQUIRED for dense 128K; fp16 does not fit). MTP4 speculative decode, 1-layer draft. gpu-memory-utilization 0.88, scheduler 8192, max-num-seqs 64, prefix cache on with zero hit delta. tokSOut = client post-first rate at p512/g128, median n=5 (not engine-native decode); tokSPrefill = actual input tokens / client TTFT at p2048/g1, median n=5 (not llama-bench pp). 230W configured cap. E2 provisional self-reported; independent reproduction pending. Measured 2026-08-09 Run 31.

Reactions

Submitted

Aug 9, 2026, 10:57 PM

Last edited

Qwen3.6-27B

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int469.31754.6
Show all details for Qwen3.6-27B

Model

Qwen/Qwen3.6-27B

Display name

Qwen3.6-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

69.3

Prefill tok/s

1754.6

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

0.26.1rc1.dev457+gc810e5ee9.xpu

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.88

Max running seqs

64

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 131072 --gpu-memory-utilization 0.88 --kv-cache-dtype fp8 --max-num-seqs 64 --max-num-batched-tokens 8192 --enable-prefix-caching --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":4}

Extra flags

--max-model-len 131072 --language-model-only --speculative-config {method:mtp,num_speculative_tokens:4}

Notes

Single-stream C1. vLLM XPU pinned nightly (vllm/vllm-openai-xpu@sha256:2c427ef477da092eb6f2cdbbbd24950b5fa171565b916db69d4c7bb10e68ca97), vllm-xpu-kernels 0.1.12. Model: llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GPTQ-Int4 (dense 27B, GPTQ-INT4 g128 sym, preserved BF16 MTP head). Patches: patch_mtp_nightly.py + patch_mtp_boundary.py. fp8 KV cache (REQUIRED for dense 128K; fp16 does not fit). MTP4 speculative decode, 1-layer draft. gpu-memory-utilization 0.88, scheduler 8192, max-num-seqs 64, prefix cache on with zero hit delta. tokSOut = client post-first rate at p512/g128, median n=5 (not engine-native decode); tokSPrefill = actual input tokens / client TTFT at p2048/g1, median n=5 (not llama-bench pp). 230W configured cap. E2 provisional self-reported; independent reproduction pending. Measured 2026-08-09 Run 31.

Reactions

Submitted

Aug 9, 2026, 10:57 PM

Last edited

Qwen3.6-27B

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int469.31754.6
Show all details for Qwen3.6-27B

Model

Qwen/Qwen3.6-27B

Display name

Qwen3.6-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

69.3

Prefill tok/s

1754.6

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

0.26.1rc1.dev457+gc810e5ee9.xpu

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.88

Max running seqs

64

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 131072 --gpu-memory-utilization 0.88 --kv-cache-dtype fp8 --max-num-seqs 64 --max-num-batched-tokens 8192 --enable-prefix-caching --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":4}

Extra flags

--max-model-len 131072 --language-model-only --speculative-config {method:mtp,num_speculative_tokens:4}

Notes

Single-stream C1. vLLM XPU pinned nightly (vllm/vllm-openai-xpu@sha256:2c427ef477da092eb6f2cdbbbd24950b5fa171565b916db69d4c7bb10e68ca97), vllm-xpu-kernels 0.1.12. Model: llmfan46/Qwen3.6-27B-uncensored-heretic-v2-Native-MTP-Preserved-GPTQ-Int4 (dense 27B, GPTQ-INT4 g128 sym, preserved BF16 MTP head). Patches: patch_mtp_nightly.py + patch_mtp_boundary.py. fp8 KV cache (REQUIRED for dense 128K; fp16 does not fit). MTP4 speculative decode, 1-layer draft. gpu-memory-utilization 0.88, scheduler 8192, max-num-seqs 64, prefix cache on with zero hit delta. tokSOut = client post-first rate at p512/g128, median n=5 (not engine-native decode); tokSPrefill = actual input tokens / client TTFT at p2048/g1, median n=5 (not llama-bench pp). 230W configured cap. E2 provisional self-reported; independent reproduction pending. Measured 2026-08-09 Run 31.

Reactions

Submitted

Aug 9, 2026, 10:57 PM

Last edited

Qwen3.6-35B-A3B

3B MoE · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int41139.88715.0
Show all details for Qwen3.6-35B-A3B

Model

Qwen/Qwen3.6-35B-A3B

Display name

Qwen3.6-35B-A3B

Base model

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

1139.8

Prefill tok/s

8715

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

64

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

vllm

Engine version

0.21.1.dev18+gec426505f XPU

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

Max running seqs

64

Scheduler delay

Num parallel

64

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 16384 --max-num-seqs 64 --max-num-batched-tokens 8192 --served-model-name Qwen3.6-35B-A3B-MTP-Preserved-GPTQ-Int4 --language-model-only

Extra flags

--max-model-len 16384 --language-model-only

Notes

CONCURRENCY MAX: 64 concurrent users. Aggregate 1967.7 t/s wall-agg, generation 1139.8 t/s, diverse 512-token prompts, median TPOT 56ms. No-MTP native int4 v4 (MTP+concurrency blocked by GDN kernel). max-num-seqs 64. 165W. Multi-user aggregate throughput, not single-stream decode. 2026-08-07.

Reactions

🔥

Submitted

Aug 7, 2026, 12:19 PM

Last edited

Qwen3.6-35B-A3B

3B MoE · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4204.68715.0
Show all details for Qwen3.6-35B-A3B

Model

Qwen/Qwen3.6-35B-A3B

Display name

Qwen3.6-35B-A3B

Base model

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

204.6

Prefill tok/s

8715

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

vllm

Engine version

0.21.1.dev18+gec426505f XPU

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

Max running seqs

64

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 16384 --max-num-seqs 64 --max-num-batched-tokens 8192 --served-model-name Qwen3.6-35B-A3B-MTP-Preserved-GPTQ-Int4 --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":4}

Extra flags

--max-model-len 16384 --language-model-only --speculative-config {method:mtp,num_speculative_tokens:4}

Notes

Single-stream MTP4 num_spec=4. vLLM 0.21 XPU int4moe + 2 patches: native int4 int8 store, BF16 MTP draft. Recurrent single-layer MTP emits 4 draft tokens/step. short/32 decode 204.6 t/s (grid, prefix-cache-off, warmup discarded). Prefill p4k 8715 t/s @230W. Decode @165W. --max-num-batched-tokens 8192. Model: Qwen3.6-35B-A3B-MTP-Preserved-GPTQ-Int4. 2026-08-07.

Reactions

Submitted

Aug 7, 2026, 12:19 PM

Last edited

Qwen3.6-27B

28B · Qwen

Intel Arc Pro B70 32GB
llama.cppQ4_K_M23.01007.00k
Show all details for Qwen3.6-27B

Model

Qwen/Qwen3.6-27B

Display name

Qwen3.6-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

23

Prefill tok/s

1007

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

1935

Output tokens

128

Prefill tokens

0

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b10255+ (071327508) SYCL

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf -ngl 99 -fa on -ctk q8_0 -ctv q4_1 -c 16384 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0

Extra flags

--ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0

Notes

Single-stream dense 27B. ThinkingCap-Qwen3.6-27B Q4_K_M GGUF. engine rate timings.predicted_per_second. q8_0 K + q4_1 V KV, FA on. 230W (dense scales with power; 79°C peak). vLLM FP8 blocked on this card (no XPU kernel, KeyError PlatformEnum.XPU). Measured 2026-08-06 Run 19.

Reactions

Submitted

Aug 6, 2026, 3:05 PM

Last edited

Qwen3.6-35B-A3B

3B MoE · Qwen

Intel Arc Pro B70 32GB
llama.cppQ4_K_XL69.01498.00k
Show all details for Qwen3.6-35B-A3B

Model

Qwen/Qwen3.6-35B-A3B

Display name

Qwen3.6-35B-A3B

Base model

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

69

Prefill tok/s

1498

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

1935

Output tokens

128

Prefill tokens

0

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b10255+ (071327508) SYCL

Quantization

Q4_K_XL

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Qwen3.6-35B-A3B-UD-Q4_K_XL.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 16384 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0

Notes

Single-stream SYCL llama-server. engine rate timings.predicted_per_second (per AGENTS.md §9.4). Qwen3.6-35B-A3B-UD-Q4_K_XL GGUF. q8_0 K + q4_1 V KV, FA on, -ngl 99 -ncmoe 0. 150W (MoE self-limits to ~140W). MoE baseline for vLLM MTP comparison. Measured 2026-08-06 Run 19.

Reactions

Submitted

Aug 6, 2026, 3:05 PM

Last edited

Qwen3.6-35B-A3B

3B MoE · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4132.97535.00k
Show all details for Qwen3.6-35B-A3B

Model

Qwen/Qwen3.6-35B-A3B

Display name

Qwen3.6-35B-A3B

Base model

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

132.9

Prefill tok/s

7535

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

55

Output tokens

32

Prefill tokens

0

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

vllm

Engine version

0.21.1.dev18+gec426505f XPU

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

auto

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.92

Max running seqs

1

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 16384 --gpu-memory-utilization 0.92 --max-num-seqs 1 --max-num-batched-tokens 8192 --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":1} --cudagraph-capture-sizes 1 2 4 8 16 32

Extra flags

--max-model-len 16384 --language-model-only --speculative-config {method:mtp,num_speculative_tokens:1} --cudagraph-capture-sizes 1

Notes

Single-stream. vLLM 0.21 XPU + 4 in-container patches (github.com/SergiioB/intel-arc-pro-b70-inference-cookbook): native int4 dtype (uint8->int8), BF16 MTP draft, XpuFusedMoe kwarg strip, GDN spec assert->warning. MTP speculative 1 layer num_spec=1. --max-num-batched-tokens 8192 clears chunking cap. Model: llmfan46/Qwen3.6-35B-A3B-uncensored-heretic-Native-MTP-Preserved-GPTQ-Int4. FLASH_ATTN PIECEWISE graphs. 150W. tokSOut=short/g32 best; tokSPrefill=pp2048 (1945-token prompt). Byte-identical greedy replays verified; KL audit vs eager pending. Measured 2026-08-06 Run 20.

Reactions

Submitted

Aug 6, 2026, 3:05 PM

Last edited

ThinkingCap-Qwen3.6-27B

28B · Qwen

Intel Arc Pro B70 32GB
llama.cppQ6_K27.9936.0
Show all details for ThinkingCap-Qwen3.6-27B

Model

bottlecapai/ThinkingCap-Qwen3.6-27B

Display name

ThinkingCap-Qwen3.6-27B

Base model

Qwen3.6-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

27.9

Prefill tok/s

936

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b10255+ SYCL (master 071327508)

Quantization

Q6_K

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

draft_mtp

Spec model

Spec draft model

Spec tokens

4

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m ThinkingCap-Qwen3.6-27B-Q6_K-MTP.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75

Notes

ThinkingCap-Qwen3.6-27B Q6_K-MTP with MTP-4 spec decoding (draft-mtp n-max 4). Decode = engine rate timings.predicted_per_second, 4 diverse prompts x 3 reps steady-state (reps 2-3), 200W, 128K ctx (MTP VRAM ceiling for dense). Prefill = llama-bench pp4096 (Run 9, same build). q8_0 K + q4_1 V, FA on.

Reactions

Submitted

Aug 5, 2026, 7:59 AM

Last edited

ThinkingCap-Qwen3.6-27B

28B · Qwen

Intel Arc Pro B70 32GB
llama.cppQ5_K_M26.3936.0
Show all details for ThinkingCap-Qwen3.6-27B

Model

bottlecapai/ThinkingCap-Qwen3.6-27B

Display name

ThinkingCap-Qwen3.6-27B

Base model

Qwen3.6-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

26.3

Prefill tok/s

936

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b10255+ SYCL (master 071327508)

Quantization

Q5_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

draft_mtp

Spec model

Spec draft model

Spec tokens

4

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m ThinkingCap-Qwen3.6-27B-Q5_K_M-MTP.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75

Notes

ThinkingCap-Qwen3.6-27B Q5_K_M-MTP with MTP-4 spec decoding (draft-mtp n-max 4). Decode = engine rate timings.predicted_per_second, 4 diverse prompts x 3 reps steady-state (reps 2-3), 200W, 128K ctx (MTP VRAM ceiling for dense). Prefill = llama-bench pp4096 (Run 9, same build). q8_0 K + q4_1 V, FA on.

Reactions

Submitted

Aug 5, 2026, 7:58 AM

Last edited

Qwen3.6-35B-A3B

3B MoE · Qwen

Intel Arc Pro B70 32GB
llama.cppQ4_K_XL72.62128.0
Show all details for Qwen3.6-35B-A3B

Model

Qwen/Qwen3.6-35B-A3B

Display name

Qwen3.6-35B-A3B

Base model

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

72.6

Prefill tok/s

2128

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

32768

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b10255+ SYCL (master 071327508)

Quantization

Q4_K_XL

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-bench -m MODEL.gguf -ngl 99 -ncmoe 0 -fa 1 -ctk q8_0 -ctv q4_1 -p 512,4096,8192,32768 -n 128 -r 3 -b 8192 -ub 4096 -t 8 -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --p 512,4096,8192,32768 --r 3 --b 8192 --dev SYCL0

Notes

llama-bench tg128 + pp4096 (Run 9). q8_0 K + q4_1 V, FA on, -ngl 99 -ncmoe 0, 150W. Master 0804 build (#25874 quantized-KV XMX FA). Ctx = bench ctx (max prompt 32K).

Reactions

Submitted

Aug 5, 2026, 6:49 AM

Last edited

ThinkingCap-Qwen3.6-27B

28B · Qwen

Intel Arc Pro B70 32GB
llama.cppQ4_K_M27.5621.30k
Show all details for ThinkingCap-Qwen3.6-27B

Model

bottlecapai/ThinkingCap-Qwen3.6-27B

Display name

ThinkingCap-Qwen3.6-27B

Base model

Qwen3.6-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

27.5

Prefill tok/s

621.3

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

4014

Output tokens

350

Prefill tokens

0

Context length

204800

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b10053+PR25690

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

KV cache size

Prefix caching

no

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

mtp

Spec model

Spec draft model

Spec tokens

4

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

yes

MTP draft layers

Temperature

0.2

Top P

0.95

Top K

20

Min P

0

Repeat penalty

Mirostat

Command

llama-server -m ThinkingCap-Qwen3.6-27B-Q4_K_M.gguf -dev SYCL0 --parallel 1 -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 204800 -t 8 -b 8192 -ub 4096 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75

Extra flags

SYCL/Level Zero; asymmetric KV K=q8_0,V=q4_1; prefill temperature=0; reasoning enabled

Notes

Fresh single-B70 ThinkingCap benchmark on 2026-07-22. Independent phases with prompt cache disabled: 4,014-token prefill at temperature 0 and separate 350-token reasoning/decode task at temperature 0.2. Two warm-ups plus five measured repetitions per phase. Prefill samples=[622.0612,621.6988,621.3770,620.9613,620.5751], CV=0.094%. Decode samples=[26.6528,27.2787,27.2420,28.2763,28.0818], CV=2.421%. MTP-4 decode acceptance=1016/1154 (88.04%). Full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096; 200K context; 165 W. Build: llama.cpp b10053 0dc74e332 plus PR #25690 0bd0ec609. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw artifact SHA-256: 91eba0e2c34bf90b2d895b7d6e1819c763aadb9ed154dbb0da48b94b368d41d2.

Reactions

Submitted

Jul 22, 2026, 8:24 PM

Last edited

gemma-4-31B-it

33B · Gemma

Intel Arc Pro B70 32GB
llama.cppQ4_K_M26.6384.80k
Show all details for gemma-4-31B-it

Model

google/gemma-4-31B-it

Display name

gemma-4-31B-it

Base model

gemma-4-31B

Revision

main

Family

Gemma

Parameters

33B

Active params

MoE

no

Output tok/s

26.6

Prefill tok/s

384.8

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

4021

Output tokens

350

Prefill tokens

0

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b9853

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

mtp

Spec model

Spec draft model

mtp-gemma-4-31B-it.gguf

Spec tokens

4

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

yes

MTP draft layers

Temperature

0.2

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m gemma-4-31B-it-Q4_K_M.gguf -dev SYCL0 --parallel 1 -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -t 8 -b 8192 -ub 4096 --spec-type draft-mtp --spec-draft-model mtp-gemma-4-31B-it.gguf --spec-draft-n-max 4 --spec-draft-p-min 0.75

Extra flags

SYCL/Level Zero; asymmetric KV K=q8_0,V=q4_1; prefill temperature=0

Notes

Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 180 W. MTP-4 enabled; decode draft acceptance 225/297 (75.76%).

Reactions

Submitted

Jul 22, 2026, 10:55 AM

Last edited

gemma-4-31B-it

33B · Gemma

Intel Arc Pro B70 32GB
llama.cppQ4_K_M24.8361.20k
Show all details for gemma-4-31B-it

Model

google/gemma-4-31B-it

Display name

gemma-4-31B-it

Base model

gemma-4-31B

Revision

main

Family

Gemma

Parameters

33B

Active params

MoE

no

Output tok/s

24.8

Prefill tok/s

361.2

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

4021

Output tokens

350

Prefill tokens

0

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b9853

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

mtp

Spec model

Spec draft model

mtp-gemma-4-31B-it.gguf

Spec tokens

4

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

yes

MTP draft layers

Temperature

0.2

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m gemma-4-31B-it-Q4_K_M.gguf -dev SYCL0 --parallel 1 -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -t 8 -b 8192 -ub 4096 --spec-type draft-mtp --spec-draft-model mtp-gemma-4-31B-it.gguf --spec-draft-n-max 4 --spec-draft-p-min 0.75

Extra flags

SYCL/Level Zero; asymmetric KV K=q8_0,V=q4_1; prefill temperature=0

Notes

Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 165 W. MTP-4 enabled; decode draft acceptance 223/298 (74.83%).

Reactions

Submitted

Jul 22, 2026, 10:50 AM

Last edited

gemma-4-31B-it

33B · Gemma

Intel Arc Pro B70 32GB
llama.cppQ4_K_M16.4332.90k
Show all details for gemma-4-31B-it

Model

google/gemma-4-31B-it

Display name

gemma-4-31B-it

Base model

gemma-4-31B

Revision

main

Family

Gemma

Parameters

33B

Active params

MoE

no

Output tok/s

16.4

Prefill tok/s

332.9

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

4021

Output tokens

350

Prefill tokens

0

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b9853

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

0.2

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m gemma-4-31B-it-Q4_K_M.gguf -dev SYCL0 --parallel 1 -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -t 8 -b 8192 -ub 4096

Extra flags

SYCL/Level Zero; asymmetric KV K=q8_0,V=q4_1; prefill temperature=0

Notes

Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 150 W. MTP disabled.

Reactions

Submitted

Jul 22, 2026, 10:45 AM

Last edited

Qwen3.6-27B

28B · Qwen

Intel Arc Pro B70 32GB
llama.cppQ5_K_M25.1613.00k
Show all details for Qwen3.6-27B

Model

Qwen/Qwen3.6-27B

Display name

Qwen3.6-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

25.1

Prefill tok/s

613.0

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

4014

Output tokens

350

Prefill tokens

0

Context length

204800

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 7840HS w/ Radeon 780M Graphics

RAM

25.8GB

OS

Windows 11 Pro 10.0.26100 build 26100

Power

230W

Engine

llama.cpp

Engine version

b9853

Quantization

Q5_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

yes

Spec method

mtp

Spec model

Spec draft model

Spec tokens

4

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

yes

MTP draft layers

Temperature

0.2

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Qwen3.6-27B-MTP-Q5_K_M.gguf -dev SYCL0 --parallel 1 -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 204800 -t 8 -b 8192 -ub 4096 --spec-type draft-mtp --spec-draft-n-max 4 --spec-draft-p-min 0.75

Extra flags

SYCL/Level Zero; asymmetric KV K=q8_0,V=q4_1; prefill temperature=0

Notes

Single Intel Arc Pro B70 clean-suite result from 2026-07-16. Independent phases: prefill uses the recorded ~4K prompt at temperature 0; decode uses a separate 350-token task at temperature 0.2. Batch/concurrency 1; full GPU offload; Flash Attention on; K=q8_0,V=q4_1; b=8192,ub=4096. Engine: llama.cpp SYCL b9853 (7af4279f4), IntelLLVM 2026.0.0. Actual host: AMD Ryzen 7 5700X3D, 32 GB RAM, Ubuntu 26.04. LocalMaxxing currently reuses an older B70 hardware profile for this account; its structured CPU/RAM/OS/power fields are stale and should be ignored. Raw source SHA-256: 6d4f82df37f469d3f0aba0e1cc3e5642919b1f9c50d2ef5d5c37e4750f20ba4c. Power cap 165 W. MTP-4 enabled; decode draft acceptance 199/216 (92.13%).

Reactions

Submitted

Jul 22, 2026, 10:40 AM

Last edited