Get startedLeaderboardDecode calculatorModelsReportsHardwareBenchmarksMarketplaceRentalsProAPI Docs
Language
SergiioB

SergiioB

@SergiioB · Member since 2026

← Leaderboard
Mainly llama.cpp 74 approved runs

Submissions

74

74 approved public runs

Models

17

unique models

Avg output

91.8

Best 1139.8 tok/s

Avg prefill

2062

Best 10371.0 tok/s

Avg total

248.6

Best 2033.5 tok/s

Avg TTFT

4587.2ms

Best 49ms

Hardware

2

distinct setups

Reactions

2

across all runs

Personal bests

Personal bests17

Fastest approved result for each model.

3B MoE9 runs🔥
Output1139.8 tok/s
Prefill8715.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

vllm · GPTQ-Int4 · 3w ago

28B11 runs
Output224.2 tok/s
Prefill4.8 tok/s
TTFT15419ms
Depth
Intel Arc Pro B70 32GB

vllm · GPTQ-Int4 · 1w ago

3B MoE9 runs
Output186.6 tok/s
Prefill7160.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

vllm · GPTQ-INT4-G64-sym-local+DFlash-BF16-local · 2w ago

07 runs
Output164.5 tok/s
Prefill383.4 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q5_K_M · 1mo ago

01 run
Output145.9 tok/s
Prefill
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q5_K_M · 1mo ago

28B5 runs
Output112.7 tok/s
Prefill1696.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

vllm · GPTQ-Int4+draft-INT4-RTN · 1w ago

3B MoE3 runs
Output108.4 tok/s
Prefill9072.9 tok/s
TTFT320ms
Depth
Intel Arc Pro B70 32GB

vllm · GPTQ-Int4 · 1w ago

3B MoE1 run
Output87.4 tok/s
Prefill766.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q4_0 · 2w ago

35B MoE2 runs
Output73.2 tok/s
Prefill416.1 tok/s
TTFT77ms
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q4_K_M · 2mo ago

28B7 runs
Output69.3 tok/s
Prefill1754.6 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

vllm · GPTQ-Int4 · 3w ago

35B2 runs
Output62.8 tok/s
Prefill1682.4 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q4_K_XL · 1mo ago

02 runs
Output50.3 tok/s
Prefill553.9 tok/s
TTFT58ms
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q6_K · 2mo ago

28B2 runs
Output31.1 tok/s
Prefill389.6 tok/s
TTFT190ms
Depth
Intel Arc Pro B70 32GB ×2

vllm · FP8 · today

30B3 runs
Output29.2 tok/s
Prefill1301.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q4_K_XL · 2w ago

28B3 runs
Output27.9 tok/s
Prefill936.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q6_K · 3w ago

33B6 runs
Output26.6 tok/s
Prefill384.8 tok/s
TTFT
Depth0k
Intel Arc Pro B70 32GB

llama.cpp · Q4_K_M · 1mo ago

28B1 run
Output25.1 tok/s
Prefill613.0 tok/s
TTFT
Depth
Intel Arc Pro B70 32GB

llama.cpp · Q5_K_M · 1mo ago

Hardware breakdown

Hardware breakdown2

Distinct rigs used across approved submissions.

Intel Arc Pro B70 32GB

72 runs · best 1139.8 tok/s

Intel Arc Pro B70 32GB ×2

2 runs · best 31.1 tok/s

Engines used

Engines used2

Runtime mix and top quantizations.

llama.cpp43Q4_K_M · Q4_0 · Q4_K_XL
vllm31FP8 · GPTQ-Int4 · GPTQ-Int4+draft-INT4-RTN
All runs

All runs74

Full approved benchmark history · page 1 of 3.

Qwen3.8-27B-FP8

28B · Qwen

31.1

tok/s

Hardware

Intel Arc Pro B70 32GB ×2

Engine

vllm · FP8

TTFT

190ms

Context

2k · today

Show all run details

Model

Qwen/Qwen3.8-27B-FP8

Display name

Qwen3.8-27B-FP8

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

31.1

Prefill tok/s

389.6

Total tok/s

39.4

TTFT

189.9ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB ×2

GPU slots

GPU count

2

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

FP8

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8920 servedModel: qw38-fp8

Extra flags

Notes

Reactions

Submitted

Aug 30, 2026, 6:06 PM

Last edited

Qwen3.8-27B-FP8

28B · Qwen

28.6

tok/s

Hardware

Intel Arc Pro B70 32GB ×2

Engine

vllm · FP8

TTFT

205ms

Context

2k · today

Show all run details

Model

Qwen/Qwen3.8-27B-FP8

Display name

Qwen3.8-27B-FP8

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

28.6

Prefill tok/s

360.8

Total tok/s

36.2

TTFT

205.1ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB ×2

GPU slots

GPU count

2

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

FP8

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8920 servedModel: qw38-fp8

Extra flags

Notes

Reactions

Submitted

Aug 30, 2026, 4:04 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

320ms

Context

2k · 1w ago

Show all run details

Model

SergiioB/Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Display name

Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Base model

Ornith-1.5-35B-A3B

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

108.4

Prefill tok/s

9072.9

Total tok/s

2033.5

TTFT

320.3ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

2906

Output tokens

128

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: ornith15

Extra flags

Notes

Reactions

Submitted

Aug 21, 2026, 10:36 AM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

297ms

Context

2k · 1w ago

Show all run details

Model

SergiioB/Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Display name

Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Base model

Ornith-1.5-35B-A3B

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

69.9

Prefill tok/s

9780

Total tok/s

1435.2

TTFT

297.1ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

2906

Output tokens

128

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: ornith15

Extra flags

Notes

Reactions

Submitted

Aug 21, 2026, 10:18 AM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

49ms

Context

2k · 1w ago

Show all run details

Model

SergiioB/Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Display name

Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Base model

Ornith-1.5-35B-A3B

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

94.1

Prefill tok/s

654.7

Total tok/s

104.4

TTFT

48.9ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

32

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

150W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: ornith15

Extra flags

Notes

Reactions

Submitted

Aug 21, 2026, 10:14 AM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

24360ms

Context

131k · 1w ago

Show all run details

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

107.8

Prefill tok/s

1693.9

Total tok/s

70.8

TTFT

24360ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

5

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

5

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Single Intel Arc Pro B70 32GB. Stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching ON. Realistic 5-user coding serving at the Qwen3.8-27B model-card RECOMMENDED non-thinking sampling (temperature=0.7, top_p=0.80, top_k=20, presence_penalty=1.5). 5 concurrent multi-turn coding sessions, ~8K session prompts, 512-token generations, 45/45 OK, real completions in history. tokSOut = sum of per-stream rates (5 x 21.6 median per-user tok/s). Per-turn Sigma: t1 134.4 / t2 106.3 / t3 100.5. TTFT 21-26s per turn (prefix cache lands 0-38% of shared tokens at C5 on this build). MTP acceptance 34-41% under concurrency. Greedy sibling: 127.4 Sigma.

Reactions

Submitted

Aug 19, 2026, 1:39 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

337ms

Context

131k · 1w ago

Show all run details

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

102.6

Prefill tok/s

1693.9

Total tok/s

70.8

TTFT

337ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Single Intel Arc Pro B70 32GB. Stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching ON. Qwen3.8-27B model-card RECOMMENDED non-thinking sampling: temperature=0.7, top_p=0.80, top_k=20, presence_penalty=1.5. C1 client post-first, n=5 median 102.61 (95.7-106.7), calibrated real-world Pi prompts, MTP acceptance 91.8%. Greedy sibling on the same stack: 106.7 (93% acceptance) - recommended sampling costs ~4% here. Cold input 1694 tok/s p8192/g1.

Reactions

Submitted

Aug 19, 2026, 1:39 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

22850ms

Context

131k · 1w ago

Show all run details

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

127.4

Prefill tok/s

1693.9

Total tok/s

70.8

TTFT

22850ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

5

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. Realistic 5-user coding serving, v2 with REAL completions kept in session history (supersedes the stubbed same-day 138.3 record). Stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching ON. 5 concurrent sessions, ~8K start, 3 turns, g512, 60/60 OK. tokSOut = sum of per-stream generation rates (5 x 25.5 median per-user tok/s, all turns). Per-turn: t1 165.5, t2 127.9, t3 114.2 aggregate. TTFT 22.6-25.0s every turn: prefix cache lands 0-38% of shared tokens at C5 on this build (vs 91% at C1), so ~45-53K session tokens re-prefill per turn. MTP acceptance 43-56% under concurrency. Short-prompt C5 (203.8) is a separate record; C1 is 106.7.

Reactions

Submitted

Aug 19, 2026, 12:59 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

335ms

Context

131k · 1w ago

Show all run details

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

106.7

Prefill tok/s

1693.9

Total tok/s

70.8

TTFT

335ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. Current full stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching ON (zero hits, unique prompts). C1 client post-first, n=5 median 106.7 (103.2-111.3), calibrated real-world Pi prompt set; MTP acceptance 89-96%; cold input 1694 tok/s at p8192/g1. Supersedes same-day 100.2 (same config, run-to-run variance). Cache-off sibling record: 112.65. Degenerate INDEX filler measures ~71 (44% acceptance) - prompt family matters.

Reactions

Submitted

Aug 19, 2026, 12:59 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

22350ms

Context

131k · 1w ago

Show all run details

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

138.3

Prefill tok/s

1693.9

Total tok/s

248.6

TTFT

22350ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

5

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

5

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. Realistic 5-user coding serving on the current stack (vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching on). Workload: 5 concurrent multi-turn coding sessions, ~8K-token session prompts, 512-token generations, 60/60 OK, 0 crashes. tokSOut = sum of per-stream generation rates (5 x 27.7 median per-user tok/s). TTFT is per-turn with ~45K session tokens prefilling per wave; turn-completion wall aggregate incl. prefill = 54 tok/s. Short-prompt C5 aggregate (203.8) is a separate record. MTP acceptance 47% under concurrency.

Reactions

Submitted

Aug 19, 2026, 12:17 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

22350ms

Context

131k · 1w ago

Show all run details

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

138.3

Prefill tok/s

1693.9

Total tok/s

248.6

TTFT

22350ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

5

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

5

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. Realistic 5-user coding serving on the current stack (vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching on). Workload: 5 concurrent multi-turn coding sessions, ~8K-token session prompts, 512-token generations, 60/60 OK, 0 crashes. tokSOut = sum of per-stream generation rates (5 x 27.7 median per-user tok/s). TTFT is per-turn with ~45K session tokens prefilling per wave; turn-completion wall aggregate incl. prefill = 54 tok/s. Short-prompt C5 aggregate (203.8) is a separate record. MTP acceptance 47% under concurrency.

Reactions

Submitted

Aug 19, 2026, 12:16 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

369ms

Context

131k · 1w ago

Show all run details

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

100.2

Prefill tok/s

1693.9

Total tok/s

70.8

TTFT

369ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. Current full stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1 overlay, prefix caching ON (zero hits, unique prompts). C1 client post-first, n=5 median (92.9-105.1), calibrated real-world Pi prompt set, same prompt file as the 2026-08-18 record. p8192/g128 median 100.0; cold input 1694 tok/s at p8192/g1. Prompt-content sensitive: degenerate INDEX filler measures ~71 (44% MTP acceptance) vs 93-96% on realistic text. Cache-off record on the same patches: 112.65.

Reactions

Submitted

Aug 19, 2026, 12:13 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

15419ms

Context

2k · 1w ago

Show all run details

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

224.2

Prefill tok/s

4.8

Total tok/s

285.7

TTFT

15418.9ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

32

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. vLLM XPU nightly digest f01e24f6, MTP4 speculative decoding, GDN mixed-split v5, draft-INT4 S+M1 overlay, prefix caching on, fp8 KV cache, 230W configured cap. Aggregate C32 throughput via lmx remote harness (256 output tokens, 3 iterations). Single-stream C1 is a separate record.

Reactions

Submitted

Aug 19, 2026, 11:32 AM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

8077ms

Context

2k · 1w ago

Show all run details

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

200.6

Prefill tok/s

9.2

Total tok/s

253.2

TTFT

8077.1ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

16

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. vLLM XPU nightly digest f01e24f6, MTP4 speculative decoding, GDN mixed-split v5, draft-INT4 S+M1 overlay, prefix caching on, fp8 KV cache, 230W configured cap. Aggregate C16 throughput via lmx remote harness (256 output tokens, 3 iterations). Single-stream C1 is a separate record.

Reactions

Submitted

Aug 19, 2026, 11:32 AM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

414ms

Context

2k · 1w ago

Show all run details

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

203.8

Prefill tok/s

178.9

Total tok/s

248.6

TTFT

413.7ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

5

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. vLLM XPU nightly digest f01e24f6, MTP4 speculative decoding, GDN mixed-split v5, draft-INT4 S+M1 overlay, prefix caching on, fp8 KV cache, 230W configured cap. Aggregate C5 throughput via lmx remote harness (256 output tokens, 3 iterations). Single-stream C1 is a separate record.

Reactions

Submitted

Aug 19, 2026, 11:32 AM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

171ms

Context

2k · 1w ago

Show all run details

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

56.8

Prefill tok/s

433.4

Total tok/s

70.8

TTFT

170.7ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. vLLM XPU nightly digest f01e24f6, MTP4 speculative decoding, GDN mixed-split v5, draft-INT4 S+M1 overlay, prefix caching on, fp8 KV cache, 230W configured cap. Aggregate C1 throughput via lmx remote harness (256 output tokens, 3 iterations). Single-stream C1 is a separate record.

Reactions

Submitted

Aug 19, 2026, 11:32 AM

Last edited

Qwen3.8-27B

28B · Qwen

112.7

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4+draft-INT4-RTN

TTFT

Context

131k · 1w ago

Show all run details

Model

Qwen/Qwen3.8-27B

Display name

Qwen3.8-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

112.7

Prefill tok/s

1696

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

0.27.2rc1.dev77+gac7509e2b XPU

Quantization

GPTQ-Int4+draft-INT4-RTN

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.88

Max running seqs

64

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

B70_DRAFT_LMHEAD_INT4=1 B70_DRAFT_MTP_INT4=1 vllm serve /model --quantization gptq --dtype float16 --max-model-len 131072 --gpu-memory-utilization 0.88 --kv-cache-dtype fp8 --max-num-seqs 64 --max-num-batched-tokens 8192 --no-enable-prefix-caching --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":4}

Extra flags

--max-model-len 131072 --no-enable-prefix-caching --language-model-only --speculative-config {method:mtp,num_speculative_tokens:4}

Notes

NEW identity vs cmsur82fz06svms01ga1f0z83 (BF16 draft 83.7). Same 1x B70, same image vllm/vllm-openai-xpu@sha256:f01e24f6c7ff01f1e0662234255a1372297d1dbd89d003cf13c8fad3eab1ba4f, kernels 0.1.12.3, same target SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16. Extra patches after Qwen MTP nightly+boundary: patch_gdn_mixed_split_v5.py + patch_draft_lmhead_int4.py + patch_draft_mtp_int4.py (runtime RTN INT4 of draft LM head + 5 MTP linears only; target verify stays BF16). tokSOut = client post-first at p512/g128, median n=5, C1, cache OFF (zero prefix_cache_hits_total delta): 112.65 (range 111.58-117.73). Matched same-harness BF16-draft arm was 81.20, not the older Run40 83.7 card. tokSPrefill = actual input/TTFT at p8192/g1 MTP4 n=5: 1696 (flat vs 1691). Accept 510/540=94.44% vs 95.86%. Short agentic +32.8%; 8K/16K long-ctx agentic +37%/+26%. 230 W configured cap; measured median draw 205 W at p512/g128. Speed-only: no token/KL/task-quality parity vs BF16 draft. E2 self-reported. Measured 2026-08-18.

Reactions

Submitted

Aug 19, 2026, 6:31 AM

Last edited

Qwen3.8-27B

28B · Qwen

83.7

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-Int4

TTFT

Context

131k · 2w ago

🚀
Show all run details

Model

Qwen/Qwen3.8-27B

Display name

Qwen3.8-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

83.7

Prefill tok/s

1774

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

0.27.2rc1.dev77+gac7509e2b XPU

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.88

Max running seqs

64

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 131072 --gpu-memory-utilization 0.88 --kv-cache-dtype fp8 --max-num-seqs 64 --max-num-batched-tokens 8192 --no-enable-prefix-caching --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":4}

Extra flags

--max-model-len 131072 --no-enable-prefix-caching --language-model-only --speculative-config {method:mtp,num_speculative_tokens:4}

Notes

Single-stream C1. vLLM XPU nightly vllm/vllm-openai-xpu@sha256:f01e24f6c7ff01f1e0662234255a1372297d1dbd89d003cf13c8fad3eab1ba4f, vllm-xpu-kernels 0.1.12.3. Model SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16 (Qwen3_5ForConditionalGeneration, dense 27B, GPTQ-INT4 g128 sym desc_act=false, preserved BF16 MTP head). Patches: patch_mtp_nightly.py + patch_mtp_boundary.py. fp8 KV cache, gpu-memory-utilization 0.88, scheduler 8192, max-num-seqs 64, prefix cache OFF (--no-enable-prefix-caching, zero cache-query delta). MTP4 speculative decode, 1-layer draft. tokSOut = client post-first rate at p512/g128, median n=5 (not engine-native decode); tokSPrefill = actual input tokens / client TTFT at p8192/g1, median n=5, no-spec mode (not llama-bench pp). 230 W configured cap; measured campaign-window mean draw 196 W (MTP4 mode). E2 self-reported; independent reproduction pending. Measured 2026-08-15. Claims-audited: 83.7 tok/s MTP4 p512/g128 n=5 range 78.3-84.4; exact-128K full-context p130944/g128 56.3 tok/s n=1. 2000 t/s prefill NOT reached; max cold-input median 1851 t/s (no-spec p2048/g1).

Reactions

🚀

Submitted

Aug 15, 2026, 7:13 PM

Last edited

Qwen3.8-27B

28B · Qwen

22.3

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_M

TTFT

Context

131k · 2w ago

Show all run details

Model

Qwen/Qwen3.8-27B

Display name

Qwen3.8-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

22.3

Prefill tok/s

1121.4

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

llama.cpp

Engine version

b10255+ SYCL (071327508) build-sycl-0804

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Qwen3.8-27B-Q4_K_M.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0

Notes

Qwen3.8-27B (dense 27B, hybrid GDN linear+full attention, MTP head not in GGUF quant). Community quant unsloth/Qwen3.8-27B-GGUF Q4_K_M (SHA-256 7b2aec3b...cc89f1b) -- weights released 2026-08-14T15:00Z, benched same day. 230W cap (stock). llama-bench tg128 mean 22.30 t/s (+/-0.01, n=5, C1, cold cache). Cold input pp8192 mean 1121.40 t/s (+/-1.89, n=5); pp512 778.80 / pp2048 1115.99 / pp4096 1137.02. Load at -c 131072 verified, 10.2 GiB VRAM free after load. KV q8_0 K + q4_1 V, FA on. Correctness: coherent smoke 5/5, deterministic replay pass, token exactness pass; reference parity not run. Cross-cap note: at matched 150W cap this model runs tg128 mean 17.78 (n=5, separate submission); prior Qwen3.6-27B Q4_K_M ran tg128 21.3 mean (n=3, 150W, Run 9) -- comparisons across caps are directional only. Dense power scaling 150->230W measured +25% decode on this model.

Reactions

Submitted

Aug 14, 2026, 3:50 PM

Last edited

Qwen3.8-27B

28B · Qwen

22.3

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_M

TTFT

Context

131k · 2w ago

Show all run details

Model

Qwen/Qwen3.8-27B

Display name

Qwen3.8-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

22.3

Prefill tok/s

1121.4

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

llama.cpp

Engine version

b10255+ SYCL (071327508) build-sycl-0804

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Qwen3.8-27B-Q4_K_M.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0

Notes

Qwen3.8-27B (dense 27B, hybrid GDN linear+full attention, MTP head not in GGUF quant). Community quant unsloth/Qwen3.8-27B-GGUF Q4_K_M (SHA-256 7b2aec3b...cc89f1b) -- weights released 2026-08-14T15:00Z, benched same day. 230W cap (stock). llama-bench tg128 mean 22.30 t/s (+/-0.01, n=5, C1, cold cache). Cold input pp8192 mean 1121.40 t/s (+/-1.89, n=5); pp512 778.80 / pp2048 1115.99 / pp4096 1137.02. Load at -c 131072 verified, 10.2 GiB VRAM free after load. KV q8_0 K + q4_1 V, FA on. Correctness: coherent smoke 5/5, deterministic replay pass, token exactness pass; reference parity not run. Cross-cap note: at matched 150W cap this model runs tg128 mean 17.78 (n=5, separate submission); prior Qwen3.6-27B Q4_K_M ran tg128 21.3 mean (n=3, 150W, Run 9) -- comparisons across caps are directional only. Dense power scaling 150->230W measured +25% decode on this model.

Reactions

Submitted

Aug 14, 2026, 3:50 PM

Last edited

Qwen3.8-27B

28B · Qwen

17.8

tok/s

Hardware

Intel Arc Pro B70 32GB

Engine

llama.cpp · Q4_K_M

TTFT

Context

131k · 2w ago

Show all run details

Model

Qwen/Qwen3.8-27B

Display name

Qwen3.8-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

17.8

Prefill tok/s

787.5

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

150W

Engine

llama.cpp

Engine version

b10255+ SYCL (071327508) build-sycl-0804

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Qwen3.8-27B-Q4_K_M.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0

Notes

Qwen3.8-27B (dense 27B, hybrid GDN linear+full attention, MTP head not in GGUF quant). Community quant unsloth/Qwen3.8-27B-GGUF Q4_K_M (SHA-256 7b2aec3b...cc89f1b) -- weights released 2026-08-14T15:00Z, benched same day. llama-bench tg128 mean 17.78 t/s (+/-0.17, n=5, C1, cold cache). Cold input pp8192 mean 787.45 t/s (+/-1.74, n=5). Load at -c 131072 verified, 10.2 GiB VRAM free after load. KV q8_0 K + q4_1 V, FA on. Correctness: coherent smoke 5/5, deterministic replay pass, token exactness pass; reference parity not run. At matched 150W/build/quant/flags the prior Qwen3.6-27B Q4_K_M ran tg128 21.3 mean (n=3, Run 9); 3.8 mean 17.78 (n=5) -- ~16% slower decode. Comparison directional: different checkpoints, n mismatch (3 vs 5).

Reactions

Submitted

Aug 14, 2026, 3:41 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-INT4-G64-sym-local+DFlash-BF16-local

TTFT

Context

16k · 2w ago

Show all run details

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

186.6

Prefill tok/s

7160

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU)

Quantization

GPTQ-INT4-G64-sym-local+DFlash-BF16-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

auto

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.9

Max running seqs

1

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /mnt/models/nemotron-lightning-gptq-sym64-20260812 --quantization gptq --dtype float16 --max-model-len 16384 --max-num-seqs 1 --max-num-batched-tokens 8192 --gpu-memory-utilization 0.90 --no-enable-prefix-caching --language-model-only --async-scheduling --speculative-config '{"method":"dflash","model":"/mnt/models2/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-DFlash-BF16","num_speculative_tokens":7}' --host 127.0.0.1 --port 8001

Extra flags

--max-model-len 16384 --no-enable-prefix-caching --language-model-only --async-scheduling --speculative-config {"method":"dflash","model":"/mnt/models2/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-DFlash-BF16","num_speculative_tokens":7}

Notes

Self-reported E2, not independently reproduced. C1 (max_num_seqs=1), prefix cache explicitly off (--no-enable-prefix-caching; enable_prefix_caching=False; zero prefix_cache_hits delta). Timing is client monotonic SSE: tokSOut is median client post-first decode at exact p2048/g128 n=5 = 186.61 t/s (range 174.60-201.83; families research/rag/tool/document/assistant). Additional decode cells on the same warm server: p512/g128 median 194.61 (range 140.20-220.01, acceptance 45.1% — wide family spread, not the representative scalar); p8192/g128 median 157.92 (143.50-170.25, acceptance 53.0%). tokSPrefill is the p8192/g1 n=5 median COLD INPUT RATE from client TTFT (7160 t/s, range 7117-7226), which includes scheduling and first-token work — NOT isolated engine prefill. Also p2048/g1 cold input median 6456 t/s. Spec: method=dflash n_spec=7; window acceptance 1830/3521 = 52.0%. Target is a LOCAL symmetric GPTQ INT4 G64 conversion of the named BF16 repo (not an official HF quant). Draft is a LOCAL NVFP4 E2M1→BF16 reconstruction of NVIDIA DFlash, not a published BF16 draft. Stack: vllm/vllm-openai-xpu@sha256:1da0a95485455f08588c11080b9718992fd7d434c6a965d74654903a9d999c57 plus local det image patches (native grouped-topk v2, SSU B8/W4, at::zeros grouped-GEMM); VLLM_XPU_ENABLE_XPU_GRAPH=1; PIECEWISE+FULL graphs; --async-scheduling; --quantization gptq --dtype float16 --max-model-len 16384 --gpu-memory-utilization 0.90. Configured cap 150 W; measured cell-window averages ~149-160 W; peak interval-average 179.3 W (0.5 s energy1_input on p8192/g128); pkg max 68.0 C. Deterministic raw-completion replay smoke exact_match=true (scope: smoke only, not logit/KL or task-quality). n=3 screen 20260813T080234Z is superseded and was not used. Matched-except-speculation vs prior no-spec n=5 graph campaign (p8192/g128 87.25 t/s): 1.81× on that cell only.

Reactions

Submitted

Aug 13, 2026, 8:40 AM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-INT4-G64-sym-local

TTFT

Context

16k · 2w ago

Show all run details

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

92.7

Prefill tok/s

10349

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU) + deterministic grouped-GEMM fix

Quantization

GPTQ-INT4-G64-sym-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

default

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.9

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /model --dtype bfloat16 --quantization gptq --max-model-len 16384 --gpu-memory-utilization 0.90 --no-enable-prefix-caching --async-scheduling

Extra flags

--max-model-len 16384 --no-enable-prefix-caching --async-scheduling

Notes

Self-reported. C1 client post-first decode, exact p512/g128 n=3 median (92.65 t/s, range 92.61-92.65), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). GPTQ conversion published at SergiiioB/Nemotron-3.5-Lightning-30B-A3B-GPTQ-INT4-G64-sym. tokSPrefill = cold input rate at p8192 from median client TTFT (10,349 tok/s; TTFT median 0.7916s, n=3, per-rep 10347-10357) — client-observed rate, NOT isolated engine prefill. NO speculative decoding. Deterministic image with grouped-GEMM at::zeros atomic-buffer fix. Deterministic replay verified (8/8 identical temp-0 requests). 150W cap; measured draw ~89-90W. Claims-auditor verified decode; prefill corrected from earlier 10371 (rounding error).

Reactions

Submitted

Aug 12, 2026, 11:54 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-INT4-G64-sym-local

TTFT

Context

16k · 2w ago

Show all run details

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

92.7

Prefill tok/s

10371

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU) + deterministic grouped-GEMM fix

Quantization

GPTQ-INT4-G64-sym-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

default

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.9

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /model --dtype bfloat16 --quantization gptq --max-model-len 16384 --gpu-memory-utilization 0.90 --no-enable-prefix-caching --async-scheduling

Extra flags

--max-model-len 16384 --no-enable-prefix-caching --async-scheduling

Notes

Self-reported. C1 client post-first decode, exact p512/g128 n=3 median (92.65 t/s, range 92.61-92.65), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion of BF16 checkpoint (published at SergiiioB/Nemotron-3.5-Lightning-30B-A3B-GPTQ-INT4-G64-sym). tokSPrefill is cold input rate at p8192 from median client TTFT (10,371 tok/s; TTFT 0.791s, n=3) - client-observed rate, NOT isolated engine prefill. NO speculative decoding (native MTP tested: 0% draft acceptance on this stack; N-gram: non-deterministic; DFlash: unsupported). Deterministic image vllm-openai-xpu:1da0a954-det0123 with grouped-GEMM at::zeros atomic-buffer fix (eliminates the XPU FP race). DETERMINISTIC REPLAY VERIFIED: 8 identical temperature-0 requests produce byte-identical output. 150W cap; measured draw ~89-90W.

Reactions

Submitted

Aug 12, 2026, 10:37 PM

Last edited

Hardware

Intel Arc Pro B70 32GB

Engine

vllm · GPTQ-INT4-G64-sym-local

TTFT

Context

16k · 2w ago

Show all run details

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

93

Prefill tok/s

8368

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU)

Quantization

GPTQ-INT4-G64-sym-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

default

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

0.95

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /mnt/models/nemotron-lightning-gptq-sym64-20260812 --host 0.0.0.0 --port 8001 --no-enable-prefix-caching --async-scheduling --gpu-memory-utilization 0.95 --max-model-len 16384 --enforce-eager false --trust-remote-code

Extra flags

--no-enable-prefix-caching --async-scheduling --max-model-len 16384 --enforce-eager false

Notes

Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). tokSPrefill is the cold input rate at exact p8192/g128 from median client TTFT (8,368 tok/s, TTFT median 0.9789s, n=3) - a client-observed input rate including scheduling and first-token work, NOT an isolated engine prefill measurement. NO speculative decoding in this config: native MTP was tested and rejected on this stack (0% draft acceptance), N-gram rejected (temperature-0 nondeterminism), DFlash not supported for this model on this stack. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent.

Reactions

Submitted

Aug 12, 2026, 9:08 PM

Last edited

ModelHardwareEnginetok/s outprefilltok/s totalTTFTDepthShare
Qwen3.8-27B-FP8

28B · Qwen

Intel Arc Pro B70 32GB ×2
vllmFP831.1389.639.4190ms
Show all details for Qwen3.8-27B-FP8

Model

Qwen/Qwen3.8-27B-FP8

Display name

Qwen3.8-27B-FP8

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

31.1

Prefill tok/s

389.6

Total tok/s

39.4

TTFT

189.9ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB ×2

GPU slots

GPU count

2

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

FP8

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8920 servedModel: qw38-fp8

Extra flags

Notes

Reactions

Submitted

Aug 30, 2026, 6:06 PM

Last edited

Qwen3.8-27B-FP8

28B · Qwen

Intel Arc Pro B70 32GB ×2
vllmFP828.6360.836.2205ms
Show all details for Qwen3.8-27B-FP8

Model

Qwen/Qwen3.8-27B-FP8

Display name

Qwen3.8-27B-FP8

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

28.6

Prefill tok/s

360.8

Total tok/s

36.2

TTFT

205.1ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB ×2

GPU slots

GPU count

2

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

FP8

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8920 servedModel: qw38-fp8

Extra flags

Notes

Reactions

Submitted

Aug 30, 2026, 4:04 PM

Last edited

Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

3B MoE · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4108.49072.92033.5320ms
Show all details for Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Model

SergiioB/Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Display name

Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Base model

Ornith-1.5-35B-A3B

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

108.4

Prefill tok/s

9072.9

Total tok/s

2033.5

TTFT

320.3ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

2906

Output tokens

128

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: ornith15

Extra flags

Notes

Reactions

Submitted

Aug 21, 2026, 10:36 AM

Last edited

Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

3B MoE · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int469.99780.01435.2297ms
Show all details for Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Model

SergiioB/Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Display name

Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Base model

Ornith-1.5-35B-A3B

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

69.9

Prefill tok/s

9780

Total tok/s

1435.2

TTFT

297.1ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

2906

Output tokens

128

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: ornith15

Extra flags

Notes

Reactions

Submitted

Aug 21, 2026, 10:18 AM

Last edited

Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

3B MoE · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int494.1654.7104.449ms
Show all details for Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Model

SergiioB/Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Display name

Ornith-1.5-35B-A3B-GPTQ-Int4-sym-G128-MTP-BF16-MixedCal-v2

Base model

Ornith-1.5-35B-A3B

Revision

main

Family

Qwen

Parameters

36B

Active params

3B

MoE

yes

Output tok/s

94.1

Prefill tok/s

654.7

Total tok/s

104.4

TTFT

48.9ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

32

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

150W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: ornith15

Extra flags

Notes

Reactions

Submitted

Aug 21, 2026, 10:14 AM

Last edited

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4107.81693.970.824360ms
Show all details for Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

107.8

Prefill tok/s

1693.9

Total tok/s

70.8

TTFT

24360ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

5

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

5

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Single Intel Arc Pro B70 32GB. Stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching ON. Realistic 5-user coding serving at the Qwen3.8-27B model-card RECOMMENDED non-thinking sampling (temperature=0.7, top_p=0.80, top_k=20, presence_penalty=1.5). 5 concurrent multi-turn coding sessions, ~8K session prompts, 512-token generations, 45/45 OK, real completions in history. tokSOut = sum of per-stream rates (5 x 21.6 median per-user tok/s). Per-turn Sigma: t1 134.4 / t2 106.3 / t3 100.5. TTFT 21-26s per turn (prefix cache lands 0-38% of shared tokens at C5 on this build). MTP acceptance 34-41% under concurrency. Greedy sibling: 127.4 Sigma.

Reactions

Submitted

Aug 19, 2026, 1:39 PM

Last edited

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4102.61693.970.8337ms
Show all details for Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

102.6

Prefill tok/s

1693.9

Total tok/s

70.8

TTFT

337ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Single Intel Arc Pro B70 32GB. Stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching ON. Qwen3.8-27B model-card RECOMMENDED non-thinking sampling: temperature=0.7, top_p=0.80, top_k=20, presence_penalty=1.5. C1 client post-first, n=5 median 102.61 (95.7-106.7), calibrated real-world Pi prompts, MTP acceptance 91.8%. Greedy sibling on the same stack: 106.7 (93% acceptance) - recommended sampling costs ~4% here. Cold input 1694 tok/s p8192/g1.

Reactions

Submitted

Aug 19, 2026, 1:39 PM

Last edited

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4127.41693.970.822850ms
Show all details for Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

127.4

Prefill tok/s

1693.9

Total tok/s

70.8

TTFT

22850ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

5

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. Realistic 5-user coding serving, v2 with REAL completions kept in session history (supersedes the stubbed same-day 138.3 record). Stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching ON. 5 concurrent sessions, ~8K start, 3 turns, g512, 60/60 OK. tokSOut = sum of per-stream generation rates (5 x 25.5 median per-user tok/s, all turns). Per-turn: t1 165.5, t2 127.9, t3 114.2 aggregate. TTFT 22.6-25.0s every turn: prefix cache lands 0-38% of shared tokens at C5 on this build (vs 91% at C1), so ~45-53K session tokens re-prefill per turn. MTP acceptance 43-56% under concurrency. Short-prompt C5 (203.8) is a separate record; C1 is 106.7.

Reactions

Submitted

Aug 19, 2026, 12:59 PM

Last edited

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4106.71693.970.8335ms
Show all details for Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

106.7

Prefill tok/s

1693.9

Total tok/s

70.8

TTFT

335ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. Current full stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching ON (zero hits, unique prompts). C1 client post-first, n=5 median 106.7 (103.2-111.3), calibrated real-world Pi prompt set; MTP acceptance 89-96%; cold input 1694 tok/s at p8192/g1. Supersedes same-day 100.2 (same config, run-to-run variance). Cache-off sibling record: 112.65. Degenerate INDEX filler measures ~71 (44% acceptance) - prompt family matters.

Reactions

Submitted

Aug 19, 2026, 12:59 PM

Last edited

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4138.31693.9248.622350ms
Show all details for Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

138.3

Prefill tok/s

1693.9

Total tok/s

248.6

TTFT

22350ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

5

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

5

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. Realistic 5-user coding serving on the current stack (vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching on). Workload: 5 concurrent multi-turn coding sessions, ~8K-token session prompts, 512-token generations, 60/60 OK, 0 crashes. tokSOut = sum of per-stream generation rates (5 x 27.7 median per-user tok/s). TTFT is per-turn with ~45K session tokens prefilling per wave; turn-completion wall aggregate incl. prefill = 54 tok/s. Short-prompt C5 aggregate (203.8) is a separate record. MTP acceptance 47% under concurrency.

Reactions

Submitted

Aug 19, 2026, 12:17 PM

Last edited

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4138.31693.9248.622350ms
Show all details for Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

138.3

Prefill tok/s

1693.9

Total tok/s

248.6

TTFT

22350ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

5

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

5

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. Realistic 5-user coding serving on the current stack (vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1, prefix caching on). Workload: 5 concurrent multi-turn coding sessions, ~8K-token session prompts, 512-token generations, 60/60 OK, 0 crashes. tokSOut = sum of per-stream generation rates (5 x 27.7 median per-user tok/s). TTFT is per-turn with ~45K session tokens prefilling per wave; turn-completion wall aggregate incl. prefill = 54 tok/s. Short-prompt C5 aggregate (203.8) is a separate record. MTP acceptance 47% under concurrency.

Reactions

Submitted

Aug 19, 2026, 12:16 PM

Last edited

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4100.21693.970.8369ms
Show all details for Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

100.2

Prefill tok/s

1693.9

Total tok/s

70.8

TTFT

369ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

yes

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. Current full stack: vLLM XPU nightly f01e24f6, MTP4, GDN mixed-split v5, draft-INT4 S+M1 overlay, prefix caching ON (zero hits, unique prompts). C1 client post-first, n=5 median (92.9-105.1), calibrated real-world Pi prompt set, same prompt file as the 2026-08-18 record. p8192/g128 median 100.0; cold input 1694 tok/s at p8192/g1. Prompt-content sensitive: degenerate INDEX filler measures ~71 (44% MTP acceptance) vs 93-96% on realistic text. Cache-off record on the same patches: 112.65.

Reactions

Submitted

Aug 19, 2026, 12:13 PM

Last edited

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4224.24.8285.715419ms
Show all details for Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

224.2

Prefill tok/s

4.8

Total tok/s

285.7

TTFT

15418.9ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

32

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. vLLM XPU nightly digest f01e24f6, MTP4 speculative decoding, GDN mixed-split v5, draft-INT4 S+M1 overlay, prefix caching on, fp8 KV cache, 230W configured cap. Aggregate C32 throughput via lmx remote harness (256 output tokens, 3 iterations). Single-stream C1 is a separate record.

Reactions

Submitted

Aug 19, 2026, 11:32 AM

Last edited

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4200.69.2253.28077ms
Show all details for Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

200.6

Prefill tok/s

9.2

Total tok/s

253.2

TTFT

8077.1ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

16

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. vLLM XPU nightly digest f01e24f6, MTP4 speculative decoding, GDN mixed-split v5, draft-INT4 S+M1 overlay, prefix caching on, fp8 KV cache, 230W configured cap. Aggregate C16 throughput via lmx remote harness (256 output tokens, 3 iterations). Single-stream C1 is a separate record.

Reactions

Submitted

Aug 19, 2026, 11:32 AM

Last edited

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4203.8178.9248.6414ms
Show all details for Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

203.8

Prefill tok/s

178.9

Total tok/s

248.6

TTFT

413.7ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

5

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. vLLM XPU nightly digest f01e24f6, MTP4 speculative decoding, GDN mixed-split v5, draft-INT4 S+M1 overlay, prefix caching on, fp8 KV cache, 230W configured cap. Aggregate C5 throughput via lmx remote harness (256 output tokens, 3 iterations). Single-stream C1 is a separate record.

Reactions

Submitted

Aug 19, 2026, 11:32 AM

Last edited

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int456.8433.470.8171ms
Show all details for Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Model

SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Display name

Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16

Base model

Qwen3.8-27B

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

56.8

Prefill tok/s

433.4

Total tok/s

70.8

TTFT

170.7ms

Peak VRAM

Power draw

Hardware cost

Prompt tokens

74

Output tokens

256

Prefill tokens

Context length

2048

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

vllm

Engine version

Quantization

GPTQ-Int4

Backend

cuda

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

KV cache size

Prefix caching

Attention backend

Flash attention

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

Concurrency

1

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

# Remote endpoint: http://127.0.0.1:8000 servedModel: qwen38

Extra flags

Notes

Self-reported. Single Intel Arc Pro B70 32GB. vLLM XPU nightly digest f01e24f6, MTP4 speculative decoding, GDN mixed-split v5, draft-INT4 S+M1 overlay, prefix caching on, fp8 KV cache, 230W configured cap. Aggregate C1 throughput via lmx remote harness (256 output tokens, 3 iterations). Single-stream C1 is a separate record.

Reactions

Submitted

Aug 19, 2026, 11:32 AM

Last edited

Qwen3.8-27B

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int4+draft-INT4-RTN112.71696.0
Show all details for Qwen3.8-27B

Model

Qwen/Qwen3.8-27B

Display name

Qwen3.8-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

112.7

Prefill tok/s

1696

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

0.27.2rc1.dev77+gac7509e2b XPU

Quantization

GPTQ-Int4+draft-INT4-RTN

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.88

Max running seqs

64

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

B70_DRAFT_LMHEAD_INT4=1 B70_DRAFT_MTP_INT4=1 vllm serve /model --quantization gptq --dtype float16 --max-model-len 131072 --gpu-memory-utilization 0.88 --kv-cache-dtype fp8 --max-num-seqs 64 --max-num-batched-tokens 8192 --no-enable-prefix-caching --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":4}

Extra flags

--max-model-len 131072 --no-enable-prefix-caching --language-model-only --speculative-config {method:mtp,num_speculative_tokens:4}

Notes

NEW identity vs cmsur82fz06svms01ga1f0z83 (BF16 draft 83.7). Same 1x B70, same image vllm/vllm-openai-xpu@sha256:f01e24f6c7ff01f1e0662234255a1372297d1dbd89d003cf13c8fad3eab1ba4f, kernels 0.1.12.3, same target SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16. Extra patches after Qwen MTP nightly+boundary: patch_gdn_mixed_split_v5.py + patch_draft_lmhead_int4.py + patch_draft_mtp_int4.py (runtime RTN INT4 of draft LM head + 5 MTP linears only; target verify stays BF16). tokSOut = client post-first at p512/g128, median n=5, C1, cache OFF (zero prefix_cache_hits_total delta): 112.65 (range 111.58-117.73). Matched same-harness BF16-draft arm was 81.20, not the older Run40 83.7 card. tokSPrefill = actual input/TTFT at p8192/g1 MTP4 n=5: 1696 (flat vs 1691). Accept 510/540=94.44% vs 95.86%. Short agentic +32.8%; 8K/16K long-ctx agentic +37%/+26%. 230 W configured cap; measured median draw 205 W at p512/g128. Speed-only: no token/KL/task-quality parity vs BF16 draft. E2 self-reported. Measured 2026-08-18.

Reactions

Submitted

Aug 19, 2026, 6:31 AM

Last edited

Qwen3.8-27B

28B · Qwen

Intel Arc Pro B70 32GB
vllmGPTQ-Int483.71774.0
Show all details for Qwen3.8-27B

Model

Qwen/Qwen3.8-27B

Display name

Qwen3.8-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

83.7

Prefill tok/s

1774

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

0.27.2rc1.dev77+gac7509e2b XPU

Quantization

GPTQ-Int4

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

fp8

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.88

Max running seqs

64

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

vllm serve /model --quantization gptq --dtype float16 --max-model-len 131072 --gpu-memory-utilization 0.88 --kv-cache-dtype fp8 --max-num-seqs 64 --max-num-batched-tokens 8192 --no-enable-prefix-caching --language-model-only --speculative-config {"method":"mtp","num_speculative_tokens":4}

Extra flags

--max-model-len 131072 --no-enable-prefix-caching --language-model-only --speculative-config {method:mtp,num_speculative_tokens:4}

Notes

Single-stream C1. vLLM XPU nightly vllm/vllm-openai-xpu@sha256:f01e24f6c7ff01f1e0662234255a1372297d1dbd89d003cf13c8fad3eab1ba4f, vllm-xpu-kernels 0.1.12.3. Model SergiioB/Qwen3.8-27B-GPTQ-Int4-sym-G128-MTP-BF16 (Qwen3_5ForConditionalGeneration, dense 27B, GPTQ-INT4 g128 sym desc_act=false, preserved BF16 MTP head). Patches: patch_mtp_nightly.py + patch_mtp_boundary.py. fp8 KV cache, gpu-memory-utilization 0.88, scheduler 8192, max-num-seqs 64, prefix cache OFF (--no-enable-prefix-caching, zero cache-query delta). MTP4 speculative decode, 1-layer draft. tokSOut = client post-first rate at p512/g128, median n=5 (not engine-native decode); tokSPrefill = actual input tokens / client TTFT at p8192/g1, median n=5, no-spec mode (not llama-bench pp). 230 W configured cap; measured campaign-window mean draw 196 W (MTP4 mode). E2 self-reported; independent reproduction pending. Measured 2026-08-15. Claims-audited: 83.7 tok/s MTP4 p512/g128 n=5 range 78.3-84.4; exact-128K full-context p130944/g128 56.3 tok/s n=1. 2000 t/s prefill NOT reached; max cold-input median 1851 t/s (no-spec p2048/g1).

Reactions

🚀

Submitted

Aug 15, 2026, 7:13 PM

Last edited

Qwen3.8-27B

28B · Qwen

Intel Arc Pro B70 32GB
llama.cppQ4_K_M22.31121.4
Show all details for Qwen3.8-27B

Model

Qwen/Qwen3.8-27B

Display name

Qwen3.8-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

22.3

Prefill tok/s

1121.4

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

llama.cpp

Engine version

b10255+ SYCL (071327508) build-sycl-0804

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Qwen3.8-27B-Q4_K_M.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0

Notes

Qwen3.8-27B (dense 27B, hybrid GDN linear+full attention, MTP head not in GGUF quant). Community quant unsloth/Qwen3.8-27B-GGUF Q4_K_M (SHA-256 7b2aec3b...cc89f1b) -- weights released 2026-08-14T15:00Z, benched same day. 230W cap (stock). llama-bench tg128 mean 22.30 t/s (+/-0.01, n=5, C1, cold cache). Cold input pp8192 mean 1121.40 t/s (+/-1.89, n=5); pp512 778.80 / pp2048 1115.99 / pp4096 1137.02. Load at -c 131072 verified, 10.2 GiB VRAM free after load. KV q8_0 K + q4_1 V, FA on. Correctness: coherent smoke 5/5, deterministic replay pass, token exactness pass; reference parity not run. Cross-cap note: at matched 150W cap this model runs tg128 mean 17.78 (n=5, separate submission); prior Qwen3.6-27B Q4_K_M ran tg128 21.3 mean (n=3, 150W, Run 9) -- comparisons across caps are directional only. Dense power scaling 150->230W measured +25% decode on this model.

Reactions

Submitted

Aug 14, 2026, 3:50 PM

Last edited

Qwen3.8-27B

28B · Qwen

Intel Arc Pro B70 32GB
llama.cppQ4_K_M22.31121.4
Show all details for Qwen3.8-27B

Model

Qwen/Qwen3.8-27B

Display name

Qwen3.8-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

22.3

Prefill tok/s

1121.4

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

230W

Engine

llama.cpp

Engine version

b10255+ SYCL (071327508) build-sycl-0804

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Qwen3.8-27B-Q4_K_M.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0

Notes

Qwen3.8-27B (dense 27B, hybrid GDN linear+full attention, MTP head not in GGUF quant). Community quant unsloth/Qwen3.8-27B-GGUF Q4_K_M (SHA-256 7b2aec3b...cc89f1b) -- weights released 2026-08-14T15:00Z, benched same day. 230W cap (stock). llama-bench tg128 mean 22.30 t/s (+/-0.01, n=5, C1, cold cache). Cold input pp8192 mean 1121.40 t/s (+/-1.89, n=5); pp512 778.80 / pp2048 1115.99 / pp4096 1137.02. Load at -c 131072 verified, 10.2 GiB VRAM free after load. KV q8_0 K + q4_1 V, FA on. Correctness: coherent smoke 5/5, deterministic replay pass, token exactness pass; reference parity not run. Cross-cap note: at matched 150W cap this model runs tg128 mean 17.78 (n=5, separate submission); prior Qwen3.6-27B Q4_K_M ran tg128 21.3 mean (n=3, 150W, Run 9) -- comparisons across caps are directional only. Dense power scaling 150->230W measured +25% decode on this model.

Reactions

Submitted

Aug 14, 2026, 3:50 PM

Last edited

Qwen3.8-27B

28B · Qwen

Intel Arc Pro B70 32GB
llama.cppQ4_K_M17.8787.5
Show all details for Qwen3.8-27B

Model

Qwen/Qwen3.8-27B

Display name

Qwen3.8-27B

Base model

Revision

main

Family

Qwen

Parameters

28B

Active params

MoE

no

Output tok/s

17.8

Prefill tok/s

787.5

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

131072

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

AMD Ryzen 7 5700X3D (8 threads)

RAM

30GB

OS

Linux

Power

150W

Engine

llama.cpp

Engine version

b10255+ SYCL (071327508) build-sycl-0804

Quantization

Q4_K_M

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

99

Split mode

KV cache dtype

q8_0/q4_1

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

4096

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

8

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

llama-server -m Qwen3.8-27B-Q4_K_M.gguf -ngl 99 -ncmoe 0 -fa on -ctk q8_0 -ctv q4_1 -c 131072 -b 8192 -ub 4096 -t 8 --no-mmap -dev SYCL0

Extra flags

--ncmoe 0 --ctk q8_0 --ctv q4_1 --b 8192 --dev SYCL0

Notes

Qwen3.8-27B (dense 27B, hybrid GDN linear+full attention, MTP head not in GGUF quant). Community quant unsloth/Qwen3.8-27B-GGUF Q4_K_M (SHA-256 7b2aec3b...cc89f1b) -- weights released 2026-08-14T15:00Z, benched same day. llama-bench tg128 mean 17.78 t/s (+/-0.17, n=5, C1, cold cache). Cold input pp8192 mean 787.45 t/s (+/-1.74, n=5). Load at -c 131072 verified, 10.2 GiB VRAM free after load. KV q8_0 K + q4_1 V, FA on. Correctness: coherent smoke 5/5, deterministic replay pass, token exactness pass; reference parity not run. At matched 150W/build/quant/flags the prior Qwen3.6-27B Q4_K_M ran tg128 21.3 mean (n=3, Run 9); 3.8 mean 17.78 (n=5) -- ~16% slower decode. Comparison directional: different checkpoints, n mismatch (3 vs 5).

Reactions

Submitted

Aug 14, 2026, 3:41 PM

Last edited

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

3B MoE

Intel Arc Pro B70 32GB
vllmGPTQ-INT4-G64-sym-local+DFlash-BF16-local186.67160.0
Show all details for NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

186.6

Prefill tok/s

7160

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU)

Quantization

GPTQ-INT4-G64-sym-local+DFlash-BF16-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

auto

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

8192

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.9

Max running seqs

1

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /mnt/models/nemotron-lightning-gptq-sym64-20260812 --quantization gptq --dtype float16 --max-model-len 16384 --max-num-seqs 1 --max-num-batched-tokens 8192 --gpu-memory-utilization 0.90 --no-enable-prefix-caching --language-model-only --async-scheduling --speculative-config '{"method":"dflash","model":"/mnt/models2/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-DFlash-BF16","num_speculative_tokens":7}' --host 127.0.0.1 --port 8001

Extra flags

--max-model-len 16384 --no-enable-prefix-caching --language-model-only --async-scheduling --speculative-config {"method":"dflash","model":"/mnt/models2/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-DFlash-BF16","num_speculative_tokens":7}

Notes

Self-reported E2, not independently reproduced. C1 (max_num_seqs=1), prefix cache explicitly off (--no-enable-prefix-caching; enable_prefix_caching=False; zero prefix_cache_hits delta). Timing is client monotonic SSE: tokSOut is median client post-first decode at exact p2048/g128 n=5 = 186.61 t/s (range 174.60-201.83; families research/rag/tool/document/assistant). Additional decode cells on the same warm server: p512/g128 median 194.61 (range 140.20-220.01, acceptance 45.1% — wide family spread, not the representative scalar); p8192/g128 median 157.92 (143.50-170.25, acceptance 53.0%). tokSPrefill is the p8192/g1 n=5 median COLD INPUT RATE from client TTFT (7160 t/s, range 7117-7226), which includes scheduling and first-token work — NOT isolated engine prefill. Also p2048/g1 cold input median 6456 t/s. Spec: method=dflash n_spec=7; window acceptance 1830/3521 = 52.0%. Target is a LOCAL symmetric GPTQ INT4 G64 conversion of the named BF16 repo (not an official HF quant). Draft is a LOCAL NVFP4 E2M1→BF16 reconstruction of NVIDIA DFlash, not a published BF16 draft. Stack: vllm/vllm-openai-xpu@sha256:1da0a95485455f08588c11080b9718992fd7d434c6a965d74654903a9d999c57 plus local det image patches (native grouped-topk v2, SSU B8/W4, at::zeros grouped-GEMM); VLLM_XPU_ENABLE_XPU_GRAPH=1; PIECEWISE+FULL graphs; --async-scheduling; --quantization gptq --dtype float16 --max-model-len 16384 --gpu-memory-utilization 0.90. Configured cap 150 W; measured cell-window averages ~149-160 W; peak interval-average 179.3 W (0.5 s energy1_input on p8192/g128); pkg max 68.0 C. Deterministic raw-completion replay smoke exact_match=true (scope: smoke only, not logit/KL or task-quality). n=3 screen 20260813T080234Z is superseded and was not used. Matched-except-speculation vs prior no-spec n=5 graph campaign (p8192/g128 87.25 t/s): 1.81× on that cell only.

Reactions

Submitted

Aug 13, 2026, 8:40 AM

Last edited

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

3B MoE

Intel Arc Pro B70 32GB
vllmGPTQ-INT4-G64-sym-local92.710349.0
Show all details for NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

92.7

Prefill tok/s

10349

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU) + deterministic grouped-GEMM fix

Quantization

GPTQ-INT4-G64-sym-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

default

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.9

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /model --dtype bfloat16 --quantization gptq --max-model-len 16384 --gpu-memory-utilization 0.90 --no-enable-prefix-caching --async-scheduling

Extra flags

--max-model-len 16384 --no-enable-prefix-caching --async-scheduling

Notes

Self-reported. C1 client post-first decode, exact p512/g128 n=3 median (92.65 t/s, range 92.61-92.65), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). GPTQ conversion published at SergiiioB/Nemotron-3.5-Lightning-30B-A3B-GPTQ-INT4-G64-sym. tokSPrefill = cold input rate at p8192 from median client TTFT (10,349 tok/s; TTFT median 0.7916s, n=3, per-rep 10347-10357) — client-observed rate, NOT isolated engine prefill. NO speculative decoding. Deterministic image with grouped-GEMM at::zeros atomic-buffer fix. Deterministic replay verified (8/8 identical temp-0 requests). 150W cap; measured draw ~89-90W. Claims-auditor verified decode; prefill corrected from earlier 10371 (rounding error).

Reactions

Submitted

Aug 12, 2026, 11:54 PM

Last edited

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

3B MoE

Intel Arc Pro B70 32GB
vllmGPTQ-INT4-G64-sym-local92.710371.0
Show all details for NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

92.7

Prefill tok/s

10371

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU) + deterministic grouped-GEMM fix

Quantization

GPTQ-INT4-G64-sym-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

default

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

gptq

SGLang quant

GPU mem util

0.9

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /model --dtype bfloat16 --quantization gptq --max-model-len 16384 --gpu-memory-utilization 0.90 --no-enable-prefix-caching --async-scheduling

Extra flags

--max-model-len 16384 --no-enable-prefix-caching --async-scheduling

Notes

Self-reported. C1 client post-first decode, exact p512/g128 n=3 median (92.65 t/s, range 92.61-92.65), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion of BF16 checkpoint (published at SergiiioB/Nemotron-3.5-Lightning-30B-A3B-GPTQ-INT4-G64-sym). tokSPrefill is cold input rate at p8192 from median client TTFT (10,371 tok/s; TTFT 0.791s, n=3) - client-observed rate, NOT isolated engine prefill. NO speculative decoding (native MTP tested: 0% draft acceptance on this stack; N-gram: non-deterministic; DFlash: unsupported). Deterministic image vllm-openai-xpu:1da0a954-det0123 with grouped-GEMM at::zeros atomic-buffer fix (eliminates the XPU FP race). DETERMINISTIC REPLAY VERIFIED: 8 identical temperature-0 requests produce byte-identical output. 150W cap; measured draw ~89-90W.

Reactions

Submitted

Aug 12, 2026, 10:37 PM

Last edited

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

3B MoE

Intel Arc Pro B70 32GB
vllmGPTQ-INT4-G64-sym-local93.08368.0
Show all details for NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Model

nvidia/NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Display name

NVIDIA-Nemotron-3.5-Lightning-30B-A3B-BF16

Base model

Revision

main

Family

Parameters

32B

Active params

3B

MoE

yes

Output tok/s

93

Prefill tok/s

8368

Total tok/s

TTFT

Peak VRAM

Power draw

Hardware cost

Prompt tokens

0

Output tokens

0

Prefill tokens

Context length

16384

Batch size

1

Hardware class

DISCRETE_GPU

Hardware

Intel Arc Pro B70 32GB

GPU slots

GPU count

1

VRAM

32GB

Chip vendor

Chip family

Chip variant

Unified memory

NPU TOPS

CPU

RAM

OS

Linux

Power

Engine

vllm

Engine version

v0.26.1rc1.dev668+g3ee2df303 (XPU)

Quantization

GPTQ-INT4-G64-sym-local

Backend

xpu

Tensor parallel

Pipeline parallel

GPU layers

Split mode

KV cache dtype

default

KV cache size

Prefix caching

Attention backend

Flash attention

yes

Chunked prefill

Prefill chunk

Continuous batching

CPU offload

CPU layers

Rope scaling

Rope scale

Yarn ext factor

Engine quant

SGLang quant

GPU mem util

0.95

Max running seqs

Scheduler delay

Num parallel

1

Concurrency

Spec decoding

no

Spec method

Spec model

Spec draft model

Spec tokens

Spec ngram

Spec draft TP

Spec draft window

MTP enabled

no

MTP draft layers

Temperature

Top P

Top K

Min P

Repeat penalty

Mirostat

Command

VLLM_XPU_ENABLE_XPU_GRAPH=1 vllm serve /mnt/models/nemotron-lightning-gptq-sym64-20260812 --host 0.0.0.0 --port 8001 --no-enable-prefix-caching --async-scheduling --gpu-memory-utilization 0.95 --max-model-len 16384 --enforce-eager false --trust-remote-code

Extra flags

--no-enable-prefix-caching --async-scheduling --max-model-len 16384 --enforce-eager false

Notes

Self-reported. C1 client post-first decode, exact p512/g128 n=5 median (93.00 t/s, range 92.96-93.03), prefix cache off, XPU graphs (PIECEWISE+FULL compiled). Local symmetric GPTQ INT4 G64 conversion (not an official quant). tokSPrefill is the cold input rate at exact p8192/g128 from median client TTFT (8,368 tok/s, TTFT median 0.9789s, n=3) - a client-observed input rate including scheduling and first-token work, NOT an isolated engine prefill measurement. NO speculative decoding in this config: native MTP was tested and rejected on this stack (0% draft acceptance), N-gram rejected (temperature-0 nondeterminism), DFlash not supported for this model on this stack. 150W configured cap; measured decode draw ~89-90W. KNOWN CAVEAT: temperature-0 deterministic replay does not hold on this stack (XPU compiled-kernel FP race at contested tokens); outputs remain coherent.

Reactions

Submitted

Aug 12, 2026, 9:08 PM

Last edited