MulaiPapan peringkatDecode calculatorModelReportsHardwareTolok UkurMarketplaceSewaProDokumen API
Bahasa
Actual Computer — Every computer, one endpoint
Community field report

Strata (Qwen3.8-Flash-Next IQ3_S) on a single RTX 5080

125B MoE model served locally on one 16 GB RTX 5080 plus 92 GB RAM: ~80-99 tok/s decode, prefill scaling to 140k tok/s at 32K context, MTP speculative decoding at ~80% acceptance.

A 125B-parameter mixture-of-experts model served entirely locally on one consumer GPU plus system RAM. Measured with the LocalMaxxing CLI against Strata's OpenAI-compatible endpoint.

Setup

  • Model: Qwen3.8-Flash-Next (original), IQ3_S size, served natively by the Strata engine
  • Engine: Strata 0.1.38 (self-compiled for Blackwell / compute 12.0; commit 99f3dbd)
  • GPU: NVIDIA GeForce RTX 5080, 15.9 GB VRAM, driver 580.159.04
  • CPU: AMD Ryzen 7 9800X3D (AVX-512), 16 threads
  • RAM: 92 GB (the engine loads ~55 GB of expert weights into RAM at startup)
  • OS: Linux (kernel 7.0), CUDA 13.1
  • Disk: model files on a separate NVMe (Strata-data); ~84 GB installed

Relevant engine settings (strata-iq3_s.json):

bash
strata generate \
  --pack <Strata-data>/packs/iq3_s \
  --native <Strata-data>/models/IQ3_S/...-00001-of-00002.gguf \
  --ple-gguf <Strata-data>/models/IQ3_S/...-00002-of-00002.gguf \
  --mtp <Strata-data>/mtp/rt \
  --max-context 65536 --kv int8 --kv-resident 32768 \
  --spec 4 --spec-min-p 0.5 --expert-cache auto --prefill auto

Startup: weights are read from disk into RAM at about 1.5 GiB/s (~51 s), then the GPU's expert cache is filled (4,217 experts, ~8 GiB of VRAM). The engine reports 257 MiB of VRAM free with everything loaded.

Methodology

  • Remote speed test via lmx speed-test run custom --mode remote, one warm-up plus three timed iterations; the median is reported.
  • Speculative decoding (MTP draft head) is on for every run; it is recorded as specMethod: MTP.
  • Decode throughput is measured over the inter-token streaming window, not including prefill. TTFT is the time from request to the first streamed token.
  • The context sweep (lmx kvcache run) resends the full prefix at each depth, so each point is a cold prefill plus its decode. The engine does not expose a persistent-KV-cache session API, so cache reuse between requests could not be measured through this path (the sweep flags this).
  • Quantization IQ3_S; the label comes from the loaded GGUF filename.

Results

Short-prompt run (253 prompt tokens, 512 generated):

MetricValue
Decode96.6 tok/s (median 98.8)
TTFT561 ms

Cold-context sweep (128 generated tokens per point):

ContextDecode (tok/s)Prefill (tok/s)TTFT (ms)
4,09680.619,200213
16,09693.283,485193
32,09699.4140,423229

Engine-side counters during these runs (from strata-iq3_s.log):

  • Speculative acceptance: 78-81% of MTP drafts accepted.
  • Decode expert-cache hit rate: 77-90% (VRAM hits / lookups).
  • KV streaming: 96-99% of block reads hit VRAM.

The engine's own end-to-end prefill timings for the cold sweep requests were 1,973 tok/s (4K), 3,463 tok/s (16K), and 3,791 tok/s (32K) - i.e. reading a fresh long prompt costs about 2-3.8k tokens/s, while the higher tokSPrefill figures above reflect the shorter incremental read once the bulk of the prefix is cached.

Observations

  • Finding: The model is usable interactively on a single 16 GB card - decode holds ~80-99 tok/s and TTFT stays under ~250 ms out to 32K context. Prefill throughput scales cleanly with prompt length (tokens/s nearly doubling from 4K to 32K), which is the expected signature of a compute-bound prefill.
  • Finding: Speculative decoding via the MTP draft head is doing real work: ~78-81% acceptance, consistent with roughly 2x effective output speed.
  • Caveat: The KV-cache context path is not an Apple-to-apple cached-context benchmark - Strata has no portable persistent-session API, so the sweep measures cold prefill. Real conversational reuse should show lower effective prefill cost.
  • Caveat: Runs are unverified on LocalMaxxing: the OpenAI-compatible client path does not capture the prompt SHA-256, the raw engine timings blob, or the explicit MTP specDraftTokens/specAcceptedTokens counts, so the site marks them verifiedRun: false. Numbers are measured, not model-generated.
  • Caveat: Peak VRAM was not captured; the engine reported ~257 MiB free with everything resident, implying near-full VRAM use on the 5080.
  • Note: The disk footprint is tight - IQ3_S needs ~84-90 GB and the first download filled the target drive before completing.

Reproduction

bash
# speed test (short prompt / long output)
lmx speed-test run custom --mode remote \
  --base-url http://127.0.0.1:8080 \
  --hf-id Qwen/Qwen3.8-Flash-Next --served-model strata \
  --quantization IQ3_S --spec-method mtp \
  --max-tokens 512 --hardware hardware.json

# context sweep
lmx kvcache run custom --mode remote \
  --base-url http://127.0.0.1:8080 \
  --hf-id Qwen/Qwen3.8-Flash-Next --served-model strata \
  --levels 4000,16000,32000 --output-tokens 128 --hardware hardware.json
Attached evidenceLocalMaxxing runs referenced by this report — open any run to see the full submission
Discussion

2 comments

Questions, reproduction notes, and follow-up results.

S
@soulrider4ever

This is really amazing, great job - I tried this on my MI210 (not supported) and got worse numbers, this is really game changing though! Can’t wait for Qwen4.0

S
@soulrider4ever

https://www.localmaxxing.com/en/runs/cmv2mmxf300j9mr01tvgevk08 was my iq3_s run