A 125B-parameter mixture-of-experts model served entirely locally on one consumer GPU plus system RAM. Measured with the LocalMaxxing CLI against Strata's OpenAI-compatible endpoint.
Setup
- Model: Qwen3.8-Flash-Next (original), IQ3_S size, served natively by the Strata engine
- Engine: Strata 0.1.38 (self-compiled for Blackwell / compute 12.0; commit
99f3dbd) - GPU: NVIDIA GeForce RTX 5080, 15.9 GB VRAM, driver 580.159.04
- CPU: AMD Ryzen 7 9800X3D (AVX-512), 16 threads
- RAM: 92 GB (the engine loads ~55 GB of expert weights into RAM at startup)
- OS: Linux (kernel 7.0), CUDA 13.1
- Disk: model files on a separate NVMe (
Strata-data); ~84 GB installed
Relevant engine settings (strata-iq3_s.json):
strata generate \
--pack <Strata-data>/packs/iq3_s \
--native <Strata-data>/models/IQ3_S/...-00001-of-00002.gguf \
--ple-gguf <Strata-data>/models/IQ3_S/...-00002-of-00002.gguf \
--mtp <Strata-data>/mtp/rt \
--max-context 65536 --kv int8 --kv-resident 32768 \
--spec 4 --spec-min-p 0.5 --expert-cache auto --prefill auto
Startup: weights are read from disk into RAM at about 1.5 GiB/s (~51 s), then the GPU's expert cache is filled (4,217 experts, ~8 GiB of VRAM). The engine reports 257 MiB of VRAM free with everything loaded.
Methodology
- Remote speed test via
lmx speed-test run custom --mode remote, one warm-up plus three timed iterations; the median is reported. - Speculative decoding (MTP draft head) is on for every run; it is recorded as
specMethod: MTP. - Decode throughput is measured over the inter-token streaming window, not including prefill. TTFT is the time from request to the first streamed token.
- The context sweep (
lmx kvcache run) resends the full prefix at each depth, so each point is a cold prefill plus its decode. The engine does not expose a persistent-KV-cache session API, so cache reuse between requests could not be measured through this path (the sweep flags this). - Quantization
IQ3_S; the label comes from the loaded GGUF filename.
Results
Short-prompt run (253 prompt tokens, 512 generated):
| Metric | Value |
|---|---|
| Decode | 96.6 tok/s (median 98.8) |
| TTFT | 561 ms |
Cold-context sweep (128 generated tokens per point):
| Context | Decode (tok/s) | Prefill (tok/s) | TTFT (ms) |
|---|---|---|---|
| 4,096 | 80.6 | 19,200 | 213 |
| 16,096 | 93.2 | 83,485 | 193 |
| 32,096 | 99.4 | 140,423 | 229 |
Engine-side counters during these runs (from strata-iq3_s.log):
- Speculative acceptance: 78-81% of MTP drafts accepted.
- Decode expert-cache hit rate: 77-90% (VRAM hits / lookups).
- KV streaming: 96-99% of block reads hit VRAM.
The engine's own end-to-end prefill timings for the cold sweep requests were
1,973 tok/s (4K), 3,463 tok/s (16K), and 3,791 tok/s (32K) - i.e. reading a fresh
long prompt costs about 2-3.8k tokens/s, while the higher tokSPrefill figures
above reflect the shorter incremental read once the bulk of the prefix is cached.
Observations
- Finding: The model is usable interactively on a single 16 GB card - decode holds ~80-99 tok/s and TTFT stays under ~250 ms out to 32K context. Prefill throughput scales cleanly with prompt length (tokens/s nearly doubling from 4K to 32K), which is the expected signature of a compute-bound prefill.
- Finding: Speculative decoding via the MTP draft head is doing real work: ~78-81% acceptance, consistent with roughly 2x effective output speed.
- Caveat: The KV-cache context path is not an Apple-to-apple cached-context benchmark - Strata has no portable persistent-session API, so the sweep measures cold prefill. Real conversational reuse should show lower effective prefill cost.
- Caveat: Runs are unverified on LocalMaxxing: the OpenAI-compatible client
path does not capture the prompt SHA-256, the raw engine timings blob, or the
explicit MTP
specDraftTokens/specAcceptedTokenscounts, so the site marks themverifiedRun: false. Numbers are measured, not model-generated. - Caveat: Peak VRAM was not captured; the engine reported ~257 MiB free with everything resident, implying near-full VRAM use on the 5080.
- Note: The disk footprint is tight - IQ3_S needs ~84-90 GB and the first download filled the target drive before completing.
Reproduction
# speed test (short prompt / long output)
lmx speed-test run custom --mode remote \
--base-url http://127.0.0.1:8080 \
--hf-id Qwen/Qwen3.8-Flash-Next --served-model strata \
--quantization IQ3_S --spec-method mtp \
--max-tokens 512 --hardware hardware.json
# context sweep
lmx kvcache run custom --mode remote \
--base-url http://127.0.0.1:8080 \
--hf-id Qwen/Qwen3.8-Flash-Next --served-model strata \
--levels 4000,16000,32000 --output-tokens 128 --hardware hardware.json

This is really amazing, great job - I tried this on my MI210 (not supported) and got worse numbers, this is really game changing though! Can’t wait for Qwen4.0
https://www.localmaxxing.com/en/runs/cmv2mmxf300j9mr01tvgevk08 was my iq3_s run