Aan de slagLeaderboardDecode calculatorModellenReportsHardwareBenchmarksMarktplaatsVerhuurProAPI Docs
Taal
Actual Computer — Every computer, one endpoint

Verdachte runs

These submissions report an output speed more than 10× the most generous theoretical decode ceiling for their hardware, model and quantization — including full-acceptance speculative decoding. They are excluded from the leaderboards and model pages until proven.

If one of these is yours and you can reproduce it, contact the admins with a canonical-prompt run captured by the lmx CLI (engine timings, prompt/output hashes, acceptance stats). Approved appeals move the run back onto the leaderboard. Back to the leaderboard

#1Qwen3.8-27B-NVFP4Google Cloud TPU v6e 256 Host Intel Xeon Platinumvllm · 1-bit7,516,192,768,000,000,000 tok/s

Need more proof of what's happening here

anonymous · 9/4/2026 · batch 1 · 2.000.000.000 output tokens

#2LFM2-350M-NVFP4A16RTX 5090 32GBsglang · NVFP417,682,412,430,045,650 tok/s

Need to see outputs and further proof

anonymous · 9/4/2026 · batch 1 · 512 output tokens

#3LFM2-350M-NVFP4A16RTX 5090 32GBsglang · NVFP4 · DFlash217,682,412,430,045,650 tok/s

dflash draft tokens 2000000000

anonymous · 9/4/2026 · batch 1 · 2.000.000.000 output tokens

#4LFM2-350M-NVFP4A16RTX 5090 32GBsglang · NVFP4 · DFlash24,184,615,291,759,722 tok/s

draft tokens 2000000000

anonymous · 9/3/2026 · batch 1 · 2.000.000.000 output tokens

#5LFM2-350M-NVFP4A16RTX 5090 32GBsglang · NVFP4 · DFlash267,721,558,877,207.3 tok/s

draft tokens 335539200

anonymous · 9/3/2026 · batch 1 · 335.539.200 output tokens

#6LFM2-350M-NVFP4A16RTX 5090 32GBsglang · NVFP4 · DFlash23,414,359,934,193.1 tok/s

draft tokens 16711680

anonymous · 9/3/2026 · batch 1 · 16.711.680 output tokens

#7LFM2-350M-NVFP4A16RTX 5090 32GBsglang · NVFP4 · DFlash222,149,160,757.7 tok/s

draft tokens 131072

anonymous · 9/3/2026 · batch 1 · 131.072 output tokens

#8Qwen3.8-27B-NVFP4Google Cloud TPU v6e Host Intel Xeon Platinumvllm · BitNet b1.5810,361,459.6 tok/s

VRAM not updated, and need more proof of what's happening here

anonymous · 9/4/2026 · batch 64 · 512 output tokens

#9GLM-5.3-BF16Google Cloud TPU v6e Host Intel Xeon Platinumvllm · BF16 · DFlash6,843,561.6 tok/s

Need further outputs to confirm

anonymous · 9/4/2026 · batch 32 · 512 output tokens

#10LFM2-350M-NVFP4A16RTX 5090 32GBsglang · NVFP4 · dflash4,508,232.9 tok/s

draft tokens 16384, requires large output token budget for accurate speed calcs

anonymous · 9/3/2026 · batch 1 · 16.384 output tokens

#11LFM2-350M-NVFP4A16RTX 5090 32GBsglang · NVFP4 · dflash1,854,541.6 tok/s

requires large output token budget for accurate speed calcs

anonymous · 9/3/2026 · batch 1 · 16.384 output tokens

#12LFM2-350M-NVFP4A16RTX 5090 32GBsglang · NVFP4 · dflash1,684,543 tok/s

requires large output token budget for accurate speed calcs

anonymous · 9/3/2026 · batch 1 · 8.192 output tokens

#13LFM2-350M-NVFP4A16RTX 5090 32GBsglang · NVFP4 · dflash1,107,027 tok/s

requires large output token budget for accurate speed calcs

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#14LFM2-350M-NVFP4A16RTX 5090 32GBsglang · NVFP4 · dflash71,381.1 tok/s

requires large output token budget for accurate speed calcs

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#15Qwen3.8-27B-NVFP4RTX 5090 32GBsglang · NVFP4 · ngram4,765.7 tok/s

No output sample, prefill tok too few, and draft details not found

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#16Qwen3.8-27B-NVFP4RTX 5090 32GBsglang · NVFP4 · ngram4,743.5 tok/s

No output sample, prefill tok too few, and draft details not found

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#17Qwen3.8-27B-NVFP4RTX 5090 32GBsglang · NVFP4 · dflash4,726.6 tok/s

No output sample, prefill tok too few, and draft details not found

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#18Qwen3.8-27B-NVFP4RTX 5090 32GBsglang · NVFP4 · dflash4,723.4 tok/s

prefill tok too few, no output or prompt samples

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#19Qwen3.8-27B-NVFP4RTX 5090 32GBsglang · NVFP4 · dflash4,691.3 tok/s

no output or input samples, dflash draft 512, too few input tok

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#20Qwen3.8-27B-NVFP4RTX 5090 32GBsglang · NVFP4 · dflash4,666.8 tok/s

512 dflash draft, no output or input samples, too few input/output tokens

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#21Qwen3.8-27B-NVFP4RTX 5090 32GBsglang · NVFP4 · dflash4,654.9 tok/s

512 dflash draft, no output or input samples, too few input/output tokens

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#22Qwen3.8-27B-NVFP4RTX 5090 32GBsglang · NVFP4 · dflash4,621.9 tok/s

512 dflash draft, no output or input samples, too few input/output tokens

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#23Qwen3.8-27B-NVFP4RTX 5090 32GBsglang · NVFP4 · ngram4,613.8 tok/s

No output sample, prefill tok too few, and draft details not found

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#24Qwen3.8-27B-NVFP4RTX 5090 32GBsglang · NVFP44,546.8 tok/s

Output sample doesn't make sense and draft details not found

anonymous · 9/4/2026 · batch 1 · 512 output tokens

#25Qwen3.8-27B-NVFP4RTX 5090 32GBsglang · NVFP4 · dflash2,103.6 tok/s

no input or output samples, high dflash draft number

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#26GLM-5.3-FlashRTX PRO 6000 Blackwell 96GB ×2sglang · NVFP4 · dflash1,493.6 tok/s

high dflash draft number, no output or input samples

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#27GLM-5.3-FlashRTX PRO 6000 Blackwell 96GB ×2sglang · NVFP4 · dflash1,205.7 tok/s

high dflash draft number 64, no input or output samples

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens

#28Qwen3.8-27B-NVFP4RTX 5090 32GBsglang · NVFP4 · dflash1,074.5 tok/s

high dflash draft size 24, no input or output samples

anonymous · 9/3/2026 · batch 1 · 2.048 output tokens