Decode roofline
How fast can your model decode?
Estimate model fit, per-session speed, and useful aggregate throughput from the memory system up.
Configuration
Start with an example or enter exact values.
01
Hardware memory
Use sustained bandwidth when you have it. Catalog bandwidth produces a looser ceiling.
02
Model weights
Resident weights decide fit. Active weights decide ordinary token traffic.
03
Serving target
The reservation controls memory fit; active context controls decode traffic.
Advanced architecture and KV cache
64 layers · 8 KV heads · 16-bit cache
+
Advanced architecture and KV cache
64 layers · 8 KV heads · 16-bit cache
