Get startedLeaderboardDecode calculatorModelsHardwareBenchmarksMarketplaceRentalsProAPI Docs
Language
Atomic Chat DFlash campaign banner
Decode roofline

How fast can your model decode?

Estimate model fit, per-session speed, and useful aggregate throughput from the memory system up.

Configuration

Start with an example or enter exact values.

Live calculation
01

Hardware memory

Use sustained bandwidth when you have it. Catalog bandwidth produces a looser ceiling.

02

Model weights

Resident weights decide fit. Active weights decide ordinary token traffic.

03

Serving target

The reservation controls memory fit; active context controls decode traffic.

Advanced architecture and KV cache

64 layers · 8 KV heads · 16-bit cache

+