Roofline Model Calculator
From
Why More Compute Does Not Mean Faster AI
by Mona Mishra
Chip Configuration
H100 SXM5
TPU v4
A100 80GB
H200
Peak Compute (TFLOPS, BF16)
HBM Bandwidth (GB/s)
Your Workload
Arithmetic Intensity (FLOPs/Byte)
Quick Estimate
— Select a workload —
LLM Decode (batch=1) — ~0.5
LLM Decode (batch=16) — ~8
LLM Decode (batch=128) — ~64
LLM Prefill (short seq) — ~150
LLM Prefill (long seq) — ~400
Dense training step — ~1000
Ridge Point
—
Attainable Perf
—
MXU Utilization
—
of peak compute
—
Roofline Chart