» Cedana / Capacity planning

Tokens to GPUs Calculator

Tokens have become the unit that many practitioners use to report AI growth. The trouble is that the numbers are hard to picture. Is a trillion tokens a month a lot? It depends. The same volume needs a different amount of hardware depending on the GPU, the model, and the workload (agentic coding, chat). So we built a calculator. Enter a token volume and pick a model. It tells you how many GPUs that works out to, with a range that shows how firm the estimate is. Every default links to its source, and you can swap in your own measurements.

:: 70b total / 70b active / 70 gb weights / ref h100 / moderate confidence
Sets the GPU for the headline, the calculation, and the sensitivity diagram. All GPU types stay in the chart.
An input token costs less GPU time than an output token.
Average load divided by installed capacity. Traffic peaks and spare capacity make this less than 100%.
Advanced assumptions
AssumptionLowBaseHigh
H100 tok/s
B200 ÷ H100
A100 ÷ H100
V100 ÷ H100
Input cost
The first row is output tokens per second for one GPU of the reference type at full load. Input cost is the GPU time for one input token divided by the time for one output token.
Applied when the reference GPU holds one model copy and the other GPU type does not.
» GPU count
367
H100 GPUs / base
211
Low
853
High
46
8x H100 nodes

1 trillion tokens in one month (30 days) on Llama 3.3 70B needs approximately 367 H100 GPUs. Confidence is moderate.

1101001k10k
Single GPUs
Nodes with 8 GPUs
Base estimatePlausible rangeLog axis. Select a row to see its calculation.
One model copy needs 3 V100 GPUs (77 GB against 32 GB). Treat the V100 result as theoretical.
» Calculation

How the calculator gets the H100 number

Each step changes one quantity. Read from left to right.

  1. 01
    Start
    1 trillion
    tokens in one month (30 days)
  2. 02
    ÷ 2,592,000 seconds
    385,802
    tokens per second, average
  3. 03
    × 0.475 for the token mix
    183,256
    output-equivalent tokens per second
  4. 04
    ÷ 50% utilization
    366,512
    tokens per second of installed capacity
  5. 05
    ÷ 1,000 tok/s for each H100
    367
    H100 GPUs
gpus = tokens ÷ seconds × [output share + input share × input cost] ÷ utilization ÷ tok/s for each gpu
» Sensitivity

Which assumption moves the H100 result most

Each bar shows the GPU count when one assumption moves from its best case to its worst case. The other assumptions stay at the base value. The longest bar is the assumption to measure first.

H100 speed for this modelmeasurement uncertainty
244 at 1,500 tok/s
814 at 450 tok/s
Average utilizationyour operating choice
229 at 80%
611 at 30%
Cost of an input tokenmeasurement uncertainty
251 at 0.1 of output
482 at 0.5 of output
Measurement uncertaintyYour operating choiceVertical line: base, 367 H100 GPUs before rounding
» Per-GPU detail
GPUtok/s per GPUTokens per GPU per dayGPUs per copyLowBaseHigh8-GPU nodes
A10045040.93M14368151,966102
V10014413.10M31,3382,5477,122319
H1001,00090.95M121136785346
B2004,000363.79M1479225212
» Model presets and evidence

Reference speed is tokens per second for one GPU, as low, base, and high. "Measured" values come from public benchmarks across loose and strict latency targets. "Estimated" values have no direct benchmark.

ModelTotal / activeWeightsReference speedB200 ÷ H100BasisSource
Llama 3.1 8B8B / 8B8 GBH100: 3,000, 6,000, 12,5002.3, 4, 6Estimated. One measured point for the high value.morph »
Llama 3.3 70B70B / 70B70 GBH100: 450, 1,000, 1,5002.3, 4, 6Measured on H100 and B200.inferencex »
cerebrium »
gpt-oss 120B117B / 5.1B65 GBH100: 740, 1,400, 2,6006, 9, 12Measured on H100 and B200. Weight size is not verified.inferencex »
DeepSeek V4 Flash284B / 13B160 GBB200: 1,500, 4,000, 9,0005, 9, 14Estimated. No benchmark found. Scaled from V4 Pro and gpt-oss by active parameters.datacamp »
MiniMax M3 428B428B / n/a428 GBH100: 157, 240, 5373.5, 4.5, 5.3Measured on H100 and B200. Weight size assumes 8-bit.inferencex »
DeepSeek R1 / V3671B / 37B671 GBH100: 23, 75, 26612, 16, 21Measured on H100 and B200. Weight size assumes 8-bit.inferencex »
GLM-5.3753B / 40B755 GBB200: 1,400, 3,000, 7,0004, 8, 16Measured on B200 only. Sources disagree: GLM-5 at FP4 gives 1,417 to 2,935, GLM-5.3 gives 4,329 to 10,243. H100 ratio is estimated.inferencex 5.3 »
inferencex 5 fp4 »

size »
DeepSeek V4 Pro1600B / 49B865 GBB200: 468, 1,000, 2,8006, 14, 21Measured on B200 only. H100 ratio is estimated from DeepSeek R1.inferencex »
size »
» Other assumptions
AssumptionValueBasisSource
A100 ÷ H1000.35, 0.45, 0.60H100 measured at 1.8x to 2.9x the A100 on 70B modelshyperstack »
perplexity »
V100 ÷ H1000.10, 0.18, 0.25Specifications only: 125 against 312 TFLOPS, 900 against 2,039 GB/s. No LLM serving benchmark.spheron »
Input cost0.1, 0.3, 0.5Derived from 128:128 and 2024:128 runs on 4x H100 (result 0.33)e2e networks »
GPU memory32, 80, 80, 180 GBV100, A100, H100, B200. The fit check adds 10% to the weight size.inferencex »

Values read on 2 October 2026.

» Limits of this model
  • The result is a planning estimate. Expect an error of 2x in each direction for measured presets, and more for estimated presets. Measure your model on your hardware before you buy.
  • Speed changes with the latency target. The low and high values of each preset come from strict and loose latency targets.
  • For large mixture-of-experts models, the B200 lead over the H100 is 4x to 21x. FP4 support and memory size cause this. The ratio is specific to each model.
  • A100 and V100 results for the large models are theoretical. One model copy needs more GPUs than one node contains.
  • The V100 value has no measured LLM serving data. It comes from hardware specifications.
  • The range combines independent errors as a root sum of squares in log space. It is not a statistical confidence interval.
  • The calculator does not include prompt caching, speculative decoding, reasoning-token overhead, failures, or regional duplication.