» Cedana / Capacity planning
Tokens to GPUs Calculator
Tokens have become the unit that many practitioners use to report AI growth. The trouble is that the numbers are hard to picture. Is a trillion tokens a month a lot? It depends. The same volume needs a different amount of hardware depending on the GPU, the model, and the workload (agentic coding, chat). So we built a calculator. Enter a token volume and pick a model. It tells you how many GPUs that works out to, with a range that shows how firm the estimate is. Every default links to its source, and you can swap in your own measurements.
1 trillion tokens in one month (30 days) on Llama 3.3 70B needs approximately 367 H100 GPUs. Confidence is moderate.
How the calculator gets the H100 number
Each step changes one quantity. Read from left to right.
- 01Start1 trilliontokens in one month (30 days)
- 02÷ 2,592,000 seconds385,802tokens per second, average
- 03× 0.475 for the token mix183,256output-equivalent tokens per second
- 04÷ 50% utilization366,512tokens per second of installed capacity
- 05÷ 1,000 tok/s for each H100367H100 GPUs
Which assumption moves the H100 result most
Each bar shows the GPU count when one assumption moves from its best case to its worst case. The other assumptions stay at the base value. The longest bar is the assumption to measure first.
| GPU | tok/s per GPU | Tokens per GPU per day | GPUs per copy | Low | Base | High | 8-GPU nodes |
|---|---|---|---|---|---|---|---|
| A100 | 450 | 40.93M | 1 | 436 | 815 | 1,966 | 102 |
| V100 | 144 | 13.10M | 3 | 1,338 | 2,547 | 7,122 | 319 |
| H100 | 1,000 | 90.95M | 1 | 211 | 367 | 853 | 46 |
| B200 | 4,000 | 363.79M | 1 | 47 | 92 | 252 | 12 |
Reference speed is tokens per second for one GPU, as low, base, and high. "Measured" values come from public benchmarks across loose and strict latency targets. "Estimated" values have no direct benchmark.
| Model | Total / active | Weights | Reference speed | B200 ÷ H100 | Basis | Source |
|---|---|---|---|---|---|---|
| Llama 3.1 8B | 8B / 8B | 8 GB | H100: 3,000, 6,000, 12,500 | 2.3, 4, 6 | Estimated. One measured point for the high value. | morph » |
| Llama 3.3 70B | 70B / 70B | 70 GB | H100: 450, 1,000, 1,500 | 2.3, 4, 6 | Measured on H100 and B200. | inferencex » cerebrium » |
| gpt-oss 120B | 117B / 5.1B | 65 GB | H100: 740, 1,400, 2,600 | 6, 9, 12 | Measured on H100 and B200. Weight size is not verified. | inferencex » |
| DeepSeek V4 Flash | 284B / 13B | 160 GB | B200: 1,500, 4,000, 9,000 | 5, 9, 14 | Estimated. No benchmark found. Scaled from V4 Pro and gpt-oss by active parameters. | datacamp » |
| MiniMax M3 428B | 428B / n/a | 428 GB | H100: 157, 240, 537 | 3.5, 4.5, 5.3 | Measured on H100 and B200. Weight size assumes 8-bit. | inferencex » |
| DeepSeek R1 / V3 | 671B / 37B | 671 GB | H100: 23, 75, 266 | 12, 16, 21 | Measured on H100 and B200. Weight size assumes 8-bit. | inferencex » |
| GLM-5.3 | 753B / 40B | 755 GB | B200: 1,400, 3,000, 7,000 | 4, 8, 16 | Measured on B200 only. Sources disagree: GLM-5 at FP4 gives 1,417 to 2,935, GLM-5.3 gives 4,329 to 10,243. H100 ratio is estimated. | inferencex 5.3 » inferencex 5 fp4 » size » |
| DeepSeek V4 Pro | 1600B / 49B | 865 GB | B200: 468, 1,000, 2,800 | 6, 14, 21 | Measured on B200 only. H100 ratio is estimated from DeepSeek R1. | inferencex » size » |
| Assumption | Value | Basis | Source |
|---|---|---|---|
| A100 ÷ H100 | 0.35, 0.45, 0.60 | H100 measured at 1.8x to 2.9x the A100 on 70B models | hyperstack » perplexity » |
| V100 ÷ H100 | 0.10, 0.18, 0.25 | Specifications only: 125 against 312 TFLOPS, 900 against 2,039 GB/s. No LLM serving benchmark. | spheron » |
| Input cost | 0.1, 0.3, 0.5 | Derived from 128:128 and 2024:128 runs on 4x H100 (result 0.33) | e2e networks » |
| GPU memory | 32, 80, 80, 180 GB | V100, A100, H100, B200. The fit check adds 10% to the weight size. | inferencex » |
Values read on 2 October 2026.
- The result is a planning estimate. Expect an error of 2x in each direction for measured presets, and more for estimated presets. Measure your model on your hardware before you buy.
- Speed changes with the latency target. The low and high values of each preset come from strict and loose latency targets.
- For large mixture-of-experts models, the B200 lead over the H100 is 4x to 21x. FP4 support and memory size cause this. The ratio is specific to each model.
- A100 and V100 results for the large models are theoretical. One model copy needs more GPUs than one node contains.
- The V100 value has no measured LLM serving data. It comes from hardware specifications.
- The range combines independent errors as a root sum of squares in log space. It is not a statistical confidence interval.
- The calculator does not include prompt caching, speculative decoding, reasoning-token overhead, failures, or regional duplication.