TL;DR
- Every published rule for owning GPUs instead of renting them turns on one number: how much of the time you expect to keep them busy. The guides put the break-even at 50 to 70% sustained utilization for training or self-hosting, and at 2 million to 50 million tokens a day for inference.
- Every one of those rules asks you to forecast demand. None asks how many of the hours you already pay for end in finished work. On a real fleet that number is lower than the dashboard says, because GPUs spend paid hours cold-starting, redoing work a failure destroyed, and sitting where no job can use them.
- The on-prem formulas price a GPU as if it works whenever it is switched on, and a node bills every hour whether or not it serves. So the number to put in the calculation is the one you can measure on your own stack: the share of paid GPU-hours that became finished work. Which number you pick decides which side of the break-even you land on.
- In this piece we walk through what the calculators say, how the on-prem formula prices an hour, what renting has cost against owning, and the dated rates on offer today. The second half covers the four numbers that set your own break-even, why a node bills while idle, where the missing hours go, and what changes when a job can be saved and moved.
What the calculators tell you the threshold is
At what point does owning GPUs become cheaper than renting them? An r/LocalLLaMA thread asking exactly that has more than 120 comments, and on r/learnmachinelearning a poster says the rent-versus-own math is starting to feel broken for teams doing under 50 hours a month of GPU work.
Different guides give you different thresholds. Spheron, in an April 2026 guide, puts it at "at under 70% GPU utilization, cloud wins on total cost of ownership". Introl, writing in February 2026, states that "Self-hosted breakeven requires 50%+ GPU utilization for 7B models". Between them, the guides work in a range of 50 to 70% sustained utilization.
For inference, the threshold comes in tokens rather than hours. Spheron's April 2026 figures are "At 16M+ tokens/day on H100 on-demand or 22M+ on H200, self-hosting on GPU cloud is cheaper". The AI Engineer puts the floor lower: "Self-hosting inference pays off in two cases: you push past roughly two million tokens a day, or your data legally cannot leave your network". Across the guides, the crossover runs from 2 million to 50 million tokens a day.
Every one of those rules asks you for a forecast: how busy the fleet will be, or how many tokens a day you will send it. None asks how many of the hours you pay for produce finished work.
The on-prem formula prices every hour the node exists
Spheron publishes the owning side as a formula with every term named: hardware depreciation, electricity, cooling folded into the power usage effectiveness (PUE) figure, networking, and staff. Its worked example prices an 8x H100 SXM5 server at approximately $350,000 and spreads it over 36 months at 720 hours a month, which is $13.50 an hour for the node, or $1.69 per GPU-hour. Electricity adds $0.21 per GPU-hour, from 0.7 kW times 1.80 for server overhead, times 1.4 for PUE, at $0.12 per kWh.
With a staffing allocation on top, the page reaches "approximately $4.75/hr per GPU, before networking, maintenance contracts, and facility lease costs". On the same page, Spheron's own on-demand H100 rate is $2.90 an hour, below the owned figure it has just worked out, and the page tells you to put in your own electricity rate, PUE, staff allocation and server price.
The number to look at is the 720 hours a month. That is every hour the node exists, not the hours it runs a job, so the formula prices an owned GPU as though it works whenever it is switched on. Divide the same capital by the hours the node delivered work and the cost per delivered hour goes up. A Reddit poster totalling the cost of self-hosting in France counts that in: the full cost, they write, has to include "hardware depreciation, upfront capex, replacement parts, cooling, noise, internet, storage, and the fact that my machines are not generating tokens 24/7 like a commercial service would".
Renting the same hardware has cost about twice as much, for training runs
Epoch AI worked the comparison both ways for frontier training runs. In "How much does it cost to train frontier AI models?", published 3 June 2024, it prices each run as amortized hardware and energy for hardware the lab owns, then again as rented cloud hardware, and finds the rented estimates "around twice as high on average".
That number covers less than it seems to. It is the final training run of a frontier model, priced from disclosed or estimated hardware and energy use, not a live rental quote and not an inference fleet. It also prices a run that finished, and nothing in the method counts one that was interrupted and started again.
GPU-hour prices move, so the rate you use needs a date
The price of a GPU-hour moves month to month. The one-year H100 contract index that SemiAnalysis tracks rose from $1.70 to $2.35 between October 2025 and March 2026, and one of the list prices in the table below is a promotion that ends on 30 September 2026. So any rate you put in the calculation is only true on the day it was read, and the date belongs next to it.
| Source | Hardware | Rate per GPU-hour | Date the rate was read |
|---|---|---|---|
| GetDeploying, median across 40 providers | H100, on demand | $3.38 | 6 September 2026, as the page states |
| GMI Cloud list price | H100 | from $2.00 | September 2026 |
| GMI Cloud list price | B200 | from $4.00 | September 2026 |
| Lambda, 8x on-demand instance | H100 SXM | $3.99 | September 2026 |
| Lambda, 8x on-demand instance | B200 SXM6 | $6.69 | September 2026 |
| Together AI list price | HGX H100 | $5.49 | September 2026 |
| Together AI promotion, "valid until 09/30/26" | HGX H100 | $3.99 | September 2026 |
| SemiAnalysis one-year rental contract index | H100 | $1.70 rising to $2.35 | October 2025 and March 2026, reported April 2026 |
GetDeploying sums up the spread in one line: "As of September 6, 2026, the median on-demand price is $3.38 per GPU per hour across 40 providers". One vendor lists the same H100 at $2.00 and another at $5.49 on the same day, and one of the two is a promotion that its own page says ends on 30 September 2026.
The arithmetic below uses the SemiAnalysis contract figures. They are April 2026 reporting on an index sold by subscription, so the break-even it produces uses the March 2026 rate.
Break-even utilization depends on four numbers you already have
The guides above give you a threshold as a buyer. If you own the fleet you are also the seller, selling hours to your own teams, so your break-even comes from the seller's arithmetic. For an operator selling GPU-hours, break-even utilization is the share of the fleet's hours that must be sold and delivered for revenue to cover the fixed cost of owning it. Four inputs set it, and all four are settled before the first hour sells.
- The contract rate per GPU-hour the fleet signs.
- The depreciation life it books, in years.
- The annual financing cost.
- The annual site cost, meaning what it costs to house, power, and run the fleet.
The depreciation life is the one you pick with the least outside discipline, and it draws the most scrutiny: SemiAnalysis reports that financial analysts criticized any neocloud or hyperscaler that used a 6-year depreciation period for its GPU compute assets. Because all four inputs differ by operator, there is no industry break-even figure to quote.
The calculation is short. Add up the annual fixed cost of depreciation, financing, and the site. Divide it by the fleet's GPU-hours in the year, which is the GPU count times 8,760. Divide again by the contract rate per GPU-hour. The result is the fraction of the fleet's annual GPU-hours that must be sold and delivered.
To see what one point of utilization is worth, take a fleet of 10,000 GPUs. It has 87,600,000 GPU-hours in a year, so one point is 876,000 hours, and at $2.35 per GPU-hour, the one-year H100 contract rate SemiAnalysis reported for March 2026, that point is worth about $2.1 million a year. Put in your own count and rate for your own figure.
A node bills every hour, whether or not it serves
With an inference API, the bill follows the traffic: your monthly token volume times your provider's published price per million tokens. A node you own or rent bills differently. Its hourly rate runs against every hour in the month, serving or idle. Eight GMI Cloud H200s at "from $2.60 /GPU-hour" is $20.80 an hour, and 730 hours, the average month in a 365-day year, is about $15,184 whether the node served anything or not.
Lambda states the rule for its instances. It bills "for as long as they're running, regardless if they're actively being used", so an hour spent loading a model or re-running work after a failure is billed like any other hour.
That is the asymmetry the utilization threshold measures. A quiet month lowers the per-token bill and leaves the node bill where it was, so your self-hosted cost per token is a fixed bill divided by however many tokens the node produced. And the number of tokens produced appears on no price list, because it depends on your batch sizes, your context lengths, your traffic pattern, and the hours the node spends with nothing to do.
A Hacker News commenter working out what a token costs on hardware you own puts the assumption in the open: "A single 1kW B200 GPU will set you back $50k, and can do 125 tokens per second with LLama4. Let's imagine you can use it for 36 months, at a DC, cooling and electricity price of 20 cents per kWh". That is the right calculation to run. It assumes 125 tokens per second for all 36 months, and how close your hardware comes depends on how much of those months it spends working.
The missing points are cold starts, failures, and immobility
A fleet pays for three kinds of hour that deliver nothing: hours spent cold-starting, hours spent redoing work a failure destroyed, and hours capacity spends stuck where no job can use it.
A cold start is a GPU loading weights and starting a serving engine before it can answer anything, plus the capacity you hold ready so that no request has to wait for that load. Based on our experience and our discussions with providers, over-provisioning ranges from 20 to 80% depending on the model portfolio hosted. A fleet hosting many models holds more of it, because swapping a model in is itself a cold start. Capacity held that way is powered, depreciating, and, on a fleet that sells hours, unsold.
A failure costs the hours between the last saved state and the failure, billed twice and delivered once. The restart re-runs work already paid for, so you buy the same piece of work twice.
Immobility is capacity in the wrong place or the wrong shape: GPUs that are powered and healthy and cannot take the job that is waiting, because the unused capacity is scattered across the fleet in pieces too small for it. Once a job starts, the resources it landed on stay bound to it until it ends, whatever changes around it.
None of this shows up as an incident. What shows up is a lower count of delivered hours against the same fixed costs.
What changes when a job can be saved and moved
Those three kinds of hour have one root. Once a job is running, its state is stuck where it is, so a failure throws the work away and a badly placed job cannot move. Saving the complete state of a running GPU workload, so that it can be brought back later on the same machine or a different one, is called checkpointing, and doing it automatically is what we do at Cedana. A job saved that way does not start over after a failure, and it can leave the node it landed on before it finishes.
A move is not free. It costs a restore plus the work the job did since its last checkpoint. Restore time follows the size of the checkpoint, meaning how many bytes have to move, not the parameter count. Cedana does not decide where a job goes. An operator's policy or an agent starts the move.
A move has two limits. The destination has to match the origin. A checkpoint records the GPU, driver, engine and model versions it was taken against. If any of them changes, the checkpoint is invalid and the workload cold-starts instead. And the job has to fit on one node. The single-GPU and multi-GPU-on-a-single-node tiers ship in production today. Multi-node, where a single workload spans hundreds or thousands of GPUs across many nodes, is in design partnership with leading enterprises and neoclouds.
The rate, the depreciation life, the financing, and the site cost are all fixed before the first hour sells. How many of the paid hours become finished work is the one term still open once the hardware is racked, and it is the term we work on. Cedana is automated GPU checkpointing and migration infrastructure that increases the useful work your GPUs deliver. We save the state of a running job and bring it back on compatible hardware, as of its last checkpoint. We do not change the price of an hour, the depreciation schedule, or the demand. We change only which of the paid hours turn into finished work. Measure that share on the fleet you already run before you put a number in the calculation.
Related:
- How to put a dollar figure on the GPU-hours that produced nothing
- How much utilization improvement do you need to break even on GPU checkpointing?
- Tasks per dollar: the economics of running your own coding models
- Cost per token and revenue per megawatt are the same number
- The fleet's real yield metric is tokens per gigabyte of VRAM
- Sovereignty has a bill: the transferred duties and the availability math
Common questions
At what point does it make more sense to own GPUs outright than to rent them by the hour or pay per token through an API?
Spheron and Introl put the break-even for owning instead of renting at roughly 50 to 70% sustained utilization for training or self-hosting, and other guides put the inference crossover between 2 million and 50 million tokens a day. Each of those thresholds asks for a forecast of demand, not a measurement of what your own stack already does with the hours you pay for, and reading the wrong one changes which side of the calculation you land on.
What does a GPU-hour cost right now, and why does the price swing so much?
GetDeploying reports a median on-demand H100 rate of $3.38 per GPU-hour across 40 providers as of 6 September 2026, while list prices in September 2026 ranged from $2.00 to $5.49 a GPU-hour depending on the provider, and one of those was a promotion dated to expire on 30 September 2026. A SemiAnalysis index of one-year H100 rental contracts rose from $1.70 to $2.35 per GPU-hour between October 2025 and March 2026, which is why every rate in this calculation needs its date attached to it.
What does it cost, per token, to run a model yourself once you count hardware, power and utilization?
A node you own or rent bills its hourly rate whether or not it is serving anything, so the self-hosted cost per token is that fixed bill divided by however many tokens the node produced, and a quiet month raises it while a busy month lowers it. An API's per-token price does not carry that swing, because the bill follows the traffic instead of the hour.
How much is idle, queued, or unused GPU capacity costing us?
A GPU loading weights, re-running work a failure destroyed, or holding capacity in the wrong place is billed at the same rate as one serving traffic, so a utilization dashboard can read high while most of that busy time produces nothing. Cold starts, failures, and immobility are the three kinds of paid hour that show up only as a lower count of delivered hours against the same fixed costs, without ever tripping an alarm.


