Maximize revenue per megawatt.
Cedana restores frontier models in about a minute: a 1.6T model in 70 seconds. Run the fleet hotter, keep GPUs earning through failures and upgrades, with zero changes to your code or stack.
In production trials with leading AI infrastructure companies.
Native cold start grows with the model. Cedana restore doesn't.
Scale deployments for frontier inference.
8× NVIDIA B200 · 1.7TB memory · CUDA 12.9 · official sglang cookbook recipes · models ordered by parameter count
Key frontier models, measured.
| Model | Params | Checkpoint (GiB) | Native | Cedana | Speedup |
|---|---|---|---|---|---|
| MiniMax-M2.7MoE | 229B | 244 | 564s | 57s | 9.9× |
| GLM-5.2-FP8MoE · Multimodal | 753B | 734 | 1322s | 61s | 21.7× |
| Kimi-K2.6MoE · Multimodal | 1100B | 670 | 1217s | 63s | 19.3× |
| DeepSeek-V4-ProMoE | 1600B | 873 | 2051s | 70s | 29.3× |
Three problems solved. More ways to monetize your fleet.
More models on the same GPUs. Higher utilization.
Start multi-GPU frontier models in seconds to reduce overprovisioning.
Failures become fast, stateful resumes.
Automatically resume multi-GPU inference workloads through failures, without losing customer work.
Defragment, rebalance, monetize across GPUs.
Ship a spot-priced preemptible tier that preserves long-running batch inference.
One always-warm 8-GPU frontier replica costs ~$175K per year before it serves a token (at $2.50/GPU-hour; plug in your rate). Scale-to-zero returns it.
What vMotion did for the datacenter, Cedana does for GPUs.
Cedana operates at the CUDA driver level, beneath your inference engines and orchestrators. Seamlessly integrates with K8s. No code, stack, or config changes.
More revenue per megawatt from the fleet you already financed.
Long cold starts and model diversity result in overprovisioning. Sub-minute restore collapses the warm pool.
Recovery downtime is unserviced debt on financed GPUs. Resuming execution state turns long outages into seconds.
Sub-minute restore and stateful resiliency extend the revenue life of prior-generation silicon by making it viable inference capacity.
Inference is becoming stateful. State is becoming the bottleneck.
Agents run longer.
Long-horizon agents accumulate session state worth real money. Every eviction, failure, or cold start throws it away.
Models keep growing.
Frontier open-weight models push cold starts from seconds to half an hour, and warm pools stop being affordable.
More models, longer contexts.
Multi-model routing means constant swapping; growing context windows mean gigabytes of KV and session state per workload. More rebuilds, each one more expensive.
Questions we hear first.
Cedana is the automation control plane for stateful inference. It checkpoints, restores, and migrates live GPU workloads across the fleet, without rebuilding state from scratch. By turning execution state into a schedulable object, Cedana lets fleets produce more tokens with higher reliability and lower cost.