Maximize tokens per GPU.
Cedana restores frontier models in about a minute: a 1.6T model in 70 seconds. Collapse the warm pool, hold SLO through failures and upgrades, with zero changes to your engines or schedulers.
In production trials with leading AI infrastructure companies.
Native cold start grows with the model. Cedana restore doesn't.
Scale deployments for frontier inference.
8× NVIDIA B200 · 1.7TB memory · CUDA 12.9 · official sglang cookbook recipes · models ordered by parameter count
Key frontier models, measured.
| Model | Params | Checkpoint (GiB) | Native | Cedana | Speedup |
|---|---|---|---|---|---|
| MiniMax-M2.7MoE | 229B | 244 | 564s | 57s | 9.9× |
| GLM-5.2-FP8MoE · Multimodal | 753B | 734 | 1322s | 61s | 21.7× |
| Kimi-K2.6MoE · Multimodal | 1100B | 670 | 1217s | 63s | 19.3× |
| DeepSeek-V4-ProMoE | 1600B | 873 | 2051s | 70s | 29.3× |
Restore time scales with checkpoint size, not parameter count: MoE and quantization keep it sublinear as models grow.
Three taxes on your margin, removed.
More models on the same GPUs. Higher utilization.
Start multi-GPU frontier models in seconds to reduce overprovisioning.
Failures become fast, stateful resumes.
Automatically resume multi-GPU inference workloads through failures, without losing customer work.
Drain, upgrade, rebalance without killing workloads.
Migrate live inference workloads off nodes for maintenance and engine upgrades, and rebalance fragmented capacity across the fleet.
The primitive your scheduler is missing.
Your stack including scheduler, router, cache tiering stays the same. Cedana sits transparently beneath them, capturing and restoring live execution state. Inference and engine agnostic.
No code, stack, or config changes.
Widen the spread between cost per GPU-hour and revenue per GPU-hour.
Cold starts: Reduce overprovisioning of warm replicas. The full catalog and every dedicated endpoint stay seconds away instead of always loaded.
Failures: Stop paying for failures twice. Stateful resume keeps multi-GPU workloads serving through faults, so an incident costs seconds, not SLA credits and lost customer work.
Immobility: Run the fleet hotter. Consolidate and rebalance live workloads, and drain nodes for upgrades and maintenance without taking capacity offline.
Inference is becoming stateful. State is becoming the bottleneck.
Agents run longer.
Long-horizon agents accumulate session state worth real money. Every eviction, failure, or cold start throws it away.
Models keep growing.
Frontier open-weight models push cold starts from seconds to half an hour, and warm pools stop being affordable.
More models, longer contexts.
Multi-model routing means constant swapping; growing context windows mean gigabytes of KV and session state per workload. More rebuilds, each one more expensive.
Questions we hear first.
Cedana is the missing primitive for stateful inference. It checkpoints, restores, and migrates live GPU workloads across the fleet, without rebuilding state from scratch. By turning execution state into a schedulable object, Cedana lets fleets produce more tokens with higher reliability and lower cost.