Run inference elastically.
On your infrastructure.
One primitive: GPU workload portability.
Cedana checkpoints running inference workers, capturing the entire state — memory, KV cache, CUDA context, container — and resumes them on a different GPU.
Inside the loop: the policy engine.
What's happening inside the chart above. Cedana's policy engine elastically resizes warm GPU buffers to match live demand — scaling capacity up before a spike, releasing it after. You set the policy. Cedana provisions to track actual load.
Unlock operational leverage.
Resilience for mission-critical workloads
When a worker is lost to failure or preemption, Cedana resumes it automatically — keeping valuable KV cache and in-flight sessions intact.
Less overprovisioning
Match provisioning closer to demand with 2–10× faster cold starts.
Idle GPUs get to work
Training and inference share the same GPUs. Long jobs immediately yield to inference spikes, with no lost work.
Right workload, right GPU
Dynamically migrate running inference workers across nodes without losing state. Put latency-critical traffic on faster GPUs, batch work on cheaper ones, and rebalance whenever cost or demand shifts.
Seamless integration.
Integrates with a Helm chart in minutes. No code or config changes. Supports NVIDIA Dynamo.
More from cedana automation.
Command your
compute.
Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.