Live migration for
CPU and GPU containers.
Save. Migrate. Resume.
One API across CPU, GPU, distributed jobs, Kubernetes pods, and SLURM batch.
Policy-based automation at scale.
Stateful reliability
Survive node loss, preemption, and hardware faults — workloads resume from the exact step.
Price / performance
Run on spot, ride out reclaim, and consolidate fleet utilization beyond static scheduling.
Advanced orchestration
Move running jobs based on live signals — failures, demand, priorities, or maintenance.
Elastic scaling
Resize compute across instances, clusters, regions, and clouds without losing progress.
More of the cedana platform.
What automation unlocks.
Reliability
Automatically continue workloads from catastrophic failures without losing progress or restarting.
Productivity
Automatically migrate workloads to eliminate idle GPUs and increase throughput.
Operations
Perform maintenance without losing workload progress or manual re-submission.
Command your
compute.
Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.