cedana / product · how it works

Live migration for
CPU and GPU containers.

Save a workload, migrate it across nodes or clusters, and resume right where it left off — with zero code changes and no manual checkpoint hooks.
Mechanism

Save. Migrate. Resume.

One API across CPU, GPU, distributed jobs, Kubernetes pods, and SLURM batch.

cedana lifecycleSAVE
step 01
save
checkpoint state
step 02
migrate
rdma transfer
step 03
resume
restored on new node
no code changes · okk8s + slurm · nativecpu + gpu · ok
What you get

Policy-based automation at scale.

01

Stateful reliability

Survive node loss, preemption, and hardware faults — workloads resume from the exact step.

02

Price / performance

Run on spot, ride out reclaim, and consolidate fleet utilization beyond static scheduling.

03

Advanced orchestration

Move running jobs based on live signals — failures, demand, priorities, or maintenance.

04

Elastic scaling

Resize compute across instances, clusters, regions, and clouds without losing progress.

Automation · Status: LIVE

What automation unlocks.

01

Reliability

Automatically continue workloads from catastrophic failures without losing progress or restarting.

02

Productivity

Automatically migrate workloads to eliminate idle GPUs and increase throughput.

03

Operations

Perform maintenance without losing workload progress or manual re-submission.

~ / cedana / deploy ready

Command your
compute.

Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.

deployK8s helm chart · SLURM plug-in
first migration<30 min
code changes0
workloadstraining · inference · HPC
01
AWS
02
Google Cloud
03
NVIDIA
04
K8s
05
SLURM
06
Nvidia Dynamo