cedana / use cases · distributed training
Unbreakable
multi-node training.
Seamless and transparent. Resume training runs from the exact step — even after mid-epoch failures on large multi-node clusters.
Mechanism
Train. Fail. Resume from the same step.
Cedana captures full distributed state across ranks and resumes from where the job left off — no rewinding to the last manual checkpoint.
distributed training · 8× rank · nccl ring● RUNNING
■rank-0
■rank-1
■rank-2
■rank-3
■rank-4
■rank-5
■rank-6
■rank-7
nccl ring
step14,000 / 50,000
steps lost · 0downtime · —nodes · 2 / 2resume · same step
What you get
Production training, without the firefighting.
01
Train effectively
Deploy training jobs on SLURM or Kubernetes using KubeFlow, Ray, and other components of choice. Keep your configurations.
02
Failure mitigation
System-level checkpoint/restore ensures no lost work even during mid-epoch failures on large multi-node clusters.
03
Seamless multi-node
Safely spin up and down training runs without needing to reconstruct state.
04
Planet-scale compute
Manage training runs across clusters, both on-prem and in the cloud. Resume from system-level checkpoints on GPUs anywhere.
Other use cases
More from cedana automation.
~ / cedana / deploy● ready
Command your
compute.
Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.
deployK8s helm chart · SLURM plug-in
first migration<30 min
code changes0
workloadstraining · inference · HPC
01
AWS
02
Google Cloud
03
NVIDIA
04
K8s
05
SLURM
06
Nvidia Dynamo