cedana / use cases · distributed training

Unbreakable
multi-node training.

Seamless and transparent. Resume training runs from the exact step — even after mid-epoch failures on large multi-node clusters.
Mechanism

Train. Fail. Resume from the same step.

Cedana captures full distributed state across ranks and resumes from where the job left off — no rewinding to the last manual checkpoint.

distributed training · 8× rank · nccl ringRUNNING
rank-0
rank-1
rank-2
rank-3
rank-4
rank-5
rank-6
rank-7
nccl ring
step14,000 / 50,000
steps lost · 0downtime · nodes · 2 / 2resume · same step
What you get

Production training, without the firefighting.

01

Train effectively

Deploy training jobs on SLURM or Kubernetes using KubeFlow, Ray, and other components of choice. Keep your configurations.

02

Failure mitigation

System-level checkpoint/restore ensures no lost work even during mid-epoch failures on large multi-node clusters.

03

Seamless multi-node

Safely spin up and down training runs without needing to reconstruct state.

04

Planet-scale compute

Manage training runs across clusters, both on-prem and in the cloud. Resume from system-level checkpoints on GPUs anywhere.

~ / cedana / deploy ready

Command your
compute.

Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.

deployK8s helm chart · SLURM plug-in
first migration<30 min
code changes0
workloadstraining · inference · HPC
01
AWS
02
Google Cloud
03
NVIDIA
04
K8s
05
SLURM
06
Nvidia Dynamo