cedana / use cases · automatic spot

Automatic,
unbreakable spot capacity.

Run long-running workloads on spot reliably, without interruption or losing work. Save up to 80% with zero code or config changes.
Caltech · Status: SHIPPING

Caltech's CompBio group saved 80% with 2× faster results.

Scientists use Cedana to run end-to-end training, inference, and GROMACS evaluation loops on spot capacity. Automatically.

spot fleet · auto-recovery on reclaimSTABLE
spot-00running
spot-01running
spot-02running
spot-03running
spot-04running
spot-05running
spot-06running
·spot-07idle
··· awaiting events ···
cost saved · 74%work lost · 0reclaims handled · auto
Capabilities

Unbreakable, stateful reliability.

Long-running, stateful workloads automatically resume on a new instance through revocations or failures. Your workload doesn't lose progress, and you don't waste time babysitting jobs.

01

Automated

Live-migrate GPU workloads before failures happen. System-level checkpoint/restore ensures no lost-work even during mid-epoch failures — on multi-node clusters.

02

Job-level SLAs

Assign individual jobs SLAs for reliability, costs, and other criteria — required for efficiently sharing compute across users and groups.

03

No code changes

Checkpointing is transparently and continuously performed with no impact on performance. No need to manage checkpoints.

04

Seamless install

Just add a few lines to your Helm chart and you're ready to go.

Pain points · Status: SOLVED

From broken defaults to automated.

COST
Skyrocketing hosting costs
Save up to 80% with zero code or config changes.
BATCH
Long batch jobs stuck on on-demand
Migrate to spot, finish reliably — no restarts.
TUNE
Endless Karpenter / autoscaler tuning
Eliminate guesswork; let the workload move instead.
BABYSIT
Manually restarting evicted jobs
Transparently checkpoints, migrates, and resumes.
~ / cedana / deploy ready

Command your
compute.

Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.

deployK8s helm chart · SLURM plug-in
first migration<30 min
code changes0
workloadstraining · inference · HPC
01
AWS
02
Google Cloud
03
NVIDIA
04
K8s
05
SLURM
06
Nvidia Dynamo