Automatic,
unbreakable spot capacity.
Caltech's CompBio group saved 80% with 2× faster results.
Scientists use Cedana to run end-to-end training, inference, and GROMACS evaluation loops on spot capacity. Automatically.
Unbreakable, stateful reliability.
Long-running, stateful workloads automatically resume on a new instance through revocations or failures. Your workload doesn't lose progress, and you don't waste time babysitting jobs.
Automated
Live-migrate GPU workloads before failures happen. System-level checkpoint/restore ensures no lost-work even during mid-epoch failures — on multi-node clusters.
Job-level SLAs
Assign individual jobs SLAs for reliability, costs, and other criteria — required for efficiently sharing compute across users and groups.
No code changes
Checkpointing is transparently and continuously performed with no impact on performance. No need to manage checkpoints.
Seamless install
Just add a few lines to your Helm chart and you're ready to go.
From broken defaults to automated.
More from cedana automation.
Command your
compute.
Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.