use case: high performance computing
Cedana supercharges HPC: faster insight, higher throughput, lower cost.
Cedana automatically migrates workloads intelligently across your HPC cluster. Our intelligent control plane automates reliability, maintenance windows, resource overallocation, and underutilization.
production workloads● live · running on cedana
01
GROMACS
02
Molecular dynamics
03
Simulation codes
04
Training / Inference
05
Batch HPC
↳ in production with Caltech Computational Biology Lab● 5/5 workloads · stateful
How it works in practice
01RELIABILITY
Automated stateful reliability
Workloads automatically resume through failures (GPU, OOM, and more) without losing progress. No code changes needed.
failover · cluster● FAILING
lost · 0● sla preserved
02PRIORITIZATION
Prioritize without loss
Free a long running job for an urgent one, then put it back exactly where it was.
priority queue● LONG JOB
lost · 0● auto-pause · resume
↳ No checkpointing code. No management.
03OPERATIONS
Automate maintenance windows
When a node needs patching, draining, or a reboot, Cedana checkpoints the running workloads, lets the window proceed, and brings them back automatically afterward.
node · maintenance window● RUNNING
downtime · 0● auto · stateful
Fits your scheduler
SLURM
Kubernetes
Cedana plugs into the scheduler you already run. No new orchestration layer to adopt.
Other use cases