cedana / how we work

How we work.

The problems Cedana removes, the difference against restart-based tooling, stateful reliability, and how it scales.
The Problem → The Fix

AI workloads cannot move once running.
Cedana makes them liquid.

Four ways stranded compute drains budgets — and how Cedana resolves each, live and in place.

— the problemstatus: degraded
→ with cedanastatus: ok
cluster · 64 gpus · 25% utilized● 75% IDLE
active 16 idle 48
01ERR.IDLE

Idle GPUs

Valuable compute remains stranded while critical work is delayed.

75% idle
cluster · 64 gpus · 88% utilized● 88% ACTIVE
active 56 idle → reassigned
01RESOLVED

Maximize throughput

Workloads shift to idle GPUs, reclaiming stranded capacity and maximizing cluster throughput.

up to 88% utilization
training · 50 000 stepsrunning…
░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░░
0%compute saved: 0
02FATAL.RESTART

Expensive failures

Failures and preemptions force workloads to restart from scratch — up to 65% of compute wasted.

65% compute lost
NODE-A · FAULT✗ checkpoint step 14 832migrate()NODE-B · OKresumedstep 14 832 / 50 000● Δ 0.42s↳ live migration · sla preserved
02RESOLVED

Automated reliability

Workloads automatically migrate to healthy infrastructure and resume after failures.

0 lost progress
capacity · 10 nodes30% buffer · idle
active 7 safety buffer 3
03WASTE.30%

Overprovisioned GPUs

Capacity is routinely over-provisioned by 10–50% just to maintain reliability and hit SLAs.

10–50% overprovisioned
capacity · 10 nodes● 0% buffer · full
active 10 reclaimed 3
03RESOLVED

Eliminate overprovisioning

Automatic migration and recovery remove the need for large safety buffers to meet SLAs.

0% safety buffer
scheduler · static routing● ROUTE FAULT
SCHEDULERw1w2w3w4
1 lane down · scheduler can't reroutew3 stuck
04STUCK.SCHEDULE

Rigid infrastructure

Schedulers cannot dynamically adapt to failures, demand, or changing priorities.

cannot adapt
scheduler · adaptive routing● REROUTED
SCHEDULERw1w2w3w4
load reroutes in real timeall lanes healthy
04RESOLVED

Adaptive infrastructure

Kubernetes and SLURM adapt workloads in real time to failures and demand.

real-time
The Difference

The Cedana diff.

— without migrationSTATUS · DEGRADED
Expensive failures
Up to 65% compute lost
Over-provisioned GPUs
10–50% capacity buffers
Idle GPUs
Stranded compute while jobs wait
Rigid infrastructure
Schedulers cannot adapt
+ with cedanaSTATUS · OK
Automated reliability
Workloads resume automatically
Eliminate overprovisioning
SLAs without safety buffers
Maximize throughput
Workloads migrate to idle GPUs
Adaptive infrastructure
Workloads adjust in real time
AWSGoogle CloudNVIDIASLURMKubernetesLambdaCoreWeaveAWSGoogle CloudNVIDIASLURMKubernetesLambdaCoreWeave
~ / cedana / deploy ready

Command your
compute.

Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.

deployK8s helm chart · SLURM plug-in
first migration<30 min
code changes0
workloadstraining · inference · HPC
01
AWS
02
Google Cloud
03
NVIDIA
04
K8s
05
SLURM
06
Nvidia Dynamo