cedana / product · orchestration

Real-time compute
orchestration.

Schedulers predict. Cedana adapts. Resize, migrate, and preempt jobs in real time — across instances, clusters, regions, and clouds — without losing progress.
Stateful reliability · durable compute

Run forever. Pay only when active.

Continuous automatic snapshots let workloads survive infrastructure failures and pause indefinitely without burning compute or losing progress. Resume when pricing or performance lines up.

orchestrator · live re-pack● ACTIVE · worker-1
worker-1
35%
worker-2
49%
worker-3
47%
worker-4
90%
elastic · okpreempt-safe · oksteps lost · 0
Elastic + dynamic scaling

Scale across instances, clusters, regions, clouds.

01

Live resize

Migrate jobs to fewer or bigger machines on demand — no restart, no warmup penalty.

02

Preempt-safe

Continuous checkpointing means preemption costs no work. Spot becomes a first-class tier.

03

Cross-cluster

Move a running job to another cluster — different region, different provider — same workload state.

04

Priority-aware

High-priority workloads bump lower-priority ones aside; the bumped jobs come back exactly where they were.

Automation · Status: LIVE

What automation unlocks.

01

Reliability

Automatically continue workloads from catastrophic failures without losing progress or restarting.

02

Productivity

Automatically migrate workloads to eliminate idle GPUs and increase throughput.

03

Operations

Perform maintenance without losing workload progress or manual re-submission.

~ / cedana / deploy ready

Command your
compute.

Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.

deployK8s helm chart · SLURM plug-in
first migration<30 min
code changes0
workloadstraining · inference · HPC
01
AWS
02
Google Cloud
03
NVIDIA
04
K8s
05
SLURM
06
Nvidia Dynamo