cedana / product · performance

Improve performance.
Reduce costs.

Eliminate wasteful compute by automatically suspending and resuming workloads based on real demand. Reduce fragmentation, drop safety buffers, and run hotter without losing reliability.
Utilization · before / after

Maximize utilization.
No idle resources.

Suspend and resume workloads on demand. Reclaim the safety buffer most clusters can't.

fleet utilization · 28 nodes● live
before30%
after85%
Cold starts · measured

Lightning-fast cold starts
— even at 405B.

hugging-quants/Meta-Llama-3.1-405B-Instruct-AWQ-INT4 on sglang · 8× H100 GPUs · 596 GB checkpoint. Native cold-start finishes in 550s; Cedana resume in 241s — a 2.28× speedup.

cold start · meta-llama 3.1 405b awq-int4 · sglang · 8× h100● measuring
native0 s
cedana0 s
speedup · 2.28×model · hugging-quants/Meta-Llama-3.1-405B-Instruct-AWQ-INT4topology · 8× H100 · 596 GB ckpt
Hardware support

CPU and GPU.

Save, migrate, and resume transparently across CPU (Intel, ARM) and GPU (NVIDIA) workloads.

01

Intel

x86_64 · checkpoint/restore
02

ARM

aarch64 · checkpoint/restore
03

NVIDIA

cuda 12.x · h100 / b200 / gh200
Automation · Status: LIVE

What automation unlocks.

01

Reliability

Automatically continue workloads from catastrophic failures without losing progress or restarting.

02

Productivity

Automatically migrate workloads to eliminate idle GPUs and increase throughput.

03

Operations

Perform maintenance without losing workload progress or manual re-submission.

~ / cedana / deploy ready

Command your
compute.

Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.

deployK8s helm chart · SLURM plug-in
first migration<30 min
code changes0
workloadstraining · inference · HPC
01
AWS
02
Google Cloud
03
NVIDIA
04
K8s
05
SLURM
06
Nvidia Dynamo