cedana / throughput · v1.4 · build 2826
GPU job migration
infrastructure.
Cedana checkpoints, migrates, and resumes live GPU jobs across your fleet — raising throughput, reliability, and AI revenue per MW.
~ / cedana / migrate · livenode-A ─▸ node-B
The Problem → The Fix
AI workloads cannot move once running.
Cedana makes them liquid.
Four ways stranded compute drains budgets — and how Cedana resolves each, live and in place.
~ / cedana / problem ─▸ fix · live● 4 resolved
✕Idle GPUs:: 75% idle└▸Maximize throughput:: 88% util✕Expensive failures:: 65% lost└▸Automated reliability:: 0 lost✕Overprovisioned GPUs:: 50% over└▸Eliminate overprovisioning:: 0% buffer✕Rigid infrastructure:: no adapt└▸Adaptive infrastructure:: real-time
— the problemstatus: degraded
→ with cedanastatus: ok
01ERR.IDLE
Idle GPUs
75% idle
01RESOLVED
Maximize throughput
up to 88% utilization
02FATAL.RESTART
Expensive failures
65% compute lost
02RESOLVED
Automated reliability
0 lost progress
03WASTE.30%
Overprovisioned GPUs
10–50% overprovisioned
03RESOLVED
Eliminate overprovisioning
0% safety buffer
04STUCK.SCHEDULE
Rigid infrastructure
cannot adapt
04RESOLVED
Adaptive infrastructure
real-time
Live Proof · Console
Cold-Start Race
Same model. Same hardware.
21.7× faster to first token.
GLM-5.2-FP8 · 753B MoE · Multimodal · vLLM v0.19.0 · TP=8 · 8× B200 · 734 GiB ckpt · elapsed 00:00.00s
NATIVE · vLLM v0.19.0 · TP=8
● COLD START
00:00.00
$ vllm serve zai-org/GLM-5.2-FP8 --tp=8INFO 00:00:03 spawning 8 workers · EP=8INFO 00:00:40 loading shards 12/94 …INFO 00:05:00 loading shards 40/94 …INFO 00:12:00 loading shards 78/94 …INFO 00:16:20 loading shards 94/94 …INFO 00:18:40 building cuda graph (×8)INFO 00:20:30 warming kv cache · vision towerINFO 00:21:30 jit compile attention
CEDANA · RESTORE
▲ WARM RESUME
00:00.00
$ cedana resume glm-5.2-fp8 --from=snap.az-1[+] fetch snapshot :: ok (734 GiB · memlock)[+] verify hash :: ok (sha256: c4e1…9a)[+] restore cuda ctx :: ok (8× driver attached)[+] repopulate gpu mem :: ok (weights + KV)[+] thaw connections :: ok (sockets)[+] register endpoint :: ok (port 8000)READY 00:01:01 first token
NATIVE
0.0s
CEDANA
0.0s
Advanced · Status: SCALING
Built for the hardest workloads.
01 · ADVANCED WORKLOADS
Distributed by default.
Works transparently with MPI & NCCL.
02 · SCALE
From a single node to an AI factory.
The Platform
unlock() your scheduler.
01
Work with what you have
No rip-and-replace. No code changes. No disruption to your teams.
02
Kubernetes & SLURM
Built for AI and HPC. Native support for SLURM and Kubernetes.
03
First migration in <30 min
K8s helm chart or SLURM plug-in. No changes to your config.
~ / cedana / deploy● ready
Command your
compute.
Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.
deployK8s helm chart · SLURM plug-in
first migration<30 min
code changes0
workloadstraining · inference · HPC
01
AWS
02
Google Cloud
03
NVIDIA
04
K8s
05
SLURM
06
Nvidia Dynamo