cedana / throughput · v1.4 · build 2826
GPU Migration Infrastructure
for Inference.
Cedana checkpoints, migrates, and resumes running inference across your fleet, any engine, any model, any orchestrator, raising throughput, reliability, and AI revenue per MW.
~ / cedana / migrate · livenode-A ─▸ node-B
The Problem → The Fix
AI workloads cannot move once running.
Cedana makes them liquid.
Four ways stranded compute drains budgets — and how Cedana resolves each, live and in place.
~ / cedana / problem ─▸ fix · live● 4 resolved
✕Idle GPUs:: 75% idle└▸Maximize throughput:: 88% util✕Expensive failures:: 65% lost└▸Automated reliability:: 0 lost✕Overprovisioned GPUs:: 50% over└▸Eliminate overprovisioning:: 0% buffer✕Rigid infrastructure:: no adapt└▸Adaptive infrastructure:: real-time
— the problemstatus: degraded
→ with cedanastatus: ok
01ERR.IDLE
Idle GPUs
75% idle
01RESOLVED
Maximize throughput
up to 88% utilization
02FATAL.RESTART
Expensive failures
65% compute lost
02RESOLVED
Automated reliability
0 lost progress
03WASTE.30%
Overprovisioned GPUs
10–50% overprovisioned
03RESOLVED
Eliminate overprovisioning
0% safety buffer
04STUCK.SCHEDULE
Rigid infrastructure
cannot adapt
04RESOLVED
Adaptive infrastructure
real-time
Live Proof · Console
Cold-Start Race
Same model. Same hardware.
21.7× faster to first token.
GLM-5.2-FP8 · 753B MoE · Multimodal · vLLM v0.19.0 · TP=8 · 8× B200 · 734 GiB ckpt · elapsed 00:00.00s
NATIVE · vLLM v0.19.0 · TP=8
● COLD START
00:00.00
$ vllm serve zai-org/GLM-5.2-FP8 --tp=8INFO 00:00:03 spawning 8 workers · EP=8INFO 00:00:40 loading shards 12/94 …INFO 00:05:00 loading shards 40/94 …INFO 00:12:00 loading shards 78/94 …INFO 00:16:20 loading shards 94/94 …INFO 00:18:40 building cuda graph (×8)INFO 00:20:30 warming kv cache · vision towerINFO 00:21:30 jit compile attention
CEDANA · RESTORE
▲ WARM RESUME
00:00.00
$ cedana resume glm-5.2-fp8 --from=snap.az-1[+] fetch snapshot :: ok (734 GiB · memlock)[+] verify hash :: ok (sha256: c4e1…9a)[+] restore cuda ctx :: ok (8× driver attached)[+] repopulate gpu mem :: ok (weights + KV)[+] thaw connections :: ok (sockets)[+] register endpoint :: ok (port 8000)READY 00:01:01 first token
NATIVE
0.0s
CEDANA
0.0s
Advanced · Status: SCALING
Built for the hardest workloads.
01 · ADVANCED WORKLOADS
Distributed by default.
Works transparently with MPI & NCCL.
02 · SCALE
From a single node to an AI factory.
The Platform
unlock() your scheduler.
01
Work with what you have
No rip-and-replace. No code changes. No disruption to your teams.
02
Kubernetes & SLURM
Built for AI and HPC. Native support for SLURM and Kubernetes.
03
First migration in <30 min
K8s helm chart or SLURM plug-in. No changes to your config.
~ / cedana / deploy● ready
Command your
compute.
Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.
deployK8s helm chart · SLURM plug-in
first migration<30 min
code changes0
workloadstraining · inference · HPC
01
AWS
02
Google Cloud
03
NVIDIA
04
K8s
05
SLURM
06
Nvidia Dynamo