cedana / throughput · v1.4 · build 2826

GPU Migration Infrastructure
for Inference.

Cedana checkpoints, migrates, and resumes running inference across your fleet, any engine, any model, any orchestrator, raising throughput, reliability, and AI revenue per MW.
~ / cedana / migrate · livenode-A ─▸ node-B
Live Proof · Console
Cold-Start Race

Same model. Same hardware.
21.7× faster to first token.

GLM-5.2-FP8 · 753B MoE · Multimodal · vLLM v0.19.0 · TP=8 · 8× B200 · 734 GiB ckpt · elapsed 00:00.00s
NATIVE · vLLM v0.19.0 · TP=8
● COLD START
00:00.00
$ vllm serve zai-org/GLM-5.2-FP8 --tp=8
INFO 00:00:03 spawning 8 workers · EP=8
INFO 00:00:40 loading shards 12/94 …
INFO 00:05:00 loading shards 40/94 …
INFO 00:12:00 loading shards 78/94 …
INFO 00:16:20 loading shards 94/94 …
INFO 00:18:40 building cuda graph (×8)
INFO 00:20:30 warming kv cache · vision tower
INFO 00:21:30 jit compile attention
CEDANA · RESTORE
▲ WARM RESUME
00:00.00
$ cedana resume glm-5.2-fp8 --from=snap.az-1
[+] fetch snapshot :: ok (734 GiB · memlock)
[+] verify hash :: ok (sha256: c4e1…9a)
[+] restore cuda ctx :: ok (8× driver attached)
[+] repopulate gpu mem :: ok (weights + KV)
[+] thaw connections :: ok (sockets)
[+] register endpoint :: ok (port 8000)
READY 00:01:01 first token
NATIVE
0.0s
CEDANA
0.0s
delta · 1261s savedspeedup · 21.7×verdict · cedana wins
Advanced · Status: SCALING

Built for the hardest workloads.

01 · ADVANCED WORKLOADS

Distributed by default.
Works transparently with MPI & NCCL.

02 · SCALE

From a single node to an AI factory.

The Platform

unlock() your scheduler.

01

Work with what you have

No rip-and-replace. No code changes. No disruption to your teams.

02

Kubernetes & SLURM

Built for AI and HPC. Native support for SLURM and Kubernetes.

03

First migration in <30 min

K8s helm chart or SLURM plug-in. No changes to your config.

~ / cedana / deploy ready

Command your
compute.

Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.

deployK8s helm chart · SLURM plug-in
first migration<30 min
code changes0
workloadstraining · inference · HPC
01
AWS
02
Google Cloud
03
NVIDIA
04
K8s
05
SLURM
06
Nvidia Dynamo