DeepSeek-V4-Pro (1600B, MoE)
Native cold start: 2051 seconds (34.2 minutes). Cedana restore: 70 seconds from an 873 GiB checkpoint, a 29.3× speedup. The largest model in the set and the largest win.
Measured cold start times for key frontier open-weight models on 8× NVIDIA B200. Native engine cold start grows with model size; Cedana restores every model in 57–70 seconds, including a 1.6T-parameter model that natively takes 34 minutes.
Scale deployments for frontier inference.
8× NVIDIA B200 · 1.7TB memory · CUDA 12.9 · official sglang cookbook recipes · models ordered by parameter count
| Model | Params | Checkpoint (GiB) | Native | Cedana | Speedup |
|---|---|---|---|---|---|
| MiniMax-M2.7MoE | 229B | 244 | 564s | 57s | 9.9× |
| GLM-5.2-FP8MoE · Multimodal | 753B | 734 | 1322s | 61s | 21.7× |
| Kimi-K2.6MoE · Multimodal | 1100B | 670 | 1217s | 63s | 19.3× |
| DeepSeek-V4-ProMoE | 1600B | 873 | 2051s | 70s | 29.3× |
Restore time scales with checkpoint size, not parameter count: MoE and quantization keep it sublinear as models grow.
Native cold start: 2051 seconds (34.2 minutes). Cedana restore: 70 seconds from an 873 GiB checkpoint, a 29.3× speedup. The largest model in the set and the largest win.
Native cold start: 1217 seconds (20.3 minutes). Cedana restore: 63 seconds from a 670 GiB checkpoint, a 19.3× speedup. Restores faster than the smaller GLM-5.2 because its checkpoint is smaller: restore tracks checkpoint size, not parameter count.
Native cold start: 1322 seconds (22 minutes). Cedana restore: 61 seconds from a 734 GiB checkpoint, a 21.7× speedup.
Native cold start: 564 seconds (9.4 minutes). Cedana restore: 57 seconds from a 244 GiB checkpoint, a 9.9× speedup.
All runs on a single node with 8× NVIDIA B200 (Blackwell) and 1.7TB of system memory, CUDA 12.9. Models were served with sglang using the official sglang cookbook recipes for each model, unmodified.
Native cold start is the time from engine launch to ready-to-serve, including weight loading and full engine initialization. Cedana restore is the time to restore the same fully-initialized engine from a Cedana checkpoint to ready-to-serve. Checkpoint sizes are as reported in the table.
Cedana captures at the OS and CUDA driver level, so results are independent of the serving framework; equivalent behavior applies across vLLM and SGLang. Accelerated cold starts are also validated on H100 and A100 systems.