» Cedana / benchmarks / model switching
Coding Portfolio: RTXPro 6000 Results
We ran 40 tasks from SWE-bench Verified against four open-weight coding models on a single node with 8 RTX PRO 6000 Blackwell GPUs. Our serving stack was Cedana, NVIDIA Dynamo and NeMo Switchyard. Switchyard enables us to automatically route tasks to models, and Cedana enables us to fit multiple models on a single node by switching them 10‑30x faster.
Read how the stack was built »
We used 4 models (GLM-4.5, Devstral, Qwen2.5 Coder and Qwen3.6) that need 14 GPUs between them, so the node swapped models in and out 29 times. Each swap was a Cedana restore from checkpoint, and restores averaged 33 to 70 seconds depending on the model.
In the video, the top lane is with Cedana and the bottom is without. Cedana finishes the same work in 152 minutes with 40 tasks passed. Without Cedana the same task, on the same GPU takes 317 minutes, 2.1 times as long, and costs $84 against $41 at an assumed $16 per node-hour.
Same verified task set: 2.1× faster, 52% lower compute cost.
| Measure | Nativemodeled | Cedanarecorded | Advantage |
|---|---|---|---|
| Successful tasks | |||
| Tasks passed when Cedana finishes, at 152 minutes | 13 | 40 | 3.1× |
| Speed and hardware | |||
| Wall-clock hours to finish | 5.3 h | 2.5 h | 2.1× |
| GPU-hours used | 42.2 | 20.3 | 2.1× |
| GPU-hours spent loading models | 19.3 | 1.2 | 16.7× |
| Nodes to keep all four models loaded | 2 | 1 | 2× |
| Cost at $16 per node-hour, on demand | $84 | $41 | 2.1× |
| Token serving efficiency | |||
| Input tokens served per wall-clock hour | 15.4M | 32.1M | 2.1× |
| Input tokens served when Cedana finishes | 37.5M | 81.5M | 2.2× |
- Recorded: actual run performed Sept 21, 2026. 145 tasks from SWE-bench Verified, 8 GPUs, 29 restores, 40 tasks passed, 81.5M input tokens.
- Cost assumes $16 per node-hour for the 8-GPU node, on demand.