» Cedana / benchmarks / model switching

Coding Portfolio: RTXPro 6000 Results

We ran 40 tasks from SWE-bench Verified against four open-weight coding models on a single node with 8 RTX PRO 6000 Blackwell GPUs. Our serving stack was Cedana, NVIDIA Dynamo and NeMo Switchyard. Switchyard enables us to automatically route tasks to models, and Cedana enables us to fit multiple models on a single node by switching them 10‑30x faster.

Read how the stack was built »

We used 4 models (GLM-4.5, Devstral, Qwen2.5 Coder and Qwen3.6) that need 14 GPUs between them, so the node swapped models in and out 29 times. Each swap was a Cedana restore from checkpoint, and restores averaged 33 to 70 seconds depending on the model.

In the video, the top lane is with Cedana and the bottom is without. Cedana finishes the same work in 152 minutes with 40 tasks passed. Without Cedana the same task, on the same GPU takes 317 minutes, 2.1 times as long, and costs $84 against $41 at an assumed $16 per node-hour.

~ / cedana / swe-bench / cedana vs native● 30 s
» Results · same work, both lanes

Same verified task set: 2.1× faster, 52% lower compute cost.

MeasureNativemodeledCedanarecordedAdvantage
Successful tasks
Tasks passed when Cedana finishes, at 152 minutes13403.1×
Speed and hardware
Wall-clock hours to finish5.3 h2.5 h2.1×
GPU-hours used42.220.32.1×
GPU-hours spent loading models19.31.216.7×
Nodes to keep all four models loaded212×
Cost at $16 per node-hour, on demand$84$412.1×
Token serving efficiency
Input tokens served per wall-clock hour15.4M32.1M2.1×
Input tokens served when Cedana finishes37.5M81.5M2.2×
29 Cedana restores took 21 min 36 s in total. 11 native cold starts took 4 h 30 min.
  1. Recorded: actual run performed Sept 21, 2026. 145 tasks from SWE-bench Verified, 8 GPUs, 29 restores, 40 tasks passed, 81.5M input tokens.
  2. Cost assumes $16 per node-hour for the 8-GPU node, on demand.