Secure and productive frontier coding capabilities
Enterprises and startups increasingly host their own coding models to keep code inside their infrastructure and per-token costs under their control. However, frontier-level capabilities require a portfolio of models. Coding work spans autocomplete, small edits, complex debugging, and long-horizon repo-wide planning. No single model is optimal across all of them. Serving a portfolio takes a router that matches each task to the right model and escalates when a smaller model fails. Coding agents connect to a single API endpoint for every task.
Coding-agent traffic is heavy-tailed. In TraceLab, a public trace of about 4,300 Claude Code and Codex sessions, the median request finishes in 38 seconds, but the 6% of requests that run past 10 minutes use 69% of all request time. A resident mid-size model serves most requests, and the frontier model is needed only for the tail. Table 1 lists the portfolio built for that demand shape, the tier each model serves, and its checkpoint footprint on an 8x NVIDIA B200 node.
| Task type | Description | Model | VRAM |
|---|---|---|---|
| Long-horizon planning | Repo-wide reasoning, task decomposition, architecture decisions | DeepSeek-V4-Pro-NVFP4 | ~873 GiB |
| Frontier / hard tasks | Complex debugging, SWE-bench Pro-class | Kimi-K2.6-NVFP4 / GLM-5.2 | 670-734 GiB |
| Medium complexity | Feature work, refactors, tests, PR-scale changes | MiniMax-M2.7 | ~244 GiB |
| Small tasks | Quick edits, docstrings, lint fixes | MiniMax-M2.7 | 0 (reuses resident model) |
| Autocomplete / FIM | Sub-200 ms inline completion | Qwen3-Coder-30B-A3B FP8 | ~40 GiB |
| Batch inference | Nightly evals, bulk codegen; preemptible | Any portfolio member | 0 (reuses resident model) |
| Fine-tuning (v2+) | LoRA on executor trajectories | MiniMax-M2.7 base | 0 (reuses resident model) |
| Portfolio total | Distinct footprints exceed the node. ~1-min switching makes it fit | Portfolio: 1,891 GiB | Node: 1,341 GiB |
Table 1: Portfolio of Coding Models
A portfolio of frontier models exceeds the VRAM of a single node, and distributing models across dedicated nodes leaves expensive capacity idle when the model isn't in use.
Serving a portfolio on a single node solves the economics, but hits two walls: 1) switching frontier models results in long cold start times (9-34 mins), and 2) pre-empting a long-running coding job to escalate an urgent task destroys its in-flight sessions, resulting in workflow bottlenecks, lost compute, and wasted time.
NVIDIA Dynamo, NeMo Switchyard and Cedana
In this post we show how NVIDIA Dynamo, NeMo Switchyard, and Cedana integrate to host a portfolio of frontier open weight coding models on a single 8x B200 node.
This makes model switching cheap enough to escalate by default, based on four core capabilities:
01 · Intelligent inference: Dynamo serves the full portfolio behind one OpenAI-compatible endpoint across SGLang, vLLM, and TensorRT-LLM, and its KV-aware router sends each request to the worker already holding its prefix cache, cutting recomputation for coding agents that resend repo context every turn.
02 · Intelligent model routing: NeMo Switchyard, NVIDIA's open source model router, picks the model for each request at runtime. Its escalation strategy starts tasks on a cheap model and promotes to a larger one after failed turns. That strategy is only affordable when the larger model does not cost a cold start to reach.
03 · Fast frontier model cold starts: Cedana restores multi-GPU frontier models including serving state, in 57 to 70 seconds, between 10-29x faster than a native cold start (full benchmarks).
04 · Stateful inference reliability: Cedana checkpoints multi-GPU inference workloads during in-flight sessions, and resumes those sessions intact. Workloads can automatically resume through catastrophic failures.
Fast switching delivered
The following shows the results of the combined system in switching across models on an 8xB200.
| Model | Params | Checkpoint (GiB) | Native | Cedana | Speedup |
|---|---|---|---|---|---|
| MiniMax-M2.7 · MoE | 229B | 244 | 564s | 57s | 9.9× |
| GLM-5.2-FP8 · MoE · Multimodal | 753B | 734 | 1322s | 61s | 21.7× |
| Kimi-K2.6 · MoE · Multimodal | 1100B | 670 | 1217s | 63s | 19.3× |
| DeepSeek-V4-Pro · MoE | 1600B | 873 | 2051s | 70s | 29.3× |
Table 2: 10-29x faster resumes enable efficient and fast switching
Escalate fast and effectively
Fast switching makes progressive escalation practical: Switchyard tries the task on a small model, and when it fails, it escalates to a larger one. Without fast restore, each escalation to DeepSeek-V4-Pro costs 34-minute cold start, reducing frequency and flexibility. At 70 seconds, escalation becomes cheap enough to become the default: the swap costs about 3.4% of the cold start it replaces.
Fast, effective routing
NVIDIA NeMo Switchyard routes across model tiers reducing costs. With the NeMo Switchyard escalation router, LangChain measured a 74% cost reduction on 145 multi-turn agent tasks compared with a frontier-only baseline, sending only 7% of calls to the frontier model at about a 6-point accuracy tradeoff. That test routed between warm API endpoints. A self-hosted node has two options, and both are expensive. Option 1: keep the frontier model resident, and its GPUs sit idle between escalations. Option 2: load the frontier model on demand, and each escalation costs a 9 to 34-minute cold start, while the model being swapped out loses its in-flight sessions. Cedana restores a frontier model in about a minute and checkpoints the outgoing model with its sessions intact, which makes escalation affordable on one node.
With Cedana, a Switchyard route change triggers a checkpoint of the running model and a restore of the target, and the swap completes in 57 to 70 seconds (Table 2). The outgoing model keeps its sessions: a repo-wide refactor or fine-tuning job that has run for hours can be checkpointed when an urgent workload needs the node and resumed later from where it stopped, including its in-flight requests.
The Checkpoint Mechanism
Cedana checkpoints the entire inference state at the OS and CUDA driver level: weights, KV cache, allocator state, CUDA graphs, and in-flight state, the full container including host and GPU. Checkpoints are agnostic to the engine (SGLang, vLLM, TensorRT-LLM).
Integration and deployment guide
The full stack deploys with a single Helm chart in minutes. The chart installs Cedana as a Kubernetes DaemonSet that integrates seamlessly with Dynamo, accelerating cold starts and providing stateful reliability across every worker. When Switchyard routes a task to a new model, Cedana checkpoints the running model and restores the target in its place. No code modifications or config changes required. Cedana includes an optional UI and control plane, giving teams a single place to define their model portfolio, set Switchyard routing policies, and monitor the full stack in real time.

Diagram 1: architecture
Refer to the Cedana and Dynamo Deployment Guide for the Helm chart and the model-switching walkthrough, and to the 8x B200 cold-start benchmark for methodology and per-model results.
Forthcoming
In our next blog post we'll provide performance and efficiency benchmarks.


