Writing from the
cedana team.
insights from the team.
Field reports, engineering deep-dives, and benchmarks from the Cedana team.

One frontier restart burns half a monthly 99.9% error budget
Calculate how model restart time consumes an inference error budget, what replicas change, and how to compare a measured restore with your availability SLO.
read →
The recovery time you can put in a filing
Define and test LLM recovery time from failure to serving again. Separate restore benchmarks from the detection, placement and recovery your team must document.
read →
When the node dies under a Jupyter session
Learn what a notebook file cannot recover after a GPU node fails, how process checkpoints preserve kernel state, and which restoration limits still apply.
read →
One GPU fails and 71 healthy GPUs wait
Understand how one failure stalls a tensor-parallel NVL72 workload, how recovery consumes healthy GPU-hours, and where published measurements stop.
read →
Five fleet signals, five policies: from alert to automatic action
Connect GPU health, thermal, reclaim, fragmentation and maintenance signals to workload-preserving actions. See which triggers ship and which remain designs.
read →
GPU failure frequency scales with the fleet, not the on-call rotation
Use published cluster studies to understand how GPU job interruptions scale, distinguish measurements from projections, and estimate your own fleet's rate.
read →
What to do when vLLM or SGLang stops responding and nothing crashed
Separate a stalled serving engine from a hung GPU. Understand health-check limits and why recovery needs a checkpoint from before the engine stopped.
read →
What to do when DCGM flags a GPU that has a job running on it
Interpret DCGM and Xid alerts, distinguish repair from workload recovery, and decide which checkpoint to restore before draining a degraded GPU node.
read →
Is it worth acting on a GPU failure prediction?
Assess GPU failure predictions using precision, warning time and the cost of acting. Understand what checkpoints change and which failures give no warning.
read →
What happens to a training job when a GPU fails
Understand Xid faults, GPU resets and the state a training job loses. Learn why recovery depends on a checkpoint taken before the hardware fails.
read →
One node failed and the whole training job died
Learn why one failed rank stops a distributed training job, what NCCL timeouts mean, and how checkpoint state determines what a restart can recover.
read →
What torchrun, torchft and Ray Train restart from when a node dies
Compare how torchrun, torchft, Ray Train and Kubeflow recover after a node fails, including saved state, code changes and checkpoint intervals.
read →Showing 12 of 13 posts