cedana / blog · field reports & engineering deep-dives

Writing from the
cedana team.

Engineering notes, benchmarks, and field reports covering how we build the automation layer for AI factories.
Featured
All posts · 98

insights from the team.

Field reports, engineering deep-dives, and benchmarks from the Cedana team.

Cedana. Blueprint of a bordered panel holding four rows. Three rows, marked with a blue plus and labelled cluster, scheduler and workload, carry blue chips reading plus helm chart one daemonset, plus three lines in slurm.conf, and plus CEDANA_ENABLE equals one. The fourth row, marked with a grey minus and labelled your code, carries a grey chip reading unchanged.
· Cedana Editorial

What actually changes on my cluster when I install Cedana?

Review the Helm install, Slurm plugins and workload opt-in settings Cedana adds to a cluster, with configuration examples and the application-code boundary.

read →
Cedana. Blueprint of a key drawn in blue, its bow a saved-state icon and its blade cut with four teeth of different heights. A blue line leads from it to a node box outlined in blue whose bottom edge is cut with four matching notches and a filled blue tick. A grey dashed line leads to a second node box below, outlined in ink, whose third notch is red and shallower; the line ends in a red cross, and a grey circular restart arrow sits inside the box.
· Cedana Editorial

Driver, CUDA, and engine upgrades with workloads running, and the one limit

Plan rolling GPU driver, CUDA and engine upgrades around compatible capacity. Learn why existing checkpoints cannot carry a workload across a version change.

read →
Cedana. Blueprint of three grey job bars at staggered heights all ending just short of a red dashed vertical line under a red calendar icon marked patch date, each with a blue save icon at its end marked checkpoint. A narrow dashed column with a wrench icon stands just past the line, and blue bars continue from the far side of it.
· Cedana Editorial

Patch the GPU cluster on the security calendar, not the job calendar

Plan GPU security patches around the bulletin deadline. Move compatible workloads before maintenance and account for the last nodes crossing a driver upgrade.

read →
Cedana. Blueprint of two tracks of one job ending at the same red line marked reclaim, with a pale red band just before it. On the top track the tick marks are far apart and a red hatched block marked work lost runs back the whole interval. On the bottom track the blue ticks are dense and the hatched loss is a sliver.
· Cedana Editorial

What happens to a training job when a spot instance is reclaimed

Compare spot interruption windows and recovery costs for GPU training. Learn what must be saved before reclaim and when spot remains worth using.

read →
Cedana. Blueprint of one pod at the centre holding a blue block of GPU memory, with six lines converging on it from small circles at the left and right edges, some solid and some dashed, each ending in an arrowhead. Beneath the pod its floor is a red dashed line over a red cross. Lower right, a dashed outline of a replacement pod with an empty slot, and a blue saved-state icon pointing into it.
· Cedana Editorial

What happens to a GPU pod when Kubernetes ends it

Compare the ways Kubernetes ends GPU pods, the warning each path provides, and what checkpoint recovery needs after eviction or spot-node termination.

read →
Cedana. Blueprint of a dashed, struck-out instance on the left with three lines leaving it to the right: two grey dashed lines to an empty machine outline and to a small stack of documents, and one solid blue line, through a saved-state icon, to a machine holding the job in blue as it was running.
· Cedana Editorial

What SkyPilot, Ray Train, SageMaker and MemVerge bring back after a spot reclaim

Compare spot recovery tools by the machines, files or running state they restore, the checkpoint code they require and their published limits.

read →
Cedana. Blueprint of three panels headed by an icon and a word: a red cross marked kill, a pause glyph marked pause, a blue save icon marked checkpoint. Each panel is a machine's memory stacked to a dashed limit with a process column beside it: emptied and struck out under kill, still dark and full under pause, and emptied with its top block rising in blue to the save icon under checkpoint.
· Cedana Editorial

Can Linux pause a process instead of killing it when memory runs out?

Learn why pausing a process does not free RAM, what earlyoom and systemd-oomd can do, and when checkpointing must act to preserve running work.

read →
Cedana. Blueprint of the same node drawn three times in a row: on the left holding three blue session blocks, in the middle standing empty inside a dashed grey window with a nut-and-spanner mark above it, on the right holding the three blue blocks again. Three saved-state icons sit above the empty node, and blue dashed arrows carry the sessions up from the first node into them and down from them into the third.
· Cedana Editorial

Draining a GPU node in Kubernetes without losing the work on it

Understand what a Kubernetes drain does to GPU pods, how disruption budgets affect it, and how to plan checkpoint and restore around maintenance.

read →
Cedana. Blueprint of three columns of three nodes. The left column is tinted blue and empty, freshly upgraded. The middle and right columns are outlined in ink and hold grey work blocks, with a blue block beside each grey one in the middle column where work arrived from the left along blue dashed arrows. A red dashed line runs between the left column and the rest, and a grey arrow trying to cross it back ends in a red cross.
· Cedana Editorial

Why an NVIDIA GPU Operator upgrade waits for your workloads

Learn why GPU Operator upgrades wait for active workloads, what causes driver pods to stall, and how compatibility limits shape a rolling upgrade.

read →

Showing 9 of 9 posts