Writing from the
cedana team.
insights from the team.
Field reports, engineering deep-dives, and benchmarks from the Cedana team.

What actually changes on my cluster when I install Cedana?
Review the Helm install, Slurm plugins and workload opt-in settings Cedana adds to a cluster, with configuration examples and the application-code boundary.
read →
Driver, CUDA, and engine upgrades with workloads running, and the one limit
Plan rolling GPU driver, CUDA and engine upgrades around compatible capacity. Learn why existing checkpoints cannot carry a workload across a version change.
read →
Patch the GPU cluster on the security calendar, not the job calendar
Plan GPU security patches around the bulletin deadline. Move compatible workloads before maintenance and account for the last nodes crossing a driver upgrade.
read →
What happens to a training job when a spot instance is reclaimed
Compare spot interruption windows and recovery costs for GPU training. Learn what must be saved before reclaim and when spot remains worth using.
read →
What happens to a GPU pod when Kubernetes ends it
Compare the ways Kubernetes ends GPU pods, the warning each path provides, and what checkpoint recovery needs after eviction or spot-node termination.
read →
What SkyPilot, Ray Train, SageMaker and MemVerge bring back after a spot reclaim
Compare spot recovery tools by the machines, files or running state they restore, the checkpoint code they require and their published limits.
read →
Can Linux pause a process instead of killing it when memory runs out?
Learn why pausing a process does not free RAM, what earlyoom and systemd-oomd can do, and when checkpointing must act to preserve running work.
read →
Draining a GPU node in Kubernetes without losing the work on it
Understand what a Kubernetes drain does to GPU pods, how disruption budgets affect it, and how to plan checkpoint and restore around maintenance.
read →
Why an NVIDIA GPU Operator upgrade waits for your workloads
Learn why GPU Operator upgrades wait for active workloads, what causes driver pods to stall, and how compatibility limits shape a rolling upgrade.
read →Showing 9 of 9 posts