Writing from the
cedana team.
insights from the team.
Field reports, engineering deep-dives, and benchmarks from the Cedana team.

Stateful AI has the same boundary problem as long HPC runs
See why inference, fine-tuning and agent workloads face the same allocation limits as HPC simulations, but carry different state and recovery obligations.
read →
A 90-day wall-time PoC: what to measure before you change policy
Design a Slurm checkpointing proof of concept using accounting history, application review and agreed thresholds for completion time, queue impact and exceptions.
read →
GROMACS shows both the value and the limit of application checkpointing
Use GROMACS maxh, cpt and cpi to understand checkpointing across allocation limits, and see what other applications must build to offer the same recovery path.
read →
Measuring the cost of wall-time termination from your sacct data
Use sacct to count Slurm TIMEOUT records, calculate exposed node-hours and identify repeat jobs. Separate accounting evidence from proof that work was lost.
read →
Why everyone over-requests memory on a shared cluster
Understand why cluster users request extra memory, how Slurm limits and sampled peaks affect sizing, and what changing an allocation costs.
read →
Rebooting Slurm nodes for a kernel update without losing the running jobs
Compare Slurm reboot and reservation workflows, account for the idle time before maintenance, and plan checkpoint recovery around a kernel update.
read →
Your Slurm job was cancelled due to time limit. What to do now
Confirm a Slurm TIMEOUT, check what progress survived, and compare checkpoints, requeue and job chains for runs that exceed the wall-time limit.
read →
Slurm has no suspend option for GPU jobs. What preemption without killing the job looks like
Learn why Slurm suspend keeps GPUs allocated, how requeue and grace time work, and where checkpointing can preserve a preempted job's progress.
read →
DMTCP, application checkpoints, workflow managers and system-level checkpointing: what each covers on a Slurm cluster
Compare application checkpoints, workflow resume, DMTCP and system-level GPU checkpoints for Slurm jobs, including setup responsibilities and limits.
read →Showing 9 of 9 posts