Writing from the
cedana team.
insights from the team.
Field reports, engineering deep-dives, and benchmarks from the Cedana team.

Which GPU workloads should I checkpoint first, and which should I leave alone?
Choose a small first workload set for GPU checkpointing. Review opt-in controls, cheap retries, real-time inputs, external side effects and untested job servers.
read →
How do I run a proof of concept for GPU checkpointing?
Plan a GPU checkpointing evaluation with a matched baseline, real interruptions, correctness checks and agreed success criteria before expanding the install.
read →
How to read a checkpoint benchmark
Evaluate checkpoint benchmarks using nine disclosures, from capture scope and clock boundaries to storage, repeated runs and performance after restoration.
read →
Does multi-node checkpointing ship today?
Distinguish a workload moving between nodes from one spanning nodes. See Cedana's shipped single-node coverage and the multi-node tier still in design partnership.
read →
Hot standby versus checkpoint recovery on the same hardware
Compare a spare GPU node with checkpoint recovery by cost, outage tolerance and saved session state. Choose per service tier and include time to find capacity.
read →
What cannot be checkpointed in a GPU workload?
Understand GPU checkpoint boundaries, unsupported resources, version constraints and external side effects your application must handle after a restore.
read →
Version mismatch and the compatibility matrix
Check the GPU, driver, engine and model versions a checkpoint records. Learn why a supported driver range does not mean checkpoints restore across versions.
read →
System-level vs application-level GPU checkpointing: the category and the bar
Compare application and system-level GPU checkpoints by saved state, code changes, overhead and six criteria for evaluating a checkpointing claim.
read →
Which checkpointing approach brings back the state your job is holding?
Compare GPU checkpointing approaches by the state they save, their limits and the integration work needed to decide whether to build or buy.
read →
What CRIU and cuda-checkpoint do when you wire them together yourself
See how CRIU and cuda-checkpoint save GPU workloads, where their support stops, and what your team must build around the open-source tools.
read →
Does Kubernetes, Docker, Slurm or Nextflow already restore a running job?
Compare what Kubernetes, Docker, Podman, Slurm and Nextflow preserve, what they restart, and where restoring a running GPU job needs additional tooling.
read →Showing 13–23 of 23 posts