Engineering notes and field reports from the Cedana team.
01Posts · 05
The Wall-Time Limit forces an expensive tradeoff in HPC.
Jobs reach their wall-time limits and lose hours or days of in-memory progress. However, the limit itself is not the problem.
read →Roofline Analysis and the Inference Value Chain
Once you stop renting intelligence, you own performance.
read →The Era of Stateful Inference: how to improve cost per token.
We are entering the stateful inference era, driven by frontier models with longer context windows, longer in-flight sessions, and single instances spanning 8, 16, or more GPUs.
read →The Utilization Ceiling: Why AI and HPC Schedulers Hit 30% and How to Fix This
GPU utilization across AI and HPC workloads is fundamentally capped at 30% because schedulers cannot migrate running jobs. Cedana's CPU and GPU migration capability surpasses this limitation, unlocking near-full utilization.
read →Using Cedana to Live-Migrate Stateful Workloads Between Spot Instances
Save, migrate and resume a running XGBoost workload in six steps.
read →