Writing from the
cedana team.
insights from the team.
Field reports, engineering deep-dives, and benchmarks from the Cedana team.
The Wall-Time Limit Forces an Expensive Tradeoff in HPC.
Jobs reach their wall-time limits and lose hours or days of in-memory progress. However, the limit itself is not the problem.
read →Roofline Analysis and the Inference Value Chain
Once you stop renting intelligence, you own performance.
read →The Era of Stateful Inference: How to Improve Cost per Token.
We are entering the stateful inference era, driven by frontier models with longer context windows, longer in-flight sessions, and single instances spanning 8, 16, or more GPUs.
read →The Utilization Ceiling: Why AI and HPC Schedulers Hit 30% and How to Fix This
GPU utilization across AI and HPC workloads is fundamentally capped at 30% because schedulers cannot migrate running jobs. Cedana's CPU and GPU migration capability surpasses this limitation, unlocking near-full utilization.
read →The Fork in the Road Post GPT-5
On GPT-5 being underwhelming, scaling laws and where we can get reasonable progress from. Also, should AGI still be the goal?
read →Exploring new frontiers of Reinforcement Learning with Cedana
Cedana <3 RL !
read →Supercharging Scientific Computing with Cedana
How we leverage open-source tech (like Kueue) to push scientific compute and accelerate time to insights.
read →Using Cedana to Live-Migrate Stateful Workloads Between Spot Instances
Save, migrate and resume a running XGBoost workload in six steps.
read →Showing 8 of 8 posts