cedana / blog · field reports & engineering deep-dives
Writing from the
cedana team.
Engineering notes, benchmarks, and field reports covering how we build the automation layer for AI factories.
All posts · 08
insights from the team.
Field reports, engineering deep-dives, and benchmarks from the Cedana team.
The Wall-Time Limit Forces an Expensive Tradeoff in HPC.
Jobs reach their wall-time limits and lose hours or days of in-memory progress. However, the limit itself is not the problem.
read →Roofline Analysis and the Inference Value Chain
Once you stop renting intelligence, you own performance.
read →The Era of Stateful Inference: How to Improve Cost per Token.
We are entering the stateful inference era, driven by frontier models with longer context windows, longer in-flight sessions, and single instances spanning 8, 16, or more GPUs.
read →The Utilization Ceiling: Why AI and HPC Schedulers Hit 30% and How to Fix This
GPU utilization across AI and HPC workloads is fundamentally capped at 30% because schedulers cannot migrate running jobs. Cedana's CPU and GPU migration capability surpasses this limitation, unlocking near-full utilization.
read →Exploring new frontiers of Reinforcement Learning with Cedana
Cedana <3 RL !
read →Showing 5 of 5 posts