How to choose a GPU for LLM inference
Build an inference GPU shortlist using your traffic, memory needs, precision, latency, interconnect and rack cost, then measure startup and recovery on your path.
read →What roofline analysis cannot see: the costs outside the steady state
Add startup, recovery and movement time to your inference measurements. Understand what steady-state roofline analysis explains and which fleet costs sit outside it.
read →What Cedana saves, and what it does not
See which GPU, process, file, network and scheduler state Cedana saves, what remains outside the checkpoint, and what a compatible restore requires.
read →Sliding windows, state space, and the cost of remembering
Examine how attention layouts bound session memory, distinguish windowing from cache compression, and identify what one published configuration can establish.
read →The accuracy budget is part of the performance budget
Evaluate weight, KV-cache and expert-math precision changes against task quality, latency, throughput and memory on the same workload before changing production.
read →What an SGLang serve command commits your deployment to
Read SGLang presets as coupled choices about latency, throughput, memory and parallelism. Understand the tradeoffs and keep example values tied to their version.
read →Why more GPU is not a performance plan
Use compute, memory bandwidth, capacity and interconnect readings to diagnose inference limits before buying more GPUs. Separate busy hardware from useful output.
read →Out of bandwidth or out of capacity: memory pressure and the quantization levers
Distinguish slow token generation from a KV cache that cannot admit more work. Match weight or cache quantization to the limit and test the resulting quality.
read →Long context changes the economics of a fast first token
Size long-context serving around first-token latency, cache residency and concurrency. Track prefix misses and tail latency instead of relying on a median alone.
read →The two ceilings: why the ridge point is the GPU number that matters
Calculate a GPU's ridge point from dense compute and memory bandwidth. Compare H100, H200 and B200 on a consistent basis before diagnosing your workload.
read →One model, two workloads: prefill, decode, and why disaggregation exists
Understand prefill and decode resource needs, the KV-cache transfer cost of separating them, and how your traffic changes the right balance between serving pools.
read →MoE serving is a network problem, and more GPUs will not fix it
Diagnose compute, memory and interconnect limits in MoE inference. See why activated parameter counts and extra GPUs cannot replace communication measurements.
read →Where checkpoints live, how long they stay, and who can read them
Choose where GPU checkpoints live, how long they remain and who can read them. Identify the encryption and key-management answers a security review still needs.
read →What actually changes on my cluster when I install Cedana?
Review the Helm install, Slurm plugins and workload opt-in settings Cedana adds to a cluster, with configuration examples and the application-code boundary.
read →How much storage does checkpointing a GPU cluster need?
Estimate checkpoint storage from workload memory, retained copies and job count. Review quota controls, storage bandwidth and unanswered lifecycle behaviors.
read →When retrying the job is the right call
Decide whether a GPU job needs checkpoints using interruption risk, work at stake, rebuild cost and deadline slack. Keep application idempotency in either path.
read →How to know a restored GPU workload is correct
Test whether a restored GPU workload continues the same computation. Compare repeatable runs, check multi-GPU boundaries and record the exact environment.
read →Does moving a job change its results?
Understand what a checkpoint preserves when a GPU job moves, which versions must match, and how live inputs and ordinary GPU variation affect reproducibility.
read →Cost per token and revenue per megawatt are the same number
Calculate cost per delivered token from fleet costs and serving logs, then see how cold starts, failures and stranded capacity also affect revenue per megawatt.
read →The fleet's real yield metric is tokens per gigabyte of VRAM
Measure delivered tokens against installed VRAM. See how memory bandwidth, warm replicas, failures and stranded GPUs affect the output of an inference fleet.
read →KV-cache offload versus a saved worker: what each survives
Compare the state preserved by KV-cache offload and a full worker checkpoint. Learn what survives a crash, what must reload, and where prefix reuse still helps.
read →Tasks per dollar: the economics of running your own coding models
Compare coding APIs and self-hosted GPUs using cost per completed task. Account for idle node hours, model swaps and task completion on your own traffic.
read →Demand response for GPU fleets: what the programs require
Review two GPU demand response trials, the workloads they could flex, and why deeper power cuts require saved job state. Separate trial results from design targets.
read →Stateful AI has the same boundary problem as long HPC runs
See why inference, fine-tuning and agent workloads face the same allocation limits as HPC simulations, but carry different state and recovery obligations.
read →What is inside a GPU checkpoint?
See the weights, KV cache, CUDA context and process state inside a GPU checkpoint, how they affect its size, and what a restart has to rebuild.
read →One frontier restart burns half a monthly 99.9% error budget
Calculate how model restart time consumes an inference error budget, what replicas change, and how to compare a measured restore with your availability SLO.
read →The recovery time you can put in a filing
Define and test LLM recovery time from failure to serving again. Separate restore benchmarks from the detection, placement and recovery your team must document.
read →Every era of computing needed a migration primitive. GPUs are next
Trace migration through operating systems, virtual machines, containers and databases to understand why GPU fleets need portable running state.
read →How Cedana works: checkpoint, restore, and migration below the serving engine
Explore Cedana's daemon, CRIU and GPU capture layers, the state they restore, and the policies, storage paths and compatibility limits around them.
read →What is GPU checkpointing? A plain explanation
Learn what GPU checkpointing saves, how checkpoint, snapshot, restore and migration differ, and why saving model weights alone cannot resume a workload.
read →What a GPU checkpointing layer costs while the workload runs
Understand steady-state GPU checkpointing overhead, why driver-call patterns matter, and what Cedana's published single-GPU measurements leave unanswered.
read →Which GPU workloads should I checkpoint first, and which should I leave alone?
Choose a small first workload set for GPU checkpointing. Review opt-in controls, cheap retries, real-time inputs, external side effects and untested job servers.
read →How do I run a proof of concept for GPU checkpointing?
Plan a GPU checkpointing evaluation with a matched baseline, real interruptions, correctness checks and agreed success criteria before expanding the install.
read →How long do agentic sessions run, and where does their state live?
Explore published coding-agent traces to see how turns, cache reuse and pauses shape session affinity, GPU memory use and the cost of a lost worker.
read →When the node dies under a Jupyter session
Learn what a notebook file cannot recover after a GPU node fails, how process checkpoints preserve kernel state, and which restoration limits still apply.
read →Sovereignty has a bill: the transferred duties and the availability math
Size the availability responsibilities of running inference on your own hardware, including spare capacity, recovery time and a single node's outage allowance.
read →Your KV cache now rivals your weights
Calculate KV cache bytes from model configuration, precision and context length, then estimate how many live sessions fit after the weights are loaded.
read →How to read a checkpoint benchmark
Evaluate checkpoint benchmarks using nine disclosures, from capture scope and clock boundaries to storage, repeated runs and performance after restoration.
read →The engine layer: batching, caching, and configuration as a commitment
Understand how continuous batching, paged KV memory and prefix caching affect serving performance, and why changing a worker's regime means launching a new one.
read →Does disaggregated serving remove the need to move GPU workers?
See why separating prefill and decode does not make serving workers stateless, and what moving a worker means for its KV cache, sessions and startup cost.
read →H100 for prefill, H200 for decode: the mismatch you pay for twice
Compare H100 and H200 memory and compute for prefill and decode. Account for cache-transfer cost and why moving across GPU models still requires a cold start.
read →Expert placement in a mixture-of-experts deployment is a systems decision
Separate a model router's expert choices from the deployment layout you control. Review expert placement, attention parallelism and communication overlap.
read →One GPU fails and 71 healthy GPUs wait
Understand how one failure stalls a tensor-parallel NVL72 workload, how recovery consumes healthy GPU-hours, and where published measurements stop.
read →Five fleet signals, five policies: from alert to automatic action
Connect GPU health, thermal, reclaim, fragmentation and maintenance signals to workload-preserving actions. See which triggers ship and which remain designs.
read →Does multi-node checkpointing ship today?
Distinguish a workload moving between nodes from one spanning nodes. See Cedana's shipped single-node coverage and the multi-node tier still in design partnership.
read →GPU failure frequency scales with the fleet, not the on-call rotation
Use published cluster studies to understand how GPU job interruptions scale, distinguish measurements from projections, and estimate your own fleet's rate.
read →Sharing GPUs without fixed MIG slices
Compare MIG, time-slicing, MPS and serial sharing through checkpoints. Understand isolation, fixed slice sizes and the save-and-restore cost of switching jobs.
read →Hot standby versus checkpoint recovery on the same hardware
Compare a spare GPU node with checkpoint recovery by cost, outage tolerance and saved session state. Choose per service tier and include time to find capacity.
read →What cannot be checkpointed in a GPU workload?
Understand GPU checkpoint boundaries, unsupported resources, version constraints and external side effects your application must handle after a restore.
read →Driver, CUDA, and engine upgrades with workloads running, and the one limit
Plan rolling GPU driver, CUDA and engine upgrades around compatible capacity. Learn why existing checkpoints cannot carry a workload across a version change.
read →Version mismatch and the compatibility matrix
Check the GPU, driver, engine and model versions a checkpoint records. Learn why a supported driver range does not mean checkpoints restore across versions.
read →System-level vs application-level GPU checkpointing: the category and the bar
Compare application and system-level GPU checkpoints by saved state, code changes, overhead and six criteria for evaluating a checkpointing claim.
read →Patch the GPU cluster on the security calendar, not the job calendar
Plan GPU security patches around the bulletin deadline. Move compatible workloads before maintenance and account for the last nodes crossing a driver upgrade.
read →A 90-day wall-time PoC: what to measure before you change policy
Design a Slurm checkpointing proof of concept using accounting history, application review and agreed thresholds for completion time, queue impact and exceptions.
read →Which checkpointing approach brings back the state your job is holding?
Compare GPU checkpointing approaches by the state they save, their limits and the integration work needed to decide whether to build or buy.
read →GROMACS shows both the value and the limit of application checkpointing
Use GROMACS maxh, cpt and cpi to understand checkpointing across allocation limits, and see what other applications must build to offer the same recovery path.
read →Measuring the cost of wall-time termination from your sacct data
Use sacct to count Slurm TIMEOUT records, calculate exposed node-hours and identify repeat jobs. Separate accounting evidence from proof that work was lost.
read →What CRIU and cuda-checkpoint do when you wire them together yourself
See how CRIU and cuda-checkpoint save GPU workloads, where their support stops, and what your team must build around the open-source tools.
read →Does Kubernetes, Docker, Slurm or Nextflow already restore a running job?
Compare what Kubernetes, Docker, Podman, Slurm and Nextflow preserve, what they restart, and where restoring a running GPU job needs additional tooling.
read →How much utilization improvement do you need to break even on GPU checkpointing?
Calculate the utilization gain needed to cover a GPU checkpointing fee using your operating cost, paid hours, restore cost and recoverable work.
read →Which utilization number goes into your own-versus-rent calculation?
Compare owning GPUs, renting nodes and paying per token using delivered GPU-hours, operating costs and utilization instead of demand forecasts alone.
read →How to put a dollar figure on the GPU-hours that produced nothing
Calculate the cost of idle GPUs, cold starts and recompute using your own hourly rate. Separate paid busy time from work your fleet actually keeps.
read →Why an idle notebook keeps its GPU until you kill it
See why idle notebooks and Ray actors keep their GPUs, how shutdown tools release them, and what a checkpoint must preserve before a session ends.
read →What NVIDIA Dynamo Snapshot and Modal GPU memory snapshots restore today
Compare Dynamo Snapshot, Modal, InferX and Cedana by captured state, restore compatibility, supported workloads and the limits behind their benchmarks.
read →How much of your GPU utilization was work you kept?
Distinguish GPU activity, MFU and useful output. Use job accounting and startup timings to estimate how much paid GPU time produced work you kept.
read →How to fill idle GPUs without killing the job that fills them
Compare how GPU schedulers lend and reclaim idle capacity, what preemption costs borrowers, and where checkpointing can preserve their work.
read →Where the minutes go when an LLM worker cold starts
Trace LLM cold starts through weight loading, compilation and warmup. Compare caching fixes with checkpoint restore and understand the storage limits.
read →Can LLM inference scale to zero without paying for warm replicas?
Understand when LLM inference can scale to zero, what warm replicas cost, and how restore time and request latency determine the capacity you keep ready.
read →How to swap models on one GPU without a cold start
Compare GPU model swaps using vLLM sleep mode, SGLang and checkpoint restore. Learn where parked state lives and what each swap still costs.
read →The Wall-Time Limit Forces an Expensive Tradeoff in HPC.
Jobs reach their wall-time limits and lose hours or days of in-memory progress. However, the limit itself is not the problem.
read →Why a cluster with free GPUs still cannot place an 8-GPU job
See why scattered free GPUs cannot fit a large job, what bin packing and consolidation change, and why moving running state matters for defragmentation.
read →What would a GPU scheduler do differently if it could move running jobs?
Compare how GPU schedulers handle placement, preemption and time limits, and what checkpointing could change once a job has already started.
read →Why a healthy inference worker can still be in the wrong place
Understand how GPU topology, decode load and hardware fit affect inference workers, and why admission-time placement cannot rebalance running sessions.
read →What to do when vLLM or SGLang stops responding and nothing crashed
Separate a stalled serving engine from a hung GPU. Understand health-check limits and why recovery needs a checkpoint from before the engine stopped.
read →What to do when DCGM flags a GPU that has a job running on it
Interpret DCGM and Xid alerts, distinguish repair from workload recovery, and decide which checkpoint to restore before draining a degraded GPU node.
read →Is it worth acting on a GPU failure prediction?
Assess GPU failure predictions using precision, warning time and the cost of acting. Understand what checkpoints change and which failures give no warning.
read →What happens to a training job when a GPU fails
Understand Xid faults, GPU resets and the state a training job loses. Learn why recovery depends on a checkpoint taken before the hardware fails.
read →One node failed and the whole training job died
Learn why one failed rank stops a distributed training job, what NCCL timeouts mean, and how checkpoint state determines what a restart can recover.
read →What torchrun, torchft and Ray Train restart from when a node dies
Compare how torchrun, torchft, Ray Train and Kubeflow recover after a node fails, including saved state, code changes and checkpoint intervals.
read →What happens to a training job when a spot instance is reclaimed
Compare spot interruption windows and recovery costs for GPU training. Learn what must be saved before reclaim and when spot remains worth using.
read →What happens to a GPU pod when Kubernetes ends it
Compare the ways Kubernetes ends GPU pods, the warning each path provides, and what checkpoint recovery needs after eviction or spot-node termination.
read →What SkyPilot, Ray Train, SageMaker and MemVerge bring back after a spot reclaim
Compare spot recovery tools by the machines, files or running state they restore, the checkpoint code they require and their published limits.
read →What happens when a cluster job runs out of memory
Identify which memory limit killed your cluster job, where to find the evidence, and why recovery requires state saved before the OOM kill.
read →Why everyone over-requests memory on a shared cluster
Understand why cluster users request extra memory, how Slurm limits and sampled peaks affect sizing, and what changing an allocation costs.
read →Can Linux pause a process instead of killing it when memory runs out?
Learn why pausing a process does not free RAM, what earlyoom and systemd-oomd can do, and when checkpointing must act to preserve running work.
read →Draining a GPU node in Kubernetes without losing the work on it
Understand what a Kubernetes drain does to GPU pods, how disruption budgets affect it, and how to plan checkpoint and restore around maintenance.
read →Rebooting Slurm nodes for a kernel update without losing the running jobs
Compare Slurm reboot and reservation workflows, account for the idle time before maintenance, and plan checkpoint recovery around a kernel update.
read →Why an NVIDIA GPU Operator upgrade waits for your workloads
Learn why GPU Operator upgrades wait for active workloads, what causes driver pods to stall, and how compatibility limits shape a rolling upgrade.
read →Your Slurm job was cancelled due to time limit. What to do now
Confirm a Slurm TIMEOUT, check what progress survived, and compare checkpoints, requeue and job chains for runs that exceed the wall-time limit.
read →Slurm has no suspend option for GPU jobs. What preemption without killing the job looks like
Learn why Slurm suspend keeps GPUs allocated, how requeue and grace time work, and where checkpointing can preserve a preempted job's progress.
read →DMTCP, application checkpoints, workflow managers and system-level checkpointing: what each covers on a Slurm cluster
Compare application checkpoints, workflow resume, DMTCP and system-level GPU checkpoints for Slurm jobs, including setup responsibilities and limits.
read →Roofline Analysis and the Inference Value Chain
Once you stop renting intelligence, you own performance.
read →The Era of Stateful Inference: How to Improve Cost per Token.
We are entering the stateful inference era, driven by frontier models with longer context windows, longer in-flight sessions, and single instances spanning 8, 16, or more GPUs.
read →The Utilization Ceiling: Why AI and HPC Schedulers Hit 30% and How to Fix This
GPU utilization across AI and HPC workloads is fundamentally capped at 30% because schedulers cannot migrate running jobs. Cedana's CPU and GPU migration capability surpasses this limitation, unlocking near-full utilization.
read →Using Cedana to Live-Migrate Stateful Workloads Between Spot Instances
Save, migrate and resume a running XGBoost workload in six steps.
read →