Checkpoint a warmed inference worker once, restore it on demand, and scale GPU capacity to live traffic instead of peak-of-peak.