cedana / product · reliability

Unbreakable AI
and HPC.

Live GPU workload migration. Consolidate fragmented capacity, preempt safely, and survive failures across regions and clouds — without changing a line of code.
Global utilization

Unify the fleet. Shared, not fragmented.

Pool GPU and CPU resources across a distributed fleet into one logical system. Idle capacity moves to where it's needed — across clusters, regions, and workloads — eliminating fragmentation and static reservations.

global fleet · auto-restore on failRUNNING
NODE-A · us-east-1c · gpu-7
RUNNING
─ idle ─
NODE-B · us-east-1d · gpu-3
STANDBY
········
steps lost · 0downtime · SLA · preserved
Reliability + resilience

Safely preemptable. Always.

01

Preventive maintenance

Live-migrate GPU workloads off nodes before hardware failures hit — zero work lost.

02

Job-level SLAs

Assign reliability targets and cost ceilings per job. Share clusters across teams without contention.

03

Datacenter-ready

Support for confidential computing containers and VMs for security-sensitive deployments.

04

Planet-scale fault tolerance

Train across multiple clusters — on-prem + cloud — and resume from system-level checkpoints on any GPU globally.

Automation · Status: LIVE

What automation unlocks.

01

Reliability

Automatically continue workloads from catastrophic failures without losing progress or restarting.

02

Productivity

Automatically migrate workloads to eliminate idle GPUs and increase throughput.

03

Operations

Perform maintenance without losing workload progress or manual re-submission.

~ / cedana / deploy ready

Command your
compute.

Your compute, liquid — checkpoint, migrate, and resume live GPU jobs across the fleet.

deployK8s helm chart · SLURM plug-in
first migration<30 min
code changes0
workloadstraining · inference · HPC
01
AWS
02
Google Cloud
03
NVIDIA
04
K8s
05
SLURM
06
Nvidia Dynamo