A breakdown of what 99%, 99.9%, and 99.99% uptime actually mean for AI inference providers, mapping each reliability tier to the specific failure domains it must survive. 99% covers node-level failures (GPU faults, driver crashes), 99.9% requires surviving a full data center outage with live multi-DC traffic routing, and 99.99% demands multi-region deployment with reserved failover capacity. The post details real failure modes across compute, network, storage, and software layers, explains why infrastructure ownership matters when things break at 3 a.m., and provides a list of pointed questions to ask any inference provider about their architecture, failover testing, and SLA measurement methodology.
Table of contents
The problem with reliability numbersWhat each nine actually requiresWhat to ask before you commit to a providerOur numbers definedCome see the architecture548 Impressions