Google presents a cluster-level reliability framework for TPU superpods designed for trillion-parameter AI model training. Rather than measuring instance-level uptime, the framework evaluates aggregate availability of 'cubes' (64-chip units) within a superpod using a binomial distribution model. For Ironwood (7th-gen TPU) superpods with 144 cubes (9,216 chips), the model guarantees 130 cubes available 95% of the time, providing 8,320 fully interconnected chips for large-scale training runs. The framework integrates with a three-layer reliability stack covering infrastructure, frameworks (JAX/Pathways), and application-level fault tolerance (auto-checkpointing), all aimed at maximizing ML training goodput.

7m read timeFrom cloud.google.com
Post cover image
Table of contents
Reliability for AI supercomputersDeep dive: The mathematics of availability at scale
123 Impressions