On August 14, 2024, about 0.4% of Neon customer projects in us-east-1 experienced up to 2 hours of downtime after an EC2 instance hosting a pageserver failed without warning. The incident review details the timeline: alerts fired 10 minutes after failure, migration of 17,037 projects began 30 minutes in, and full recovery took ~2 hours. Key delays included human-in-the-loop alerting, a semi-manual migration script, and stuck projects requiring manual intervention. Neon's new Storage Controller service — built with a reconciliation-loop model and autonomous heartbeat-based failure detection — can respond to node failures in seconds without human intervention. The fix is already in production for large projects (>64GiB) and Neon is accelerating rollout to all paying customers.