Reliability work is typically framed as deterministically eliminating a fixed percentage of incidents, but this framing is misleading. Historical incident data has too much variance and systems change over time, making point estimates of reliability improvements essentially meaningless beyond persuading leadership. A more honest framing is that all reliability mechanisms — load shedding, autoscaling, canary deployments, etc. — improve the *odds* of a system staying up rather than guaranteeing outcomes. This probabilistic lens is especially useful for valuing resilience work like improving incident responder skills, which has no quantifiable impact estimate but clearly improves the odds of handling future unforeseen incidents well.
Table of contents
Share this:1.9K Impressions