Outages often don't announce themselves — they appear as slowly rising latency or creeping error rates while no one is watching. Effective monitoring starts with establishing baselines so alerts fire on meaningful deviations, not static thresholds. Monitoring outcomes from the user's perspective (synthetic transactions like login or checkout flows) beats checking individual components in isolation. Data from 1.8 million outages shows ~68% start outside business hours, and while the median outage lasts under two minutes, a long tail of multi-hour failures demands alert design that accounts for rare, prolonged events. Alert fatigue is addressed through weekly alert review sessions to deliberately reduce noise, and anomaly detection should require a duration threshold before paging anyone.