Every monitoring tool ships with the same implicit promise: visibility equals safety. Turn on all the dashboards, wire up all the alerts, and nothing will surprise you.
In practice, the opposite happens. Teams that monitor everything respond to nothing well. Alert fatigue is not a buzzword — it is a measurable failure mode. When your on-call engineer gets 40 pages in a week, the 41st gets the same tired glance as the rest. The signal that matters drowns in the noise that doesn't.
We learned this the hard way. Here is the framework we use now to decide what earns a page, what earns a weekly glance, and what we deliberately stop watching.
The first filter is simple: does this signal correspond to something a customer would notice?
A spike in CPU on a background worker is interesting. A spike in response latency on an endpoint a customer hits every few seconds is urgent. Both might eventually matter, but they do not deserve the same treatment.
We sort every metric into one of three buckets:
The hard part is not sorting. The hard part is being honest about which bucket something belongs in and resisting the urge to promote everything to bucket one "just in case."
Every unnecessary alert carries a cost. The obvious cost is time — someone stops what they are doing, context-switches, investigates, and concludes nothing is wrong. The less obvious cost is trust. After enough false pages, engineers start assuming alerts are noise. They check slower. They silence more aggressively. When a real incident arrives, the muscle memory is already degraded.
We started tracking two numbers: pages per week and pages that led to an actual intervention. When the ratio got uncomfortable, we knew we had a noise problem, not a coverage problem.
Trimming alerts feels dangerous. Nobody wants to be the person who turned off the warning that would have caught the next outage. So we set rules to make the decision less emotional:
No alert without a defined response. If the runbook says "look at it and see if it's a problem," that alert should not page anyone. It belongs in a weekly review dashboard. An alert is only worth waking someone up for if there is a concrete action to take when it fires.
Thresholds must reflect real impact, not theoretical ceilings. A queue backing up to 500 items might sound alarming. But if the system clears 500 items in under a minute during normal operation, that threshold is just generating noise. Set thresholds based on observed customer impact, not on round numbers that feel tidy.
Every alert gets a review date. When we create an alert, we attach a review date — usually 30 or 60 days out. On that date, someone checks whether the alert has fired, whether those fires led to real action, and whether the threshold still makes sense. Alerts that have never fired in 90 days either need a lower threshold or need to be retired.
Silence is a valid conclusion. Removing an alert is not the same as ignoring a problem. It means you have decided the risk of that failure mode does not justify the ongoing cost of page-level urgency.
After applying this framework, our weekly page count dropped significantly. More importantly, the pages that remained were almost always real. Engineers started treating every alert as credible again, which shortened response times without any tooling changes.
We did not lose coverage. The metrics we demoted from pages to weekly reviews are still recorded. If one trends badly over days, we catch it in review. We just stopped pretending that every wobble required an immediate human response.
Monitoring everything equally is not diligence. It is an abdication of judgment. It pushes the decision about what matters from the calm moment of system design into the chaotic moment of an alert firing at 3 AM.
Choosing what not to alert on forces you to think clearly about what actually matters to the people who depend on your system. That thinking, done in advance, is worth more than any dashboard.
Be the first to comment.
0 comments
Loading comments...