Alert fatigue: measuring alert quality and tuning thresholds with symptoms, SLOs and burn rates
For technicians, MSP NOC staff and systems administrators who already respond to monitoring alerts and set basic thresholds, and whose on-call is now noisy. It builds on 'Monitoring and alert response' and 'RMM fundamentals: monitoring, alerts and patching'. You'll sort alerts into symptoms and causes and decide what pages, what becomes a ticket and what stays on a dashboard; measure alert precision and page load; calculate error budgets and burn rates for a service level objective; choose tuning changes knowing what they do to detection and reset time; and keep an alert review record. It uses Google's Site Reliability Engineering books and the Prometheus documentation; the ideas carry over to RMM and other monitoring tools, whose settings differ. Your client agreements and on-call policy decide which services get SLOs and who may change alert rules. Ends with a supervisor-graded alert review.
- Level
- Intermediate
- Length
- About 95 minutes
- Contents
- 5 lessons · final exam
- Status
- Published · updated 10 Oct 2026
Skills you'll practise
- Classify alerts as symptom-based or cause-based and decide whether each should page, open a ticket or appear only on a dashboard
- Calculate alert precision and page load per shift from an alert review, and identify the alerts that cause alert fatigue
- Calculate an error budget and burn rate, and choose burn-rate alert thresholds and windows for a service level objective
- Choose a tuning change for a noisy alert (window, duration, prediction, grouping, inhibition, silence, automation or removal) and explain its effect on detection and reset time
- Write an alert review record that tracks page load and records a decision, owner and re-check for each noisy alert
Course outline
- 1.Classify alerts by symptom or cause, and decide what pagesLesson · 16 min
- 2.Calculate alert precision and page load, and find the fatigueLesson · 14 min
- 3.Calculate an error budget and burn rate, and choose burn-rate alertsLesson · 18 min
- 4.Choose a tuning change and know what it does to detection and reset timeLesson · 16 min
- 5.Write an alert review recordLesson · 12 min
- 6.Alert fatigue: measuring alert quality and tuning thresholds with symptoms, SLOs and burn rates: knowledge checkKnowledge check · 16 questions
- 7.Alert fatigue: measuring alert quality and tuning thresholds with symptoms, SLOs and burn rates: practical exerciseKnowledge check · 1 question
- 8.Final exam10 questions · passing it completes the course, so people who already know the material can test out
Sources it draws on
The lessons and questions are written from these references, so learners can go back to the original.
- Google SRE book, chapter 6: Monitoring Distributed Systems (Rob Ewaschuk)
- Google SRE book, chapter 11: Being On-Call
- Google SRE workbook, chapter 5: Alerting on SLOs
- Prometheus documentation: Alerting (best practices)
- Prometheus documentation: Alerting rules (for and keep_firing_for)
- Prometheus documentation: Alertmanager (grouping, inhibition, silences)
- Prometheus documentation: Query functions (predict_linear)
See it with your own jobs and topics
Tell us about your team and we'll walk you through setup, from choosing jobs to your first skills check.