Prometheus Alert Rule Builder
Build Prometheus alerting rules for availability, latency, saturation and SLO burn rate, with the for clause, labels and annotations that make an alert actionable rather than noisy.
Last reviewed by the Radiatus Cloud team
Want this automated for your stack?
We build CI/CD, Kubernetes & IaC pipelines that scale.
Alert on symptoms, not on causes
The most durable guidance from Google's SRE practice is to alert on what users experience rather than on every internal condition that might cause it. High CPU is not a problem if requests are still fast. A full disk on a node with no workloads is not urgent. Alerting on causes produces a page for every anomaly, and once responders learn that most pages are not real, the one that is gets ignored too. A small set of symptom alerts, tied to error rate, latency and availability, catches nearly everything that matters and pages far less often.
The for clause is what removes flapping
An expression that is true for one scrape interval fires immediately without a for clause, so a single slow scrape becomes a page. The for duration requires the condition to hold continuously before the alert becomes firing, which filters transients at the cost of delaying detection by that duration. Five minutes is a common compromise for a page; a warning that feeds a dashboard rather than a phone can afford much longer.
Multi window burn rate is the better SLO alert
A naive SLO alert fires when the error budget is being consumed faster than the allowed rate, which pages on any brief spike. The multi window multi burn rate approach, described in the SRE workbook, requires a fast burn over a short window and confirms it over a longer one, so a genuine incident pages within minutes while a one minute blip does not page at all. It is more configuration for a substantially better signal to noise ratio.
Related tools
- CI/CD Security Gap Analyzer — Checklist based analyzer for CI/CD pipeline security gaps.
- Docker Security Scanner — A new tool extracted from the codebase.
- Terraform Scanner — A new tool extracted from the codebase.
- SQL Formatter — Format and indent SQL queries for readability. Handles joins, subqueries and CTEs, supports common dialects, and runs entirely in your browser.
Frequently Asked Questions
What should page versus what should just be recorded?
Page for symptoms a user can feel: elevated error rate, latency past the objective, or a service that is down. Record everything else as a warning that appears on a dashboard or in a ticket. A page that does not require action within minutes should not be a page.
What for duration should I use?
Long enough that transients do not fire and short enough that detection is useful. Five minutes is a common default for a paging alert. Warnings can sit at fifteen or thirty minutes. Zero means a single bad scrape pages someone.
What is multi window burn rate alerting?
It fires when the error budget is burning fast over a short window and confirms over a longer one. The short window gives quick detection, the long window suppresses brief spikes. It is the approach the SRE workbook recommends over a single threshold.
Why use rate() with a five minute window?
rate() needs at least four samples to be reliable, so the window should be at least four times the scrape interval. Five minutes on a fifteen or thirty second scrape is comfortable. A shorter window on a slow scrape produces noisy and sometimes empty results.
How do I avoid alerts firing on absent metrics?
An expression over a metric that stops being reported evaluates to no data rather than to zero, so the alert silently stops firing. Use absent() or absent_over_time() as a separate alert to catch a target that has disappeared entirely.
Privacy & Security
Everything runs in your browser; nothing is uploaded.
How to Use
Choose an alert type and set the thresholds to generate a ready to paste alerting rule.
Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.