DevOps

PromQL Query Builder

Build PromQL queries for rates, error ratios, histogram percentiles, top-k and aggregation, with the counter and gauge rules that make the difference between a correct query and a plausible one.

Last reviewed by the Radiatus Cloud team

Query appears here.

Want this automated for your stack?

We build CI/CD, Kubernetes & IaC pipelines that scale.

Talk to an engineer

Counters and gauges need different functions

A counter only increases and resets to zero when the process restarts. Its raw value is meaningless, because it depends on how long the process has been running; what you want is the rate of change, which is what rate and increase compute while correctly handling the reset. A gauge goes up and down and its current value is the thing of interest, so applying rate to it produces a number that means nothing. Using the wrong function is the most common PromQL error and it produces a graph that looks entirely plausible.

Aggregate before you divide, and never average an average

An error ratio must sum the numerator and the denominator across the relevant series before dividing. Dividing series by series and then averaging weights a low traffic instance identically to a busy one, so a single request failing on an idle pod can drag the ratio to a value the aggregate never reached. The same applies to histogram quantiles: the buckets must be summed with sum by (le) before histogram_quantile is applied, because a quantile computed per series and then averaged is not a quantile of anything.

The range window has a floor

rate needs at least two samples and is unreliable below about four, so the range should be at least four times the scrape interval. On a thirty second scrape, a rate over one minute produces jumpy or empty results while five minutes is stable. The upper bound is responsiveness: a fifteen minute window smooths a spike into invisibility. Five minutes is the common default because it is comfortably above the floor for most scrape intervals.

Related tools

Frequently Asked Questions

When do I use rate versus increase versus irate?

rate gives a per second average over the window and is what you want for graphs and alerts. increase is rate multiplied by the window, useful for "how many in the last hour". irate uses only the last two samples and is far too noisy for alerting, though it can be useful for debugging a spike.

Why is my error ratio wrong?

Almost always because it divides before aggregating. Sum the numerator and denominator across series first, then divide. Dividing per series and averaging gives an idle instance the same weight as a busy one, so one failure on a quiet pod can dominate the result.

How short can the rate window be?

At least four times the scrape interval. rate needs several samples to be meaningful, so a one minute window on a thirty second scrape has two samples and produces jumpy or empty results. Five minutes is the usual safe default.

Why does my alert stop firing during an outage?

Because a metric that stops being reported evaluates to no data rather than to a bad value, so the expression returns nothing and the alert resolves. Pair every threshold alert with an absent_over_time alert on the same job.

How do I compute a percentile correctly?

Sum the histogram buckets with sum by (le) over the rate of the bucket counter, then apply histogram_quantile. Precision is limited by the bucket boundaries, so a p99 between a 0.5s and a 1s bucket is interpolated rather than measured.

Privacy & Security

Everything runs in your browser; nothing is uploaded.

Data: None
Client-side-Side
Active
v1.0

How to Use

Choose a query pattern and fill in the metric and labels to get a correct PromQL expression.

Disclaimer: This tool is provided "as is" without warranty of any kind. Results are for educational and utility purposes.