Prometheus Recording Rules and Alerting Rules in Practice - 夜莺博客

Prometheus Recording Rules and Alerting Rules in Practice

Prometheus rules come in two flavours and people routinely conflate them. Recording rules precompute an expensive expression and store the result as a new time series; alerting rules evaluate an expression and produce alerts. Both live in the same rule file, both are evaluated per group at a fixed interval, and both are invisible to your dashboards until you get the naming and aggregation right. This article covers the syntax, the naming convention that makes rules maintainable at scale, and the two timing keywords that determine alert quality.

Rule Groups and Where They Live

Rules are defined in YAML files referenced from the rule_files field of the Prometheus configuration. Groups are evaluated sequentially, and rules inside a group run in order at the same evaluation time. Files can be reloaded at runtime with SIGHUP, but only if every rule file parses — one syntax error means nothing reloads, so validate before signalling.

groups:
  - name: node-aggregations
    interval: 30s
    limit: 0
    labels:
      team: platform
    rules:
      - record: instance:node_cpu:rate5m
        expr: rate(node_cpu_seconds_total{mode!="idle"}[5m])
        labels:
          tier: compute

Recording Rule Syntax

# The output time series name. Must be a valid metric name.
record: <string>

# Evaluated every cycle; result stored under the name in 'record'.
expr: <string>

# Labels to add or overwrite on the result.
labels:
  [ <labelname>: <labelvalue> ]

Precomputing matters most for dashboards, which re-run the same expression every refresh. If a panel takes two seconds to render because it is computing a five-minute rate over ten thousand series, a recording rule reduces that to a single series lookup.

The Naming Convention

Prometheus recommends the general form level:metric:operations:

  • level — the aggregation level and labels of the output.
  • metric — the metric name, unchanged except for stripping _total from counters used with rate().
  • operations — the operations applied, newest first. _sum is omitted when other operations are present.
# Rate of HTTP requests per job, averaged over 5 minutes
job:http_requests:rate5m

# Ratio of failures to total, using _per_ and a ratio suffix
job:http_failures_per_requests:ratio_rate5m

Two aggregation rules that prevent most bad rules:

  • Always specify a without clause listing the labels you are aggregating away, so other labels such as job survive and you do not create a collision.
  • Never average a ratio or an average of averages. Aggregate numerator and denominator separately, then divide.

Alerting Rules

groups:
  - name: availability
    rules:
      - alert: InstanceDown
        expr: up == 0
        for: 5m
        keep_firing_for: 5m
        labels:
          severity: page
        annotations:
          summary: "Instance {{ $labels.instance }} down"
          description: "{{ $labels.instance }} of job {{ $labels.job }} down for 5 minutes."

for is the debounce: the alert is pending until the expression has been true for that long, which is how you stop a single scrape failure from paging someone. keep_firing_for is the mirror image — it keeps an alert firing for a period after the condition clears, which suppresses resolution/re-fire flapping when a metric is intermittently unavailable.

Alerting rule names must be valid label values. Labels and annotations both support templating, with {{ $labels.x }} for the alert's labels and {{ $value }} for the evaluated value.

Reliability Details

  • Rule evaluation gaps. If a group has not finished evaluating before the next cycle starts, the next evaluation is skipped and rule_group_iterations_missed_total increments. A climbing counter means your rule group is too expensive for its interval.
  • Limiting series. A per-group limit caps how many series a rule or alert may produce. Exceeding it discards the series and clears the alerts, and the event is logged as an evaluation error.
  • query_offset shifts a group's evaluation timestamp backwards, which is useful when rules depend on metrics that arrive with a known delay.
  • Alerting rules do not send notifications. That is Alertmanager's job — grouping, inhibition, silencing and routing all live there.

Validate Before You Ship

promtool check rules /etc/prometheus/rules/*.yml
promtool test rules unit_test.yml
curl -X POST http://localhost:9090/-/reload     # if --web.enable-lifecycle

Related reading on this site: Prometheus snmp_exporter: Monitor Switches and Routers for scraping switches and routers, Thanos: Long-Term Prometheus Storage and Global Queries for retention beyond the local Prometheus, and Alertmanager Routing, Grouping and Silences That Work for the routing and inhibition layer these rules feed.

原文链接:https://prometheus.io/docs/prometheus/latest/configuration/recording_rules/