blackbox_exporter: ICMP, TCP and HTTP Probes in Prometheus - 夜莺博客

blackbox_exporter: ICMP, TCP and HTTP Probes in Prometheus

blackbox_exporter answers the question host-based exporters cannot: is this endpoint actually reachable and responding the way a user experiences it? It probes HTTP, HTTPS, DNS, TCP and ICMP from outside the target, which makes it the right tool for SLI monitoring, for validating that a path works rather than that an interface is up, and for catching the difference between "the service is running" and "the service answers in under a second". This guide covers module definitions, the relabeling pattern required by the multi-target exporter model, and the permission fix that trips up almost every first ICMP probe.

The multi-target pattern

blackbox_exporter is a multi-target exporter: Prometheus passes the target as a URL parameter rather than the exporter being configured with a static target list. The scrape configuration therefore needs relabeling to move the target from the Prometheus target address into a __param_target parameter, and to point the actual scrape at the exporter.

scrape_configs:
  - job_name: 'blackbox'
    metrics_path: /probe
    params:
      module: [http_2xx]      # look for an HTTP 200 response
    static_configs:
      - targets:
          - http://prometheus.io
          - https://prometheus.io
          - http://example.com:8080
    relabel_configs:
      - source_labels: [__address__]
        target_label: __param_target
      - source_labels: [__param_target]
        target_label: instance
      - target_label: __address__
        replacement: 127.0.0.1:9115    # the blackbox exporter's real host:port

The three relabel rules do three things: copy the configured target into __param_target, promote it to the instance label so you can tell probes apart in a graph, and rewrite the scrape address to the exporter itself. Miss the third rule and Prometheus tries to scrape your website with the exporter's metrics path.

Probes can also be targeted through DNS service discovery, which is useful for "probe every A record of this name":

  - job_name: blackbox_all
    metrics_path: /probe
    params:
      module: [http_2xx]
    dns_sd_configs:
      - names: [example.com, prometheus.io]
        type: A
        port: 443
    relabel_configs:
      - source_labels: [__address__]
        target_label: __param_target
      - source_labels: [__param_target]
        replacement: https://$1/
        target_label: __param_target
      - source_labels: [__param_target]
        target_label: instance
      - target_label: __address__
        replacement: 127.0.0.1:9115
      - source_labels: [__meta_dns_name]
        target_label: __param_hostname   # sets the Host header and TLS SNI

__param_hostname is the detail that makes virtual-host probing work: it sets the HTTP Host header and the TLS SNI so you can probe an IP address with the correct name attached.

Module definitions in blackbox.yml

modules:
  http_2xx:
    prober: http
    timeout: 5s
    http:
      valid_status_codes: []        # defaults to 2xx
      method: GET
      preferred_ip_protocol: ip4

  http_post_2xx:
    prober: http
    timeout: 5s
    http:
      method: POST
      headers:
        Content-Type: application/json
      body: '{}'

  tcp_connect:
    prober: tcp
    timeout: 5s
 
  icmp_check:
    prober: icmp
    timeout: 5s
    icmp:
      preferred_ip_protocol: ip4

  dns_example:
    prober: dns
    timeout: 5s
    dns:
      query_name: www.example.com
      query_type: A
      valid_rcodes: [NOERROR]

The HTTP prober supports several checks worth using for real service validation rather than simple liveness: fail_if_ssl and fail_if_not_ssl, fail_if_matches_regexp and fail_if_not_matches_regexp. That last pair is how you catch the classic failure where the load balancer returns HTTP 200 with an error page — assert on content, not just on the status code. The TCP prober can also do banner and protocol conversations with query_response (send and expect pairs), letting you probe something like an IRC or SMTP banner rather than just port openness.

Timeouts, permissions and reloading

  • Each probe's timeout is derived from the Prometheus scrape_timeout, slightly reduced to allow for network delay. It can be capped lower in the module, and the default ceiling is 120 seconds if neither is set.
  • ICMP needs privileges. On Linux the exporter needs root, the CAP_NET_RAW capability, or a user inside net.ipv4.ping_group_range. The clean fix is setcap cap_net_raw+ep blackbox_exporter; alternatively set net.ipv4.ping_group_range = 0 2147483647 in sysctl. Without this the ICMP module fails silently in a way that looks like a network problem.
  • Reload configuration without a restart: send SIGHUP or POST to /-/reload. An invalid file is rejected and the previous configuration stays active — check the log if a change appears not to apply.
  • Verify a probe by hand before trusting the scrape: curl 'http://localhost:9115/probe?target=example.com&module=icmp_check' and read probe_success. Add debug=true to get a full trace of the probe.

What to alert on

probe_success is the obvious signal, but the useful ones are the durations: probe_duration_seconds for total probe time, probe_http_duration_seconds for the phases across an HTTP request, and per-phase TLS timing. Alerting on a slow phase rather than a failed probe catches degradation before users do. Route those alerts through the notification paths in Prometheus SNMP exporter monitoring switches, so a probe failure reaches a human with the target and the module in the payload. For router and switch instrumentation, combine this with the SNMP polling described in Prometheus Alertmanager routing and silences — reachability probes and interface counters answer different halves of the same question.

原文链接:https://github.com/prometheus/blackbox_exporter