Prometheus snmp_exporter: Monitoring Switches and Routers - 夜莺博客

Prometheus snmp_exporter: Monitoring Switches and Routers

Prometheus cannot speak SNMP. The bridge is snmp_exporter, a small Go service that
translates an HTTP scrape into SNMP polls using a mapping file that ties OIDs to metric names.
The design is elegant — one exporter polls thousands of devices because Prometheus passes the
target as a URL parameter — but the generator, the relabeling and the counter semantics are
where people get stuck. This article covers all three, plus the alert rules that turn SNMP
counters into useful pages.

Why an Exporter and Not a Collector

Every other Prometheus integration pulls metrics from an endpoint that already speaks
Prometheus. Devices do not, and writing a custom poller means reimplementing retries, timeouts
and MIB parsing. snmp_exporter already does it and is maintained. If you would rather have a
poller writing directly to a time-series database, the TIG alternative is documented in
SNMP graphing with the TIG stack.

Step 1: Generate snmp.yml

The shipped example config is not enough for real devices; generate your own from the MIBs
you actually load.

git clone https://github.com/prometheus/snmp_exporter.git
cd snmp_exporter/generator

# Put vendor MIBs where the generator can find them
#   mibs/   — copy IF-MIB, CISCO-*, HUAWEI-*, JUNIPER-* as needed
#   generator.yml describes which modules to build

# Build the generator (Docker is easiest)
docker build -t snmp-generator .
docker run -v "${PWD}:/opt/" snmp-generator generate
#   produces snmp.yml — mount it into the exporter container

A generator.yml for a typical switch module:

modules:
  if_mib:
    walk:
      - 1.3.6.1.2.1.2                     # IF-MIB interfaces table
      - 1.3.6.1.2.1.31                    # IF-MIB extended (ifName, ifHCInOctets)
    lookups:
      - source_indexes: [ifIndex]
        lookup: 1.3.6.1.2.1.2.2.1.2       # ifDescr
        drop_source_indexes: false
    overrides:
      ifAlias:
        type: DisplayString
      ifHCInOctets:
        type: counter
      ifHCOutOctets:
        type: counter
  cisco_cpu:
    walk:
      - 1.3.6.1.4.1.9.9.109.1.1.1        # CISCO-PROCESS-MIB cpmCPUTotalTable

Correct types are what make Prometheus useful: counter feeds
rate() properly, DisplayString stays a label.

Step 2: Run the Exporter and Scrape Many Devices

docker run -d --name snmp_exporter -p 9116:9116 \
  -v /etc/snmp_exporter/snmp.yml:/etc/snmp_exporter/snmp.yml \
  prom/snmp-exporter

curl 'http://localhost:9116/snmp?target=10.0.0.20&module=if_mib' | head
# prometheus.yml
scrape_configs:
  - job_name: snmp
    scrape_interval: 60s
    scrape_timeout: 20s
    metrics_path: /snmp
    params:
      module: [if_mib]
    static_configs:
      - targets:
          - 10.0.0.20
          - 10.0.0.21
    relabel_configs:
      - source_labels: [__address__]
        target_label: __param_target
      - source_labels: [__param_target]
        target_label: instance
      - target_label: __address__
        replacement: 127.0.0.1:9116

The relabeling is the whole trick: __address__ is rewritten to the exporter, the
real device address is moved into __param_target, and instance is set
back to the device so dashboards and alerts read naturally.

Use a file-based service discovery source for anything beyond a handful of devices:

  - job_name: snmp
    file_sd_configs:
      - files: ['/etc/prometheus/targets/switches.yml']

Step 3: SNMPv3 Everywhere

# snmp.yml — generated auth stanza
auths:
  snmpv3_user:
    version: 3
    security_level: authPriv
    username: monitoring
    password: AuthPass123
    auth_protocol: SHA
    priv_protocol: AES
    priv_password: PrivPass123
    context_name:

# In prometheus.yml
    params:
      module: [if_mib]
      auth: [snmpv3_user]

SNMPv2c community strings travel in clear text. There is no deployment in which that is
acceptable on a routed network.

Step 4: Alert Rules That Earn Their Keep

groups:
- name: network
  rules:
  - alert: InterfaceDown
    expr: ifOperStatus{job="snmp"} == 2 and on(instance, ifName) (ifAdminStatus == 1)
    for: 2m
    labels: {severity: warning}
    annotations:
      summary: "{{ $labels.instance }} {{ $labels.ifName }} is down"

  - alert: InterfaceFlapping
    expr: changes(ifOperStatus{job="snmp"}[15m]) > 4
    for: 0m
    labels: {severity: warning}

  - alert: InterfaceErrorsIncreasing
    expr: rate(ifInErrors{job="snmp"}[5m]) + rate(ifOutErrors{job="snmp"}[5m]) > 1
    for: 5m
    labels: {severity: critical}

  - alert: DeviceRebooted
    expr: sysUpTime < 3600
    for: 0m
    labels: {severity: info}

  - alert: SNMPScrapeFailing
    expr: up{job="snmp"} == 0
    for: 5m
    labels: {severity: warning}
    annotations:
      summary: "SNMP polling failed for {{ $labels.instance }}"

The SNMPScrapeFailing alert is the one everyone forgets and everyone needs: a
device that stops answering SNMP disappears from your dashboards silently, and a silent
monitoring gap is indistinguishable from a healthy device.

Combine ifAdminStatus with ifOperStatus or you will page on
deliberately shut ports. And keep 64-bit counters — 32-bit octet counters wrap quickly and
rate() on a wrapped counter produces a fantasy spike.

Scaling Notes

  • One exporter handles thousands of targets, but it is a single point of failure. Run two
    behind a load balancer or split by site.
  • Set scrape_timeout below the device's real SNMP latency; a timeout produces a
    up == 0 rather than partial data.
  • Keep a gold snmp.yml in version control and generate it in CI. Hand-edited
    exporter configs drift within a month.
  • For modern devices, prefer gNMI streaming telemetry over polling where it is available —
    see gNMI streaming telemetry on IOS-XE — and keep SNMP for everything that predates it.
  • Route alerts through a routing tree with silences for maintenance windows:
    Alertmanager routing and notification configuration covers the configuration.

原文链接:https://snmp-monitoring.info/guides/grafana