SNMP Traps vs Polling: Designing Network Alerting - 夜莺博客

SNMP Traps vs Polling: Designing Network Alerting

Every network monitoring design has to answer one question: do you learn about a fault by asking devices, or by devices telling you? The answer is both, with a clear division of labour - and getting that division wrong produces the two failure modes every NOC knows: alerts that arrive minutes late, and alert storms that bury the real fault.

What Each Method Is Actually Good At

Polling Traps / informs
Detects State at poll time (up/down, counters, thresholds) The moment an event occurs
Latency Up to one poll interval Immediate (UDP, may be lost; informs are acknowledged)
Reliability High: measured directly Lower: UDP, no retransmission unless using SNMPv2c informs or v3 confirmed
Answers 'is it up now?' Yes - authoritative No - a trap only reports a transition
History / trending Yes No
Misses Events shorter than the interval Devices that cannot send, or whose trap is lost

The design rule that follows: polling is the source of truth for current state; traps are for speed and for events that do not persist. Never let a trap alone be the only evidence that something is broken, and never rely on a 5-minute poll to catch a 3-second flap.

Poll Intervals That Make Sense

  • 30-60 s for device reachability and critical link state on core and distribution gear.
  • 60-300 s for interface counters on access ports, storage and non-critical infrastructure.
  • 5-15 min for inventory, configuration drift and software versions.

Push intervals out when the collection server is the bottleneck, not when you wish the alerts were quieter. If a poll every 60 seconds is too noisy, the problem is the alarm definition, not the interval.

Which Events Belong in Traps

# Cisco IOS: a deliberately small trap set
snmp-server enable traps snmp linkdown linkup
snmp-server enable traps bgp
snmp-server enable traps config
snmp-server enable traps entity
snmp-server host 10.10.10.90 version 3 priv zbx
!
! do NOT enable traps coldstart/warmstart on every device unless you like 400 alerts
! when a distribution switch reloads

The trap set worth having: link state on uplinks, routing adjacency changes, configuration changes (a change you did not make is always interesting), hardware faults and power/PSU events. The trap set worth suppressing: everything emitted per-port per-second during a loop or a broadcast storm.

Deduplication and Dependencies - the Two Things That Kill Alert Storms

  1. Dependencies. Every interface alarm depends on the parent device being reachable. Every access-port alarm depends on the uplink. Configure this in the monitoring tool and a switch reload produces one alert instead of two hundred.
  2. Deduplication. A flapping link generates a down/up/down sequence. Group them into a single event with a flap count, and alert on the count rather than on each transition.
  3. Suppression windows. During maintenance, suppress the device - and verify afterwards that suppression ended, because a forgotten maintenance window is how real outages get missed.
  4. Rate limiting at the receiver. A trap receiver with no rate limit will be flooded exactly when the network is failing, and lose the trap you needed.

Making Traps Reliable

  • Use SNMPv3 with authPriv, or v2c informs (acknowledged) rather than unacknowledged traps, where the platform supports it.
  • Source traps from the loopback or management interface so the receiver can identify the device regardless of link state.
  • Send traps over the management network where possible; on a failing production path you will lose them.
  • Monitor the monitoring: track trap-receiver input rate and alert if it drops to zero for a device that normally reports. Silence is not health.

The Alert You Actually Want

For each alarm, be able to write one sentence: what broke, what it means for users, and what the responder should do first. If you cannot, the alarm is telemetry, not alerting - put it on a dashboard and keep it out of the pager. This single discipline reduces alert volume more than any technical tuning.

Related reading: Zabbix SNMP templates for switches, Icinga 2 configuration, and snmpwalk and OID troubleshooting when polling returns nothing at all.

原文链接:https://datatracker.ietf.org/doc/html/rfc3417