Zabbix SNMP Switch Monitoring: Templates and Traps - 夜莺博客

Zabbix SNMP Switch Monitoring: Templates and Traps

Zabbix is the most common open-source replacement for a commercial NMS, and network devices are its weakest out-of-the-box case: a host template designed for a Linux server does not poll interfaces, does not discover ports and does not receive traps. This is the minimum viable setup for switches and routers, in the order you should build it.

Step 1 - Get SNMP Right on the Device First

Zabbix cannot fix a half-configured device. On every switch, decide three things before adding a host:

  • Version. SNMPv3 with authPriv, or SNMPv2c with a source restriction. Never v2c community 'public' on a management network.
  • Source restriction. An ACL limiting polling to the Zabbix server only.
  • Contact/location. Set them; they are what turns an alert into an actionable one.
# Cisco IOS example
snmp-server community RO_POLL ro ZBX_ACL
snmp-server location DC1-RackA12
snmp-server contact noc@example.com
snmp-server enable traps snmp linkdown linkup
snmp-server enable traps entity
snmp-server host 10.10.10.50 version 2c TRAP_STRING

snmp-server group ZBX v3 priv
snmp-server user zbx ZBX v3 auth sha <auth> priv aes 128 <priv>
snmp-server host 10.10.10.50 version 3 priv zbx

Verify from the Zabbix server before blaming the template: snmpwalk -v3 -l authPriv -u zbx -a SHA -A auth -x AES -X priv 10.10.20.1 sysDescr. If the walk fails, nothing downstream will work.

Step 2 - Add the Host and the Right Interface

  • Create the host with the device's management IP, and set the SNMP interface (not the agent interface) to the same address and port 161.
  • Attach a link template first (for example Cisco IOS by SNMP or SNMP Generic), then a device-specific template only if one exists for that exact platform.
  • Set the macro for the SNMP credentials at host level ({$SNMP_COMMUNITY}) or globally so you never paste credentials into templates.

Step 3 - Interface Discovery Is the Whole Point

A device template that does not use low-level discovery (LLD) will not monitor your ports. The two items that drive it:

  • Discovery rule on ifDescr / ifName to find interfaces.
  • Item prototypes on ifHCInOctets, ifHCOutOctets, ifOperStatus, ifHighSpeed - the 64-bit HC counters, always, because 32-bit counters wrap in minutes on a 10G link.
walk ifName            : .1.3.6.1.2.1.31.1.1.1.1
ifHCInOctets           : .1.3.6.1.2.1.31.1.1.1.6
ifHCOutOctets          : .1.3.6.1.2.1.31.1.1.1.10
ifOperStatus           : .1.3.6.1.2.1.2.2.1.8
ifHighSpeed            : .1.3.6.1.2.1.31.1.1.1.15
ifAlias (description)  : .1.3.6.1.2.1.31.1.1.1.18

Two practical adjustments to the stock templates: use the interface alias/description in the item name so an alert reads 'Gi1/0/24 - uplink-to-core', and add a discovery filter that excludes Null, Vlan1 and stacked-port pseudo-interfaces so your dashboards are not flooded with noise.

Step 4 - Triggers That Do Not Cry Wolf

  • Link down: trigger on ifOperStatus change, but add a dependency so SFP/Ethernet ports covered by an aggregate do not each alert when an uplink dies.
  • Errors: alert on the rate of ifInErrors, not the absolute counter, and require sustained breach (two consecutive checks) before firing.
  • Device reachable: one host-level ping/SNMP trigger per device, with all interface triggers depending on it. Without that, a device reboot produces 200 alerts.
avg(/device/net.if.in.errors[ifHCInErrors.1],5m)>10 and
diff(/device/net.if.in.errors[ifHCInErrors.1])=1

Step 5 - Receive Traps as Well as Poll

Polling misses fast events (a 3-second link flap, a routing adjacency reset). Configure the Zabbix server as a trap receiver and create a trapper item per device, or use the built-in SNMP trap item type with a matching string such as linkDown. Keep the trap path for events and the polling path for trend and threshold data - do not try to make either one do both jobs.

Operational Hygiene

Once polling works: keep a maintenance window object per device so that a planned firmware reload does not page anyone; export templates to your Git repository; and add the polling server's own capacity to your monitoring. A Zabbix server that is 40% behind on its queue produces false 'device unreachable' alerts, which is how monitoring loses credibility.

Related: snmpwalk and OID troubleshooting for when a walk returns nothing, and sFlow on OS10 when you need per-flow rather than per-interface data.

原文链接:https://www.zabbix.com/documentation/current/en/manual/config/items/itemtypes/snmp