ASR9000 SNMP Timeouts: Architecture and Troubleshooting - 夜莺博客

ASR9000 SNMP Timeouts: Architecture and Troubleshooting

SNMP timeouts on an ASR 9000 are usually not an SNMP problem. IOS XR is a distributed operating system: line cards hold the counters, the route processor switch processor answers the request, and the answer has to be fetched across the inter-process communication path between them. When an NMS times out against a router whose CPU is comfortably idle, the cause is normally the cost of that assembly work, or a control-plane protection mechanism deciding your polling looks like an attack. Understanding the packet path explains both.

How an SNMP request moves inside the box

  1. The request arrives in band or out of band depending on your management-plane protection configuration, and is punted to the control plane on the RSP.
  2. It is handed to NETIO — the IOS-XR equivalent of IOS's IP input process, handling process-level switching.
  3. If the request is "for me", it goes to the SNMP daemon (snmpd), which evaluates it and dispatches the required work.
  4. For counters that live on a line card, the request is relayed to that node and the replies are gathered and assembled.
  5. The assembled response is returned to the NMS.

Every one of those hops is a place where latency accumulates, and step 4 is why polling an XR router is not equivalent to polling an IOS router with the same number of interfaces. On a chassis with many line cards and a large MIB walk, the assembly time dominates the response time — so a timeout value tuned on a fixed IOS switch fails on the chassis.

Control-plane protection: the other half of the story

IOS XR protects the control plane with LPTS (Local Packet Transport Services) and MPP (Management Plane Protection). LPTS polices which punted traffic is allowed to reach the control plane and at what rate; MPP restricts management traffic to specified interfaces. Two consequences for monitoring:

  • SNMP overload control deliberately limits the impact SNMP can have on the device. If an NMS polls aggressively — especially a MIB walk that generates a burst of requests — overload control rate-limits or drops them, which the NMS reports as a timeout.
  • If MPP is configured, SNMP arriving on an interface that is not permitted is dropped before it ever reaches snmpd. The device looks healthy, the community string is correct, and nothing answers.

What to check on the router

show snmp
show snmp users
show snmp context
show snmp engineid
show running-config snmp-server
show control-plane management-plane
show lpts pifib hardware police location 0/0/CPU0
show processes cpu | include snmp

Check the SNMP process CPU rather than total CPU: the RSP can be at 15 percent overall while snmpd is saturated by a MIB walk, which is exactly the scenario that produces intermittent timeouts correlated with the polling window.

What to fix on the NMS

Cisco's guidance is explicit about the default behaviour of most NMS platforms being wrong for a distributed chassis:

  • Use dynamic timeout where the NMS supports it. If not, multiply the default timeout by the number of applications simultaneously polling the agent on the ASR 9000.
  • Use dynamic retries where available; otherwise derive the retry count from testing rather than accepting the vendor default.
  • Stagger polling intervals so multiple management applications do not converge on the same minute. Five tools each polling 200 interfaces at *:00 is a self-inflicted outage.
  • Replace bulk MIB walks with targeted get requests on the OIDs you actually trend. A walk on the interface table of a chassis-length MIB is the single most common cause of "SNMP is broken on our new router".

Feature support notes for XR

Feature Availability
SNMP informs (not inform proxy) From 4.1
Full AES encryption for SNMPv3 4.1
IPv6 support for the SNMP engine transport 4.2
VRF-aware SNMP engine 3.3
Warm standby on the SNMP agent Yes (relevant to RP switchover behaviour)
Management Plane Protection / SNMP overload control Yes

Standards-based MIB coverage includes ENTITY-MIB, IF-MIB, IP routing MIBs, protocol MIBs (BGP, OSPF, IS-IS), MPLS/pseudowire/VPLS MIBs, and IEEE 802.x MIBs for LAG, CFM and OAM. Where you cannot find an OID, check the ASR9K MIB guide before assuming a hardware limitation — cross-platform capability files on XR are not exhaustively documented, which is why the platform-specific MIB guide exists.

A short diagnostic order

  1. Can the NMS reach the router with a single snmpget of sysUpTime from the same source address? If not, it is MPP, routing or ACL, not SNMP.
  2. Does show snmp show the expected engine ID, users and context?
  3. Is snmpd CPU saturated during the polling window?
  4. Are timeouts correlated with other management tools' schedules? If yes, stagger and increase timeout.
  5. Is the failing request a walk? Replace it with targeted gets and retest before escalating.

For the fabric and interface side of XR troubleshooting see ASR 9000 IOS XR show commands, and for hardware-level verification see ASR 9000 environment, power and temperature.

原文链接:https://community.cisco.com/t5/service-providers-knowledge-base/asr9000-xr-understanding-snmp-and-troubleshooting/ta-p/3144621