Troubleshoot Input Drops in Cisco IOS XR - 夜莺博客

Troubleshoot Input Drops in Cisco IOS XR

Introduction

This document describes how to troubleshoot input drops on the interface on Cisco IOS XR routers. It covers ASR 9000 series routers, CRS series routers, and GSR 12000 series routers.

Background Information

In Cisco IOS, an input drop was due to the interface input queue getting full. Too many packets were punted to the CPU for process switching and it was not able to handle them fast enough. The input queue built up until full, causing drops.

On the ASR 9000, the Network Processor (NP) of the line card realizes that the router misses a load info and increments an NP drop counter, which is uploaded to the interface input drop counter.

The important mental shift is that "input drop" is a symptom category, not a mechanism. On classic IOS the mechanism was a full input queue. On IOS XR the interface input drop counter is a roll-up that collects drops detected all over the data path - the ingress port, the network processor, the CEF lookup stage, the netio node graph and the punt path to the control plane. The same interface counter can therefore be incremented by half a dozen unrelated root causes, and the job of troubleshooting is to narrow down which counter actually moved.

Problem: Increment in the Input Drop

Unknown Destination MAC Address or dot1q VLAN

When a Cisco IOS router sends Ethernet keepalives on its GigabitEthernet interface, those keepalives increment the input drops on the XR router because they do not have the destination MAC address of the XR router.

On the ASR 9000, the Input drop invalid DMAC and Input drop other counters in the controller stats are not incremented. To recognize these drops on the ASR 9000, find the NP handling the interface with the input drops:

show controllers np ports all location 0/<x>/CPU0
show contr np counters np<y> location 0/<x>/CPU0

For packets received with a dot1q VLAN not configured on a subinterface of the ASR 9000, the Input drop unknown 802.1Q counter is not incremented in show controllers gigabitEthernet 0/0/0/30 stats. This mismatch between the interface counter and the controller counter is the single most confusing part of the whole procedure: the interface input drop climbs, but every counter you would normally check on the port stays flat because the drop happened further into the pipeline than the port statistics can see.

Packets Dropped Due to Unrecognized Upper-Level Protocol

This category of drops is described in the Cisco Support Community article "ASR9000/XR: Drops for Unrecognized Upper-Level Protocol Error". These are packets whose EtherType or protocol field does not resolve to a handler the line card knows about - for example a proprietary or unregistered EtherType that nobody has asked the interface to process. The packet reaches the ingress and finds no protocol handler, so it is discarded and counted as an input drop.

NP Drops on the ASR 9000

Drop counters in the NP of the ASR 9000 are reported as input drops when they apply to a packet received on an interface and dropped. This does not happen for Packet Switch Engine (PSE) drops on the CRS and the XR 12000 — they are not counted as input drops.

If input drops appear on an ASR 9000 and do not match one of the reasons above, run the commands shown previously to find the NP handling the interface, then check its counters. For example, IPV4_PLU_DROP_PKT in the NP counters means the CEF/PLU entry says the packet must be dropped — for instance, there is no default route and unreachables are disabled, so packets not matching a more specific route hit a drop entry in the default CEF handler.

If a drop counter explains the input drops but the counter name is not self-explanatory, refer to "ASR9000/XR: Troubleshoot Packet Drops and Understanding NP Drop Counters". Note that first-generation (Trident) line cards use certain counter names, while new-generation (Typhoon) line cards have new counter names; find the similar counter name based on the description.

Netio

Collect a show netio idb <interface> to see the interface input drop and the netio node drop counters, which helps isolate where in the data path the packets are being dropped.

A Repeatable Troubleshooting Sequence

Knowing the individual causes is not the same as having an order to try them in. The sequence below moves from cheapest and least disruptive to most, and stops as soon as a counter matches a known cause.

  1. Confirm the drop is real and local. Run show interfaces <interface> accounting and compare the rate of input drops against the input rate. A flat counter is not your problem; a climbing counter that tracks a specific traffic type is.
  2. Check the controller statistics for the port. If Input drop unknown 802.1Q, Input drop invalid DMAC or the other named input drop counters moved, you already have your answer and you can stop.
  3. If the controller counters are flat but the interface input drop climbed, the drop happened after the MAC stage. Find the NP with show controllers np ports all location 0/<x>/CPU0.
  4. Read the NP counters with show contr np counters np<y> location 0/<x>/CPU0 and look for a single dominant drop counter rather than a spread across many.
  5. Run show netio idb <interface> to see which node in the netio graph owns the dropped packets.
  6. Only then escalate to a TAC case with the captured counters, because the counters are what let the TAC engineer name the root cause without a packet capture.

Collecting Evidence Before You Escalate

Most input drop cases that stall in a TAC queue do so because the evidence bundle was incomplete. Capture the following before you open the case, ideally while the drops are actively climbing:

  • show interfaces <interface> and show interfaces <interface> accounting - the rate context.
  • show controllers gigabitEthernet 0/0/0/30 stats - the named port-level drop counters.
  • show controllers np ports all location 0/<x>/CPU0 and the matching show contr np counters - the NP attribution.
  • show netio idb <interface> - the node-level attribution.
  • The exact line-card generation (Trident versus Typhoon) from show inventory, because the counter names differ.

With that bundle in hand, the counter name alone usually identifies the cause: a dot1q drop against an unconfigured subinterface, an invalid DMAC drop against an Ethernet keepalive from an IOS neighbour, or an IPV4_PLU_DROP_PKT drop against a missing default route with unreachables disabled. Each of those has a one-line fix - add the subinterface, stop the foreign keepalives, or restore the default route - and the counters confirm the fix took effect without a packet capture.

Understanding Which Counter Moved

The confusion in input drop cases almost always comes from looking at the wrong counter first. IOS XR has three separate places where a drop can be recorded for the same packet, and only one of them is the interface input drop counter.

  • Port and controller counters - show controllers gigabitEthernet 0/0/0/30 stats reports the named physical-layer drops such as Input drop unknown 802.1Q and Input drop invalid DMAC. These see drops that happen at the MAC stage.
  • NP counters - show contr np counters np<y> location 0/<x>/CPU0 reports drops that happen once the packet is inside the network processor, including the PLU and CEF mismatch drops. These are invisible to the port statistics.
  • Netio node counters - show netio idb <interface> reports drops attributed to a specific node in the internal data-path graph, which is the closest thing XR has to "where in the pipeline did this die".

When the interface input drop counter climbs but the port counters are flat, the drop is at the NP or the netio node, not at the port - which immediately rules out the dot1q and invalid-DMAC causes and sends you to the NP counters. When the port counters do move, you can usually stop after step two of the sequence, because the named counter already tells you the cause.

Verifying the Fix

Drops that stop climbing are the only acceptable proof. After you apply a fix - adding the missing subinterface, silencing the foreign keepalives, or restoring the default route - clear or baseline the counters and watch them for long enough to cover the traffic pattern that was triggering the drops.

clear counters gigabitEthernet 0/0/0/30
show interfaces gigabitEthernet 0/0/0/30 accounting
show controllers np counters np0 location 0/1/CPU0 | include drop

If the interface input drop stays flat across a business-cycle's worth of traffic, the fix took. If it resumes at the same rate and against the same named counter, the cause you identified was a symptom of something else - most commonly a second traffic type hitting the same interface - and you should return to step three of the sequence with the counter names from the second capture.

One last point on baselines: a small, steady trickle of input drops is normal on a busy interface and is not worth chasing. The signal is not the absolute number but the correlation - drops that track a specific peer, a specific VLAN, or a specific time of day. Capture a baseline before you change anything, because once you have paused a service or shifted traffic to fix the drops, the baseline is gone and you can no longer prove the counter moved because of your change rather than because the traffic did.

Products Covered

  • ASR 9000 Series Aggregation Services Routers
  • IOS XR Software

Related Reading on This Site

For a deeper walk through the NP counter names and the data-path stages that produce them, see IOS XR input drops and NP counters troubleshooting. If the drops are paired with physical-layer errors, IOS XR interface CRC errors troubleshooting covers the optics and cabling side. When the cause turns out to be scheduling or policy rather than a drop entry, IOS XR QoS class-map and policy shaping configuration explains the egress path, and IOS XR line card diagnostics with show diag and OIR covers the hardware path when a card is implicated.

原文链接:https://www.cisco.com/c/en/us/support/docs/ios-nx-os-software/ios-xr-software/213967-troubleshoot-input-drops-in-ios-xr.html