NX-OS Troubleshooting Tools: Consistency Checkers and More - 夜莺博客

NX-OS Troubleshooting Tools: Consistency Checkers and More

On NX-OS, the question that separates a five-minute fix from a multi-day investigation is whether the hardware and the software agree. Interface counters can look clean while the forwarding tables disagree with the control plane, and no amount of show interface will reveal it. NX-OS ships a set of consistency checkers for exactly this, plus a packet-analysis tool inside the switch. This article covers the tools worth knowing before you escalate, and the order in which to use them.

Consistency checkers: the first tool to reach for

A consistency checker validates the software state against the hardware state and logs the result as PASSED or FAILED. It exists to support root cause analysis and fault isolation when a symptom cannot be explained by configuration.

switch# show consistency-checker copp
switch# show consistency-checker gwmacdb
switch# show consistency-checker l3-interface interface ethernet 1/1 brief
switch# show consistency-checker vpc
switch# show consistency-checker multicast nlb cluster-ip <cluster-ip>
switch# show consistency-checker span
switch# show consistency-checker sflow

Coverage is platform- and release-dependent, so check the guide for your release before assuming a checker exists on your hardware. The useful ones in practice:

Checker What it proves
copp Control-plane policing is programmed in hardware as intended
gwmacdb Gateway MAC address database is consistent between hardware and software
l3-interface Layer 3 settings of SVI and routed interfaces are correctly programmed
vpc vPC inconsistencies, including LACP individual (I) state members lacking an egress mask
span, sflow, itd Mirroring, telemetry and ITD programming consistency

A checker run before and after a change is more valuable than a single run, because FAILED is only meaningful against a known-good baseline.

Ethanalyzer: packet capture on the switch itself

Ethanalyzer is a Wireshark-compatible capture running on the supervisor, which removes the need to mirror traffic to an external analyser for control-plane problems:

switch# ethanalyzer local interface mgmt capture-filter "udp port 1812" limit-captured-frames 500
switch# ethanalyzer local interface inband display-filter "bgp" limit-captured-frames 200
switch# ethanalyzer local interface inband capture-filter "host 10.0.0.5" write bootflash:cap.pcap

Typical uses: proving whether a RADIUS/TACACS request left the box, seeing whether BGP keepalives are arriving, and catching the actual control-plane packet behind an intermittent adjacency failure. Display filters use Wireshark syntax, so existing knowledge transfers.

Processes, resources and onboard failure logging

switch# show processes cpu sort | head
switch# show system resources
switch# show system error-id 0x401e0008
switch# show obfl error-stats
switch# show logging logfile | last 50

show system error-id translates an error code printed in a syslog message into a facility and description — for example an autocopy failure to a standby supervisor with the reason "standby disk may be full". This is one of the highest-value commands on the platform because it turns opaque bootvar and sysmgr messages into readable causes. OBFL records low-level hardware events across reboots, which is how you prove a problem pre-dates your change window.

Monitoring and telemetry tools

  • sFlow for sampled interface traffic and top-talker analysis; show sflow and its consistency checker confirm the agents are actually running.
  • SPAN for traditional mirroring, with the SPAN consistency checker to verify programming.
  • SNMP and RMON where a poller-based model is in place, including the PCAP SNMP parser for debugging odd counter values.
  • Embedded Event Manager for event-driven diagnostics — trigger a capture or a show-tech automatically when an interface flaps rather than waiting to notice.
  • Thermal monitoring and congestion detection commands for hardware-level and buffer-level evidence.

An escalation-ready order

  1. show system error-id against every error code in the log.
  2. The relevant consistency checker for the failing feature.
  3. Ethanalyzer capture for any control-plane protocol problem.
  4. OBFL to establish whether the fault predates the change.
  5. show processes cpu sort and show system resources for resource contention.
  6. Only then collect a full tech-support and escalate, with the timeline and the checker output attached.

For the vPC-specific checker and its failure signatures see Nexus vPC peer-link troubleshooting and type-1 mismatch; for an IOS-XR equivalent of the same methodology see ASR 9000 IOS XR show commands for troubleshooting.

Capturing evidence on the standby and off-box

Two habits shorten escalation considerably. First, save rather than print: write bootflash:cap.pcap on an Ethanalyzer capture and show tech-support > bootflash:ts.txt give support a file instead of a scroll-back. Second, when the fault follows a supervisor switchover, collect from the standby as well — a consistency checker that passes on the active supervisor and fails on the standby is a strong signal of a programming divergence between the two, which is exactly the class of fault a single-supervisor investigation misses.

switch# show tech-support > bootflash:ts-$(date +%F).txt
switch# show consistency-checker l3-interface interface ethernet 1/1 detail
switch# show system internal flash
switch# show system health

Choosing the right tool for the symptom

Symptom First tool
Counters clean, forwarding wrong Consistency checker for the affected feature
Protocol adjacency will not stay up Ethanalyzer capture on inband
Unexplained resets or pre-existing faults OBFL error statistics
Error code in the syslog show system error-id
Control plane flooded by one host CoPP checker plus Ethanalyzer on the punting path
Intermittent loss under load sFlow or SPAN with buffer and queue counters

原文链接:https://www.cisco.com/c/en/us/td/docs/dcn/nx-os/nexus9000/106x/configuration/troubleshooting/cisco-nexus-9000-series-nx-os-troubleshooting-guide-106x/m-troubleshooting-tools-and-methodology.html