Cisco Nexus 9000 Troubleshooting Cheat Sheet: Tools - 夜莺博客

Cisco Nexus 9000 Troubleshooting Cheat Sheet: Tools

When traffic slows down on a Nexus 9000, the fastest path to root cause is knowing which built-in tool matches which symptom. Cisco's official Nexus cheat sheet for beginners organizes exactly that: Ethanalyzer for CPU-bound traffic, SPAN and DMirror for port mirroring, ELAM and the N9K Packet Tracer for hardware forwarding analysis, PACL/RACL/VACL for selective capture, plus OBFL and event-histories for post-mortem diagnostics. This article condenses that cheat sheet into a practical reference every NX-OS engineer should keep at hand.

The single most common mistake when troubleshooting NX-OS is reaching for the wrong layer. A port that shows 100% line-rate utilisation but no drops is a capacity problem, not a forwarding problem. A port that shows drops with negligible utilisation is a forwarding or buffer problem, and needs ELAM or the packet tracer, not tcpdump. Choosing the tool from the symptom instead of from habit is what turns a two-hour investigation into a ten-minute one.

Background: CPU Traffic Versus ASIC Traffic

NX-OS gives you no visibility into ASIC-forwarded traffic from the Linux side of the box, because that traffic never touches the CPU. Any tool that runs as a process — Ethanalyzer being the prime example — can only see traffic that is punted to the supervisor: control protocols, exception traffic, and packets with TTL expiry or unsupported options.

Hardware-forwarded traffic, which is the overwhelming majority, must be inspected with hardware-aware tools: ELAM to see a real packet at a pipeline stage, the packet tracer to follow a synthetic or real packet through the ASIC, or SPAN to copy traffic out of the box to an external analyser. Severity levels, logs, OBFL and event-histories then tell you what happened after the fact.

Tool Selection Matrix

Symptom First tool Why
Control protocol flapping, CPU pegged Ethanalyzer Protocol traffic is punted to the CPU
Need the actual packets to an external analyser SPAN / ERSPAN Copies hardware-forwarded traffic off-box
One flow is dropped, counters inconclusive ELAM Inspects a real packet mid-pipeline
Unsure which path a packet takes N9K Packet Tracer Traces the full ASIC decision path
Intermittent loss, cannot be present for it PACL / RACL / VACL capture Counters prove arrival or departure
Switch rebooted or hardware faulted OBFL, event-history Persistent and in-memory event records
Suspecting a bad ASIC or port Diagnostics (online tests) Exercises hardware, not software

Ethanalyzer: Capturing CPU-Directed Traffic

Ethanalyzer captures traffic destined to or from the CPU and is excellent for slowness, congestion and latency issues. It works like tcpdump and can be run with detail options to show full packet headers, similar to Wireshark dissection, directly in the terminal.

That CPU-only scope is both its strength and its limitation. It is exactly the right tool for OSPF hello mismatches, BGP keepalive problems, DHCP relay failures, ARP storms, and any protocol that flaps; it will show you nothing about a data-plane flow that is being dropped in the ASIC, because that packet is never copied to the supervisor.

The syntax mirrors tcpdump, with a capture filter applied before buffering and a display filter applied on the captured set:

Nexus9000# ethanalyzer local interface inband display-filter "host 10.1.1.1 and icmp"
Nexus9000# ethanalyzer local interface inband capture-filter "arp" limit-captured-frames 100
Nexus9000# ethanalyzer local interface mgmt capture-filter "tcp port 22" limit-captured-frames 50
Nexus9000# ethanalyzer local interface inband write bootflash:capture.pcap
Nexus9000# ethanalyzer local interface inband display-filter "ospf" decode-internal

Practical notes that save time: always cap the capture with limit-captured-frames, because an unbounded Ethanalyzer session on a busy control plane will consume supervisor memory and can destabilise the switch. Write to bootflash and analyse the pcap off-box if you need to correlate with a real analyser, and remember that detail gives you a hex plus decoded view while the default view is a one-line summary per packet.

SPAN and DMirror

SPAN mirrors selected interfaces or VLANs to a destination port:

Nexus9000(config)# monitor session 1
Nexus9000(config-monitor)# source interface ethernet 1/1
Nexus9000(config-monitor)# destination interface ethernet 1/5
Nexus9000(config-monitor)# no shut

DMirror provides the same capture capability for CPU-directed traffic on Broadcom-based Nexus devices.

A local SPAN session is a copy of hardware-forwarded traffic, which means it is the only way to hand real data-plane packets to an external wire analyser. Use the direction keyword deliberately — both doubles the mirrored volume and can oversubscribe the destination port, whereas rx or tx halves it. Filter by VLAN when you only care about one broadcast domain, and verify sessions with:

Nexus9000(config)# monitor session 2
Nexus9000(config-monitor)# source interface ethernet 1/10 both
Nexus9000(config-monitor)# filter vlan 100
Nexus9000(config-monitor)# destination interface ethernet 1/20
Nexus9000(config-monitor)# no shut

Nexus9000# show monitor session 1
Nexus9000# show monitor session all

The classic SPAN failure mode is a destination port that cannot absorb the mirrored rate: if Ethernet1/20 is a 10G port and the source is a saturated 40G port, you will see the SPAN destination drop frames and wrongly conclude the source is losing traffic. Always size the destination for the mirrored peak. For captures that must cross a routed boundary or a data centre, ERSPAN is the answer rather than stretching a Layer 2 SPAN across the fabric.

ELAM and N9K Packet Tracer

ELAM (Embedded Logic Analyzer Module) captures a single packet as it traverses the ASIC pipeline - ideal for verifying forwarding decisions and checking packet alterations. The Nexus 9000 Packet Tracer detects the path a packet takes through the hardware, useful for packet flow and forwarding issues. Both are non-intrusive but require understanding of the specific ASIC architecture.

Because ELAM is per-ASIC and per-pipeline-stage, you must first identify which ASIC and which instance the ingress port belongs to. The general N9K sequence is:

Nexus9000# show platform internal tah interface ethernet 1/1
Nexus9000# debug platform internal tah elam asic 0
Nexus9000(debug-platform-utils)# trigger init in-select 15 out-select 0
Nexus9000(debug-elam)# set outer ipv4 src_ip 10.1.1.1 dst_ip 10.2.2.2
Nexus9000(debug-elam)# start
Nexus9000(debug-elam)# report

The report output tells you the result code at that stage: whether the packet was forwarded, dropped, and if dropped, which lookup failed. Common results map to recognisable faults — a source-MAC lookup miss points at MAC learning or a VLAN mismatch, an ACL result points at a policy, and a TTL or MTU result points at the packet itself. ELAM keywords vary by ASIC family and platform revision, so confirm the exact trigger syntax against the NX-OS platform documentation for your chassis before running it in production.

The packet tracer is the friendlier of the two and is usually the first choice:

Nexus9000# debug platform packet-trace packet 16 fia-trace
Nexus9000# debug platform packet-trace enable
! generate or wait for the traffic, then:
Nexus9000# show platform packet-trace summary
Nexus9000# show platform packet-trace packet 16
Nexus9000# debug platform packet-trace reset

Always reset the trace when finished; a lingering packet-trace session consumes buffers and can mask the counters you were originally investigating.

PACL/RACL/VACL for Intermittent Packet Loss

For intermittent traffic loss, PACL/RACL/VACL capture can confirm whether packets arrive or leave a certain port or VLAN. Apply an access-list and attach it in the right direction:

Nexus9000(config)# ip access-list CAPTURE
Nexus9000(config-acl)# permit ip any any capture
Nexus9000(config)# interface ethernet 1/1
Nexus9000(config-if)# ip access-group CAPTURE in

The trick is to place the same capture ACL at two points and read the counters as a subtraction. Put it on the ingress port and on the egress port of the candidate path; if the ingress counter increments and the egress counter does not, the packet was dropped inside the switch — and the remaining question becomes which stage dropped it, which is exactly what ELAM answers. Read and clear the counters with:

Nexus9000# show ip access-lists CAPTURE
Nexus9000# clear ip access-list counters
Nexus9000# show running-config interface ethernet 1/1

Remember that permit ... capture still permits the traffic; it adds a copy and a counter, it does not filter. The cost is CPU and buffer overhead, which is why these capture ACLs should be removed as soon as the investigation closes.

OBFL and Event-Histories

OBFL (Onboard Failure Logging) records hardware and environment events in non-volatile storage:

Nexus93180(config)# show logging onboard module 1 ?

The useful queries are the environmental and error-statistics views, which survive a reboot and therefore tell you what the switch saw even if you were not watching:

Nexus93180# show logging onboard module 1 temperature
Nexus93180# show logging onboard module 1 voltage
Nexus93180# show logging onboard module 1 error-stats
Nexus93180# show logging onboard module 1 boot-reason

Event-histories log protocol and process state transitions, invaluable for understanding what happened before an issue:

Nexus93180# show ip ospf event-history ?
Nexus93180# show bgp process event-history ?
Nexus93180# show system internal dmesg
Nexus93180# show accounting log

An event-history is a ring buffer per process, and the value is in the ordering: you can see the exact sequence of adjacency state changes, timer expiries and interface transitions that led to the failure, which the main syslog ring rarely preserves in enough detail. When capturing logs, raise the buffer first so the evidence is still there when you go looking:

Nexus9000(config)# logging logfile messages 7 size 16384
Nexus9000(config)# logging level vpc 6
Nexus9000# show logging logfile | include VPC

Debugs, EEM and Diagnostics

Targeted debugs and Embedded Event Manager (EEM) policies help automate troubleshooting. For hardware health, run show diagnostic content module all to list online tests (bootup level, per-port, disruptive and monitoring tests) available on the switch, then invoke them to verify ASIC and port health.

Debug commands on NX-OS are process-scoped rather than global, which makes them far safer than the IOS equivalent — but they still cost CPU, so scope them to a specific interface, VRF or neighbour wherever the syntax allows. EEM is the right answer when the problem is intermittent and you cannot sit on a console: define an applet triggered by a syslog pattern or a counter threshold that captures state and writes it to bootflash, so the evidence is collected even though the symptom came and went.

Nexus9000# show diagnostic content module all
Nexus9000# show diagnostic result module 1
Nexus9000# diagnostic start module 1 test 1
Nexus9000# show module
Nexus9000# show hardware

Note the disruptive classification in the diagnostic content output: some tests take ports down or reset an ASIC, so they belong in a maintenance window, not in a live investigation. Start with the non-disruptive and monitoring tests, which give you health data without touching forwarding.

Verification and TAC Escalation

Whichever tool you used, close the loop by re-running the counter or test that showed the problem. A capture that proves packets arrive and an egress counter that now increments is the evidence that the fix worked; a cleared counter that increments again is the evidence that it did not. Before opening a TAC case, collect a consistent bundle: show tech-support, the relevant event-histories, OBFL output for the affected module, the counter deltas, and any pcap you wrote to bootflash. Fixing the timestamps so they line up across the switch, the analyser and the server is what makes the case solvable in one call instead of five.

FAQ

Why can't Ethanalyzer see my data traffic? Because it taps the supervisor's own interface. Only punted traffic reaches it. Use SPAN, ELAM or the packet tracer for hardware-forwarded traffic.

Do I need a maintenance window for the packet tracer? No. ELAM and the packet tracer are read-only observers on the ASIC pipeline. Diagnostics, by contrast, can be disruptive.

How many SPAN sessions can I run? It depends on the platform ASIC; the N9K family supports multiple local sessions but shares the mirroring resources across ports, so verify against your platform's scale table before adding sessions in production.

What replaces show tech-support if collecting it disrupts the switch? Nothing does — but you can scope it. Collect show tech-support detail for a single module or feature rather than the whole box, and capture the specific event-histories you need.

Related Reading

原文链接:http://cisco.com/c/en/us/support/docs/switches/nexus-9000-series-switches/218096-troubleshoot-nexus-cheat-sheet-for-begin.pdf