Cisco NCS 1002 Troubleshooting: General Diagnostics Guide - 夜莺博客

Cisco NCS 1002 Troubleshooting: General Diagnostics Guide

The Cisco NCS 1002 is a compact optical transport platform, and when a node goes unreachable or a trunk port fails to come up, the right diagnostic order saves hours of guesswork. This guide condenses Cisco's official NCS 1002 troubleshooting chapter (IOS XR releases 6.3.x through 7.3.x) into a practical procedure set: validating software installation, isolating node and management-plane failures, using loopbacks to section optical connectivity problems, and verifying alarms and onboard failure logs. It is written for transport engineers who need the exact commands, not a theory lecture.

The NCS 1002 runs IOS XR, so the diagnostic method is the familiar IOS XR one: confirm the software state, confirm the hardware and management plane, then section the optical path. The difference from a router is that most real faults show up as optical-layer problems, so loopbacks and alarm inspection carry more of the weight.

Before You Start: Prerequisites

  • Console or management access to the node, and a known-good working baseline if one exists.
  • The exact software release in use, and the release notes for it. Command availability varies across 6.3.x, 7.0.x, 7.1.x through 7.3.x.
  • The optical plan for the node: which ports are trunk, which are client, which are breakout, and what sits at the far end.
  • A record of recent changes. Most "the node broke" incidents correlate with a configuration change, an OIR event or a power event.
  • Access to show tech-support output collection, because it will be requested by TAC.

Validate the Software Installation

Start by confirming the running software and package state:

RP/0/RP0/CPU0:ios# show version
RP/0/RP0/CPU0:ios# show install active summary

Read the two outputs against each other. show version reports the running release and uptime; show install active summary lists the packages that are actually active. A mismatch between the committed release and the active packages, or a node that has rebooted into a different release after a failed activation, explains a surprising number of "the commands I know do not exist" reports.

If a software operation was interrupted or a package shows as inactive, inspect the transaction history before doing anything else:

show install log
show install committed
show install repository
show install rollback

Use the install log to identify the last completed operation and whether it committed. If the node is running a partial or rolled-back state, resolve the software state before troubleshooting anything optical — hardware alarms raised during a failed activation are often artefacts and clear once the install completes cleanly.

Node Unreachable or Console Unresponsive

When the node is unreachable, work through in order: check the management interface state and IP configuration, verify the console connection parameters, then inspect hardware module status:

show platform
show environment
show controllers management Ethernet0/RP0/CPU0/0

Interpret each output in turn. show platform gives the state of every card and its slots, so a card stuck in a non-operational state is visible immediately. show environment reports temperature, voltage, fan and power supply status — thermal and power faults are a common cause of a node that boots then fails, or one that reboots under load. show controllers management Ethernet... reports the physical and link state of the management port itself, which distinguishes a cabling or speed/duplex problem from a routing or access-list problem.

Then verify the management plane configuration:

show running-config interface MgmtEth0/RP0/CPU0/0
show running-config hostname
show route ipv4
ping <gateway-ip>
ping <remote-management-host>
show arp
  • Confirm the interface is administratively up, has the expected address and mask, and is not in a shutdown state.
  • Confirm the default route exists. A node with an address but no gateway is reachable from its own subnet and nowhere else — a frequent cause of "the node is down" when it is in fact healthy.
  • Ping the local gateway first, then a host on the far side. The point at which pings stop localises the problem to a layer.
  • Check show arp for the gateway entry. A missing ARP entry with working pings elsewhere points to an address conflict or a VLAN mismatch.
  • If the node is reachable by console but not over the network, the fault is the management path, not the node's forwarding or optical function.

For console problems, verify the terminal settings match the platform's defaults (baud rate, data bits, parity, stop bits, flow control). A console session that displays garbage characters is nearly always a baud-rate mismatch, not a failing RP.

Optical Connectivity: Use Loopbacks

Loopbacks are the fastest way to section optical faults. A facility loopback loops the signal back at the port without transmitting downstream; a terminal (inward) loopback loops it at the card. Test each segment of the path — source node, intermediate nodes, destination node — to isolate whether the problem is in the client side, the trunk, or the line system. Useful companion checks:

  • LLDP snooping - discover the far-end device and validate the optical link is carrying frames.
  • Trail trace identifier (TTI) - confirm the expected path trace is received on each section.

The method is to bisect. Set a terminal loopback on the near-end card and check whether the local port goes up and error-free. If it does, the card and its optics are likely healthy and the fault is downstream. Clear the loopback, set one at the far end, and repeat. If a loopback at the far end brings the link up, the far-end card is transmitting correctly and the problem is in the span or the intermediate equipment. If neither loopback brings the link up, the problem is local to one of the two cards or their optics.

Two cautions. First, a loopback removes the service — schedule it and note the impact. Second, always clear loopbacks explicitly when finished:

RP/0/RP0/CPU0:ios# show controllers optics 0/0/0/0
RP/0/RP0/CPU0:ios# configure
RP/0/RP0/CPU0:ios(config)# controller optics 0/0/0/0
RP/0/RP0/CPU0:ios(config-Optics)# loopback internal
RP/0/RP0/CPU0:ios(config-Optics)# no loopback internal
RP/0/RP0/CPU0:ios(config-Optics)# commit
RP/0/RP0/CPU0:ios(config-Optics)# end

Leaving a loopback configured is a well-known cause of a circuit that "mysteriously" stops passing customer traffic after a maintenance window.

Troubleshoot Trunk and Breakout Ports

For trunk port issues, verify the controller state and configuration; for breakout ports, confirm the breakout mapping matches the optics installed and that the channel configuration is consistent on both ends. A failed commit after a config change is typically a semantic error — review with:

show configuration failed

Work through the checks in this order:

  • Controller state. Confirm the trunk controller is up and the optics are recognised. An optic that is physically present but unrecognised usually indicates a compatibility or firmware issue rather than a dead optic.
  • Breakout mapping. Breakout ports depend on the installed optic's mode; a mapping that expects 4×10G from an optic configured for 1×40G will fail to come up with no obvious error. Verify the mapping matches the optic on both ends.
  • Consistency between ends. Speed, FEC mode, and payload type must match. Mismatched FEC settings are a classic cause of a link that reports up at the physical layer but never passes traffic cleanly.
  • Client-side versus trunk-side. If the client port is up and the trunk is down, the fault is on the line side; if both are down, check power and card state first.
  • After any change, verify the commit. show configuration failed prints the rejected statements and reasons, which saves guessing at what the parser disliked.

Alarms, Logs and Failure Data

Check active alarms and pull failure logs before contacting support:

show alarms brief
show logging
show tech-support

Onboard failure logging (OBFL) and process crash dumps help identify hardware versus software faults. For optical module issues, also see our H3C 光模块故障排查手册 and the Infinera DWDM resource guide.

Additional commands that pay for themselves during an incident:

show alarms detail
show logging system
show logging onboard
show oper-status
show inventory
  • show alarms detail expands each active alarm with its description and the timestamp it was raised — the sequence of timestamps is often more informative than the alarms themselves.
  • show logging system shows system-level events including boot, card insertion and removal, and power events.
  • Onboard failure logging preserves data across reloads, so it is the best source when the node has rebooted and the live logs are gone.
  • show inventory records the exact hardware and optic part numbers, which is what you need to confirm compatibility and to order replacements.

Line Card and Hardware Diagnostics

When a card is suspect, isolate hardware from configuration before swapping anything:

show platform
show redundancy
show diag
show controllers optics 0/0/0/0
show hw-module subslot 0/0/0 status

Check show redundancy for the state of the route processor pair; an RP that is not in a redundant-ready state explains loss of management access and slow or failed commits. show diag reports per-card detail including any internal errors the card has recorded. If a card reports repeated internal errors or fails to reach the operational state after a reseat, treat it as a hardware fault and follow the safe OIR procedure for removal and replacement.

Collecting Data for a TAC Case

Escalation quality determines resolution speed. Before opening a case, collect:

  1. The exact software version from show version and the active package set from show install active summary.
  2. A full show tech-support capture, taken while the fault is present. A capture taken after the fault clears is far less useful.
  3. The alarm list with timestamps, so the failure sequence is reconstructable.
  4. OBFL data and any crash or core files present on the node.
  5. The physical topology and the results of any loopback tests already performed, including which end was looped and what the port state was during the test.
  6. Optical measurements: received power at both ends of the suspect span, compared against the commissioning baseline.

Verification Checklist After the Fix

show alarms brief
show platform
show interfaces
show controllers optics 0/0/0/0
show install active summary
show logging | include ERR|CRIT
  • All previously active alarms are cleared, and no new alarm was introduced by the change.
  • Every affected card is in the operational state in show platform, with no residual failure flags.
  • Interfaces that should be up are up, with the expected line protocol state.
  • Optical receive levels are within the expected window, not merely "no alarm".
  • The running software state is consistent and committed.
  • No new error or critical messages appeared in the log since the change.

Quick Reference and Common Questions

The node answers on console but not over the network. Check the default route and the management port link state first; then ARP for the gateway. This is a management-path fault, not an optical one.

A trunk port will not come up after installing an optic. Verify the optic is recognised by the platform, then confirm the breakout mapping and FEC settings match the far end.

A commit fails with no obvious syntax error. Run show configuration failed; the reason for rejection is printed with the offending statement.

A circuit stopped passing traffic after maintenance. Check for a loopback left configured on either end before escalating.

The node rebooted and logs are missing. Pull OBFL data — it survives reloads, unlike the live log buffer.

Commands from the documentation do not exist. Verify the running release with show version; NCS 1002 command availability varies between 6.3.x and the 7.x releases.

For related diagnostics, see Cisco IOS-XR Line Card Diagnostics: show diag and Safe OIR, Cisco IOS-XR Interface CRC Errors: How to Troubleshoot, Optical Transceiver Power: Reading Rx/Tx dBm Correctly, and Coherent Optics: Pre-FEC BER, OSNR and DSP Checks. For the alarm side, read Common OTN Alarms and How to Troubleshoot Them.

原文链接:https://www.cisco.com/c/en/us/td/docs/optical/ncs1000/63x/troubleshooting/guide/troubleshooting-guide-63x/general-troubleshooting.html