Cisco ASR 9000 IOS XR show Commands for Troubleshooting - 夜莺博客

Cisco ASR 9000 IOS XR show Commands for Troubleshooting

IOS XR on the ASR 9000 is a distributed system: every line card runs its own copy of many processes, and a fault may be visible only from one node, one CPU or one subslot. Engineers who bring Cisco IOS habits to XR - a handful of global show commands - typically miss the node-scoped output that actually identifies the problem. This article arranges the commands by the layer you are debugging, from power and fabric up to control-plane processes.

Start with hardware and environment

RP/0/RSP0/CPU0:router# show platform
RP/0/RSP0/CPU0:router# show inventory
RP/0/RSP0/CPU0:router# show environment all
RP/0/RSP0/CPU0:router# show environment power-supply
RP/0/RSP0/CPU0:router# show environment temperatures
RP/0/RSP0/CPU0:router# show redundancy

show platform tells you which cards and modules the RP has actually registered - a line card that is physically present but absent from this output has a problem before it has a configuration problem. show environment in its variants reports fans, power supplies, voltages, temperatures and altitude; a single failed PSU or a temperature alarm is often the root cause behind a card that keeps reloading. Because the platform supports redundant RSPs, always confirm which node is active with show redundancy before interpreting any other output.

FPDs and subslot-level detail

RP/0/RSP0/CPU0:router# show hw-module subslot 0/1/cpu0 status
RP/0/RSP0/CPU0:router# show hw-module subslot 0/1/cpu0 brief
RP/0/RSP0/CPU0:router# show hw-module fpd
RP/0/RSP0/CPU0:router# show fpd package

Field-programmable device versions must match the software release. A mismatch shows up as an FPD "not current" entry, and the fix is an FPD upgrade - not a reload, and not a RMA.

Interfaces and optics

RP/0/RSP0/CPU0:router# show ipv4 interface brief
RP/0/RSP0/CPU0:router# show interfaces TenGigE0/1/0/0
RP/0/RSP0/CPU0:router# show interfaces TenGigE0/1/0/0 brief
RP/0/RSP0/CPU0:router# show controllers TenGigE0/1/0/0 phy

XR prints an admin-down state separately from a link-down state, and a breakout configuration changes the interface naming entirely: a 100GE port split with hw-module location 0/0/CPU0 bay 0 port 2 breakout 10xTenGigE appears as TenGigE0/0/0/2/0 through /9 rather than as one interface. When an interface is missing rather than down, check show ipv4 interfaces brief | include Ten and the breakout state before assuming a hardware fault.

Fabric and datapath

RP/0/RSP0/CPU0:router# show controllers fabric plane 0 statistics
RP/0/RSP0/CPU0:router# show controllers fabric connectivity
RP/0/RSP0/CPU0:router# show controllers npu resources

Fabric faults present as packet loss or as cards that lose connectivity to each other while every interface stays up. Comparing per-plane statistics across line cards localises the failure to a plane, a slot or a specific link, which is far faster than reading interface counters on forty ports.

Configuration and change safety

RP/0/RSP0/CPU0:router# show configuration
RP/0/RSP0/CPU0:router(config)# commit confirmed 5
RP/0/RSP0/CPU0:router# show configuration commit list
RP/0/RSP0/CPU0:router# rollback configuration last 1

XR keeps a commit database, and the target configuration is separate from the running configuration, so an unfinished change is visible with show configuration and reversible with rollback configuration. On a remote session, commit confirmed applies the change and rolls it back automatically unless it is confirmed - the single most valuable habit for anyone editing access lists on an out-of-band path.

Control plane and processes

RP/0/RSP0/CPU0:router# show processes cpu
RP/0/RSP0/CPU0:router# show processes memory
RP/0/RSP0/CPU0:router# show route summary
RP/0/RSP0/CPU0:router# show bgp summary
RP/0/RSP0/CPU0:router# show logging

On XR each process runs per node, so process commands accept a location argument and report per-CPU consumption. A process pinned at 100 percent on a single node while the rest of the chassis is idle points at that node's card or its specific task, not at a global problem. The broader workflow, including memory accounting and restart behaviour, is covered in our IOS XR processes, CPU and memory troubleshooting article.

A workable triage order

  1. show platform and show inventory - is every expected card registered?
  2. show environment all - power, temperature, fans.
  3. show redundancy - which RP is active?
  4. show interfaces ... brief - is the fault interface-scoped?
  5. Fabric and NPU counters - is traffic being dropped in the datapath?
  6. Process and logging output - is the control plane healthy?

Related material: ASR 9000 IM and SYSDB debugging, power supply module troubleshooting and IOS XR RPL routing policy.

原文链接:https://www.cisco.com/c/en/us/td/docs/routers/asr9000/software/system_management/command/reference/b-system-managment-cr-asr9000/hardware-redundancy-and-node-administration-commands.html