Cisco IOS XR: Troubleshoot Processes, CPU and Memory - 夜莺博客

Cisco IOS XR: Troubleshoot Processes, CPU and Memory

IOS XR runs as a set of preemptive, memory-protected processes on top of a microkernel, which changes troubleshooting completely compared with classic IOS. There is no monolithic scheduler to blame for high CPU: one misbehaving process can be identified, restarted or relocated without reloading the whole router. This guide covers the process toolkit on IOS XR — CPU and memory inspection, blocked process analysis, memory leak hunting with compare reports, and the process restart procedure with its preconditions.

Start with the process list

RP/0/RSP0/CPU0:router# show processes
RP/0/RSP0/CPU0:router# show processes cpu
RP/0/RSP0/CPU0:router# show processes memory
RP/0/RSP0/CPU0:router# show processes blocked

show processes lists every executable process with JID, thread count, memory and cumulative CPU. show processes cpu sorts by utilisation — the practical starting point when the control plane feels sluggish. The output columns that matter are CPU (percentage used by the process) and HH:MM:SS (run time since last restart); a process with a large CPU share and a run time measured in seconds has restarted recently, which is a different problem from one that has simply been busy for days.

Use monitor processes for a moving picture

RP/0/RSP0/CPU0:router# monitor processes
Enter number of procs to display: 15

195 processes; 628 threads; 3375 channels, 4495 fds
CPU states: 49.0% idle, 0.9% user, 50.0% kernel
Memory: 2048M total, 1576M avail, page size 4K

    JID TIDS Chans   FDs Tmrs   MEM   HH:MM:SS   CPU  NAME
      1   27  198     2    1      0    6:11:43 50.01% kernel
     52    5  215    44    5   228K    0:00:05  0.72% devc-conaux
    293    7   31    39   11   352K    0:00:09  0.04% shelfmgr
    315    3  177    14    4     1M    0:00:11  0.03% sysdb_svr_local
    298    9  25   111    9     2M    0:00:09  0.00% snmpd

Refreshing output catches the difference between a steady load and a spike every few seconds, which a one-shot show cannot. Keep an eye on FDs as well: a file-descriptor count that climbs monotonically is a leak, even when CPU looks healthy.

Blocked processes and starvation

RP/0/RSP0/CPU0:router# show processes blocked
RP/0/RSP0/CPU0:router# show processes blocked location 0/1/CPU0

A blocked process is waiting on a resource it cannot acquire — typically memory, a semaphore or a queue that the peer process stopped draining. Blocked processes are the usual root cause of "protocols are down but CPU is low": nothing is spinning, so nothing shows as hot, yet the protocol state machine cannot advance.

Hunting memory growth

RP/0/RSP0/CPU0:router# show memory summary
RP/0/RSP0/CPU0:router# show memory compare start
RP/0/RSP0/CPU0:router# show memory compare report

show memory compare snapshots allocations per process, then reports the difference. The output ranks processes by growth and marks those that restarted during the sampling window with an asterisk.

  JID   name                 mem before   mem after    difference   mallocs  restarted
  ---   ----                 ----------   ---------    ----------   -------  ---------
   84   driver_infra_partner  577828       661492       83664        65
  279   gsp                   268092       335060       66968        396
  111   nrssvr                29152        37232        8080         60
  269   ifmgr                 539308       530652       -8656        -196        *

Retake the snapshot over an interval that reflects real traffic (an hour, not a minute) before concluding anything. A process that grows then releases is normal caching; one that only grows is a leak candidate worth a TAC case.

Restarting a process safely

RP/0/RSP0/CPU0:router# process restart nvgen
RP/0/RSP0/CPU0:router# process restart dumper location 0/1/CPU0
RP/0/RSP0/CPU0:router# show processes aborts
RP/0/RSP0/CPU0:router# show processes log

Restarting a process is far cheaper than a reload, but it is not free: a restarted routing protocol withdraws and relearns its routes, and dependent processes may reset. Check the process's dependents, confirm you are not restarting something mandatory mid-convergence, and do it in a maintenance window for anything protocol-related. show processes aborts tells you whether the process has a history of crash-restarts, which would change the plan entirely.

A practical order of operations

  1. show processes cpu and show processes blocked to see whether the problem is work or waiting.
  2. show logging filtered for process names to find a crash or restart loop.
  3. show memory compare over a realistic window if the symptom is slow degradation.
  4. Correlate with interface and fabric counters using Cisco ASR 9000 troubleshooting show commands so you can tell a control-plane symptom from a data-plane cause.
  5. If a process must be restarted, record the rollback configuration first as described in IOS XR commit and rollback configuration.

原文链接:https://www.cisco.com/c/en/us/td/docs/routers/xr12000/software/xr12k_r3-9/system_management/command/reference/yr39xr12k_chapter11.html