SONiC Troubleshooting Guide: Drops, Optics, Techsupport - 夜莺博客

SONiC Troubleshooting Guide: Drops, Optics, Techsupport

SONiC's Linux-based architecture gives you a powerful toolkit when things go wrong — but only if you know which counter means what. The official SONiC troubleshooting guide organizes diagnostics into a small set of high-yield steps: understanding the RX/TX counter families, reading optical DOM values from transceivers, generating a techsupport dump before any isolation step, and cleanly shutting down BGP sessions to take a misbehaving switch out of service. This walkthrough applies that workflow step by step, with real command output, so you can follow the same path on any SONiC distribution and any merchant silicon platform.

Why SONiC Diagnostics Look Different

A SONiC switch is a Linux server running a set of Docker containers that program a merchant silicon ASIC. That split explains nearly every quirk you will hit while troubleshooting:

  • The kernel network stack is not the data plane. Forwarding happens in the ASIC, programmed through the Switch Abstraction Interface (SAI). Commands such as ip link or ifconfig show Linux-side representations of ports; they tell you nothing about what the ASIC is forwarding or dropping.
  • Counters come from the SDK, not the kernel. show interfaces counters is a rendering of SAI port statistics read out of the ASIC. If the SDK does not expose a counter, SONiC cannot show it to you.
  • State lives in Redis. The control and data planes exchange information through databases named CONFIG_DB, APPL_DB, STATE_DB, COUNTERS_DB and ASIC_DB. When a configuration appears to be applied but has no effect, the failure is often that a value never made it from CONFIG_DB to APPL_DB.
  • Every subsystem is a container. BGP, LLDP, teamd, SNMP, the syncd/syncd-vs pair and the management framework each run separately, so their logs and their failure modes are independent.

The practical consequence is simple: for suspected data-plane faults, look at ASIC-derived counters and container logs. For suspected control-plane faults, look inside the bgp container and at the routing table. Knowing the ASIC vendor — Broadcom, Marvell, NVIDIA/Mellanox or Cisco — also matters, because counter names and available drop reasons differ between them.

Investigating Packet Drops

Start with interface counters:

admin@sonic:~$ show interfaces counters
        Iface      RX_OK   RX_RATE   RX_UTIL   RX_ERR   RX_DRP   RX_OVR
      Ethernet0  471,729,839,997  653.87 MB/s  12.77%   0   18,682   0

Counter semantics are critical:

  • RX_ERR/TX_ERR — physical layer (L2) issues: FCS errors, runt frames. Indicates link-level problems.
  • RX_DRP — ingress pipeline drops: L2/L3/ACL drops and insufficient ingress buffer.
  • TX_DRP — egress buffer drops due to congestion, including WRED.
  • RX_OVR/TX_OVR — oversized packets.

Read them as a decision tree rather than a scoreboard. Error counters that increment steadily alongside traffic point at the physical layer: a marginal optic, a dirty or over-stressed fibre connector, a bad DAC, a speed mismatch, or a FEC mismatch. Drops without errors point into the pipeline: an ACL silently discarding, a missing ARP or MAC entry, uRPF, a VLAN or MTU mismatch, or genuine congestion. Oversize counters almost always trace back to a jumbo frame configuration where one end of the link runs a larger MTU than the other.

Reading Counters Without Fooling Yourself

Absolute counters alone can mislead — a value accumulated over six months of uptime tells you little about what is happening now. Compare rates over a defined window, and clear counters deliberately before an experiment so that what you observe afterwards is attributable to the change you made:

admin@sonic:~$ show interfaces counters -p 5
admin@sonic:~$ show interfaces counters -a
admin@sonic:~$ show interfaces counters errors
admin@sonic:~$ sudo sonic-clear counters

Two habits matter. First, always record the baseline value before you clear anything, because a cleared counter that does not move is evidence that the fault is elsewhere. Second, sample twice, several seconds apart, and subtract — the difference is the only number that describes current behaviour.

Physical Link Signal and Optics

Check optical receive power with transceiver DOM (AOC/DAC cables have no DOM values); optical power should generally be above -10 dBm:

admin@sonic:~$ show interfaces transceiver eeprom Ethernet12 --dom
ChannelMonitorValues:
  RX1Power : -5.7398dBm
  RX2Power : -4.6055dBm
  RX3Power : -5.0252dBm
  RX4Power : -12.5414dBm   <-- suspect: below -10 dBm

In that sample, three of the four lanes are healthy and one is more than 6 dB below its siblings. A single weak lane on a parallel optic is the classic signature of a partial fibre break, a contaminated connector in one position, or a failing transmitter lane inside the module. Treat the delta between lanes as seriously as the absolute value: a 6 dB spread between lanes is not normal even when every reading is above the absolute threshold.

Widen the check before replacing hardware:

admin@sonic:~$ show interfaces transceiver presence
admin@sonic:~$ show interfaces transceiver info Ethernet12
admin@sonic:~$ show interfaces transceiver error-status

DOM also exposes TX power, bias current, temperature and supply voltage for most modules. Rising bias current with falling TX power on the same lane indicates an ageing laser; an unusually hot module usually means an airflow problem in the rack, not a bad optic. And remember the interoperability rule: if one end reports DOM normally and the other reports nothing, check whether the far end is an AOC or DAC, which legitimately has no channel monitor values.

Link, Speed and FEC State

admin@sonic:~$ show interfaces status
admin@sonic:~$ show interfaces status Ethernet12
admin@sonic:~$ show interfaces link-flap

Interface status distinguishes administratively down from operationally down, and shows the negotiated speed and, on most platforms, whether the link is running with RS-FEC, FC-FEC or no FEC at all. FEC is a frequent cause of a link that comes up, passes a few hundred frames and then shows growing error counters: 100G four-lane optics and their host ports must agree on the FEC mode, and a mismatch on autonegotiation between a breakout port and a single-lane optic produces exactly this pattern. When error counters climb while drops stay flat, check FEC on both ends before you swap the module.

Buffer and Queue Drops

Congestion drops are invisible in the port-level counters on some silicon, so look one level deeper:

admin@sonic:~$ show queue counters Ethernet0
admin@sonic:~$ show priority-group headroom Ethernet0
admin@sonic:~$ show buffer_pool watermark

Queue counters reveal which class of service is being discarded, and repeated drops on a single high-priority queue with an otherwise idle port is the signature of a microburst: a millisecond-scale burst that overruns the buffer while the one-second average utilisation looks harmless. Headroom counters show whether PFC is doing its job; if shared headroom for a lossless queue is exhausted, the switch starts dropping traffic it is supposed to protect, and the fix belongs in the buffer configuration rather than in the ports. If you see PFC pause frames storming the fabric, that is its own troubleshooting path — our SONiC PFC watchdog guide covers detection and mitigation in detail.

Ingress Pipeline Drops: ACLs and Layer 3

When RX_DRP climbs with zero RX_ERR and unremarkable utilisation, stop looking at the optics and start looking at the forwarding pipeline:

admin@sonic:~$ show acl table
admin@sonic:~$ show acl rule
admin@sonic:~$ show arp
admin@sonic:~$ show mac
admin@sonic:~$ show ip route

The usual suspects are an ACL that was written for one direction and applied to both, a host whose ARP entry has aged out on the switch but not on the server, a VLAN membership mismatch between a port and the SVI, or a route that points at a next hop that no longer resolves. Each of these produces clean-looking optics with silent blackholing.

Generate a Techsupport Dump

Before isolating a device, collect a dump (equivalent of 'show tech' on other NOSes):

admin@sonic:~$ show techsupport

The archive lands in /var/dump/<HOSTNAME>_YYYYMMDD_HHMMSS.tar.gz and includes interface details, routes, BGP state, transceiver info, syslog and configs. It is the single artefact every escalation path — vendor TAC, the community, your own team the following morning — will ask for first, so generate it early, while the fault is still live:

admin@sonic:~$ show techsupport --since "2 hours ago"
admin@sonic:~$ ls -lh /var/dump/
admin@sonic:~$ tar tzf /var/dump/sonic_20260914_101500.tar.gz | head -40

A dump captured after a reload or after BGP has reconverged is worth far less than one captured at the moment of failure, because it preserves the transient state that explains the incident. Keep the archive outside the switch, and note its filename in your incident record.

Logs, Events and Core Dumps

admin@sonic:~$ show logging
admin@sonic:~$ sudo tail -f /var/log/syslog
admin@sonic:~$ docker ps
admin@sonic:~$ docker logs --tail 100 bgp
admin@sonic:~$ show services status
admin@sonic:~$ show reboot-cause

Service status catches the failure mode where a container has died and taken an entire feature with it — BGP sessions that vanished because the bgp container restarted, not because the peer failed. Reboot-cause is equally important after an unexplained outage: a switch that rebooted itself because of a watchdog or an ASIC SDK crash looks exactly like a power event until you check. If the platform supports it, also look for core dumps produced by the daemons.

Check the Control Plane

admin@sonic:~$ show ip bgp summary
admin@sonic:~$ show ip route summary
admin@sonic:~$ show lldp table

A session stuck below Established, or a prefix count that has collapsed, tells you the problem is above the data plane. LLDP is a fast sanity check on physical topology: if the neighbour table does not match your cabling documentation, you are troubleshooting the wrong device. For a full command reference, our SONiC CLI cheat sheet covers the show and config families end to end.

Isolate the Device from the Network

When a SONiC switch behaves abnormally, shut down BGP sessions without touching the config:

sudo config bgp shutdown neighbor SONIC02SPINE     # by hostname
sudo config bgp shutdown neighbor 192.168.1.124    # by IP
sudo config bgp shutdown all                        # everything

This is the cleanest way to take a suspect device out of service: the routes are withdrawn so the fabric reconverges around it, but the configuration database is untouched, so bringing it back is a single command rather than a restore. Recover with the mirror-image commands, and use interface-level shutdown when you want to stop traffic on one port only:

admin@sonic:~$ sudo config bgp startup all
admin@sonic:~$ sudo config bgp startup neighbor SONIC02SPINE
admin@sonic:~$ sudo config interface shutdown Ethernet0

A Repeatable Triage Order

  1. Capture the current state: show interfaces counters and show interfaces status, recorded rather than eyeballed.
  2. Split errors from drops. Errors lead to optics and FEC; drops without errors lead to the pipeline.
  3. If errors: read DOM, compare lanes, check FEC and speed agreement on both ends.
  4. If drops: check queue and priority-group counters for congestion, then ACLs, ARP/MAC and routing.
  5. Generate the techsupport dump while the fault is live.
  6. Check service status, reboot cause and container logs for a crashed subsystem.
  7. Only then isolate: shut the BGP sessions, confirm the fabric reconverged, and continue diagnosing off-path.

Related SONiC Reading

For a focused walkthrough of the counter families above, see our SONiC packet drop troubleshooting guide. When the fault is in the optic rather than the switch, the SFP transceiver troubleshooting checklist gives you a field-ready sequence. SONiC is also covered in our Cisco 8000 XR to SONiC migration article.

原文链接:https://github.com/sonic-net/SONiC/wiki/Troubleshooting-Guide