InfiniBand Fabric Troubleshooting: ibstat and ibnetdiscover - 夜莺博客

InfiniBand Fabric Troubleshooting: ibstat and ibnetdiscover

InfiniBand fabric diagnosis is different from Ethernet diagnosis: links train to a width and a speed, ports move through defined states, and a subnet manager owns topology and routing. Two utilities from the infiniband-diags package answer most questions — ibstat for the local host's channel adapters, and ibnetdiscover for the topology of the whole fabric. This guide shows what each one tells you, how to compare discovery runs to find changes, and the port states that indicate a real problem rather than a normal transition.

Start with the local HCA: ibstat

ibstat
ibstat -l                 # list local CA names only
ibstat mlx5_0
ibstat -p                 # port-level detail

ibstat reads from the local IB driver and reports, per port: LID, SMLID, port state, physical state, active link width and link rate. Four fields decide whether the host is healthy:

  • State — should be Active for a functioning fabric connection. Init means the link is up but the subnet manager has not finished configuring it, and Down needs physical investigation.
  • Physical state — LinkUp versus Polling or Disabled. A port in Polling is physically connected but has not completed link training.
  • Rate — the negotiated per-lane rate.
  • Width — the negotiated lane count. 4x where you expect 12x means degraded cabling or a bad connector, and it will cost you most of your bandwidth while still showing the link as up.

Width and rate together are the number worth recording in a baseline. A fabric where every host suddenly reports half width is a cabling or transceiver problem, not an application problem.

Map the fabric: ibnetdiscover

ibnetdiscover
ibnetdiscover -H              # hosts (channel adapters) only
ibnetdiscover -S              # switches only
ibnetdiscover -R              # routers only
ibnetdiscover -l              # list of connected nodes
ibnetdiscover -p              # ports report: LID, port, GUID, width, speed
ibnetdiscover -f              # full information including port speed and width

ibnetdiscover performs subnet discovery and prints a human-readable topology: node GUIDs, node types, port numbers, port LIDs and NodeDescriptions. The port-level forms are the most useful during an incident, because a ports report lists every connected port with its LID, GUID, width and speed — you can see a 4x link sitting in a 12x fabric at a glance.

ibnetdiscover -p | head -40
# Switch 0x248a070300a1b2c0  lid 0x0001 ports 36
# 12  ->  0x248a070300a1b2d0[4]  lid 0x0004  "compute-04 HCA-1"  4x  HDR

Caching and diffing to find what changed

ibnetdiscover --cache /var/tmp/fabric-$(date +%F).cache
ibnetdiscover --load-cache /var/tmp/fabric-2026-09-01.cache
ibnetdiscover --diff /var/tmp/fabric-2026-09-01.cache
ibnetdiscover --diffcheck sw,ca,port,lid,nodedesc --diff /var/tmp/fabric-2026-09-01.cache

This is the capability that turns ibnetdiscover from a discovery tool into a change detector. Take a cached scan when the fabric is known good — after commissioning, and after every accepted change — then diff against it later. The diff reports differences in switches, channel adapters, routers and port connections, and --diffcheck narrows the comparison to the categories you care about, such as port connections and LIDs.

A diff that shows a switch disappearing usually means a subnet manager ownership problem or a powered-off switch rather than a broken link; a diff that shows a port connection moving indicates cabling work someone forgot to document.

The subnet manager is part of the diagnosis

ibdiagnet                  # broad fabric health scan
sminfo                     # is a subnet manager present, and which one
perfquery -x               # per-port error counters

ibdiagnet scans directed-route packets and reports duplicated GUIDs, links stuck in INIT, unresponsive links, error counters, routing checks, link width and speed anomalies, SM presence and partition key configuration in one pass. If ibstat shows Init across many hosts at once, run sminfo before touching cables: no active SM produces exactly that symptom, and the fix is on the manager, not the fabric.

A practical order of operations

  1. ibstat on the affected host — confirm state, width and rate before anything else.
  2. perfquery on that port for symbol errors: physical-layer errors that do not change link state still destroy throughput.
  3. sminfo to confirm a subnet manager is present and has the fabric under control.
  4. ibnetdiscover -p for the neighbours and their negotiated widths.
  5. ibdiagnet for a whole-fabric health report when the symptom is wider than one host.
  6. Diff against a cached baseline to see what changed since the last known-good state.

For design comparisons — when InfiniBand is worth its switching cost versus RoCEv2 over Ethernet — see InfiniBand versus RoCE versus Ethernet fabrics, and if your lossless Ethernet path is the one under test, the PFC and ECN tuning in MLNX-OS lossless RoCE PFC and ECN configuration covers the congestion-control side of the same problem.

原文链接:https://man7.org/linux/man-pages/man8/ibnetdiscover.8.html