BGP Network Troubleshooting: Best Practices and Commands - 夜莺博客

BGP Network Troubleshooting: Best Practices and Commands

BGP is a complex protocol, and effective troubleshooting requires a calm, organized approach rather than randomly executing commands across the network. This guide lays out a checklist-based workflow that starts at layer 1 and progresses upward: verify peering status, confirm connectivity between loopbacks, check route origination, inspect update exchange, and finally audit route filtering. Each step includes the exact Cisco IOS commands and realistic output so you can reproduce the workflow on your own routers.

The reason a checklist beats intuition here is the size of the search space. A single missing prefix can be caused by a physical link, a TCP session that never completed, a source-address mismatch, a next-hop that is unreachable, an outbound filter on the sender, an inbound filter on the receiver, an AS-path loop, a dampened route, a route that exists only in the Adj-RIB-In, or a maximum-prefix condition that silently tore the session down and brought it back. Guessing which of those applies usually costs an hour; walking the chain in order usually costs ten minutes, because each step eliminates a whole branch of the tree.

Step 1: Verify BGP Peering Status

Start with show ip bgp summary. When peering uses loopback interfaces, perform a loopback-to-loopback ping — a plain ping uses the outgoing interface IP, which makes iBGP fail. Then check the local configuration:

R1# show running-config | section bgp
router bgp 64500
 bgp log-neighbor-changes
 neighbor 2.2.2.2 remote-as 64500
 neighbor 3.3.3.3 remote-as 64501
 neighbor 3.3.3.3 update-source Loopback0
 address-family ipv4
  neighbor 2.2.2.2 activate
  neighbor 3.3.3.3 activate

Use debug ip tcp transactions to see which source address the peer is initiating sessions from; if it uses the physical interface IP instead of the loopback, add update-source Loopback0 on both ends.

Read the state column carefully, because it encodes the failure. Idle means the router is not even trying — usually no route to the neighbour, an administratively shut neighbour, or a session that failed and is waiting out the connect-retry timer. Active is the misleading one: the router is actively attempting a TCP connection and failing, which points at reachability, an ACL blocking TCP 179, or a TTL problem. Connect means TCP is in handshake. OpenSent and OpenConfirm mean the transport worked and the peers are negotiating capabilities such as address families, timers or authentication. Established is the steady state; a session flapping between Established and Idle at regular intervals almost always has a maximum-prefix threshold being exceeded, an authentication mismatch appearing only after capabilities are exchanged, or a hold-timer problem.

R1# show ip bgp summary
R1# show ip bgp neighbors 2.2.2.2
R1# show ip bgp neighbors 2.2.2.2 | include BGP state|Last reset|Notification
R1# show tcp brief | include 179

The Last reset and Notification fields in the neighbour detail are the single most valuable output in BGP troubleshooting. A notification code tells you whether the peer sent an administrative shutdown, a hold-timer expiry, a bad capability or a configuration mismatch, which converts a vague flap into a specific fix.

eBGP Multihop for Non-Directly-Connected Peers

eBGP defaults to TTL 1, so multihop peering fails unless configured. For a dual-homed design where the eBGP peer is two hops away:

router bgp 64500
 neighbor 3.3.3.3 remote-as 64501
 neighbor 3.3.3.3 ebgp-multihop 2
 neighbor 3.3.3.3 update-source Loopback0

Remember that the TTL must be large enough for the real path, not the perceived one. In designs where the topology changes under failure — a backup path that is one hop longer than the primary — a multihop value sized for the primary will drop the session precisely when the backup is in use. Size it for the worst case and document why.

Step 2: Verify Missing Routes and Origination

BGP only advertises locally known routes. If a network statement does not originate, check the RIB:

R2# show ip bgp | include 50.50.50.0
R2(config)# ip route 50.50.50.0 255.255.255.0 null 0

Add a static route (optionally to Null0 for blackhole aggregates) and the prefix will be originated and advertised.

A useful refinement is to check the RIB entry type. A route learned from a more specific BGP path, an IGP route and a static route all look different in show ip route, and a network statement only originates the prefix if some route for it exists and is installed. A route that is present but recursed through an unreachable next-hop, or one that is suppressed because of a redistribution loop, will not be originated even though it appears in the configuration.

Step 3: Confirm the Next Hop Is Reachable

BGP inherits next-hop reachability from the IGP, and an unreachable next hop is one of the most common causes of "the route was advertised but never appeared". Check both the prefix state and the next hop before touching any filter.

R1# show ip bgp 50.50.50.0
R1# show ip bgp 50.50.50.0 longer-prefixes
R1# show ip route 10.10.10.10
R1# show ip bgp next-hop

This step also matters for iBGP, where the next-hop is normally unchanged across the AS. If an internal router receives an iBGP route whose next hop is an external address it has no route to, the prefix is marked inaccessible and withheld from the RIB. The standard fixes are next-hop-self on the edge router announcing the route, or enabling next-hop tracking so BGP withdraws the prefix when the IGP path disappears rather than blackholing traffic.

Step 4: Update Exchange and Route Filtering Checks

R2# show ip bgp neighbor 1.1.1.1 advertised-routes
R1# show ip bgp neighbor 2.2.2.2 routes
R1# show ip bgp neighbor 2.2.2.2 received-routes

received-routes requires soft-reconfiguration inbound. When a prefix is missing, check for prefix-list, AS_PATH and community filters applied to the neighbor, then review route-maps that modify attributes.

The distinction between the three tables is the whole point of the step. Advertised-routes is the sender's opinion; received-routes is what actually arrived before inbound policy; routes is what survived inbound policy and made it into the table. If a prefix appears in advertised-routes but not in received-routes, the loss is in transit — an AS-path loop, an outbound filter at a transit provider, or a session that went down mid-update. If it appears in received-routes but not in routes, the loss is your own inbound policy, and the fix is in a prefix-list, a route-map or an as-path filter.

Two filters cause a disproportionate share of incidents. An AS-path access list intended to block long paths may also match a legitimate peer, and an outbound prefix-list missing the le keyword will drop every more-specific advertisement. Print the filter and read it as an adversary would before assuming the protocol is at fault.

Step 5: Session, Timer and Scale Health

Once routes are exchanging correctly, confirm that the session itself is not degrading under load. Maximum-prefix limits, hold timers that are too aggressive for a control plane under CPU pressure, and BGP scanner taking too long on large tables all produce intermittent problems that look like routing bugs.

R1# show ip bgp summary | include Max|PfxRcd
R1# show ip bgp neighbors 2.2.2.2 timers
R1# show ip bgp neighbor 2.2.2.2 | include Last read|Last write|Prefix activity
R1# show processes cpu | include BGP

A PfxRcd value stuck at the configured maximum is a session about to be torn down. A large gap between Last read and Last write on a quiet session is normal, but a session that is quiet on both while the IGP underneath it is flapping indicates a transport problem that BGP is masking rather than causing.

Common Root Causes, Ranked by Frequency

After enough incidents, the distribution of causes becomes predictable, and knowing it lets you test the likely hypotheses first.

  • Transport or reachability — an ACL blocking TCP 179, a missing route to the neighbour address, an MTU problem on the path, or a TTL too low for the real topology. These produce Idle or Active states rather than partial routing.
  • Source-address mismatch — one end using a physical interface while the other expects a loopback. The session may even establish and then drop unpredictably as the source address changes with the outgoing interface.
  • Next-hop accessibility — the prefix is learned, the next hop is not, and the route is withheld. Most common after an IGP change that quietly removed the path to an external next hop.
  • Policy filtering — prefix-lists, distribute-lists, AS-path access lists, communities and route-maps applied in the wrong direction or with the wrong mask-length qualifier.
  • Attribute manipulation — a route-map rewriting the local preference, MED or AS-path so that a valid route loses best-path selection and disappears from the active table.
  • Scale and resource limits — maximum-prefix thresholds, BGP scanner timeouts, memory pressure or control-plane policing dropping BGP packets under load.

Testing in this order is not arbitrary: the upper items are cheap to check and each eliminates a large fraction of the search space before you ever open a route-map.

Convergence, Stability and Monitoring

Once a session is healthy, the remaining question is whether it stays healthy under change. Measure convergence explicitly rather than assuming it: withdraw a prefix at the edge in a maintenance window and time how long the change takes to appear in the RIB several hops away. Values that are consistently worse than the design target usually indicate aggressive route damping, a scanner that cannot keep up, or an IGP whose own convergence is the real bottleneck.

For continuous visibility, export the neighbour table and the Adj-RIB-In prefix counts to your monitoring platform and alert on deltas rather than absolutes. A session that normally carries 900,000 prefixes and suddenly carries 850,000 has lost a substantial block even though the session is still Established, and that is precisely the kind of silent degradation that a state-only alarm will miss. Pair that telemetry with an external route collector when you need a third-party view of what your AS is actually announcing.

Best Practices Summary

Work one issue at a time, verify each layer before moving up, use monitoring tools such as BGPmon or ExaBGP for real-time visibility, and never run random debug commands in production. A disciplined checklist turns BGP incidents from firefighting into mechanical fixes.

Three habits separate teams that resolve BGP incidents quickly from teams that do not. First, always log neighbour changes (bgp log-neighbor-changes) and ship those logs somewhere durable, because the exact second a session dropped is often the only clue correlating a routing outage with a change window. Second, capture the relevant show output before making any change — once a session resets, the Adj-RIB-In state that explained the problem is gone. Third, prefer soft reconfiguration or route refresh over a hard reset, since a hard reset on a full-table session causes a convergence event that can itself trigger the next incident.

Continue with securing BGP with ASPA, BGP TCP-AO on Arista EOS, and IOS-XR input drop troubleshooting. Two further guides pair well with this workflow: BGP session stuck in Idle or Active for the state machine in detail, and RPKI route origin validation when the prefix is present but rejected as invalid.

原文链接:https://www.noction.com/knowledge-base/bgp-network-troubleshooting