Cisco BGP Troubleshooting: Common Issues and Fixes - 夜莺博客

Cisco BGP Troubleshooting: Common Issues and Fixes

BGP sessions that will not establish, prefixes that disappear, and neighbors that flap are the most common routing problems in service provider and enterprise networks. This Cisco-authored document condenses years of TAC experience into a practical troubleshooting flow: start with show ip bgp all summary, verify connectivity and configuration, then work through hold-timer, AFI/SAFI, next-hop, RIB and performance issues one by one. The commands and corrective actions below apply to Cisco IOS and IOS XE.

BGP Adjacency Down: First Checks

If a session is down, issue show ip bgp all summary to see the current state — IDLE or ACTIVE means the Finite State Machine has not reached Established. Then check:

  • No connectivity: verify with ping (loopback-to-loopback when peering over loopbacks), check show ip route peer_IP, layer 1 state, and any firewall or ACL blocking TCP 179.
  • Wrong AS: the log shows %BGP-3-NOTIFICATION: sent to neighbor ... 2/2 (peer in wrong AS) — correct the AS numbers.
  • Duplicate router ID: %BGP-3-NOTIFICATION ... 2/3 (BGP identifier wrong) — set unique router IDs manually with bgp router-id X.X.X.X.
  • Missing update-source: iBGP over loopbacks requires neighbor ip-address update-source interface-id.

Adjacency Bounces: Interface Flap and Hold Timer

If the neighbor continuously bounces, check for physical interface flaps with show interface and show logging, then verify that the hold timer (default 180s) is not expiring because of high CPU or packet loss. Debug with debug ip bgp and debug ip tcp transactions to see where the TCP session is being reset.

AFI/SAFI, Next-Hop and RIB Issues

Prefixes received but not installed usually indicate a next-hop reachability problem — the route's next hop must exist in the RIB or the BGP route stays hidden. Check show ip bgp for routes in the table, show ip bgp rib-failure for routes rejected by the RIB, and the best-path selection with show ip bgp bestpath.

High CPU and Slow Peer Handling

BGP scanner, router, I/O, open and event processes all consume CPU. Excessive packets in the BGP queue point at a slow peer — the slow-peer feature (e.g. neighbor X slow-peer detection threshold 120) removes the slow peer from the update path until it catches up. Memory issues are checked with show process memory and show ip bgp summary memory counters.

Reading the BGP Finite State Machine

Before running any debug command, interpret the state shown by show ip bgp summary. The BGP FSM has six states, and each one narrows the fault to a specific layer:

  • Idle. The router is not trying to peer at all. Either the neighbor is administratively shut down (neighbor x.x.x.x shutdown), the peer is unreachable because of a missing route, or a resource problem (full BGP table, memory exhaustion) prevented the session from being started.
  • Connect / Active. A TCP connection is being attempted. "Active" is misleading — it does not mean the session is working, it means the router is actively trying. Persistent Active with no movement points at a transport-layer failure: no route to the peer, an ACL or firewall dropping TCP 179, a mismatched update-source, or the remote peer not listening.
  • OpenSent. The TCP session is up and an OPEN message has been sent. Stuck here usually means the remote router is not answering, or authentication (MD5, TCP-AO) does not match.
  • OpenConfirm. Both OPEN messages have been exchanged; the router is waiting for the KEEPALIVE. A hold here often indicates an AS number mismatch, a bad BGP identifier (router ID), or an unsupported capability such as a missing address family.
  • Established. The session is up and updates are flowing. If routes are still missing at this point, the problem is policy, next-hop, or RIB-related rather than adjacency-related.

Look at the State/PfxRcd column: a number means Established and tells you how many prefixes were accepted from that peer. A state name means the session never came up. The Up/Down column tells you how long the current state has persisted, which distinguishes a one-off flap from a permanent condition.

Systematic Path for "No Connectivity"

When the state is Active or Idle, work the stack from the bottom up rather than guessing:

Router# show ip bgp summary
Router# show ip route 10.0.0.2
Router# show ip interface brief
Router# ping 10.0.0.2 source Loopback0
Router# show ip interface GigabitEthernet0/0 | include line protocol
Router# show access-lists | include 179

The logic is simple: BGP runs over TCP, so if the underlying IP path fails, BGP cannot succeed. Verify the peer address is in the routing table with the correct next hop, that the interface to that next hop is up/up, that a sourced ping (using the loopback you advertise as update-source) succeeds both ways, and that no ACL, zone-based firewall or upstream carrier filter blocks TCP port 179. Testing with telnet 10.0.0.2 179 from the router is a fast way to prove the transport is open; a refusal means the packet arrived and the peer is not listening, while a timeout means the packet never arrived.

Two configuration mistakes dominate this category. The first is a missing neighbor x.x.x.x update-source Loopback0 when peering over loopbacks — the session then forms from the physical address and collapses whenever a link changes. The second is an ebgp-multihop omission when the eBGP peer is more than one hop away; the default TTL of one kills the session.

Decoding NOTIFICATION Messages

BGP reports the reason a session failed in a NOTIFICATION message, and IOS logs it with the error code and subcode in parentheses. The most frequent ones and their fixes:

  • 2/2 (peer in wrong AS) — the remote's AS number does not match the local neighbor x remote-as. Correct the typo, or use neighbor x local-as for an intentional AS migration.
  • 2/3 (BGP identifier wrong) — duplicate router ID. Two routers are using the same BGP identifier, usually because both were left to derive it from an interface address that is now duplicated. Set explicit, unique bgp router-id values, ideally tied to loopback addresses.
  • 2/4 (bad peer AS) / 2/5 (bad BGP identifier) — capability or version mismatch during OPEN processing.
  • 4/0 (hold timer expired) — the peer stopped sending KEEPALIVEs within the negotiated hold time. Causes are highs CPU, interface flapping, severe congestion, or a one-way path failure.
  • 6/2 (administrative shutdown) — the remote peer was administratively shut. Not a fault, a change.
  • 6/4 (administrative reset) — the remote cleared the session deliberately, often a policy refresh.
  • 5/0 (connection rejected) / 2/8 (bad optional parameter) — often a firewall intervening mid-session, or a capability mismatch such as a missing address family.

Read the direction of the message. A NOTIFICATION sent to a neighbor means your router found the problem; a NOTIFICATION received from a neighbor means the remote router did. That single detail halves the search space.

Prefixes Received but Not Installed

When the session is Established and shows a plausible prefix count but the routes never reach the routing table, the fault lies in the last three stages: policy, next-hop, or RIB. Work through them in order:

Router# show ip bgp 192.0.2.0
Router# show ip bgp rib-failure
Router# show ip bgp neighbors 10.0.0.2 advertised-routes
Router# show ip bgp neighbors 10.0.0.2 routes
Router# show ip route 192.0.2.0
  • Next-hop unreachable. If the next hop of a received route is not resolvable in the RIB, the BGP route appears in show ip bgp but is flagged with no best path and never installed. On iBGP this is the classic "next hop was not carried across the AS" problem; fix it with neighbor x next-hop-self on the border routers, or by advertising the external next hop through the IGP.
  • RIB failure. show ip bgp rib-failure lists routes that lost to a more specific route already installed, or that were rejected for administrative distance reasons. For example, an OSPF internal route to the same prefix wins over iBGP by default (distance 110 versus 200).
  • Filtering. Check the inbound route-map, prefix-list and distribute-list on the neighbor. show ip bgp neighbors x routes shows what survived the inbound policy, while advertised-routes shows what you are sending out. A common surprise is an as-path access-list with an unanchored regex that accidentally matches the router's own AS.
  • Soft reconfiguration. Remember that a policy change needs clear ip bgp x soft in (with soft-reconfig or route-refresh) to take effect. Forgetting the refresh is a frequent false alarm.

Session Flapping and Route Damping

A session that comes up and goes down repeatedly is worse than one that stays down, because it injects churn into every downstream router. Diagnose the trigger before the symptom:

Router# show log | include %BGP
Router# show ip bgp summary | include Up/Down
Router# show interface GigabitEthernet0/0 | include flap
Router# show ip bgp dampening flap-statistics

If the cause is a genuinely unstable link, fix the link or add neighbor x fall-over bfd so BGP reacts to the real failure in milliseconds instead of waiting for the hold timer. If the cause is a flapping prefix rather than a flapping session — the classic case being a route learned from an unstable customer that keeps appearing and disappearing — then route flap damping is the tool. Damping assigns a penalty per withdrawal, suppresses the route above a threshold, and reuses it once the penalty decays. The default parameters are aggressive and widely regarded as harmful for large prefixes; if you deploy damping, tune the half-life, reuse and suppress values explicitly. See our write-up of damping penalty, half-life and reuse tuning before enabling the defaults.

High CPU, Slow Peers and Memory

Performance problems surface as slow convergence or as sessions that time out under load. The relevant processes are the scanner (walking the table periodically), router (best-path computation), I/O (parsing inbound updates) and the per-peering event processes.

Router# show process cpu sorted | include BGP
Router# show ip bgp summary | include neighbor|query
Router# show processes memory | include BGP

A large query count in the output of show ip bgp summary indicates that inbound updates are arriving faster than the router process can consume them, typically because one peer is sending a full table repeatedly. The slow-peer feature solves this directly: rather than letting one slow member stall the update group, the router detects it, removes it from the group and paces updates to it separately:

Router(config-router)# neighbor 10.0.0.2 slow-peer detection threshold 120
Router(config-router)# neighbor 10.0.0.2 slow-peer split-update-group dynamic

Memory issues are checked with show process memory and with the memory counters printed by show ip bgp summary. A steadily growing BGP memory footprint usually means an unstable peer is re-announcing the same prefixes, an unusually large number of unique paths is being retained, or the router is holding a full IPv4 and full IPv6 table without enough RAM. The remedy is prefix filtering — a max-prefix limit to stop the hemorrhage, and explicit inbound prefix-lists to ensure you only retain what you intend to use.

Debug Commands and Their Risks

Debug output is invaluable but dangerous: on a busy router, debug ip bgp can push CPU to 100 percent and turn a partial outage into a total one. Use these safeguards before enabling any debug:

  • Prefer show commands and clear ip bgp x soft to full debugs. show ip bgp neighbors 10.0.0.2 gives you counters, timers, capabilities and the last reset reason without touching the CPU.
  • Scope the debug to one neighbor where the syntax allows, and use an ACL with debug ip packet ... acl rather than debugging all packets.
  • Terminate output to a logging buffer or a syslog server, never to the console, and always have a second access path open. A debug storm with logging console enabled can lock out your only session.
  • Remember undebug all (u all) and keep it ready. Set a loose time limit on yourself.
Router# debug ip bgp 10.0.0.2 updates
Router# debug ip bgp 10.0.0.2 keepalives
Router# undebug all

Hardening the Sessions You Just Fixed

Once the adjacency is stable, make it harder to break by accident or attack:

  • neighbor x maximum-prefix 1000 80 restart 30 caps the damage an errant peer can do and warns before the limit is hit.
  • neighbor x password ... adds TCP MD5 signatures, or use TCP-AO on platforms that support it — a much stronger option that authenticates the whole TCP segment. Our BGP TCP-AO configuration guide walks through the key-chain side.
  • neighbor x ttl-security hops 1 makes eBGP spoofing from outside the directly connected segment far harder by requiring the TTL to match the expected hop count.
  • Filter what you accept and what you announce. Prefix-lists, AS-path access-lists and route-maps are the core tooling; see prefix-list and route-map filtering for BGP and BGP communities and as-path filtering for worked examples.
  • Validate origin with RPKI. Origin validation gives you a data-driven reason to reject a hijacked prefix, which is exactly the failure that route filtering alone cannot catch. See RPKI route origin validation deployment.

Quick Reference Checklist

Working from this list in order resolves the great majority of BGP tickets without a debug session:

  1. Session state and uptime: show ip bgp summary.
  2. Transport: ping the peer from the right source, test TCP 179, check ACLs.
  3. Configuration symmetries: AS numbers, router IDs, update-source, authentication, address families.
  4. Last reset reason: show ip bgp neighbors x.
  5. Log NOTIFICATION codes and map them to a cause.
  6. If Established, check inbound policy, next-hop reachability and rib-failure.
  7. If unstable, find the flapping link or prefix before touching timers.
  8. If slow, look at CPU, query counts and slow-peer detection.
  9. Then harden: max-prefix, authentication, TTL security, filtering, RPKI.

Also see our guides on BGP security with ASPA and BGP TCP-AO configuration on Arista EOS for additional BGP hardening and troubleshooting practice.

原文链接:https://www.cisco.com/c/en/us/support/docs/ip/border-gateway-protocol-bgp/218027-troubleshoot-border-gateway-protocol-bas.html