BGP Next-Hop Self and Third-Party Next Hop - 夜莺博客

BGP Next-Hop Self and Third-Party Next Hop

A BGP route can be in the table, marked best, and still not be used. The next hop is unreachable, the recursive lookup fails, and the prefix never reaches the routing table. This is the classic iBGP next-hop problem, and the fix is one of three commands depending on whether you are talking to a route reflector client, a peer on the same subnet, or a peer that expects the original eBGP next hop to be preserved.

How the next hop gets set

  • eBGP: the advertising router rewrites the next hop to the local interface address facing the peer, so the receiver can reach it over a directly connected link.
  • iBGP: the next hop is preserved unchanged. A route learned from an external peer keeps that external next hop as it is propagated inside the AS.

That preservation is what breaks topologies. An internal router learns 203.0.113.0/24 with next hop 198.51.100.1 — an address reachable only from the router that has the eBGP session. Without a recursive path to 198.51.100.1, the prefix stays in the BGP table and never becomes a route.

show ip bgp 203.0.113.0
!  Next Hop 198.51.100.1, from 10.0.0.9 (10.0.0.9)
!     (metric 11) from 10.0.0.9 (10.0.0.9)
!       Origin IGP, localpref 100, valid, internal, best

show ip route 198.51.100.1
! % Network not in table          <- the smoking gun

The three fixes

! 1. rewrite the next hop on the way out
router bgp 65001
 address-family ipv4 unicast
  neighbor 10.0.0.9 next-hop-self

! 2. rewrite it for everything, including reflected eBGP routes
 neighbor 10.0.0.9 next-hop-self all

! 3. or use a route-map when only some prefixes need rewriting
 route-map SET-NH permit 10
  set ip next-hop 10.0.0.1
 neighbor 10.0.0.9 route-map SET-NH out

next-hop-self under the address family is the standard fix on an eBGP border router: it gives every iBGP peer a next hop they can actually reach, at the cost of forcing traffic through this router. next-hop-self all (IOS XE) additionally rewrites the next hop on routes that were learned from iBGP peers, which matters when you are reflecting full eBGP-learned routes and want a deterministic forwarding path.

The third-party next-hop rule

There is an exception that surprises engineers in both directions. When two BGP speakers are on the same IP subnet and the next hop is also on that subnet, the original next hop is advertised unchanged — the "third-party next hop" behaviour. This is why a route reflector can reflect routes between clients that share a segment without rewriting anything, and why the same design fails the moment a client sits one hop away.

! check whether the next hop is on a shared subnet with the peer
show ip bgp neighbors 10.0.0.9 | include "BGP neighbor is|External|Internal"
show ip route 198.51.100.0 255.255.255.0

next-hop-unchanged on eBGP

! Arista EOS / NX-OS / Junos style - preserve the next hop across an AS boundary
neighbor 172.16.1.1 next-hop-unchanged
! Junos
set protocols bgp group FABRIC neighbor 172.16.1.1 next-hop-unchanged

In EVPN and large fabrics, next-hop-unchanged is what lets a spine act as a pure route reflector while the leaves keep the original VTEP next hop — essential for symmetric IRB, where rewriting the next hop to the spine would send encapsulated traffic to the wrong VTEP. If you are debugging EVPN next-hop behaviour, the type-5 route layout described in EVPN type-5 IP prefix routes is the context that makes the symptom make sense.

Diagnostic sequence

  1. show ip bgp <prefix> — confirm the route is valid and see the next hop.
  2. show ip route <next-hop> — the recursive lookup. If this fails, you have found the problem.
  3. show ip bgp <prefix> on the advertising router — is it sending a rewritten next hop?
  4. If a route-map with set ip next-hop is in play, check the order of route-map entries: a deny entry above the permit can drop the attribute change silently.

Related: the BGP best-path algorithm explains why a prefix can be best while still unusable, and route reflector cluster ID configuration covers the reflector design where next-hop behaviour is most often a surprise.

原文链接:https://www.cisco.com/c/en/us/td/docs/ios-xml/ios/iproute_bgp/configuration/xe-16/irg-xe-16-book/bgp-next-hop-self.html