BGP Confederations vs Route Reflectors: Choosing a Scale-Out - 夜莺博客

BGP Confederations vs Route Reflectors: Choosing a Scale-Out

Internal BGP requires every speaker to hear every route, and the naive implementation — full mesh — needs n × (n−1)/2 sessions. Ten routers means 45 sessions; fifty routers means 1,225; and each one needs its own configuration and its own failure modes. Two mechanisms exist to break that scaling wall: route reflectors, which selectively relax the split-horizon rule, and confederations, which split one large AS into sub-ASes that peer with each other. They solve the same problem in different ways, and the wrong choice shows up later as a design constraint you cannot undo.

The Full-Mesh Problem in Numbers

Speakers Full-mesh iBGP sessions
5 10
10 45
25 300
50 1,225
100 4,950

The problem is not just session count: every iBGP speaker must be configured with every peer, and the configuration grows quadratically with the fabric.

Route Reflectors: Relax Split Horizon Selectively

iBGP's split-horizon rule stops a router advertising a route learned from one iBGP peer to another iBGP peer, because full mesh was assumed. A route reflector is an iBGP speaker permitted to reflect routes within a cluster. Its internal peers split into client peers and non-client peers, and the reflector propagates routes between the two groups. Two optional non-transitive attributes — ORIGINATOR_ID and CLUSTER_LIST — prevent loops.

router bgp 65000
 bgp cluster-id 1.1.1.1
 neighbor 10.0.0.2 remote-as 65000
 neighbor 10.0.0.2 route-reflector-client
 neighbor 10.0.0.3 remote-as 65000
 neighbor 10.0.0.3 route-reflector-client
 neighbor 10.0.0.4 remote-as 65000        ! non-client

Clustering for redundancy is the standard approach: two (or more) reflectors in the same cluster with the same cluster-id and both configured with the same clients. Clients peer with both; if one reflector dies, the other still reflects. The CLUSTER_LIST attribute prevents a route reflected by both from looping back.

# Junos equivalent
set protocols bgp group iBGP type internal
set protocols bgp group iBGP local-address 10.0.0.1
set protocols bgp group iBGP cluster 1.1.1.1
set protocols bgp group iBGP neighbor 10.0.0.2
set protocols bgp group iBGP neighbor 10.0.0.3

Confederations: Split the AS

A confederation divides one AS into multiple member sub-ASes. Routers inside a sub-AS run iBGP with each other; sub-ASes peer with each other using a special confederation eBGP, and the whole structure appears to the outside world as a single AS number. The benefit is that you may need only partial mesh inside each sub-AS, and each sub-AS can have its own policy — which is also its cost, because policy must now be maintained in several places.

! Router in member sub-AS 65010, part of confederation 65000
router bgp 65010
 bgp confederation identifier 65000
 bgp confederation peers 65020
 neighbor 10.0.0.1 remote-as 65010        ! intra-sub-AS
 neighbor 10.0.10.2 remote-as 65020       ! confederation eBGP
 neighbor 192.0.2.9 remote-as 64500       ! true external peer

Inside a member sub-AS the rules are still iBGP rules, so a confederation usually also needs route reflectors inside each sub-AS. That is the first hint that the two techniques are complementary rather than competitive: in a large multi-region network you often end up with both.

Side by Side

Route reflector Confederation
Scaling mechanism Suppress split horizon for clients Partition the AS into sub-ASes
Attributes added ORIGINATOR_ID, CLUSTER_LIST AS_CONFED_SEQUENCE, AS_CONFED_SET
External AS path change None Confederation hops are stripped when leaving the confederation
Redundancy Cluster of reflectors with shared cluster-id Redundant confederation links between sub-ASes
Policy placement Centralised at the reflector Per sub-AS, can be delegated
Migration cost Low, incremental High; renumbering ASN and re-peering
Typical use Single-region scale-out, fast convergence Multi-region, or organisational boundaries within one AS

Choosing

Choose route reflectors for the vast majority of designs. They are simpler, well supported by every vendor, standardised in RFC 4456, and require no ASN changes. A pair of reflectors per region, clustered, with clients peering to both, scales to hundreds of speakers.

Choose a confederation when the AS itself needs to be partitioned — commonly because different regions are managed by different teams with different policies, or because a merger must combine two networks that already have iBGP designs. The confederation gives you a policy and failure boundary, at the price of a much more involved design and migration.

Design Rules Learned the Hard Way

  • Follow the IGP: reflectors must sit on the IGP topology's shortest path or you create suboptimal routing. Enable bgp next-hop-self deliberately, not by reflex.
  • Always run at least two reflectors per cluster, and make sure clients have sessions to both — a single reflector is a single point of failure for the entire control plane.
  • Do not enable additional-path selection casually: it is the fix for “only one path is visible through the reflector”, but it increases update volume and requires every participant to support it.
  • Set bgp cluster-id explicitly when you have more than one reflector in a cluster. Two reflectors without a shared cluster-id will not handle redundancy the way you expect.
  • Verify convergence by simulation, not by theory: show ip bgp summary, then break a session on purpose in a lab and measure the time to reconverge.

Related Reading

Deeper dives on the same topics from our archive:

原文链接:https://www.cisco.com/c/en/us/td/docs/ios-xml/ios/iproute_bgp/configuration/xe-16/irg-xe-16-book/configuring-internal-bgp-features.html