EVPN ESI Multihoming: All-Active vs Single-Active - 夜莺博客

EVPN ESI Multihoming: All-Active vs Single-Active

EVPN multihoming lets a server, switch or router connect to two or more Provider Edge routers with what looks like a single LAG, while the fabric treats those physical links as one Ethernet Segment. The mechanism is worth understanding in detail because the failure modes are subtle: a misconfigured ESI produces duplicated BUM traffic, and a mis-elected designated forwarder produces silent one-way black holes. This article walks through the ESI, the two EVPN route types that make multihoming work, the difference between all-active and single-active modes, and what to check when traffic takes the wrong path.

The Ethernet Segment and the ESI

An Ethernet Segment (ES) is the group of links through which a customer device attaches to more than one PE. Each ES gets a 10-octet Ethernet Segment Identifier. Every PE attached to that segment advertises the same ESI, and it is the ESI — not an IP address — that ties the PEs together into one redundancy group.

For any port that is part of a multihomed CE, you define an ESI and bind it to the interface. On an all-active deployment the connections form a link aggregation group, so the ESI rides on the bundle.

Route Type 4: Ethernet Segment Route and DF Election

Ethernet Segment routes exist so that PEs attached to the same segment can discover each other. Once an ESI is assigned, the PE advertises it with the ES-Import extended community as BGP route type 4. Any peer whose import community matches imports the route and the PEs auto-discover one another — no manual peer list required.

Those same type 4 routes drive Designated Forwarder election. DF election decides which PE is responsible for sending broadcast, unknown-unicast and multicast traffic toward the CE on a given segment, which is what prevents duplicate BUM delivery when several PEs are attached.

Route Type 1: Per-ES EAD Routes and Split Horizon

Per-ES A-D routes carry the ESI Label extended community, and that community does two jobs:

  • It signals whether the segment is configured all-active or single-active.
  • It carries the ESI label used for split-horizon filtering on the core side.

Split horizon is the mechanism that stops a frame received from a multihomed CE on one PE from being flooded back out toward the same ES by another PE. Getting the ESI label wrong is the classic cause of a loop between two PEs that are both attached to the same segment. Per-ES EAD routes are also what give fast convergence when the access-side ES fails.

All-Active vs Single-Active

All-Active  : every PE in the ES forwards traffic simultaneously.
              The access device load-balances across its links; the PEs
              load-balance toward remote PEs. Needs ESI label split-horizon.

Single-Active: exactly one PE (the DF) forwards for the segment; the rest
              stay non-DF and drop CE-sourced BUM traffic.

Single-active is the conservative choice: with one active forwarder, loop prevention is intrinsic and you avoid the split-horizon complexity. All-active is what most data centres actually deploy because it uses both uplinks and gives you real bandwidth, not just redundancy — at the cost of requiring correct ESI label handling.

Verification

show bgp l2vpn evpn route type 1
show bgp l2vpn evpn route type 4
show l2vpn evpn ethernet-segment
show l2vpn evpn ethernet-segment detail | include "ESI|DF|Topology"

The segment detail output is the fastest read: it prints the ESI value, the ES-Import RT derived from the ESI, and a Topology line showing MH, Single-active or MH, All-active alongside the operational state. If a PE shows the segment as non-operational, the other PEs will still advertise it and you get asymmetric forwarding.

Failure Modes Worth Pre-Testing

  • Duplicate BUM to the CE. DF election did not converge, or two PEs both believe they are DF. Check type 4 routes on every PE.
  • Flooding loop between PEs. Missing or mismatched ESI label — split horizon is not active.
  • Uplink down, no failover. The remote PE never learned the per-ES EAD route, so the MACs are still advertised as reachable only through the failed PE.

For background on the overlay itself, see EVPN-VXLAN Data Center Fabric: Design Guide 2026, the ESI-specific walk-through in Juniper EVPN ESI-LAG: ESI Types, LACP and Junos CLI, and our comparison of Arista EOS MLAG Configuration: Complete Guide based designs against standards-based ESI multihoming, which is the same trade-off between vendor-specific LAG control planes and BGP-signalled Ethernet Segments.

原文链接:https://www.cisco.com/c/en/us/td/docs/routers/ios-xe/mpls/mpls/m-ce-evpn-multihoming.html