VXLAN BUM Traffic: Ingress Replication vs Multicast - 夜莺博客

VXLAN BUM Traffic: Ingress Replication vs Multicast

VXLAN solves the Layer 2 scaling problem by encapsulating Ethernet frames in UDP and shipping them over a routed underlay. Unicast traffic is easy: once the destination MAC is known, the source VTEP tunnels to exactly one remote VTEP. The hard part is broadcast, unknown-unicast and multicast — collectively BUM traffic — where the VTEP does not yet know who needs a copy. There are two ways to handle that, and the choice shapes the underlay design, the control plane and how the fabric scales. This article covers both, with the failure modes each introduces.

Why BUM Traffic Is the Hard Case

An overlay VTEP learns remote MAC addresses from the control plane (BGP EVPN) or from data-plane learning. Until an address is known, a frame destined for it cannot be tunneled to a specific VTEP, so it must be delivered to every VTEP that might host the destination. That is the definition of BUM delivery, and it applies equally to genuine broadcasts such as ARP requests, to multicast, and to the first packets of any flow whose destination is still unknown.

RFC 7348 defines the VXLAN header and the UDP destination port (default 4789) and leaves the BUM replication mechanism to the implementation, which is why two deployments can both be "VXLAN" and behave very differently.

Option 1 — Multicast Underlay

Each VXLAN VNI is mapped to a multicast group in the underlay. VTEPs join the group with PIM, and a BUM frame is sent once to the group address, with the replication happening in the network rather than on the VTEP.

VTEP-A  --(one copy to 239.1.1.1)-->  underlay multicast
                                          |
                     +--------------------+--------------------+
                     |                    |                    |
                  VTEP-B               VTEP-C               VTEP-D

Strengths: the source VTEP sends a single copy regardless of how many remote VTEPs exist, so VTEP CPU and egress bandwidth do not scale with the number of members. Multicast also performs best for one-to-many applications that legitimately need group delivery.

Weaknesses: you must run PIM in the underlay, including a rendezvous point or equivalent, and multicast must be operational before any overlay traffic works. Underlay multicast in a data centre is an operational burden many teams are actively trying to remove, and a PIM failure degrades the overlay in ways that are hard to diagnose from the overlay side.

Option 2 — Ingress Replication

Each VTEP holds a list of remote VTEPs for each VNI and, when it needs to send a BUM frame, unicasts a copy to every one of them. The replication is done by the source VTEP in the data plane.

VTEP-A  --copy-->  VTEP-B
        --copy-->  VTEP-C
        --copy-->  VTEP-D

Strengths: the underlay stays a plain unicast IP fabric. No PIM, no RP, no multicast state to operate. This is why ingress replication dominates modern EVPN-VXLAN deployments.

Weaknesses: the head-end VTEP's CPU and uplink bandwidth scale linearly with the number of remote VTEPs and the BUM rate. Very large fabrics with high broadcast rates eventually need to limit BUM, often with ARP suppression.

The Control Plane Makes the Difference

The inefficiency of ingress replication is mostly an artefact of data-plane learning. When the VTEP learns remote VTEPs dynamically, all it knows is who else is participating, so it must copy to everyone. BGP EVPN changes that calculation:

  • Type 2 MAC/IP advertisement routes teach every VTEP which MAC addresses live behind which remote VTEP, so the vast majority of frames are sent as unicast.
  • ARP suppression lets the ingress VTEP answer an ARP request locally from its EVPN cache, so the ARP broadcast never enters the replication list at all.
  • The result is that in a healthy EVPN fabric, BUM becomes a small fraction of traffic — which is what makes ingress replication practical at scale.

Choosing Between Them

Factor Multicast underlay Ingress replication
Underlay requirement PIM + RP Plain unicast routing
VTEP overhead One copy per BUM frame One copy per remote VTEP
Operational complexity Higher Lower
Scales with VTEP count Well Linearly, mitigated by EVPN + ARP suppression
Best fit Multicast-heavy workloads, large member counts Standard data centre EVPN-VXLAN

Verification and Failure Modes

# Which remote VTEPs are known for the VNI
show vxlan address-table
show bgp l2vpn evpn route type 2

# Is the multicast underlay actually joined
show ip mroute
show ip pim neighbor

# ARP suppression in action
show l2vpn evpn arp-cache
  • Broadcast storm inside the overlay. BUM amplification when the control plane is down and data-plane learning floods everything. Rate-limit BUM on the VTEP as insurance.
  • One-way traffic after a VTEP restart. The recovering VTEP has lost its remote VTEP list and floods until EVPN re-converges.
  • Silent black hole with a multicast underlay. Overlay is up, PIM is not — only BUM-dependent traffic fails, which is why a multicast design should be monitored on the PIM side separately from the overlay.

For the surrounding design, see EVPN-VXLAN Data Center Fabric: Design Guide 2026, the encapsulation comparison in Geneve vs VXLAN: Overlay Encapsulation Compared, and Distributed Anycast Gateway and ARP Suppression in VXLAN for the ARP-suppression layer that keeps BUM out of the replication list.

原文链接:https://datatracker.ietf.org/doc/html/rfc7348