EVPN-VXLAN Data Center Fabric: Design Guide 2026 - 夜莺博客

EVPN-VXLAN Data Center Fabric: Design Guide 2026

EVPN-VXLAN has become the standard for enterprise and hyperscaler data centers, replacing spanning-tree-based designs and aging DCI technologies like OTV, which reaches end-of-support in April 2026. This guide covers the full stack of modern fabric design: leaf-spine Clos topology with an eBGP underlay, the critical choice between symmetric and asymmetric IRB, native multi-tenancy through VRFs and route targets, multi-site DCI, and the interoperability gotchas when mixing Arista, Cisco, Juniper and Nokia. Use it as a decision framework before you touch a single CLI.

The reason EVPN-VXLAN won is not that it is elegant. It is that it removes the two constraints that defined the previous twenty years of data center networking: the requirement to place hosts within one Layer 2 failure domain, and the requirement to run spanning tree to keep that domain loop-free. EVPN gives you Layer 2 semantics with Layer 3 behavior — any host can be reached at any rack, any subnet can span the fabric, and no link is ever blocked by STP. The cost is a control plane that you now have to design deliberately rather than inherit from defaults.

Leaf-Spine Fabric Design Principles

Every EVPN-VXLAN fabric is built on a Clos topology: every leaf connects to every spine, with no leaf-to-leaf or spine-to-spine links. ECMP provides equal-cost paths between any two leaves — typically 2-16 spines. Production sizing ranges from small fabrics (4-16 leaves, 2 spines) to hyperscale (256+ leaves, 16-32 spines). The underlay is almost always eBGP IPv4 per RFC 7938 because it scales beyond any IGP and terminates ECMP at the leaf.

Three properties of the Clos topology are worth internalizing. First, oversubscription is a property of the leaf, not the fabric: a leaf with 48 downlinks at 25G and 8 uplinks at 100G delivers a 1.5:1 ratio, and that ratio holds no matter how many leaves you add. Second, path diversity is a property of the spine count: with 4 spines and 4 uplinks per leaf, you can lose a spine and keep 3/4 of your upward bandwidth. Third, the fabric is only as deterministic as its hashing — with 4 equal-cost paths, a poorly tuned ECMP hash will still send most elephant flows down one uplink, which is why entropy-based hashing (Layer 3 + Layer 4, or better, on RoCEv2 fabrics) is not optional.

Underlay: eBGP, IGPs and Why RFC 7938 Wins

The underlay has one job: provide loopback-to-loopback reachability between every VTEP with good ECMP. Everything else is over-engineering. An IGP such as OSPF or IS-IS can do the job, and IS-IS in particular scales well, but the operational profile of eBGP is hard to beat for a fabric where the topology is a fixed two-stage Clos.

In an RFC 7938 underlay, each leaf and each spine gets its own autonomous system number, peers are established point-to-point on routed interfaces, and a single network or redistribute connected statement advertises the loopback. Because neighbors are directly connected, there is no IGP flooding domain, no designated router election and no link-state database to converge — a failure is a BGP withdrawal that propagates in milliseconds.

router bgp 65001
  router-id 10.0.0.11
  maximum-paths 4
  neighbor 10.1.1.0 remote-as 65201
    address-family ipv4 unicast
      send-community extended
  address-family ipv4 unicast
    network 10.0.0.11/32

The general rule is to keep the underlay in the default VRF, keep it IPv4-only, and resist the temptation to carry any tenant information in it. The underlay carries only loopbacks and point-to-point link subnets; tenants live entirely in the overlay.

Overlay Control Plane Options

Once the underlay is up, the overlay needs a control plane to distribute EVPN routes. Three options dominate:

  • iBGP with route reflectors on the spines — one ASN fabric-wide; spines become route reflectors for the VPNv4 and L2VPN EVPN address families. This is the classic Cisco NX-OS and Arista EOS reference design: simple policy, but the reflectors are a shared control-plane dependency.
  • eBGP leaf-to-spine with next-hop-unchanged — every leaf in its own AS, spines transparent. Scales cleanly and removes reflector failure modes, at the cost of more configuration per neighbor.
  • Controller-based (e.g. a fabric controller managing VXLAN tunnels) — less common in greenfield EVPN deployments now that BGP EVPN is universally supported.

For a fabric of any real size, the eBGP overlay option tends to age better. Route reflectors become a scaling concern once the Type-2 table is large, and a reflector outage is a fabric-wide event. With per-leaf ASNs, the failure of any single spine only removes the routes that spine was reflecting for — which in a fully meshed leaf-spine fabric is nothing, because the remaining spines carry the same information.

Addressing Plan and Loopback Design

The addressing plan is where most fabric migrations quietly go wrong. Three address pools need to be reserved up front:

  • Loopback pool — one /32 per switch, used as the VTEP source and BGP router-id. This must be stable for the life of the fabric; changing a VTEP loopback means re-establishing every tunnel.
  • Point-to-point link pool — /31 per leaf-spine link. Reserve enough that you can double the spine count without renumbering.
  • Tenant subnets — allocated per tenant VRF, with room to grow, ideally aligned so that summarization is possible at the border.

A VTEP source interface that flaps or changes address silently tears down the overlay; keep loopbacks on dedicated logical interfaces, never on a physical uplink, and put them in the default VRF.

EVPN Route Types and What They Do

EVPN carries several route types, but a fabric design really only depends on four:

  • Type-2 (MAC/IP Advertisement) — advertises individual host MACs, with optional IP and L3 VNI. Type-2 with IP is what permits ARP suppression and host-route distribution.
  • Type-3 (Inclusive Multicast Ethernet Tag) — builds the BUM replication list, telling every VTEP which peers participate in each L2 VNI.
  • Type-5 (IP Prefix) — advertises IP prefixes — used for border advertising and summarization, and the mechanism by which a fabric announces a tenant subnet as a single prefix rather than thousands of host routes.
  • Type-1/Type-4 — multihoming: auto-discovery and designated-forwarder election for ESI-based designs.

The design decision hiding in this list is whether you let Type-2 routes carry IP information. If you do, hosts get /32 routes distributed fabric-wide, which is precise but scales as O(hosts). If you summarize at the border with Type-5, you trade some optimality for a table that stays small. Most enterprises do both: Type-2 within the fabric for precision, Type-5 at the border for scale.

Asymmetric vs Symmetric IRB

Asymmetric IRB runs an SVI for every tenant VLAN on every leaf and routes on the ingress leaf using only EVPN Type-2 routes — simpler, but it does not scale past roughly 10-15 tenants. Symmetric IRB adds a dedicated L3 VNI per tenant; ingress and egress leaves each do their half of the bridging while routing happens across the L3 VNI, using both Type-2 and Type-5 routes. Rule of thumb: symmetric IRB for most new designs, asymmetric only for small, simple environments.

The distinction is easier to see at the packet level. With asymmetric IRB, the ingress leaf receives a frame on VLAN 10, routes it into VLAN 20, and transmits a frame tagged for VLAN 20 into the VXLAN tunnel using VLAN 20's L2 VNI. That means every leaf must host an SVI for both VLAN 10 and VLAN 20, so the number of SVIs per leaf equals the number of subnets in the fabric — fine at 20 subnets, unworkable at 2,000.

With symmetric IRB, the ingress leaf routes the frame into the tenant's L3 VNI, which is a single VNI per VRF. The packet crosses the fabric with the L3 VNI in its header, and the egress leaf does only the Layer 2 delivery on the destination VLAN. The ingress leaf needs an SVI only for the VLANs it actually hosts, and the L3 VNI count equals the VRF count — a number that stays in the tens even in large enterprises.

Symmetric IRB (recommended):
  Tenant subnets: 2,000      ->  SVIs per leaf: only locally hosted VLANs
  VRF count: 26              ->  L3 VNIs: 26
  Fabric-wide table: Type-2 host routes + Type-5 prefixes

Asymmetric IRB (legacy/small):
  Tenant subnets: 2,000      ->  SVIs per leaf: 2,000 (impractical)
  VRF count: 26              ->  L3 VNIs: none
  Fabric-wide table: Type-2 host routes only

There is one real advantage to asymmetric IRB: it never requires an L3 VNI, so it works in designs where the platform cannot map a VNI to a VRF. That is a shrinking constraint. Symmetric IRB should be your default.

Multi-Tenancy: The Killer Feature

Isolation is enforced in the control plane via VRF-to-L3-VNI mapping and route targets:

Tenant A:
  Route-Target: 100:100
  L3 VNI: 50100
  VLANs: VNI 10101, VNI 10102

Tenant B:
  Route-Target: 200:200
  L3 VNI: 50200
  VLANs: VNI 10201, VNI 10202

There is no way for Tenant A broadcast to reach Tenant B switches — compare that with VLAN-based isolation where one misconfigured trunk breaks the boundary instantly.

The thing that makes this genuinely different from VLAN-based isolation is that the boundary is enforced by BGP import policy rather than by forwarding-table separation. A leaf that does not import route target 100:100 will not install Tenant A's routes at all — not in a separate table, not anywhere. Isolation becomes a property of the control plane, so a misconfigured access port can at worst drop traffic; it cannot silently bridge two tenants together.

Route target design is worth thinking through before deployment. You can use a simple ASN:VNI scheme, which is easy to reason about, or a structured scheme such as ASN:tenant-id that survives VNI renumbering. Whichever you choose, document the mapping, because route-target mismatches are the single most common EVPN outage.

Multihoming: MLAG vs ESI-LAG

A server with two uplinks to two different leaves is the standard access design, and EVPN offers two ways to handle it. MLAG (Arista) or vPC (Cisco) treats the peer link as a synchronization channel and presents the two leaves as a single logical switch to the server. ESI-LAG uses EVPN multihoming: both leaves advertise the same Ethernet Segment Identifier, and one is elected designated forwarder per segment via a Type-4 route.

ESI-LAG is the cleaner long-term design because it does not require a peer link carrying production traffic, and because the failure behavior is defined by the control plane rather than by a proprietary peer-link protocol. MLAG remains extremely common and well understood, and it is not going away — but if you are building a greenfield fabric, ESI-LAG is worth the extra design effort.

MTU, Fragmentation and the 50-Byte Tax

VXLAN encapsulation adds 50 bytes of overhead. That means a 1500-byte inner frame becomes 1550 bytes on the wire, which does not fit in a 1500-byte underlay MTU. The correct fix is to raise the underlay MTU and leave the tenant-facing MTU at 1500, so that hosts never see fragmentation.

Leaf-facing access ports :  MTU 1500  (what the server sees)
Leaf uplinks and spines  :  MTU 9216  (jumbo, carries the encap)
NVE / system-level MTU   :  must be >= 1550, in practice 9216

The failure mode is subtle: small packets (pings, handshakes, DNS) succeed, and large transfers hang. If you see connectivity where a ping works and an SSH session stalls, suspect MTU before you suspect routing.

Multi-Site DCI and Interoperability

EVPN Multi-Site lets each data center run its own fabric independently, with border gateways re-originating EVPN routes and BUM suppression preventing cross-site flooding. In practice, the interoperability gotchas are almost always route-target auto-derivation (each vendor picks different defaults), ESI handling in multi-homing, and BUM replication tree selection (ingress replication vs multicast). For AI clusters, add PFC and ECN for lossless RoCEv2 plus entropy-based ECMP hashing.

The multi-site architecture is worth understanding in terms of what does not cross the DCI. The underlay stays entirely site-local; only the EVPN overlay is extended, and only between border gateways. Each site keeps its own spine layer, its own BGP topology and its own failure domain. A spine that dies in site A is invisible to site B. That independence is the whole point — it is what lets you run different hardware generations and even different vendors in each site without forcing convergence.

Interoperability, on the other hand, is where multi-vendor fabrics genuinely cost you. The specific items to test before rollout:

  • Route-target auto-derivation — Cisco, Arista, Juniper and Nokia do not all compute the same default RT from a VNI. Configure RTs explicitly in multi-vendor fabrics; never rely on auto.
  • ESI format and DF election — ESI types and the designated-forwarder algorithm must align (RFC 7432 vs the newer preference-based election).
  • BUM replication — verify that a leaf in one vendor's domain correctly builds a Type-3 list from another vendor's advertisements.
  • ARP suppression interaction — confirm both sides populate the cache from each other's Type-2 routes; a one-sided failure produces intermittent reachability.
  • MTU conventions — some platforms default to a different system MTU; check both.

Migration Strategy from a Legacy Three-Tier Network

Very few fabrics are built greenfield. The usual approach is to stand up the EVPN fabric alongside the existing three-tier network and migrate workload by workload, using a pair of border leaves as the interconnect during the transition. Key sequencing:

  1. Build and validate the underlay with a single test VLAN and no production traffic.
  2. Establish EVPN between the border leaves and the legacy core, advertising only the test subnet.
  3. Migrate one non-critical rack, verify ARP suppression, host mobility and ECMP behavior.
  4. Extend the underlay and overlay RIB, then move workloads in batches sized to your rollback window.
  5. Retire the legacy core only after the last subnet has moved and the border has been carrying production for a full change cycle.

The temptation is to migrate the fabric infrastructure faster than the workloads. Resist it — the fabric can be built and tested with almost no risk, and the risky part is always the cutover of live flows.

Failure Domains and Convergence

A well-designed fabric has no shared fate between the two forwarding planes. Losing a leaf removes its attached hosts and their routes; losing a spine removes one of N equal-cost paths. Neither should be visible to end hosts. The things that do create shared fate are worth listing explicitly:

  • Route reflectors, if you use iBGP and both reflectors are on the same pair of spines.
  • Multi-site border gateways, if the DCI is not dual-homed into each site.
  • DHCP or DNS servers attached to a single leaf.
  • The NVE source loopback, which every tunnel depends on.

Convergence targets to design toward: sub-second for link failure via BFD on the underlay, sub-5-second for leaf failure via BGP withdrawal, and under a minute for a full fabric resync after a control-plane restart.

Continue with EVPN MAC-VRF validation on ACX7000, the Arista MLAG run book, and SONiC troubleshooting.

For related design context on this site, see the EVPN multihoming vs MLAG comparison.

原文链接:https://netpilot.io/blog/evpn-vxlan-data-center-guide