Distributed Anycast Gateway and ARP Suppression in VXLAN - 夜莺博客

Distributed Anycast Gateway and ARP Suppression in VXLAN

Two features do most of the heavy lifting in a well-built VXLAN EVPN fabric: the distributed anycast gateway, which puts the default gateway on every leaf that hosts a subnet, and ARP suppression, which stops broadcast ARP from being flooded across the overlay. Deploy them together and the fabric behaves like a single large Layer 2 domain without the flooding, without HSRP/VRRP pinnings, and without the tromboning that comes from a centrally located gateway.

This article explains the mechanics, gives working configuration for NX-OS and other platforms, and lists the caveats that turn a clean design into an intermittent one.

The Problem With a Pinned Gateway

In a classic design, a subnet's default gateway lives on one pair of switches. Hosts on that subnet are configured with one gateway IP, and traffic between two hosts in the same subnet that happen to be attached to different leaves has to be bridged across the fabric or routed via the gateway pair. Add virtual machine mobility and the problem gets worse: when a VM moves to a different leaf, the ARP entry pointing at the old gateway is still valid but the path is now suboptimal, and a first-hop redundancy protocol has to be tuned carefully or the VM loses its gateway for a moment.

What the Distributed Anycast Gateway Changes

Instead of one gateway, every leaf that hosts the subnet has a gateway interface with the same IP address and the same virtual MAC address. The host's default gateway is always one hop away, wherever the host is. Because the gateway is local, a VM that migrates keeps a valid ARP cache entry - the MAC it learned for the gateway is still the correct MAC on the new leaf.

The properties that follow from this:

  • Traffic between two hosts in the same subnet is routed locally at the ingress leaf, not bridged across the fabric.
  • No first-hop redundancy protocol election is needed. HSRP and VRRP are not used for the anycast subnet.
  • The design is stateless with respect to gateway ownership, so there is no failover event to tune.
  • Subnet boundaries stop being tied to physical switch pairs, which is what makes rack-and-roll and VM mobility simple.

ARP Suppression

ARP suppression uses the EVPN control plane as the source of truth. Each VTEP keeps a cache of IP-to-MAC bindings learned locally from ARP and remotely from BGP EVPN MAC/IP advertisements. When a host ARPs for an address that is already in the cache, the ingress VTEP answers on behalf of the destination. The request never crosses the overlay.

The results: dramatically less broadcast traffic in the overlay, fewer unknown-unicast flood events, and lower ARP latency for known hosts. If an entry is not in the cache, the ARP request is flooded as normal - which is the important safety valve that keeps silent hosts working.

Configuration: NX-OS

! 1. global anycast gateway MAC - must be identical on every VTEP
fabric forwarding anycast-gateway-mac 2020.0000.00aa

! 2. the SVI becomes an anycast gateway, one per VLAN (L2 VNI)
vlan 43
  name TENANT-A-WEB
  vn-segment 30043
interface Vlan43
  no shutdown
  vrf member TENANT-A
  ip address 10.43.0.1/24
  fabric forwarding mode anycast-gateway
  no ip redirects

! 3. L3 VNI under the tenant VRF
vrf context TENANT-A
  vni 50001
  rd auto
  address-family ipv4 unicast
    route-target both auto
    route-target both auto evpn

! 4. ARP suppression on the L2 VNI (NVE interface)
interface nve1
  no shutdown
  source-interface loopback0
  host-reachability protocol bgp
  global suppress-arp
  member vni 30043
    suppress-arp
  member vni 50001
    mcast-group 239.1.1.1

ARP suppression needs hardware resources. On several Nexus platforms you must enlarge the ARP TCAM region, and that change requires a reload:

hardware access-list tcam region arp-ether size double-wide
copy running-config startup-config
# reload required for the TCAM change to take effect

Configuration: Other Platforms

The vocabulary differs but the model is identical. On ArubaOS-CX, the equivalent is a distributed L3 gateway configured on the SVI with the anycast gateway MAC set globally; on Dell OS10 the anycast gateway is applied per VLAN with a global virtual MAC. The EVPN half of the configuration - VNI-to-VLAN mapping, route targets, NVE/VXLAN tunnel source - is the same work described in this Dell OS10 VXLAN EVPN guide, and if you are still deciding between overlay encapsulation options, this Geneve vs VXLAN comparison covers the trade-offs.

Requirements That Are Not Optional

  1. Same virtual MAC everywhere. If one leaf uses a different anycast MAC, hosts that migrate see a gateway MAC change and flush their ARP cache - or worse, keep sending to a MAC that is no longer local.
  2. Anycast gateway on every VTEP hosting the VNI. Partial deployment is unsupported on the major platforms. Every VTEP that has the VNI must also have the gateway.
  3. Consistent ARP suppression policy fabric-wide. Mixed settings produce confusing asymmetry where some ARP requests are suppressed and others flooded.
  4. Symmetric IRB for routed traffic. The L3 VNI must be carried end to end, otherwise return traffic takes a different path than forward traffic and you get asymmetric routing with firewall state mismatch.
  5. MTU headroom in the underlay. The overlay adds encapsulation; if the underlay is not sized for it, large frames are dropped in ways that look like application faults. The arithmetic and the fragmentation behaviour are covered in this VXLAN MTU and jumbo frames guide.

Verification

! is the anycast MAC configured and used?
show fabric forwarding ip local-host-db
show ip arp suppression-cache detail
show ip arp summary

! is the SVI anycast?
show interface Vlan43 brief
show ip interface Vlan43

! is EVPN advertising the MAC/IP routes?
show bgp l2vpn evpn
show bgp l2vpn evpn route-type 2
show nve peers
show nve vni 30043

! on the host side - the gateway MAC must be the anycast MAC
arp -a | grep 10.43.0.1
ip neigh show 10.43.0.1

If ARP suppression is not working, check three things in order: is the destination IP actually in the local suppression cache, is the MAC/IP route present in the EVPN table from the remote VTEP, and does the TCAM still have room for the arp-ether region.

Failure Modes Worth Knowing

  • Duplicate address, split ownership. If two different subnets are configured with the same gateway IP on different leaves, hosts get unpredictable forwarding. Keep a strict IP plan and generate the SVI configuration from it.
  • Suppression hiding a silent host. A host that never sends an ARP (or traffic) after being provisioned never appears in the EVPN table, so its first inbound packet is flooded - or dropped if flooding is disabled. If you disable unknown-unicast flooding for security, you must know which hosts rely on it.
  • Anycast gateway plus MLAG on the same VLAN. Both mechanisms provide first-hop redundancy in overlapping ways. Choose one model per subnet rather than layering them. The interaction between multihoming mechanisms is dissected in this EVPN multihoming vs MLAG guide.
  • Asymmetric IRB creeping in. It usually appears when a leaf is added with the L2 VNI but not the L3 VNI. The symptom is one-way traffic for inter-subnet flows.

Deployed deliberately, the combination removes most of the reasons people give for keeping a spanning-tree-based data centre. The migration is mostly about discipline in the IP plan and the route-target scheme, not about the overlay itself.

原文链接:https://www.cisco.com/c/en/us/td/docs/dcn/nx-os/nexus9000/104x/configuration/vxlan/cisco-nexus-9000-series-nx-os-vxlan-configuration-guide-release-104x/m_configuring_vxlan_bgp_evpn.html