EVPN MAC-VRF Validation on ACX7000 Family - 夜莺博客

EVPN MAC-VRF Validation on ACX7000 Family

By Suneesh Babu — Juniper TechPost (posted 2022-10-29)

JUNOS unified way of bringing up EVPN E-LAN using the mac-vrf instance type, supporting 6,000 instances on ACX7000 with 642,000 MAC scale.

Introduction

Layer-2 VPN technologies provide multipoint services over IP/MPLS networks. The latest addition is Ethernet VPN (EVPN), where MAC learning is done via the control plane — compared to its predecessor Virtual Private LAN Service (VPLS), where MACs were learned via the data plane. RFC 7432 describes BGP MPLS Based Ethernet VPNs with three service types: VLAN-Based, VLAN-Bundle, and VLAN-Aware-Bundle.

The instance type mac-vrf is a unified way to configure EVPN E-LAN across entire Juniper platforms. The ACX7100-32C supports EVPN-MPLS E-LAN functionality with instance type mac-vrf. In this test scenario, VLAN-Based and VLAN-Bundle service types were covered for scale with the 22.2R1 Junos-EVO build (VLAN-Aware Bundle support is on the roadmap). The same scale was tested on ACX7100-48L and ACX7509.

Note: despite sharing the ACX moniker, ACX7000 products are different from ACX500/710/1000/1100/2100/2200/4000/5000/5400/6000. They use different Packet Forwarding Engines (PFE) and support different feature sets and scales.

This article documents the validation exercise end to end: the topology, the base IGP/BGP configuration, the scale configuration generation, and the verification commands that confirm the MAC-VRF instances actually came up with the expected control-plane and data-plane state. The distinguishing property of EVPN in a Metro context is that MAC reachability is distributed by BGP rather than flooded through the data plane, so a single mis-typed route target produces a silent, partial outage in which some hosts are reachable and others are not — which is precisely why structured validation matters.

Why MAC-VRF Replaced the Older Instance Types

Historically, Junos offered several different configuration containers for Layer-2 services, and the choice depended on the platform. The mac-vrf instance type unifies that model:

  • One syntax across platforms — the same routing-instances <name> instance-type mac-vrf block works for EVPN-MPLS E-LAN on ACX, MX and PTX, so template-driven automation does not need per-platform branches.
  • Explicit service-type mapping — you declare whether the instance is VLAN-based, VLAN-bundle or VLAN-aware-bundle, which makes the intent readable in the configuration rather than implied by the interface list.
  • Route-target driven isolation — the VRF target communities define which MAC-VRF instances import each other's MAC/IP routes, and this is the only knob that matters for tenant separation.
  • Scale headroom — ACX7000 hardware is sized for thousands of instances and hundreds of thousands of MACs, which is what a Metro aggregation role demands.

Service Types in Practice

The three RFC 7432 service types differ in how VLAN tags map to EVPN instances, and the choice is a traffic-engineering decision as much as a configuration one:

  • VLAN-Based — one EVPN instance per VLAN, and the VLAN ID has local significance only. Semantically clean, but it consumes one instance (and one EVI, one route-target pair) per customer VLAN, which is exactly why the tested scale of 6,000 instances matters.
  • VLAN-Bundle — a single EVPN instance carrying multiple customer VLANs, with the VLAN tag preserved across the MPLS core (typically via a label per bundle). It cuts the number of instances and route-targets dramatically, at the cost of requiring consistent VLAN numbering across the participating PEs.
  • VLAN-Aware-Bundle — a single BGP EVPN instance for the bundle with an EVPN per-VLAN broadcast domain inside it, so VLAN IDs can be re-mapped at each PE. This is the most flexible option and the most operationally demanding; on the tested ACX7000 release it was still on the roadmap.

For a Metro access network carrying many small business services, VLAN-Bundle is usually the pragmatic middle ground: fewer instances to manage, fewer route targets to misconfigure, and a scale profile closer to the platform ceiling.

Metro EVPN Topology

The test topology consists of three routers:

  • PE1: ACX7100-32C
  • PE2: ACX7100-48L
  • P1: PTX10008

Scale numbers tested: 6,000 EVPN MAC-VRF instances, with 43 local and 43 remote hosts learned per VLAN-Based MAC-VRF instance, and 64 local and 64 remote hosts per VLAN-Bundle instance — totaling 642,000 hosts. Bidirectional traffic in iMIX mode at 99.9% offer-load flows for all EVPN services; ~1 Tbps of traffic transits through the DUT.

The topology is deliberately minimal — two PEs and one core router — because the purpose is not to model a full fabric but to isolate the DUT's control-plane and forwarding behaviour. P1 provides the MPLS transport with segment routing so that the PE-to-PE label stack is realistic, and hosts are emulated behind each PE so that local and remote MAC learning both occur at full scale. iMIX with 99.9% offer-load is the standard stress profile: it keeps every queue busy without the pathological behaviour of a pure small-packet test, and any micro-burst that the PFE cannot absorb shows up immediately as traffic loss.

Base Configuration

The IGP and BGP configuration required to bring up EVPN:

regress@PE1> show configuration protocols | display inheritance no-comments
bgp {
    group IBGP {
        type internal;
        local-address 12.1.1.1;
        family inet { unicast; }
        family evpn { signaling; }
        neighbor 12.1.1.3 { export nhself; }
    }
}
mpls {
    interface ae0.0;
    interface ae1.0;
}
ospf {
    source-packet-routing {
        node-segment ipv4-index 101;
        srgb start-label 45000 index-range 2000;
    }
    area 0.0.0.0 {
        interface lo0.0 { passive; }
        interface ae0.0 { interface-type p2p; }
        interface ae1.0 { interface-type p2p; }
    }
}

Three details in that block deserve attention. First, family evpn signaling must be enabled on the BGP session — without it, the session comes up, OSPF converges, and yet no MAC/IP routes are ever exchanged, which is the single most common cause of a "service is up but hosts cannot talk" case. Second, export nhself rewrites the BGP next-hop to the local address so that the remote PE resolves the EVPN routes through the IGP; exporting without it leaves the next-hop as an unreachable address and routes are held invalid. Third, segment routing with an SRGB of 45000 over an index range of 2000 gives the transport a deterministic label space, so the EVPN service labels layer cleanly on top.

Bringing Up a Single MAC-VRF Instance

Before generating thousands of instances, validate one. A minimal VLAN-based MAC-VRF on a PE looks like this:

set routing-instances EVPN-100 instance-type mac-vrf
set routing-instances EVPN-100 service-type vlan-based
set routing-instances EVPN-100 route-distinguisher 12.1.1.1:100
set routing-instances EVPN-100 vrf-target target:65000:100
set routing-instances EVPN-100 protocols evpn encapsulation mpls
set routing-instances EVPN-100 protocols evpn interface ae0.100

set interfaces ae0 unit 100 family bridge interface-mode access
set interfaces ae0 unit 100 family bridge vlan-id 100

The route-distinguisher must be unique per instance (typically router-id:service-id), while the vrf-target is what must match on the remote PE for the two instances to exchange MAC/IP routes. Getting this backwards — unique VRFs but a shared RD, or a matching RD with a mismatched target — is the classic EVPN configuration error, and it produces a service that exists on both sides but never converges.

Verification and Validation Commands

Control-plane validation is where an EVPN deployment is won or lost. The following sequence walks from the transport up to the individual MAC:

# 1. Is the EVPN BGP session established?
show bgp summary
show bgp neighbor 12.1.1.3 | match "EVPN|NLRI|State"

# 2. Are EVPN routes being received per instance?
show route table EVPN-100.evpn.0

# 3. Instance and interface state
show evpn instance EVPN-100
show bridge mac-table instance EVPN-100

# 4. Which MACs were learned locally vs remotely, and via which PE?
show ethernet-switching table instance EVPN-100 extensive

# 5. Full MAC scale check across all instances
show evpn instance summary

# 6. Data-plane check: label bindings and forwarding
show route forwarding-table family evpn
show mpls lsp

Read the output in order. If show bgp summary shows the EVPN NLRI count as zero, the problem is transport or family configuration. If routes are received but the instance shows no interfaces, the problem is the bridge/interface binding. If MACs appear in the table as remote but traffic does not flow, the problem is the MPLS label path or an MTU mismatch on the attachment circuit — EVPN carries the customer frame plus a label stack, so an attachment circuit with a 1500-byte MTU against a 1508-byte transport requirement will pass control plane validation and drop large frames.

Generating Scale Configuration

Scale configurations were generated with a Jinja2 template driven by a Python script (create_evpn_vrf.py) that reads a YAML definition and builds the routing instances, demonstrating a practical automation approach for large-scale Metro deployments.

The YAML definition typically lists the instance name, the service type, the VLAN (or bundle), the route-target and the attachment interface, and the template expands each entry into a full routing-instances block. Two safeguards are worth building into any generator of this kind: validate that every route-distinguisher and route-target is unique before committing (a duplicate RD silently breaks reachability on exactly one service), and emit the configuration as a single load replace-friendly file so the commit is atomic. For 6,000 instances, a commit that partially applies is far worse than a commit that fails cleanly.

Common Pitfalls

  • Missing family evpn signaling on one side of the iBGP mesh — the session is up, but no MAC/IP routes are exchanged.
  • Asymmetric route targets — one PE imports, the other does not, giving one-way reachability that hides behind working ARP in one direction.
  • Duplicate route-distinguishers between instances, which can be masked by the fact that the BGP session still looks healthy.
  • Attachment-circuit MTU that does not account for the MPLS label stack.
  • Assuming ACX500/1000/2000 scale numbers apply to ACX7000 — the PFEs and the supported feature sets are different, and a design validated on one family is not a design on the other.
  • Convergence timing — MAC scale of this size means that a full restart of all sessions takes measurable time to re-advertise; design BFD and graceful restart accordingly.

Conclusion

ACX7000 Family platforms (ACX7100-32C, ACX7100-48L, ACX7509) at this tested scale are ready for Metro deployments. The EVPN services capabilities are paramount for next-generation Metro solution requirements.

Related reading on this site: EVPN-VXLAN data center fabric design, Junos EVPN-VXLAN CRB fabric configuration, and Juniper EVPN ESI-LAG multihoming. For how EVPN multihoming compares with the older MC-LAG approach, see EVPN multihoming vs MLAG: ESI and DF election.

References

  • RFC 7432: BGP MPLS-Based Ethernet VPN — https://datatracker.ietf.org/doc/html/rfc7432
  • MAC-VRF: https://www.juniper.net/documentation/us/en/software/junos/evpn-vxlan/topics/concept/mac-vrf-routing-instance-overview.html

原文链接:https://community.juniper.net/blogs/suneesh-babu/2022/10/28/evpn-mac-vrf-validation-on-acx7000