Arista EOS MLAG Configuration: Complete Guide - 夜莺博客

Arista EOS MLAG Configuration: Complete Guide

Multi-Chassis Link Aggregation (MLAG) is the technology that lets two Arista switches act as a single logical switch for dual-homed servers and switches, eliminating the 50% bandwidth penalty that Spanning Tree imposes on redundant topologies. This guide, drawn from the official Arista EOS user manual, explains how MLAG operates, the peer-link and MLAG system ID concepts, the configuration considerations that keep VLANs, LACP and STP consistent across peers, a complete working configuration example, and the exact show commands used to verify an active MLAG domain. It is an essential reference for any data center engineer running Arista EOS.

What Is MLAG and Why Use It?

Arista switches support MLAG to logically aggregate ports across two switches. Two 10-gigabit ports, one from each MLAG-configured switch, connect to a host and appear as a single 20-gigabit link. MLAG provides Layer 2 multipathing, higher bandwidth utilization, and active-active redundancy compared with traditional Spanning Tree-governed designs. It also interoperates with static LAG or LACP on the attached devices without proprietary protocols, so a server, firewall or load balancer that has no idea MLAG exists still sees one LACP neighbor with one system ID and one aggregated link.

The practical value is easy to state: in a classic three-tier design, the redundant uplink from an access switch to a second aggregation switch sits in a Spanning Tree blocking state most of the day, which means you paid for two links and forward on one. MLAG removes that idle link. Both peers forward, ECMP-style, and a failure of one peer or one peer-link member degrades capacity rather than the service.

MLAG Operation and Components

The two cooperating switches are called MLAG peer switches and communicate through an interface called the peer link. The peer link carries MLAG control information and any traffic from devices attached to only one peer. Each peer uses the peer address to form and maintain the link. The MLAG domain ID is a text string configured on each peer; the MLAG System ID (MSI) is the domain MAC address, derived automatically when MLAG forms, and is used in STP and LACP PDUs.

Three addresses matter and are easy to confuse:

  • Peer address - the IP address of the other switch's peer interface, used to build the MLAG control channel over the peer link.
  • Local interface - the SVI (typically VLAN 4094) on which each switch terminates that peer address.
  • MLAG System ID - a MAC address derived by the two peers once the control channel comes up. All MLAG port-channels present this single system ID to the attached device, which is exactly why a server's LACP bond across two different physical switches works at all.

Because the MSI is generated only when MLAG is healthy, a server's LACP bond torn across two peers that cannot talk to each other is a recipe for a split bond. That is the reason behind the dual-primary protections described later.

A Complete MLAG Configuration Example

The order of operations below matters. Bring up the peer link first, configure the MLAG control channel, and only then attach MLAG IDs to the server-facing port-channels.

switch1# configure terminal
switch1(config)# vlan 10,20,4094
switch1(config)# spanning-tree mode mstp
switch1(config)# interface Port-Channel1000
switch1(config-if-Po1000)# description MLAG-PEER-LINK
switch1(config-if-Po1000)# switchport mode trunk
switch1(config-if-Po1000)# switchport trunk allowed vlan 10,20,4094
switch1(config-if-Po1000)# no shutdown
switch1(config-if-Po1000)# exit
switch1(config)# interface Ethernet1-2
switch1(config-if-Et1-2)# channel-group 1000 mode active
switch1(config-if-Et1-2)# exit
switch1(config)# interface Vlan4094
switch1(config-if-Vl4094)# description MLAG-PEER-ADDRESS
switch1(config-if-Vl4094)# ip address 10.0.0.1/31
switch1(config-if-Vl4094)# no autostate
switch1(config-if-Vl4094)# exit
switch1(config)# mlag configuration
switch1(config-mlag)# domain-id DC1
switch1(config-mlag)# local-interface Vlan4094
switch1(config-mlag)# peer-address 10.0.0.2
switch1(config-mlag)# peer-link Port-Channel1000
switch1(config-mlag)# reload-delay mlag 300
switch1(config-mlag)# reload-delay non-mlag 330
switch1(config-mlag)# exit

Configure the second peer with the mirrored values: domain-id DC1, local-interface Vlan4094, ip address 10.0.0.2/31, peer-address 10.0.0.1 and the same peer-link port-channel number. The peer-link itself never carries an MLAG ID.

Now attach the server-facing port-channel, using the same MLAG ID on both switches:

switch1(config)# interface Port-Channel10
switch1(config-if-Po10)# description SERVER-01-BOND
switch1(config-if-Po10)# switchport mode trunk
switch1(config-if-Po10)# switchport trunk allowed vlan 10,20
switch1(config-if-Po10)# mlag 10
switch1(config-if-Po10)# exit
switch1(config)# interface Ethernet3
switch1(config-if-Et3)# channel-group 10 mode active

Repeat on switch2 with Port-Channel10, the identical switchport configuration and mlag 10, then place the second member link into channel-group 10 mode active. The two switches now present one LACP partner to the server.

Configuration Considerations

VLANs and LACP Consistency

VLAN parameters (access VLAN, switchport mode, trunk-allowed VLANs, native VLAN, trunk groups) must be configured identically on both peers for the peer-link and MLAG LAGs. LACP should be used on all MLAG interfaces including the peer link, since LACP control packets reference the MLAG system ID. Configuration discrepancies cause traffic loss in certain failure scenarios.

The subtle trap is that the switchport configuration on the MLAG port-channel is what the two peers must agree on, while the physical member interfaces inherit it. If you add VLAN 30 to Port-Channel10 on switch1 and forget switch2, traffic in VLAN 30 fails only when the server happens to hash that flow toward the switch that does not carry it - an intermittent, host-specific symptom that is hard to reason about from the server side.

Spanning Tree

STP must be configured globally and on port-channels with an MLAG ID. Port-specific STP settings (PortFast, BPDU Guard, BPDU filter) come from the switch where the port physically resides.

In show spanning-tree output, an MLAG port-channel appears a single time, and the remote peer's copy of it is displayed with a P (Peer) prefix so the topology does not look like a loop. A healthy MLAG domain therefore shows zero blocked ports that belong to the MLAG pair.

Control-Plane ACL Requirements

Any custom control-plane ACL applied to an MLAG port must include these rules, or MLAG will fail to establish:

permit tcp any any eq mlag ttl eq 255
permit udp any any eq mlag ttl eq 255
permit ip any any tracked

Do not remove permit ip any any tracked - its absence prevents MLAG from establishing. The tracked keyword keeps the ACL stateful for the MLSAG control flows that ride the peer link after the peer session comes up, and it is the single most common reason a tightly written control-plane policy silently breaks MLAG peering.

Verifying MLAG with Show Commands

switch1# show mlag
switch1# show mlag interfaces
switch1# show mlag interfaces detail
switch1# show mlag config-sanity
switch1# show spanning-tree vlan-id 3903
switch1# show spanning-tree blocked
switch1# show port-channel

Healthy peers show state: Active with peer-link status Up and MLAG ports in Active-full state. In STP output, MLAG interfaces appear as a single entry; remote interfaces are prefixed with P (Peer). A blocked-port count of zero confirms the MLAG domain created no topology loops.

switch1# show mlag
MLAG Configuration:
domain-id                          : DC1
local-interface                    : Vlan4094
peer-address                       : 10.0.0.2
peer-link                          : Port-Channel1000
heartbeat-interval                 : 4000 ms

MLAG Status:
state                              : Active
negotiation status                 : Connected
peer-link status                   : Up
local-int status                   : Up
system-id                          : 001c.73aa.bbcc
dual-primary detection             : Disabled

Read the output in this order: state tells you whether the domain is forwarding, negotiation status confirms the control channel is exchanging data with the peer, and system-id is the MAC address your server's LACP bond should be showing as its partner. If state is Active but a server bond still reports a single member link up, the problem is almost always on the server side or in the LACP timers, not in MLAG.

show mlag config-sanity is the fastest way to find the consistency mistakes described earlier: it compares the two peers' relevant configuration and lists any mismatch.

Failure Scenarios: Peer Link, Dual Primary and Orphan Ports

  • Peer-link member failure. With the peer link built as a port-channel, losing one member is invisible to MLAG. Losing every member breaks the control channel.
  • Dual primary (split brain). When the peer link fails but both switches stay powered, each may believe it owns the MLAG ports. Dual-primary detection, configured as dual-primary detection delay 100 action errdisable all-interfaces, shuts the MLAG interfaces on the designated loser instead of letting both forward.
  • Errdisabled ports after dual primary. Recover with errdisable recovery cause mlag-dual-primary and an interval, or clear them manually after the peer link is restored.
  • Orphan ports. A single-homed device attached to a non-MLAG port on one peer is reachable across the peer link, which is why the peer link must carry the same VLANs the orphan needs.

MLAG Maintenance and Split-Brain Protection

When a peer reboots, non-peer-link ports stay in errdisabled for the reload-delay period (300s on fixed switches, up to 1800s on modular platforms). Severing the peer-link cable can cause a split-brain state where each peer independently runs STP to prevent loops. MLAG ISSU lets you upgrade EOS on one peer with minimal disruption when versions are compatible.

The reload delay exists so that the rebooting switch does not come back with empty forwarding tables and black-hole traffic that the surviving peer was happily delivering. Before an EOS upgrade, confirm the peer is Active, that show mlag config-sanity is clean, and that the reload and non-MLAG delays are set to values larger than your control-plane convergence time.

Design Best Practices and Common Mistakes

  • Use a dedicated VLAN (4094 is conventional) for the peer address and set no autostate so the SVI never goes down because the last data VLAN disappeared.
  • Never reuse the peer-link port-channel number for MLAG IDs, and never assign an mlag ID to the peer link.
  • Keep the MLAG pair a two-device construct. Scaling a design by chaining MLAG pairs adds hops and failure domains; for fabric-wide multihoming, use EVPN ESI multihoming instead.
  • Mirror the peer-link VLAN list on both switches, including any transit VLAN used by orphaned devices.
  • Document the MLAG ID per port-channel; a mismatched ID is a silent misconfiguration that only appears during a peer failure.
  • Set explicit MSTP root priorities rather than letting the fabric elect a root by accident, and check show spanning-tree blocked after every change.

Frequently Asked Questions

Can MLAG peers be different switch models? They can be different port counts within the same EOS feature family, but keep both peers on the same EOS release so that the control-plane behaviour and ISUU steps match.

Do I need an IGP for the peer address? No. MLAG uses the peer address over the dedicated peer-link VLAN; it does not need to be advertised in the underlay routing protocol.

Is MLAG the same as vPC or MC-LAG? Functionally yes - all three deliver an active-active L2 multichassis bundle under a vendor-specific name. The Arista behaviour and verification commands are what this guide covers.

What happens if I delete the peer-link port-channel? The MLAG domain falls back to dual-primary handling; expect the configured action (usually errdisable on one peer) and recover the port-channel to restore normal operation.

Related Reading

For the ordered list of commands used to build a domain from scratch, see the Arista EOS MLAG Configuration: Complete Run Book. Day-to-day EOS syntax lives in the Arista EOS Configuration Cheat Sheet: Commands That Matter, and if you are scheduling traffic while you tune the fabric, the Arista EOS QoS Configuration: Class Maps to Strict Priority guide covers classification and scheduling. When a domain goes wrong after a cable change, start with Arista EOS MLAG Troubleshooting: Peer-Link and Dual-Primary.

原文链接:https://www.arista.com/en/um-eos/eos-multi-chassis-link-aggregation