Juniper MC-LAG Best Practices: Configuration and Usage - 夜莺博客

Juniper MC-LAG Best Practices: Configuration and Usage

Multi-Chassis Link Aggregation (MC-LAG) on Juniper EX and QFX switches delivers active-active redundancy, but only when the design follows the vendor's usage notes. This article summarizes the official Juniper Networks MC-LAG best practices: keeping ARP and MAC tables synchronized with arp-l2-validate, avoiding DHCP server redundancy pitfalls on MC-LAG peers, choosing the right DHCP relay agent behavior, separating ICCP and ICL across different ports and FPCs, and understanding how the LACP system ID behaves in every ICCP failure scenario. It is the practical checklist you need before deploying MC-LAG in production.

MC-LAG building blocks

Four components do all the work, and confusing them is the source of most design errors:

  • MC-AE (multi-chassis aggregated Ethernet) — the LAG bundle presented to the downstream device. The downstream switch sees a single LACP neighbour with one system ID, even though the links terminate on two different chassis.
  • ICL (interchassis link) — the physical link, or more usually an aggregated Ethernet bundle, that carries data between the two peers when traffic arrives on the wrong chassis. This is a data-plane path.
  • ICCP (Inter-Chassis Control Protocol) — a control-plane session between the two peers’ Routing Engines that synchronises state and agrees on which chassis is active. It runs over TCP port 65231 and is authenticated.
  • Backup liveness — a secondary check that lets each peer determine whether the other is still alive even when ICCP has failed, preventing both from claiming the active role.

A common mistake is to size the ICL for average traffic. In a correctly designed MC-LAG, the ICL carries only traffic that ingresses on the standby chassis, but during a link failure or a peer restart it can momentarily carry the whole bundle. Size it for the failure case, not the steady state.

Configuration walkthrough

A minimal, production-shaped two-switch MC-LAG looks like this on the primary peer, with the mirror image on the backup. Note that the ICCP interface and the ICL are separate interfaces, on separate ports.

set interfaces ae0 aggregated-ether-options lacp active
set interfaces ae0 aggregated-ether-options lacp periodic fast
set interfaces ae0 aggregated-ether-options lacp system-id 00:11:22:33:44:55
set interfaces ae0 aggregated-ether-options mc-ae mc-ae-id 1
set interfaces ae0 aggregated-ether-options mc-ae redundancy-group 1
set interfaces ae0 aggregated-ether-options mc-ae chassis-id 0
set interfaces ae0 aggregated-ether-options mc-ae mode active-active
set interfaces ae0 aggregated-ether-options mc-ae status-control active
set interfaces ae0 aggregated-ether-options mc-ae init-delay-time 240
set interfaces ae0 aggregated-ether-options mc-ae preferred

set interfaces ae1 aggregated-ether-options lacp active
set interfaces ae1 aggregated-ether-options lacp periodic fast
set interfaces ae1 description ICL

set interfaces ge-0/0/10 unit 0 family inet address 10.255.0.1/30
set protocols iccp local-ip-address 10.255.0.1
set protocols iccp peer 10.255.0.2
set protocols iccp authentication-key changeme-iccp-key

Four statements carry the behaviour: chassis-id must differ between the two peers, status-control must be active on one and standby on the other, preferred marks the box that should win the role election, and init-delay-time keeps the bundle down for a fixed period after boot so that the peer has time to establish ICCP before LACP starts negotiating. Omitting the init delay is the classic cause of a data-path black hole during a Routing Engine switchover.

ARP and MAC Synchronization

ARP and MAC address tables normally stay synchronized in MC-LAG configurations, but can get out of sync under conditions such as link flapping. To keep them in sync while the network stabilizes, enable the arp-l2-validate statement on IRB interfaces. This option validates ARP and MAC table entries and automatically applies updates when they drift apart. Note that it may impact performance in scale configurations, so it is recommended as a workaround during incidents rather than for permanent operation.

set interfaces irb unit 10 family inet address 192.0.2.1/24 arp-l2-validate

A broader operational point: the MAC table that matters for MC-LAG is the one the hardware uses, not the one the control plane believes in. After any incident in which entries were out of sync, compare the synchronised entries on both peers rather than assuming the tables converged once traffic recovered.

DHCP Design Constraints and Best Practice

Configuring the DHCP server directly on MC-LAG peers is not supported when DHCP service redundancy is required. DHCP relay with Option 82 is also not supported together with MAC address synchronization - if DHCP relay is required, configure VRRP over IRB or RVI for Layer 3 functionality instead.

In an MC-LAG active-active environment, Juniper recommends using the bootp relay agent:

set forwarding-options helpers bootp server 192.0.2.10
set forwarding-options helpers bootp interface irb.10

This avoids stale session information issues that can arise when the router uses the extended DHCP relay agent (jdhcp) process. The underlying reason is that the extended relay agent maintains per-client session state tied to the chassis that received the request; because an MC-LAG bundle can deliver a client’s subsequent requests to the other chassis, that state can end up on the wrong box. The bootp helper is stateless, so it does not care which peer handles each request.

If you need DHCP relay with Option 82 for a security or accounting policy, the design must change: terminate the relay on a pair of routers that share Layer 3 via VRRP rather than on the MC-LAG IRB itself.

ICCP and ICL Design

Use separate ports and, where possible, different Flexible PIC Concentrators (FPCs) for the interchassis link (ICL) and the Inter-Chassis Control Protocol (ICCP) interfaces. An aggregated Ethernet interface is preferred for ICCP. Configure prefer-status-control-active together with the status-control standby configuration to prevent the LACP MC-AE system ID from reverting to the LACP default system ID during an ICCP failure. This option should only be used when ICCP will never go down unless the remote peer goes down.

set interfaces ae0 aggregated-ether-options mc-ae prefer-status-control-active

That last caveat matters. prefer-status-control-active tells the standby chassis to keep pretending it is the active side when it cannot reach its peer. If ICCP is lost for a reason that is not a peer failure — a failed ICL, a routing problem between the peers’ management addresses, an ACL that blocks TCP 65231 — both chassis will believe they are active and the LACP system ID will be stable while the two halves of the bundle behave inconsistently. This is the split-brain condition the backup liveness mechanism exists to detect.

Failure Handling Scenarios

The official usage notes document the ICCP failure behavior for QFX Series switches:

  • ICCP down, ICL down or up, backup liveness not configured: LACP system ID changes to the default value.
  • ICCP down, backup liveness Active: LACP system ID changes to default for both active and standby MC-AE interfaces.
  • ICCP down, backup liveness Inactive: no change in LACP system ID.
  • ICCP up, ICL down: LACP state moves to standby and the MUX state moves to the waiting state.

The consequence of a system ID change is that the downstream switch sees a different LACP partner and tears the bundle down, which is intentional — it is safer for the access switch to have no bundle than to forward into a half-broken one. The cost is a brief outage on that bundle, which is why the design goal is to never lose ICCP in the first place.

Backup Liveness Detection

Configure the master-only statement on the IP address of the management interface for backup liveness detection on both the primary and the backup MC-LAG peers, ensuring fast and deterministic failover behavior when ICCP connectivity is lost.

set interfaces em0 unit 0 family inet address 10.10.10.1/24 master-only

With master-only, the address is only present on the Routing Engine that currently owns the master role, so a successful ping to the peer’s address proves that the peer is alive and acting as master. That is a stronger and more deterministic signal than reachability of an address that exists on both engines at once. Set the liveness interval conservatively: a liveness check that is too aggressive can declare a busy but healthy peer dead and trigger an unnecessary system ID change.

LACP roles, system ID and first-hop design

Two behavioural rules shape the Layer 2 design around MC-LAG. First, the LACP system ID presented to the access switch is a configured value, not either chassis’s MAC address, which is what allows the downstream device to treat two physical switches as one neighbour. Override it explicitly with lacp system-id rather than letting it default, because an explicit value survives a chassis replacement without forcing the access switches to renegotiate. Second, because only one chassis is active for a given MC-AE, all ingress traffic for that bundle converges on one peer and any traffic destined to the other peer’s local ports traverses the ICL.

For Layer 3, the cleanest pattern is VRRP on an IRB or RVI, with the virtual IP owned by whichever chassis is active, so that the default gateway and the MC-AE active role follow the same failure domain. Avoid running the DHCP server locally on MC-LAG peers when redundancy is required, as noted above.

Operational verification commands

show iccp
show iccp statistics
show interfaces ae0 detail
show interfaces ae0 aggregated-ether-options
show lacp interfaces ae0
show lacp interfaces ae0 extensive
show mc-ae status
show mc-ae status-interface all
show lacp statistics interfaces ae0

show iccp shows whether the control session is up, which peer is active, and the liveness state — check it first in every MC-LAG incident. show mc-ae status prints the role each chassis believes it holds for each MC-AE ID, and the synchronisation state of the LACP information. If those two disagree about which side is active, treat it as a split brain and stop, rather than continuing to debug the data path.

Testing failure scenarios before production

  • Pull a member link from the bundle on the active chassis and confirm the LACP bundle stays up and traffic keeps flowing.
  • Shut the ICL and confirm the documented behaviour — LACP moves to standby, MUX to waiting — and that no loop forms.
  • Shut ICCP while backup liveness is configured and confirm the roles settle deterministically, with the LACP system ID behaving as documented for your platform and release.
  • Power-cycle the standby chassis during traffic and confirm the init delay prevents a transient black hole.
  • Fail a Routing Engine and confirm the bundle survives without the access switch seeing a system ID change.

Do this in a maintenance window on the actual hardware and software release you will deploy. The ICCP failure behaviour has changed subtly across Junos releases, so a table read from a different version’s documentation is not a substitute for a test.

Common misconfigurations

  • Same mc-ae-id and chassis-id on both peers — the peers will not form a coherent MC-AE.
  • Running ICCP across the ICL. Losing the ICL then takes out both the data path and the control path, which is the worst possible combination.
  • No init-delay-time. Boot-time LACP negotiation begins before ICCP is up.
  • Default LACP system ID. Any chassis swap forces every access switch to renegotiate.
  • DHCP relay with Option 82 on the IRB. Not supported together with MAC synchronisation.
  • Enabling arp-l2-validate permanently on a high-scale IRB. It is an incident workaround, not a design decision.

Where MC-LAG is heading

MC-LAG is mature and widely deployed, but it remains a proprietary control protocol carrying its own failure modes. EVPN with ESI-LAG multihoming removes the ICCP/ICL dependency by using BGP to signal multi-homing and a deterministic designated-forwarder election to prevent loops. If you are building a new campus or data-centre fabric rather than extending an existing MC-LAG estate, evaluate EVPN multihoming first — see our comparison of EVPN multihoming versus MLAG, ESI and DF election.

Related Reading

原文链接:https://www.juniper.net/documentation/us/en/software/junos/mc-lag/topics/concept/best-practices-usage-notes.html