MC-LAG ICCP Failures: Liveness & LACP System ID - 夜莺博客

MC-LAG ICCP Failures: Liveness & LACP System ID

Two Juniper switches terminating one server LAG behave like a single switch only while the Inter-Chassis Control Protocol (ICCP) session between them is healthy. When that session drops, each peer has to decide on its own whether the partner is dead or merely unreachable, and that decision is expressed as a change of LACP system ID on the wire. This article walks through the documented ICCP failure matrix for MC-LAG, shows exactly which CLI output changes in each branch, and gives the timer and liveness settings that keep the downstream server stable through a peer failure.

What ICCP actually carries

ICCP runs over TCP between the peer addresses (normally loopbacks or IRB interfaces) and behaves as a transport for three client applications that Junos names in its output:

  • LACPD - exchanges the LACP state of the multichassis aggregated Ethernet (MC-AE) bundle, including the aggregate key and the system ID both peers present to the server.
  • MCSNOOPD - keeps snooping-derived state (ARP, ND, IGMP) in sync so both peers know the same set of hosts.
  • ESWD - synchronises bridging state: MAC entries and the forwarding state of MC-AE member links.

Because the shared identity a server sees is a control-plane artefact, an ICCP fault does not look like a link down to the server. The server keeps a working bundle while the two chassis quietly disagree about who should forward - the classic split-brain path to duplicate frames.

The ICCP failure matrix

Junos documents the outcome as a function of three inputs: ICCP connection state, inter-chassis link (ICL) state, and the backup liveness peer status.

ICCP ICL Backup liveness Action on MC-AE
Down Up or Down Not configured LACP system ID changes to the default value
Down Up or Down Active (peer alive) LACP system ID changes to default on both active and standby MC-AE
Down Up or Down Inactive (peer dead) No change - the survivor keeps the shared system ID
Up Down - LACP state set to standby; MUX state moves to waiting

The first row is the dangerous one: both peers give up the shared identity and revert to their own default system ID. If the two chassis were built from the same configuration template, those default system IDs can even be identical, and the server happily keeps a bundle that straddles two switches which both believe they are active.

Why the system ID change protects the server

LACP aggregates per physical link, not per port channel. A link joins the bundle only when the actor and partner identifiers and keys match, and the partner system ID (priority plus MAC) is part of that test. When one peer reverts to its default system ID and the other does not, the server sees two groups of links with different partner system IDs, marks the mismatching links as not aggregable, and drops them out of the bundle. That is the mechanism behind the statement that a system ID change "protects" the server link: it removes the ambiguity exactly where the server would otherwise have to choose between two contradictory switches.

Configuring liveness so the matrix is deterministic

Backup liveness detection gives the peers an out-of-band heartbeat over the management network that only runs while the ICCP TCP connection is operationally down. Without it, a peer cannot distinguish "ICCP lost" from "peer is dead", and the safe-but-disruptive first row of the matrix applies.

set protocols iccp local-ip-addr 10.0.0.1
set protocols iccp peer 10.0.0.2 session-establishment-hold-time 340
set protocols iccp peer 10.0.0.2 liveness-detection minimum-interval 8000
set protocols iccp peer 10.0.0.2 liveness-detection multiplier 3
set protocols iccp peer 10.0.0.2 liveness-detection backup-liveness-detection
set protocols iccp peer 10.0.0.2 backup-liveness-detection backup-peer-ip 172.16.20.2

Two rules from Juniper's usage notes matter more than the exact numbers. First, if ICCP is carried over an IRB interface, keep the liveness interval (a multihop BFD timer) at 8 seconds or more: a shorter interval mistakes the control-plane gap during graceful Routing Engine switchover (GRES) for a peer failure and forces an unnecessary system ID change. Second, put the liveness address on a master-only management IP on both Routing Engines so the heartbeat survives a GRES.

The MC-AE side has its own timers. init-delay-time holds LACP down for a configurable period after boot so the bundle does not come up before ICCP state is rebuilt, and Juniper's guidance is to keep the session-establishment hold time at least 100 seconds above the init delay (340 s hold with 240 s init delay is the documented example for QFX).

set interfaces ae0 aggregated-ether-options lacp active
set interfaces ae0 aggregated-ether-options mc-ae mc-ae-id 1
set interfaces ae0 aggregated-ether-options mc-ae chassis-id 0
set interfaces ae0 aggregated-ether-options mc-ae redundancy-group 1
set interfaces ae0 aggregated-ether-options mc-ae status-control standby
set interfaces ae0 aggregated-ether-options mc-ae prefer-status-control-active
set interfaces ae0 aggregated-ether-options mc-ae init-delay-time 120

mc-ae-id must be identical on both peers and chassis-id must differ. Adding prefer-status-control-active on top of status-control standby is the documented way to stop the system ID from reverting to default during an ICCP failure - but only apply it when ICCP cannot go down unless the remote peer is genuinely down.

Reading the failure in a live incident

Work the layers in order: control plane, then MC-AE role, then LACP neighbours, then the server.

user@qfx> show iccp
user@qfx> show iccp detail
user@qfx> show mc-ae status
user@qfx> show lacp interfaces ae0 detail | match "System"
user@qfx> show interfaces ae0 extensive | match "LACP|MC-AE"
user@qfx> show log messages | match "ICCP|mc-ae"

show iccp reports the TCP connection state, the BFD liveness state and the client applications (MCSNOOPD, LACPD, ESWD). If the session is down while the ICL is up, compare the LACP system IDs on both chassis: a mismatch is the fingerprint of this failure mode. If the two system IDs still match but the bundle is flapping, stop looking at MC-LAG and chase the physical or LACP-rate problem instead.

A useful operational habit is to alert on the three values separately - ICCP session state, ICL error counters, and LACP system ID equality between peers - because the bundle state itself stays up in both failure branches and therefore tells you nothing. This is the same pairing of control-plane and data-plane signals discussed in our ICL vs ICCP failure behaviour comparison, and the recovery ordering is covered in our backup liveness and split-brain guide.

Lab reproduction in five minutes

user@qfx> configure
user@qfx# deactivate protocols iccp
user@qfx# commit
user@qfx# run show iccp
user@qfx# run show lacp interfaces ae0 | match "System ID"

Record the system ID before and after, once with backup liveness configured and once without. If the values match while ICCP is up and diverge after the deactivation, your pair is behaving exactly as documented - and you now have a reference output to compare against the next time production surprises you. General MC-LAG design rules are collected in our MC-LAG best practices note.

原文链接:https://juniper.net/documentation/us/en/software/junos/mc-lag/topics/concept/best-practices-usage-notes.html