MC-LAG ICCP Backup Liveness Detection Config Guide - 夜莺博客

MC-LAG ICCP Backup Liveness Detection Config Guide

An MC-LAG pair looks like one switch to the downstream server only while the two peers keep talking over ICCP. When the ICCP path breaks but both chassis stay up, the pair enters a split-brain state and both peers fall back to their own LACP system ID — which makes the server drop links from whichever bundle it did not accept first. This guide covers the out-of-band safety net for that failure: ICCP backup liveness detection, how to configure it on both peers, the timers that matter, and the verification commands (show iccp, show lacp interfaces, show mc-ae status) that prove the system ID is still shared.

Why backup liveness detection exists

ICCP normally rides a single path. If only one physical link is available for ICCP, a link or FPC failure takes ICCP down while the peer is still reachable — the classic split-brain. Backup liveness detection adds a second, out-of-band channel over the management network so each peer can still confirm whether the other is alive. During a genuine split-brain both peers change their LACP system ID; with a healthy backup liveness signal the standby peer keeps the shared ID and the server sees no bundle churn.

Configuration on both peers

set interfaces ae1 aggregated-ether-options lacp active
set interfaces ae1 aggregated-ether-options lacp system-id 00:00:00:11:11:11
set interfaces ae1 aggregated-ether-options mc-ae mc-ae-id 1
set interfaces ae1 aggregated-ether-options mc-ae redundancy-group 100
set interfaces ae1 aggregated-ether-options mc-ae chassis-id 0
set interfaces ae1 aggregated-ether-options mc-ae mode active-active
set interfaces ae1 aggregated-ether-options mc-ae status-control active
set interfaces ae1 aggregated-ether-options mc-ae init-delay-time 15
set protocols iccp local-ip-addr 10.0.0.0
set protocols iccp peer 10.0.0.1 redundancy-group-id-list 100
set protocols iccp peer 10.0.0.1 liveness-detection minimum-receive-interval 300
set protocols iccp peer 10.0.0.1 liveness-detection transmit-interval minimum-interval 300
set protocols iccp peer 10.0.0.1 backup-liveness-detection backup-peer-ip 192.168.10.2

The peer IP used for ICCP and the IP in the multichassis protection configuration must be the same address, otherwise the standby peer never enters standby. Backup liveness detection must be configured on both peers, and on a router with dual Routing Engines the master-only statement is applied to the management IP so only the active RE answers the probe.

Timers and the system ID trap

Cisco-style fast BFD intervals are tempting, but if ICCP is carried over an IRB interface configure the liveness-detection interval to at least 8 seconds so graceful Routing Engine switchover keeps working. When ICCP is down and backup liveness is active, the standby peer keeps the configured LACP system ID; when backup liveness is inactive the system ID reverts to the default. The prefer-status-control-active statement alongside status-control standby stops that reversion — only use it if ICCP can never go down unless the remote peer itself is gone.

Verification

show iccp
show iccp statistics
show interfaces ae1 mc-ae
show lacp interfaces ae1
show mc-ae status
show log messages | match ICCP

Look for ICCP Connection: Up, matching LACP system IDs on both peers, and a MUX state that matches the intended active/standby role. If the two peers show different system IDs while both are reachable, backup liveness is either unconfigured on one side or being blocked in the management path — check the management VRF and any firewall between the two mgmt IPs.

Related reading: MC-LAG ICCP failure scenarios playbook, LACP hashing and load balancing, Junos commit confirmed safety net.

原文链接:https://www.juniper.net/documentation/us/en/software/junos/mc-lag/topics/concept/best-practices-usage-notes.html