MC-LAG ICCP Failure: ARP, PIM and ICL Behaviour - 夜莺博客

MC-LAG ICCP Failure: ARP, PIM and ICL Behaviour

Ask an engineer what happens when the ICCP link between two MC-LAG peers drops and most will answer "the backup peer changes its LACP system ID". That is the headline behaviour, but it is not the whole story. Junos also stops synchronising ARP and MAC state, changes how PIM neighbours are treated, and shifts which switch owns the Inter-Chassis Link for forwarding. If you only watch LACP, you can miss the reason traffic black-holes for a few hundred milliseconds - or worse, why it keeps looping. This article covers the side effects that the MC-LAG documentation describes outside the LACP section, and shows how to prove each one from a packet capture.

Recap: What MC-LAG Builds

  • ICCP: the control channel between the two peers, normally running over the ICL or over a routed path with BFD for fast failure detection.
  • ICL (Inter-Chassis Link): the forwarding path that carries traffic when one peer has no direct path to the destination - for example when a downstream MC-AE has one member link down.
  • MC-AE: the aggregated link seen by the access device, which must look like a single LAG to the server or switch below it.

ICCP itself is a TCP session between two peer addresses, and it is reachability-based rather than link-based: it lives on top of the IP route to the peer, which is why a routing change between loopbacks can take ICCP down without a single interface going down. On top of the session, the peers synchronise several independent state tables: MAC learning, ARP and neighbour discovery, IGMP, and the LACP identity and state of every MC-AE interface. Each of those tables has its own failure behaviour, and that is what the rest of this article is about.

Effect 1: LACP System ID and Link State

show mc-ae status
show lacp interfaces ae0
show iccp

The documented behaviour: if ICCP goes down, the backup MC-LAG peer changes the LACP system ID and brings down its MC-AE interfaces, so the downstream device sees a single consistent aggregator. Whether that is the right outcome depends on liveness: if the remote peer is genuinely dead, taking the local links down is correct; if it is alive but unreachable, you have just cut half your bandwidth. Note also the case where ICCP is up but the ICL is not: the standby peer's MC-AE moves to standby and its LACP multiplexer state goes to waiting, while the active peer keeps forwarding. No system ID change occurs in that scenario, which is why an ICL failure and an ICCP failure produce visibly different LACP output.

Effect 2: ARP, ND and MAC Synchronisation Stops

Peers continuously exchange ARP/ND and MAC information so either chassis can forward for the other. When ICCP fails, that synchronisation stops immediately, and entries learned by the peer are aged out locally. The practical consequence is a short window of unknown-unicast flooding and, on an MC-LAG that also serves as a first-hop gateway, a burst of ARP requests. Watch for it here:

show arp no-resolve | count
show ethernet-switching table | count
monitor traffic interface ae0 no-resolve     ! capture ARP if a black hole is suspected

Break the effect down and it has three distinct parts. First, entries that existed only because the peer learned them are no longer refreshed, so they age out on the normal ageing timer rather than being explicitly withdrawn. Second, traffic for a destination whose entry has just aged out is flooded within the VLAN until a reply repopulates the table - the classic brief unknown-unicast storm after an ICCP event. Third, if the MC-AE is also the default gateway for the attached hosts, the loss of synchronised ARP state means each host that needs to reach the gateway re-ARPes, and the two chassis answer from whichever is still active.

The ageing behaviour is the part that surprises people: because the withdrawal is implicit, a large MAC table with a long ageing timer produces a longer period of undifferentiated flooding than operators expect from a "fast" control-plane failure. Measure it rather than assuming it: take a count of the ARP and MAC tables before you break ICCP in the lab, then again at 10, 30 and 60 seconds after, and note when the counts plateau.

Effect 3: PIM and Multicast Adjacency

Juniper documents that PIM neighbour state is not torn down purely because the ICCP connection is lost - the multicast adjacency can remain, which is deliberately different from the unicast link-state behaviour. That asymmetry matters in multicast-heavy fabrics: a router can still be listed as a PIM neighbour while its unicast path has moved, so always verify the unicast RPF path rather than trusting the neighbour table.

show pim neighbors
show multicast route extensive
show route 239.1.1.1 detail

There are two independent consequences. The first is a stale adjacency: the neighbour entry stays, so a unicast RPF check against the unicast routing table can fail and silently drop multicast streams that appear, from the neighbour table, to be perfectly healthy. The second is BUM replication cost. With ICCP down, the two chassis no longer coordinate their multicast state, so the same group can be replicated twice toward the access segment. On an MC-LAG that carries IP multicast video, that duplication is visible as an effective doubling of the BUM rate on the MC-AE, and it is one of the reasons ICCP failure remediation should be treated as urgent even when unicast traffic looks unaffected.

IGMP snooping state is synchronised over the same mechanism as MAC and ARP, so a loss of ICCP means the snooping databases of the two chassis drift apart. Groups joined before the failure tend to stay on both; groups joined during the failure exist on only one. After ICCP recovers, the two databases are not automatically reconciled in a way that repairs a wrong forwarding decision, so a flood-based recovery is the practical remedy - clear the snooping entries and let joins rebuild if you suspect drift.

Effect 4: ICL Becomes the Only Path

With ICCP down, the local chassis may still have to forward traffic destined for links on the other chassis. Everything that would have crossed the ICL now crosses it twice as often, so the ICL must be sized for failure, not for steady state - the common rule is to make the ICL at least as large as the sum of the MC-AE member links it might have to carry.

Two details make the real requirement worse than the rule of thumb. First, the traffic that now traverses the ICL includes the unknown-unicast flood from effect 2 and any duplicated BUM from effect 3, both of which are demand-driven rather than steady-state. Second, crossing the ICL adds a hop, so the effective path length for those flows changes at the same moment that the fabric loses half its access bandwidth. In a design where the ICL is deliberately smaller than the MC-AE - a common cost optimisation - this is the moment the oversubscription ratio shows up as packet drops rather than as a design number.

The First Seconds After ICCP Fails

Ordering the effects in time explains most of what an operator sees on a monitoring dashboard:

  1. T+0: ICCP session drops. MAC, ARP, IGMP and LACP synchronisation all stop at the same instant.
  2. T+0 to T+1s: the failure action runs. With backup liveness detection configured, the peer status determines whether the local chassis takes over cleanly; without it, the LACP system ID changes to the default on both peers.
  3. T+1s to T+10s: the downstream device re-evaluates its aggregator, re-ARPs for its gateway, and traffic that mapped to the now-standby links is hashed elsewhere. This is the window where a monitoring system shows a short, sharp loss of packets but full recovery of throughput at half rate.
  4. T+10s to T+ageing: the flooding window of effect 2 runs until entries age out or are relearned, and the ICL carries the extra load described above.
  5. Recovery: when ICCP comes back, the configured system ID returns, the downstream device sees one partner again, and the aggregated bandwidth is restored. The transition is another LACP event for the host, so expect a second, smaller loss window.

Capturing the Evidence

The most useful debugging tool is a capture on the MC-AE, because it shows exactly what the downstream device is being told. Start with monitor traffic, which supports BPF-style filters on Junos, or with a port mirror if you need to keep the traffic on the wire.

! what the downstream host is being told: LACP PDUs only
monitor traffic interface ae0 no-resolve matching "ether proto 0x8809"

! ARP and neighbour discovery during the failure window
monitor traffic interface ae0 no-resolve matching "arp or icmp6"

! IGMP / MLD joins, if the segment carries multicast
monitor traffic interface ae0 no-resolve matching "igmp or icmp6" detail

! write it out for later analysis
monitor traffic interface ae0 write-file /var/tmp/ae0-iccp-fail.pcap

In the capture, three fields do the work. The LACPDU is a Slow Protocol frame with Ethertype 0x8809 and subtype 0x01; inside it, the Partner System field is the identity the downstream device uses to build its aggregator, and the Partner Oper Key is the admin key it sees. If the Partner System changes from your configured value to the chassis default at the same moment the ICCP session dropped, you have the documented ICCP-down behaviour on the wire. If it never changes while the failure action is supposed to be running, backup liveness detection is not answering.

# decode a collected capture on a Linux jump host
tcpdump -r ae0-iccp-fail.pcap -nn -v ether proto 0x8809 \
  -T lacp 2>/dev/null | grep -A4 "Partner"

# show only the identity changes across the whole capture
tshark -r ae0-iccp-fail.pcap -Y "slow_protocol.subtype == 0x01" \
  -T fields -e frame.time_relative -e lacp.partner_system -e lacp.partner_key

For ARP, look for the pattern rather than the individual packet: a burst of requests for the gateway address immediately after the failure, followed by replies from a single chassis. If replies arrive from both chassis for the same address, one of them is still claiming active status for the group and you have a duplicate-ARP condition, not a recovery. For multicast, the check is counting: capture for ten seconds before the failure and ten seconds after, filter to the group in question, and compare the frame counts on the MC-AE. A doubling confirms that BUM coordination has been lost.

! correlate counters at the same time as the capture
show interfaces ae0 extensive | match "packets|drops"
show interfaces ae1 extensive | match "packets|drops"   ! the ICL
show iccp peers
show mc-ae status

Three Conclusions That Are Usually Wrong

  • "The fabric is down." ICCP being down is a control-plane event; the data plane usually keeps forwarding, at reduced capacity and with flooding. Confirm with counters on the MC-AE before escalating a control-plane alarm into a service outage.
  • "The peer is dead." ICCP reachability and peer liveness are different questions. Backup liveness detection, when configured and answering, is what actually answers the second one - and it only runs while ICCP is down.
  • "PIM is fine because the adjacency is up." The adjacency surviving the ICCP failure is documented behaviour, so it proves nothing about forwarding. Check the unicast RPF path for the multicast source address instead.

Troubleshooting Order

  1. show iccp and show iccp trace - is the session down, and why (MTU, route, BFD, auth)?
  2. show mc-ae status - which state is each MC-AE in, and is liveness the reason?
  3. show lacp interfaces - confirm the system ID the downstream device actually sees.
  4. Check ARP/ND and MAC counts before and after, then confirm with a capture on the MC-AE.
  5. Check PIM neighbours, the multicast route for the affected groups, and the BUM rate on the MC-AE and the ICL.

Related reading: MC-LAG ICCP failure and LACP system ID behaviour on QFX, MC-LAG ICCP with BFD: system ID deep dive and MC-LAG failure scenarios: standby and split brain.

原文链接:Juniper: Advanced MC-LAG Concepts