MLNX-OS MLAG (sx_mlag) Configuration & Verification - 夜莺博客

MLNX-OS MLAG (sx_mlag) Configuration & Verification

MLAG on MLNX-OS is a different animal from MLAG on Arista or VLT on Dell, and sloppy configuration shows up as a dual-active pair that blackholes traffic rather than as a clean error. This guide walks the Mellanox variant end to end: enabling the required features, building the inter-peer link (IPL), assigning the shared MLAG virtual IP and system MAC, adding server downlinks, and verifying the result with the commands that actually reveal trouble.

Prerequisites and the feature set

  • Two switches running MLNX-OS with a dedicated inter-peer link (IPL) between them — ideally more than one physical link in the IPL port-channel.
  • Management IP addresses on both switches, plus a shared MLAG virtual IP.
  • Both switches must have identical feature configuration for MLAG: the IPL must not be blocked by STP.
sx01 (config) # lacp
sx01 (config) # no spanning-tree
sx01 (config) # ip routing
sx01 (config) # protocol mlag
sx01 (config) # dcb priority-flow-control enable force

Spanning tree is disabled globally in most MLAG designs because the topology is loop-free by construction: the IPL handles inter-peer forwarding, and end hosts multi-home with LAG. If you must keep STP for other ports, disable it on the MLAG member ports and the IPL instead of globally.

Step 1: Build the IPL

sx01 (config) # interface ethernet 1/21 speed 100G force
sx01 (config) # interface ethernet 1/22 speed 100G force
sx01 (config) # interface ethernet 1/21 channel-group 1 mode active
sx01 (config) # interface ethernet 1/22 channel-group 1 mode active
sx01 (config) # interface port-channel 1 ipl 1
sx01 (config) # interface port-channel 1 dcb priority-flow-control mode on force

Repeat identically on the peer. The ipl 1 keyword is what marks this port-channel as the inter-peer link — without it, MLAG comes up but the peers cannot exchange control information correctly.

Step 2: IPL addressing and the MLAG VIP

The IPL IP address must be a dedicated subnet that is not part of the management network and not routed in the data network. The MLAG virtual IP and system MAC are shared between both switches so that downstream LAGs see a single logical device.

# VLAN used for IPL inter-peer traffic
sx01 (config) # vlan 4094
sx01 (config) # interface vlan 4094
sx01 (config interface vlan 4094) # ip address 172.32.255.2 /30
sx01 (config interface vlan 4094) # ipl 1 peer-address 172.32.255.1
sx01 (config interface vlan 4094) # exit

# Shared identity for the MLAG domain
sx01 (config) # mlag-vip MLAG-PAIR ip 172.32.255.3 /29 force
sx01 (config) # mlag system-mac 00:00:5E:00:01:5D
sx01 (config) # no mlag shutdown

On the peer, invert the IPL addresses (same VLAN, same subnet, mirrored host addresses) while keeping the MLAG VIP and system MAC identical on both. A mismatched system MAC across the pair is the classic cause of a downstream LACP bundle that refuses to come up.

Step 3: Add server-facing MLAG ports

# Downlink to a dual-homed host
sx01 (config) # interface ethernet 1/1 channel-group 10 mode active
sx01 (config) # interface port-channel 10 mlag 10
sx01 (config) # vlan 100
sx01 (config vlan 100) # exit
sx01 (config) # interface port-channel 10
sx01 (config interface port-channel 10) # switchport mode trunk
sx01 (config interface port-channel 10) # switchport trunk allowed-vlan 100-200

MLAG port-channel IDs must be reserved in a range that does not collide with the IPL or with normal port-channels. The host side needs no MLAG awareness at all — it simply runs LACP to what appears to be one switch.

Verification: prove it is healthy, not just up

sx01 # show mlag
sx01 # show mlag interfaces
sx01 # show mlag-vip
sx01 # show lacp counters port-channel 10
sx01 # show interface port-channel 1 ipl

What good looks like: MLAG state is active (not standby) on the primary, the peer IP is reachable, and each MLAG port-channel reports the peer's interface as up. On the other side, verify the host's LACP view — if the host shows one member unselected, the system MAC or the LAG IDs disagree between peers.

The failure modes worth designing for

  1. IPL down, both peers think they are primary. Keep the IPL multi-link, and make sure the peers can always reach each other; a single-cable IPL is a single point of failure for the whole MLAG domain.
  2. IPL in the management network. Then a management-plane problem becomes a data-plane outage. Keep IPL addressing dedicated.
  3. STP re-enabled on MLAG ports. It will block one side of the server bond, halving bandwidth with no error anywhere.
  4. Different MTU/flow-control on the two peers. RoCE and PFC deployments break asymmetrically; keep PFC and MTU settings identical on both switches.

Configured this way, MLAG gives you an active/active Layer 2 pair with a single logical identity for hosts — and with the verification commands above, you know it is healthy rather than merely configured.

Related Reading on This Site

原文链接:https://network.nvidia.com/related-docs/solutions/TN_Nutanix_Quick_Start_Guide_on_Mellanox_SN2010_Switches_with_CLI.pdf (NVIDIA - Mellanox MLAG quick start guide)