Monitor the Reachability of ONTAP Network Ports - 夜莺博客

Monitor the Reachability of ONTAP Network Ports

Source: NetApp ONTAP 9 Documentation

Reachability monitoring is built into ONTAP 9.8 and later. Use this monitoring to identify when the physical network topology does not match the ONTAP configuration. In some cases, ONTAP can repair port reachability. In other cases, additional steps are required.

Use these commands to verify, diagnose, and repair network misconfigurations that stem from the ONTAP configuration not matching either the physical cabling or the network switch configuration.

What Reachability Monitoring Actually Checks

ONTAP groups Ethernet ports into broadcast domains. A broadcast domain is ONTAP's own model of a Layer 2 segment: a set of ports that are expected to reach each other without routing. Reachability monitoring periodically performs a Layer 2 scan across those ports and compares what the switch fabric actually delivers against what ONTAP expects. When a port can reach broadcast domains it was never assigned to, or cannot reach the ports it shares a domain with, the mismatch is reported rather than silently tolerated.

This matters because Layer 2 problems do not announce themselves. A cable patched into the wrong switch port, a VLAN pruned off a trunk, or a trunk that was never allowed for a new VLAN all produce a network that is almost entirely healthy — until a LIF fails over to a port that cannot reach the rest of its subnet, and an entire storage service disappears. Reachability monitoring turns that latent misconfiguration into an observable state you can act on before a failover event exposes it.

The scan is a Layer 2 check only. It answers "can this port see the other ports of its broadcast domain", not "can this LIF route to the NFS client" and not "is the switch port configured with the right access VLAN name". Those questions still belong to ping, traceroute and switch-side verification.

Prerequisites and Scope

  • Version: ONTAP 9.8 or later. On earlier releases the command does not exist; use broadcast domain membership and switch-side checks instead.
  • Access: cluster administrator at the admin privilege level.
  • Scope: physical Ethernet ports, VLAN ports and interface group members. Ports in the Cluster IPspace are checked the same way as data ports.
  • Not covered: IP-level reachability, LIF subnet associations, routing, and firewall behaviour between the cluster and its clients.

Procedure

1. View Port Reachability

network port reachability show

2. Use the Decision Tree and Table to Determine the Next Step

Reachability-status Description / Action
ok The port has layer 2 reachability to its assigned broadcast domain. If "ok" but there are "unexpected ports", consider merging one or more broadcast domains. If "ok" but there are "unreachable ports", consider splitting one or more broadcast domains. If no unexpected or unreachable ports, your configuration is correct.
Unexpected ports The port has L2 reachability to its assigned broadcast domain, but also to at least one other broadcast domain. Examine physical connectivity and switch configuration; the broadcast domain may need to be merged.
Unreachable ports A single broadcast domain has become partitioned into two different reachability sets — split the broadcast domain to synchronize ONTAP configuration with the physical topology (after verifying physical and switch configuration is accurate).
misconfigured-reachability The port does not have L2 reachability to its assigned broadcast domain, but does have reachability to a different broadcast domain. Repair with: network port reachability repair -node <node> -port <port>
no-reachability The port does not have L2 reachability to any existing broadcast domain. Repair with: network port reachability repair -node <node> -port <port> — the port is assigned to a new automatically created broadcast domain in the Default IPspace.
unknown Wait a few minutes and try the command again.

3. After Repairing a Port

  • Check for and resolve displaced LIFs and VLANs.
  • If the port was part of an interface group, understand what happened to that interface group.

Reading Each Status Correctly

The status vocabulary is deliberately narrow, and two of the values are easy to misread because they describe the port as healthy while reporting a problem elsewhere in the domain.

ok means the port reaches the broadcast domain it belongs to. It does not mean the domain is correctly designed. The output can still list unexpected ports — ports that are reachable but are not members of this domain — or unreachable ports — members of this domain that cannot be reached. Both are reported as extra lists alongside an ok status, and both indicate that the ONTAP model and the physical topology disagree. In the strict sense, ok with empty lists is the only fully correct result.

Unexpected ports and the multi-domain-reachability status describe a port that can see more than one Layer 2 segment. Physically, that usually means two switch ports that should be in different VLANs are actually in the same one — a trunk allowed list that is too broad, or an access port patched into the wrong VLAN. Two separate ONTAP broadcast domains that are really one segment will confuse failover groups and can cause LIFs to land on ports that cannot serve their subnet. The fix is either to correct the switch configuration, or, if the convergence is intentional, to merge the broadcast domains.

Unreachable ports describe the opposite condition: a single broadcast domain has been partitioned into two reachability sets. Commonly this is a trunk that carries the VLAN on one controller but not the other, a mislabelled patch panel, or a switch that has the VLAN created but not permitted on the inter-switch link. Because the ports are in the same ONTAP domain but not in the same physical segment, LIF failover across the partition silently breaks connectivity. Once you have verified with the switch that the physical topology is what you actually want, split the domain so ONTAP matches reality.

misconfigured-reachability and no-reachability are the two statuses that ONTAP can fix for you, and they are the reason the repair command exists. In the first case the port is live but plugged into the wrong segment; ONTAP simply moves it to the broadcast domain it can actually reach. In the second case the port reaches nothing at all; ONTAP places it into a newly created broadcast domain in the Default IPspace so the port is no longer falsely claiming membership in a domain it cannot serve.

Detailed Command Reference

Start with the summary view, then narrow to the ports that matter. The -detail flag is what turns a one-line status into a list of the exact broadcast domains a port can reach, which is what you need in order to decide between merging, splitting and repairing.

network port reachability show
network port reachability show -node node1 -port e0d
network port reachability show -detail -node node1 -port e0d
network port reachability show -instance
network port reachability show -fields node,port,reachability-status

A healthy single port looks like this:

cluster1::> network port reachability show -node node1 -port e0d
network port reachability show
node         Port     Expected Reachability        Reachability Status
-----------  -------- ---------------------------- --------------------------
node1        e0d      Default:Default              ok

Useful supporting commands when you need to reconcile the ONTAP model with the switch configuration:

network port broadcast-domain show
network port broadcast-domain show -ipspace Default -fields broadcast-domain,mtu,ports
network port show
network port vlan show
network port ifgrp show
network interface show -fields home-node,home-port,curr-node,curr-port,subnet-name

Handling Unexpected Ports: Merging Broadcast Domains

If the reachability scan shows that two domains are really one Layer 2 segment and that is what the network is supposed to be, merge them. Before merging, confirm the MTU is consistent across the domains — merging a 9000-byte domain into a 1500-byte one changes what the ports can carry and can cause traffic loss for LIFs that were relying on jumbo frames.

network port broadcast-domain show -ipspace Default

network port broadcast-domain merge -ipspace Default \
  -broadcast-domain bd-mgmt -into-broadcast-domain bd-data

After the merge, re-run network port reachability show and confirm the unexpected ports list is now empty. If a merge is not appropriate — because the two segments are genuinely separate and the switch configuration is wrong — fix the switch first and re-scan, rather than bending ONTAP to match a fault.

Handling Unreachable Ports: Splitting Broadcast Domains

Splitting is the tool for a partitioned domain. Identify the set of ports that cannot reach the rest of the domain, verify on the switch that they really are on a separate segment, then move them into a new broadcast domain.

network port broadcast-domain split -ipspace Default \
  -broadcast-domain bd-mgmt \
  -new-broadcast-domain bd-mgmt-2 \
  -ports node1:e0d,node2:e0d

Two constraints catch people out. First, if the ports belong to a failover group, all ports of that failover group must be provided in the same split. Use network interface failover-groups show to check which ports belong together. Second, if the ports already carry LIFs, those LIFs cannot be part of a subnet's ranges, and both the current port and the home port of each LIF must be included. Inspect them with:

network interface show -fields subnet-name,home-node,home-port,curr-node,curr-port
network subnet remove-ranges -subnet-name <subnet> -ranges <range> -force-update-lif-associations true

After the Repair: Displaced LIFs, VLANs and Interface Groups

Repairing a port moves it into a different broadcast domain, which means any LIF homed on that port may be re-homed elsewhere and any VLAN port built on top of it may be affected. The repair command warns you about this before it proceeds; do not confirm it blindly on a production cluster.

network interface show -fields home-node,home-port,curr-node,curr-port
network port vlan show
network port ifgrp show
network interface show -failover

Interface groups deserve particular care. If every member port of an interface group reports no-reachability, repairing each member individually removes them one at a time from the group and places each into a new broadcast domain — and once the members are gone the interface group itself is removed. The correct sequence in that situation is to fix the physical connectivity, confirm the scan reports reachability again, and only then decide whether a repair is needed at all.

What Reachability Monitoring Does Not Cover

It is a Layer 2 check, so it will be perfectly happy while an L3 problem exists. Specifically, it does not validate that a LIF's IP address sits inside its subnet, that the subnet's ranges are correct, that a client can route to the SVM, that a firewall is permitting the traffic, or that the switch has a name-correct VLAN. Pair it with the ordinary IP-level tooling:

network ping -lif <lif> -vserver <svm> <destination>
network traceroute -lif <lif> -vserver <svm> <destination>
network interface show -vserver <svm>
network route show

Operationalising the Check

Reachability is a state, not an event, so a single manual check tells you very little. Two habits make it useful. First, run network port reachability show after every change that touches physical cabling, switch VLAN configuration, or ports and interface groups — the value is highest in the window between the change and the first failover. Second, schedule it: a periodic collection of network port reachability show -fields node,port,reachability-status through the ONTAP REST API or the AutoSupport data gives you a trend and a diff, so a port that drifts from ok to misconfigured-reachability overnight is caught before the LIF that depends on it moves. ONTAP also raises EMS events for reachability changes, which can be forwarded to the same alerting pipeline you already use for disk and HA events.

Key Commands

network port reachability show
network port reachability repair -node <node> -port <port>

Troubleshooting Matrix

Symptom Likely cause First action
Status ok but unexpected ports listed Two ONTAP domains are one physical segment Verify switch VLANs, then merge or fix the switch
Status ok but unreachable ports listed Domain partitioned across switches Verify trunk allowed VLANs, then split the domain
misconfigured-reachability Port patched into the wrong segment Repair the port, then check for re-homed LIFs
no-reachability Port down, dead cable, or switch port disabled Check link state before repairing; repairing a whole interface group can delete it
unknown Scan not yet complete Wait a few minutes and re-run
LIF unreachable after a repair LIF was re-homed to a port without the required subnet network interface show -fields home-port,curr-port

Related Reading

References: network port reachability show and network port reachability repair in the ONTAP command reference.

原文链接:https://docs.netapp.com/us-en/ontap/networking/monitor_the_reachability_of_network_ports.html