Respond to Degraded ONTAP System Health - 夜莺博客

Respond to Degraded ONTAP System Health

Source: NetApp ONTAP 9 Documentation

When your system's health status is degraded, you can show alerts, read about the probable cause and corrective actions, show information about the degraded subsystem, and resolve the problem. Suppressed alerts are also shown so that you can modify them and see whether they have been acknowledged. This walkthrough follows the official procedure step by step and then adds the surrounding context — where the alerts originate, how to read the instance-level details, and how to confirm that health really has returned to OK.

You can discover that an alert was generated by viewing an AutoSupport message or an EMS event, or by using the system health commands.

How ONTAP Health Alerting Works

ONTAP models system health as a set of independent subsystems, each of which is either OK or degraded. The overall health status is a roll-up: it is OK only when every subsystem reports OK, and it becomes degraded as soon as any one of them raises an alert. That design is deliberate, because it means you never have to guess where the problem is — the subsystem name in the output is the pointer.

Subsystems you will meet most often on a production cluster include:

  • SAS-connect and SAS-exp — the storage fabric between controllers and disk shelves, covering SAS ports, expander modules and shelf connectivity.
  • Memory and NVRAM — controller memory modules and the battery- or capacitor-backed NVRAM used to protect writes.
  • Controller and HA interconnect — failover capability between the two nodes in an HA pair.
  • Service-Processor and cluster-switch — out-of-band management, and connectivity to the cluster interconnect switches.
  • Shelf and Interface — disk shelves and network interfaces that the node depends on.

Two other facts keep expectations sensible. First, alerts are informational as well as critical: a single degraded subsystem on a two-node cluster usually does not stop data serving, because the HA partner or multipath is designed to absorb exactly this kind of failure. Second, alert state is not the same as event state — the EMS event that generated the alert may have rolled out of the event log long ago while the underlying condition is still present, which is why you always confirm at the health layer rather than the log layer.

Procedure

  1. Use the system health alert show command to view the alerts that are compromising the system's health.
  2. Read the alert's probable cause, possible effect, and corrective actions to determine whether you can resolve the problem or need more information.
  3. If you need more information, use the system health alert show -instance command to view additional information available for the alert.
  4. Use the system health alert modify command with the -acknowledge parameter to indicate that you are working on a specific alert.
  5. Take corrective action to resolve the problem as described by the Corrective Actions field in the alert. The corrective actions might include rebooting the system.
  6. When the problem is resolved, the alert is automatically cleared. If the subsystem has no other alerts, the health of the subsystem changes to OK. If the health of all subsystems is OK, the overall system health status changes to OK.
  7. Use the system health status show command to confirm that the system health status is OK.
  8. If the system health status is not OK, repeat this procedure.

Step 1 — List the Degraded Subsystems

Start broad. The summary view tells you which subsystem is degraded and how severe the alert is:

cluster1::> system health alert show
Node          Subsystem   Alert ID                          Severity
------------  ----------  --------------------------------  --------
cluster1-01   SAS-connect Controller_A_to_Shelf_1_Port_A   Alert

Restrict the list when you are focused on one node or one severity level, and include acknowledged alerts when you need to see work already in progress:

cluster1::> system health alert show -node cluster1-01
cluster1::> system health alert show -severity alert
cluster1::> system health alert show -acknowledged true

It is worth noting which subsystems are clean, not just which one is degraded. system health subsystem show enumerates every subsystem and its state, so a degraded SAS-connect alongside an OK cluster-switch tells a different story from two degraded subsystems at once:

cluster1::> system health subsystem show

Step 2 — Read the Probable Cause and Corrective Actions

The alert description is not decoration. Each alert carries three fields that together tell you what is wrong, what will happen if you ignore it, and what to do about it:

cluster1::> system health alert show -instance
              Node: cluster1-01
         Subsystem: SAS-connect
          Alert ID: Controller_A_to_Shelf_1_Port_A
          Severity: Alert
   Probable Cause: The SAS port on the controller is not connected to
                   the SAS port on the shelf.
   Possible Effect: Loss of access to the disks in the shelf.
  Corrective Action: Check the SAS cabling between the controller and
                     the shelf and reseat or replace the cable.
 Acknowledge Status: false

Decide at this point whether the corrective action is something you can perform — reseating a cable, replacing a power supply, moving a workload — or whether it requires NetApp support, a maintenance window, or a failover. Do not skip to the fix just because the alert looks familiar; the Possible Effect field is what tells you how much time you have.

Step 3 — Acknowledge the Alert

Acknowledging does not fix anything. It records that a human has seen the alert and is working on it, which keeps the alert visible in tooling without repeatedly escalating it to colleagues. Be explicit about which node and which alert ID you are acknowledging, because -acknowledge without a target is one of the easier ways to silence the wrong alert:

cluster1::> system health alert modify -node cluster1-01 \
             -alert-id Controller_A_to_Shelf_1_Port_A -acknowledge true

If you hand the work to someone else, or if the condition turns out to be a false positive that you want to keep watching, clear the acknowledgement so the alert resumes its normal prominence:

cluster1::> system health alert modify -node cluster1-01 \
             -alert-id Controller_A_to_Shelf_1_Port_A -acknowledge false

Suppressed alerts appear in the same output as active ones, so during an audit it is worth listing everything rather than only the unacknowledged entries. A suppressed alert that nobody remembers suppressing is a blind spot.

Step 4 — Take Corrective Action

Work the corrective action as written. For hardware-domain alerts on a two-node cluster, the usual pattern is to resolve the fault on the node that is not currently serving, or to verify that the partner is healthy before you contemplate a takeover or a reboot. Some corrective actions explicitly include rebooting the system, and those must be scheduled as outages with the usual prerequisites: a verified cluster health check, recent Snapshot copies, and a confirmation that the partner can take over.

Supporting commands that answer the questions you will be asked before touching hardware:

cluster1::> storage failover show
cluster1::> storage shelf show
cluster1::> network port show
cluster1::> system health status show

For storage-side faults, shelf and disk views confirm whether the affected shelf is still visible to the node and whether its disks are present — a cable fault and a shelf power fault can produce similar alerts but need different responses. For alerts that point at network interfaces or interconnects, the ONTAP network troubleshooting commands are the natural next stop.

Step 5 — Confirm the Alert Cleared and Health Returned to OK

Alerts clear automatically when the condition goes away; you do not need to delete them. The sequence to verify is bottom-up:

cluster1::> system health alert show
cluster1::> system health subsystem show
cluster1::> system health status show
Health Status: OK

If the subsystem has no other alerts, its health flips to OK. If all subsystems are OK, the overall status becomes OK. An empty alert list paired with a still-degraded status is the one combination worth investigating carefully — it usually means a second alert is being raised and cleared faster than you are sampling, or that an acknowledged alert is still outstanding. If the status is not OK, repeat the procedure from the top; the second pass often surfaces a different subsystem that was masked by the first.

Give the poller time before you conclude that a fix failed. Health state is evaluated periodically, so a cable that was reseated thirty seconds ago may not be reflected yet. Re-run system health status show after a short interval rather than immediately.

Where Alerts Come From: AutoSupport and EMS

The health commands are what you run once you know there is a problem. Finding out about the problem in the first place usually happens through one of two channels:

  • AutoSupport messages — periodic and event-driven bundles sent to NetApp and, optionally, to your own mail gateway. A health alert that you did not go looking for almost always arrives first as an AutoSupport notification.
  • EMS events — the event management system that records everything the cluster considers noteworthy, including the events that generate health alerts. Filtering EMS for severity and time window is the fastest way to reconstruct what happened before an alert appeared.
cluster1::> autosupport history show
cluster1::> event log show -severity * -time >1h

Our guides on ONTAP event log and EMS filter commands and the ONTAP performance monitoring commands cover those two directions in depth — the first for understanding why an alert fired, the second for confirming that the underlying condition has not quietly degraded performance while health still reported OK.

Key Commands

system health alert show
system health alert show -instance
system health alert modify -acknowledge
system health status show

Reference: system health alert modify — https://docs.netapp.com/us-en/ontap-cli/system-health-alert-modify.html

Turning This into an Operational Routine

  1. Point AutoSupport at a mail alias someone actually reads, and make health alerts a daily check rather than an incident-driven one.
  2. List alerts, then acknowledge the ones you are actively working — with node and alert ID, never in bulk.
  3. Read the probable cause, possible effect and corrective action before choosing a maintenance window.
  4. Confirm prerequisites (storage failover show, recent Snapshot copies) before any action that reboots or fails over a node.
  5. Verify recovery at the subsystem level first, then at the system level, and leave a note in the change record.
  6. Keep a spreadsheet or dashboard mapping alert IDs to the subsystems and hardware they describe; the second time you see an alert, the diagnosis should take seconds.

Related NetApp ONTAP Reading

For the wider command set around storage administration, our NetApp ONTAP CLI cheat sheet collects the essential storage commands in one place, and the ONTAP network troubleshooting guide covers the interface and interconnect checks that many health alerts eventually require.

原文链接:https://docs.netapp.com/us-en/ontap/system-admin/respond-degraded-system-health-task.html