SnapMirror Lag Troubleshooting on ONTAP: A Checklist - 夜莺博客

SnapMirror Lag Troubleshooting on ONTAP: A Checklist

SnapMirror lag is the difference between when the snapshot was created on the source and the system time on the destination. That definition alone explains most false alarms: lag is not "time since the last successful transfer", it is snapshot age plus transfer duration, so a relationship that updates hourly will legitimately report forty-five minutes of lag immediately after a successful transfer. Before escalating a backup SLA breach, check that the number you are looking at means what you think it means.

How lag is calculated

Lag = (destination system time when the transfer completes) − (source snapshot creation timestamp). Worked examples from NetApp's own KB:

  • Snapshot created at 22:55, scheduled update runs at 23:00, transfer takes 15 minutes → lag is 20 minutes at completion. The five minutes between snapshot creation and transfer start plus the fifteen minutes of transfer.
  • Snapshot labelled sv_daily created at 17:00, the scheduled update triggers at 01:00 the next morning and transfers in 30 minutes → lag is 8 hours 30 minutes. Nothing is broken; the schedule simply replicates an eight-hour-old snapshot.
  • Snapshot at 08:00, scheduled update at 12:00, 15-minute transfer → lag 4 hours 15 minutes.

Two factors people forget: the clock and timezone on both clusters, and the duration of the transfer. If system time is wrong on either side, timestamps are wrong and lag is wrong — a destination cluster with a drifting clock can report lag that never existed. Always run date on both clusters before believing a lag alert.

The authoritative command

cluster::> snapmirror show -destination-path svm1:vol_dest -fields exported-snapshot,lag-time
cluster::> snapmirror show -fields state,status,lag-time,last-transfer-end-timestamp
cluster::> snapmirror status -l
cluster::> snapmirror lag-threshold show

Compare exported-snapshot against lag-time: the snapshot name carries the creation timestamp, so you can verify the arithmetic yourself rather than trusting the summary. The last-transfer-end-timestamp field separates "the transfer happened recently" from "the data being transferred is old".

Ordered troubleshooting path

  1. Confirm the expected lag. Derive it from the snapshot policy and schedule: expect the time between two scheduled updates plus the normal transfer duration.
  2. Check the schedule is what you think. In /etc/snapmirror.conf, inspect the minute field. A * there means the update is requested every minute — a well-known misconfiguration that queues overlapping transfers.
  3. Look for a stuck transfer. A failed or long-running transfer keeps lag climbing. The syslog will show Scheduled update failed to start. (Scheduled update was delayed because another SnapMirror operation for the relationship is in progress.) — the signature of stacked schedules.
  4. Check whether an old snapshot is in progress. If the last transferred snapshot is not the newest one, lag reflects the age of the older snapshot being copied.
  5. Validate both clocks and timezones. Incorrect destination cluster time produces inaccurate lag with no other symptom.
  6. Check bandwidth. A saturated WAN path is the classic cause of lag that grows steadily rather than jumping.
  7. Check destination resources. Disk I/O, CPU and the destination volume's free space all slow transfer throughput; a full or nearly full destination volume stalls updates entirely.

Fixing the common causes

# Stagger overlapping schedules in /etc/snapmirror.conf
# four hourly transfers -> start them 15 minutes apart
cluster::> snapmirror update -destination-path svm1:vol_dest
cluster::> snapmirror show -fields lag-time
  • Stagger schedules. If four transfers are required every hour, space their start times 15 minutes apart instead of firing them together.
  • Never schedule per minute for async mirroring. If the business needs near-synchronous replication, use Sync SnapMirror rather than an async update every minute.
  • Separate snapshot and mirror schedules. Do not let volume snapshot schedules overlap mirror/backup update windows.
  • Throttle transfers where a high-speed LAN carries many relationships; unthrottled transfers starve each other.
  • Group by transfer duration. Keep relationships with similar transfer times in the same volume, and create multiple destination volumes when primary sizes and requirements differ.

AutoSupport makes lag look worse

When you read lag from SNAPMIRROR.XML in AutoSupport, two extra factors inflate the number: the elapsed time since the last transfer and the time AutoSupport itself took to collect its sections. A twenty-minute lag in a bundle triggered at 00:00 that took twenty minutes to generate can appear as 1:40:00. Always cross-check the on-cluster value before raising an incident.

For performance rather than lag see SnapMirror and SnapVault performance troubleshooting; for adjacent ONTAP administration see CIFS/SMB share permissions.

Alerting thresholds that do not page you unnecessarily

Active IQ Unified Manager raises lag events through dedicated event types: ocumEvtMirrorVaultRelationshipLagWarning, the corresponding error, and the SnapMirror-specific ocumEvtSnapMirrorRelationshipLagWarning and ocumEvtSnapMirrorRelationshipLagError. Thresholds are configured under Settings, Event Thresholds, Relationship, as percentages of the expected interval.

Setting Threshold Effective lag for a 60-minute schedule
Warning 150% 90 minutes
Error 250% 150 minutes

The mistake to avoid is alerting on an absolute value. A one-hour schedule and a five-minute schedule cannot share the same threshold, and because lag includes normal transfer duration, an absolute threshold below the schedule interval fires on healthy relationships. Percentage-of-interval thresholds make the alert meaningful across every relationship on the estate.

Proving the transfer rather than the lag

cluster::> snapmirror show -fields last-transfer-begin-timestamp,last-transfer-end-timestamp,last-transfer-size,last-transfer-duration
cluster::> snapmirror log show -destination-path svm1:vol_dest
cluster::> event log show -severity ERROR

When the lag looks wrong but the relationship reports healthy, compare last-transfer-size against the rate of change on the source volume. A transfer that completes quickly but copies almost nothing is not a Mirror problem — it is a snapshot policy problem, because the snapshot being replicated predates the data the application is writing.

原文链接:https://kb.netapp.com/on-prem/ontap/DP/SnapMirror/SnapMirror-KBs/How_to_troubleshoot_SnapMirror_lag_issues