SnapMirror and SnapVault Performance Troubleshooting - 夜莺博客

SnapMirror and SnapVault Performance Troubleshooting

Replication performance problems split into two symptoms that need different investigations: lag is above expectation and the transfer duration misses its SLA, or the duration meets the SLA while throughput is plainly low. Chasing both with the same checklist wastes time, because the first is usually a scheduling or resource problem and the second is usually a bandwidth or layout problem. This article lays out the methodology NetApp documents, then the fixes that address each class.

Establish the expectation first

You cannot call replication slow without knowing the target. Derive the expected lag from the schedule — it is the time between two scheduled updates plus the normal transfer duration. Then take the helicopter view:

cluster::> snapmirror status -l
cluster::> snapvault status -l
cluster::> snapmirror show -fields lag-time,last-transfer-duration,state,status
cluster::> snapmirror show -fields lag-time -instance

From the output establish: how many systems are involved, how many mirror/backup services are active, which systems are simultaneously source and destination, how many relationships per system, the transfer lag, and the date and time of the last successful transfer. A cluster that is both a destination for one relationship and a source for three others has a fundamentally different performance profile from a leaf destination.

The four documented causes

Cause Signature Fix
Overloaded SnapMirror/SnapVault implementation Many relationships starting at the same minute Stagger schedules; split across destination volumes
Non-optimal space and data layout Transfers of similar size behave differently Group relationships with similar transfer times; create multiple destination volumes
High system resource utilisation Source and destination CPU, disk or protocol activity at saturation during transfer Move schedules off peak; reduce concurrent relationships
Low network bandwidth Throughput plateaus well below link capability Throttle competing transfers; check the path, not just the endpoints

Schedule discipline, in detail

Three rules eliminate most performance incidents:

  1. Stagger start times. If four transfers are required every hour, space them 15 minutes apart instead of firing all four at the top of the hour.
  2. Never schedule per minute. In /etc/snapmirror.conf, a * in the minute field means the update is requested every single minute. Overlapping requests then produce the log signature Scheduled update failed to start. (Scheduled update was delayed because another SnapMirror operation for the relationship is in progress.) If the business genuinely needs near-continuous protection, deploy Sync SnapMirror rather than an async update every minute.
  3. Separate snapshot and replication windows. Volume snapshot schedules should not coincide with mirror or vault update windows; the snapshot itself consumes the same disk resources the transfer needs.

Layout and throttling

  • Keep relationships with roughly equal transfer times in the same volume; a 2 GB nightly increment sharing a volume with a 400 GB weekly baseline produces the characteristic "sometimes fine, sometimes terrible" pattern.
  • Create multiple destination volumes when primaries differ in size and transfer requirements, so a large transfer cannot starve a small one.
  • In a high-speed LAN carrying many relationships, throttle transfers in /etc/snapmirror.conf using the bandwidth argument. Unthrottled transfers compete and all of them slow down; throttling is counter-intuitively a throughput measure, not a restriction just for WAN links.

Collecting evidence that actually isolates the bottleneck

cluster::> perfstat
cluster::> statit
cluster::> sysstat -m

Run perfstat on both source and destination (it bundles statit and sysstat output), and run statit and sysstat -m while a transfer is in flight rather than before or after. Correlation is the whole method: a destination disk at 95 percent utilisation only during the transfer window is your bottleneck; the same figure measured an hour later is meaningless.

Alongside the array data, gather the network picture: other jobs on the path, available bandwidth, any failures, expected throughput, and whether throttling is in place. A replicating pair that meets its SLA with low throughput is almost always bandwidth-constrained on a shared path or throttled deliberately — and in that case the correct action may be to adjust expectations rather than the configuration.

Diagnostic order

  1. Compute expected lag from the schedule; compare with the reported lag.
  2. Check the minute field in /etc/snapmirror.conf for a *.
  3. Look for the overlapping-schedule log message.
  4. Confirm both clusters' clocks are accurate — incorrect time produces incorrect lag with no other symptom.
  5. Measure destination utilisation during a transfer, not around it.
  6. Measure the path: if bandwidth is saturated, stagger, throttle or upgrade; if not, look at layout.
  7. Only after 1-6, consider tuning transfer parameters.

For lag specifically — including how lag is computed and why normal lag looks alarming — see SnapMirror lag troubleshooting on ONTAP, and for the file-protocol layer of the same storage, see ONTAP CIFS/SMB share permissions.

原文链接:https://kb.netapp.com/on-prem/ontap/DP/SnapMirror/SnapMirror-KBs/What_is_the_method_of_troubleshooting_SnapMirror_SnapVault_and_OSSV_performance_issues