NetApp ONTAP Commands: 50 Real-World Q&A for Admins - 夜莺博客

NetApp ONTAP Commands: 50 Real-World Q&A for Admins

NetApp ONTAP powers FAS, AFF and Select platforms in thousands of enterprises, and its command-line interface is the fastest way to diagnose storage problems when things go wrong. This collection of real-world Q&A takes you from fundamentals - volumes, aggregates, WAFL, RAID-DP - through provisioning, replication and snapshots, into the troubleshooting scenarios that actually come up in production: LUNs not visible on hosts, volumes full after file deletion, SnapMirror lag, AutoSupport failures and CIFS share access. Every answer includes the ONTAP command you would run.

Two things make ONTAP CLI work different from switch or server administration. First, almost every object is scoped to a SVM (storage virtual machine, formerly vserver), so the same command with the wrong -vserver will query an empty namespace and return a confident, wrong answer. Second, ONTAP separates the cluster shell from the SVM context: cluster1::> is the cluster admin prompt where storage objects live, and cluster1::> vserver context vs1 narrows scope for SVM-owned objects such as LUNs, shares and policies. Knowing which prompt you are at explains most "command not found" surprises.

ONTAP Fundamentals and Provisioning

cluster1::> storage aggregate show
cluster1::> volume show
cluster1::> aggr show-space
cluster1::> volume show -fields size,available,percent-used

Q. What is the relationship between an aggregate and a volume? An aggregate is the RAID layer built from physical disks; a volume is a flexible filesystem (FlexVol) or a set of constituents (FlexGroup) carved out of an aggregate. You size the aggregate for the RAID group and spares, and size the volume for the workload. storage aggregate show answers "do I have RAID capacity", volume show answers "who is using it".

Q. What is WAFL and why does it matter operationally? The Write Anywhere File Layout is ONTAP's filesystem: it writes new blocks anywhere free space exists and updates pointers, rather than overwriting in place. That pointer-based design is what makes snapshots nearly free to create and fast to restore. Operationally, it means a full volume is often a snapshot-retention problem rather than a data-growth problem.

Q. Why RAID-DP instead of RAID 5? RAID-DP adds a second parity disk per RAID group, which protects against a second failure during a rebuild - the window when an array is most vulnerable. On modern high-capacity SATA/SAS drives a rebuild can run for many hours, so double parity is the default expectation rather than an option.

Create a LUN and map it to a host with three commands:

cluster1::> lun create -vserver vs1 -path /vol/vol1/lun1 -size 100g -ostype linux
cluster1::> igroup create -vserver vs1 -igroup ig1 -protocol iscsi -ostype linux -initiator iqn.1993-08.org.debian:01
cluster1::> lun map -vserver vs1 -path /vol/vol1/lun1 -igroup ig1

Q. What has to exist before those three commands work? A SVM with an iSCSI LIF, a volume with enough free space, and the iSCSI service running on that SVM. Missing any of the three produces a failure that sounds like a LUN problem: lun create errors on space, lun map succeeds but the host never sees a target, and the initiator logs a login rejection. Check the prerequisites before the LUN:

cluster1::> volume show -vserver vs1 -volume vol1 -fields size,available
cluster1::> network interface show -vserver vs1 -data-protocol iscsi
cluster1::> iscsi service show -vserver vs1
cluster1::> igroup show -vserver vs1

Q. How do I add a second initiator to an existing igroup? igroup add -vserver vs1 -igroup ig1 -initiator <iqn>. Do this before a host rebuild, not after, and remember that adding an initiator to a mapped igroup exposes the LUN to that host - the change takes effect without any further mapping step.

NAS Access: SVMs, LIFs, Exports and Shares

cluster1::> network interface show
cluster1::> vserver show
cluster1::> volume qtree show -vserver vs1 -volume vol1
cluster1::> vserver export-policy rule show -vserver vs1
cluster1::> vserver cifs share show -vserver vs1
cluster1::> vserver cifs session show -vserver vs1 -instance

Q. A client can reach the SVM but not the export. Where do I look? Three layers, in order: is the data LIF up and reachable from the client (network interface show plus a ping or traceroute from the client), does a rule in the export policy match the client's address with the required access (NFSv3 uses the client source address, so NAT or a routed subnet breaks naive rules), and is the volume junctioned into the SVM namespace so the client's path resolves at all. The juncture is the one people forget: volume show -fields junction-path settles it in one line.

Replication, Backup and Snapshots

cluster1::> snapmirror show -fields status,lag-time
cluster1::> snapmirror initialize -destination-path DEST_SVM:DEST_VOL
cluster1::> snapmirror update -destination-path DEST_SVM:DEST_VOL
cluster1::> snapmirror break -destination-path DEST_SVM:DEST_VOL
cluster1::> volume snapshot show
cluster1::> volume snapshot delete -vserver vs1 -volume vol1 -snapshot snap_old

WAFL uses pointer-based snapshots that are space-efficient and fast. If SnapMirror fails due to a missing baseline, reinitialize; if lag increases suddenly, check WAN performance, source snapshot delays, and volume locks.

Q. SnapMirror versus SnapVault - which am I looking at? The policy name tells you. A mirror-type policy keeps a replicated copy of the source volume, which is what you use for DR failover. A vault-type policy archives snapshots on the destination at a longer retention, which is what you use for backup and compliance. snapmirror show -fields policy,type on the destination answers both questions at once.

Q. How do I know a snapshot policy is actually running? Look at the snapshot list rather than the policy definition:

cluster1::> volume snapshot policy show -vserver vs1 -policy default
cluster1::> volume snapshot show -vserver vs1 -volume vol1 -fields create-time,size
cluster1::> snapshot policy show -vserver vs1

A policy that exists but has produced no snapshots in the retention window means the schedule is not firing - commonly because the volume was moved between aggregates and the schedule was not carried over. Useful depth on both DR topologies and policy setup is in NetApp ONTAP SnapMirror: setup, monitoring and troubleshooting and NetApp ONTAP SnapMirror create: DR, SnapVault and Sync policies.

Space Efficiency and Reclamation

cluster1::> df -h
cluster1::> volume efficiency show -vserver vs1 -volume vol1
cluster1::> volume efficiency start -vserver vs1 -volume vol1 -scan-old-data true
cluster1::> volume show -vserver vs1 -volume vol1 -fields space-guarantee,autosize-mode
cluster1::> volume snapshot show -vserver vs1 -volume vol1 -sort-by size -fields size,total

Q. Users deleted a lot of files but the volume is still full. Why? This is the classic ONTAP answer in interviews, and it has four usual causes: snapshots still reference the deleted blocks, a CIFS share has an open handle on the deleted file, the volume has deduplication or compression metadata that has not been reclaimed, or the space is being held by a stale snapshot created before the deletion. df -h shows the volume-level view while volume efficiency show shows the efficiency work; sorting snapshots by size usually identifies the holding snapshot in a single command. Deleting one large stale snapshot can free more space than any amount of tuning.

Troubleshooting and Hardware Scenarios

cluster1::> storage disk show -broken
cluster1::> disk assign -disk 0a.15 -owner node1
cluster1::> lun mapping show
cluster1::> lun show -vserver vs1 -path /vol/vol1/lun1
cluster1::> system health status show
cluster1::> storage failover show

When a mapped LUN is not visible on the host, verify the igroup mapping, initiator IQNs, and that the iSCSI/FCP service is running, then rescan on the host. NFS volumes that stay full after deletion do not reclaim space instantly - check with df -h and volume efficiency show.

Q. What does "broken disk" mean and when do I assign it? storage disk show -broken lists disks that failed or were removed from an aggregate. A new spare disk in an unowned state is normal; a disk with no owner that you intend to use must be assigned with disk assign. Assigning a disk that ONTAP has marked broken does not repair it - check the disk's state and the recent event log before you assign anything, or you will add a failing disk to a healthy aggregate.

Q. How do I check whether a node is healthy overall? system health status show gives the aggregated verdict, and the cluster log gives the reason:

cluster1::> system health status show
cluster1::> system health alert show
cluster1::> event log show -severity ERROR -time-reverse -max-records 20
cluster1::> system node show -fields health,uptime

Reading the event log with -severity ERROR and a record limit is far more useful than scrolling the whole EMS backlog; the EMS message name (for example wafl.vol.full) is the searchable, supportable identifier. Filtering EMS properly is covered in NetApp ONTAP event log: EMS show and filter commands.

HA, AutoSupport and Scenario Checks

cluster1::> storage failover show
cluster1::> storage failover takeover -ofnode node1
cluster1::> storage failover giveback -ofnode node1
cluster1::> system node autosupport invoke -node node1 -type all -message "Issue with Disk X"
cluster1::> autosupport check show
cluster1::> system node autosupport show -fields state
cluster1::> vserver cifs show
cluster1::> vserver cifs session show

Generate an AutoSupport bundle before opening a support case, verify quorum during node failover, and always have real troubleshooting stories ready - understanding the why behind the command matters more than memorizing output.

Q. Takeover is refused. What are the usual reasons? The partner node must be healthy, the LIFs must be able to migrate, and the configuration must not contain something that only works on one node (a node-scoped LIF, a hardware-dependent setting). storage failover show lists the current state and the block, and the takeover planning checks behind it explain the refusal. The recovery path for the common refusal cases is written up in ONTAP takeover not possible resolution - the practical advice is to read the refusal reason rather than retrying the command.

Q. AutoSupport is not reaching NetApp. How do I debug it? autosupport check show reports connectivity and configuration checks; if those pass, confirm the node's AutoSupport state is enabled and that the mail or HTTPS transport is permitted by your network policy. Many "AutoSupport failed" tickets are firewall changes to SMTP rather than anything on the storage system. Reachability monitoring for the whole data network, including the management paths AutoSupport uses, is covered in NetApp ONTAP network troubleshooting: essential CLI commands and monitoring ONTAP network port reachability.

Ten Habits That Make ONTAP Administering Easier

  1. Always pass -vserver explicitly, even when only one SVM exists - it makes future multi-tenant mistakes visible in review.
  2. Prefer -fields and -instance over reading default output; narrower output is faster and easier to script.
  3. Check df -h and the snapshot list together before declaring a volume full problem.
  4. Read the event log by severity rather than tailing it - the newest log entry is rarely the root cause.
  5. Take an AutoSupport bundle before a failover, not after.
  6. Verify igroup membership from both sides (array and initiator) when a LUN is invisible.
  7. Keep cluster shell and SVM context commands straight; use vserver context deliberately and return with exit.
  8. Record which policy belongs to which relationship, so a later lag investigation starts from the topology instead of the symptom.
  9. Test a snapshot restore on a copy of the volume before you ever need it in production.
  10. Remember that SAN, NAS and object access to the same SVM are different protocols with different troubleshooting paths - do not mix the playbooks.

More NetApp resources on this site: ONTAP takeover not possible resolution, ONTAP ifstat port troubleshooting, and monitoring ONTAP network port reachability. For a compact command reference that complements this Q&A, see NetApp ONTAP CLI cheat sheet and NetApp ONTAP NFS export policy configuration.

原文链接:https://medium.com/@madhuprasadcse/mastering-netapp-storage-interviews-50-real-time-q-a-for-ontap-admins-ef1aa76cae58