vSAN Stretched Cluster: Witness, Sites and Site Preference - 夜莺博客

vSAN Stretched Cluster: Witness, Sites and Site Preference

A vSAN stretched cluster looks like one cluster but is really three fault domains: two data sites with hosts and storage, plus a witness host that holds no capacity and exists purely to break ties. The reason to build one is that it gives you storage-level availability across two rooms or two buildings with a single vSphere cluster and no array-based replication. The reason it goes wrong is usually the witness: where it lives, how it communicates, and which site is preferred all determine whether a link failure produces an outage or a clean failover.

Three Fault Domains, One Cluster

VMware's own framing is the clearest one: a stretched cluster is nothing more than a vSAN cluster with three fault domains — one for each data site and one for the witness. Hosts are assigned to their respective fault domain (site) during configuration.

You need a minimum of three hosts: at least one in the preferred site, at least one in the secondary site, and one host acting as the witness. In production that normally means several hosts per site for capacity, plus a small witness appliance at a third location.

The Witness Is Not a Spare Host

The witness has no storage capacity and does not hold components. Its job is to participate in the quorum calculation so that if the inter-site link fails, vSAN can determine which side should keep serving. That produces the rule that matters most in practice:

  • If the witness is reachable and only the inter-site link fails, the preferred site keeps running and the secondary site's components go unavailable.
  • If the witness is unreachable but both data sites can see each other, the cluster continues normally.
  • If a data site is lost and the witness cannot be reached either, you have lost quorum and the surviving site cannot serve.

This is why the witness location is a design decision, not an afterthought. Put it somewhere that shares as few failure modes as possible with both data sites — a third building, a small colo, or a different rack with independent power and network.

Designating the Preferred Site

Exactly one site must be designated preferred. The preferred site is the one that remains in operation when the link between sites fails, unless it is resyncing or has another problem of its own. The choice has real consequences:

  • VMs and their components should have their primary placement on the preferred site so failover does not require a resync.
  • Which site is preferred should be decided by where the majority of workloads actually run, not alphabetically.
  • Changing the preferred site later is possible but requires the cluster to be healthy — you cannot reprefer a site while one side is down, which is exactly when you might want to.

Witness Traffic Separation

By default, witness traffic flows over the vSAN network. In a stretched cluster that can be undesirable: the witness path and the inter-site data path may share hardware, so a single failure takes out both. Witness traffic separation (WTS) lets you send witness traffic over a different VMkernel interface — typically a management or dedicated link — and is applied to all hosts in the two data sites that participate in the stretched cluster.

The benefit is measurable during a link failure: without WTS a cut fibre between sites can take out witness connectivity as well, pushing the cluster into a quorum loss even though the witness host itself is perfectly healthy.

Configuration Summary

  1. Create or claim a vSAN cluster and confirm all hosts are contributing.
  2. Assign each host to its site (fault domain) — preferred and secondary.
  3. Deploy the witness host at the third location and add it to the cluster.
  4. Designate the preferred site.
  5. Optionally configure witness traffic separation on the data-site hosts.
  6. Set the storage policy so that it tolerates the failures you actually want to survive.

The storage policy is the other half of the design. A stretched cluster with a policy of Failures to tolerate = 1 places one mirror in each site; that survives a whole-site loss. A policy with no site-level redundancy can survive a disk failure but not a site failure, which defeats the point.

What to Test

  • Pull the inter-site link and confirm the preferred site keeps running and the secondary site's VMs restart on the preferred side as policy dictates.
  • Power off the witness and confirm normal operation continues while both sites can see each other.
  • Verify the resync after the link returns completes inside your maintenance window — resyncing a stretched cluster over a thin inter-site link is the most common unexpected bottleneck.

Related topics on this site: vSphere Standard vSwitch VLAN and Trunk Port Groups for the port group and VLAN plumbing behind the VMkernel interfaces, VMware NSX Segments, Transport Zones and TEPs if the stretched design also needs a stretched logical network, and iSCSI Multipathing with multipathd: Tuning Guide for the equivalent path-redundancy thinking in block storage.

原文链接:https://techdocs.broadcom.com/us/en/vmware-cis/vsan/vsan/8-0/planning-and-deployment/working-with-virtual-san-stretched-cluster/introduction-to-stretched-clusters.html