iSCSI Multipathing with multipathd: Tuning Guide - 夜莺博客

iSCSI Multipathing with multipathd: Tuning Guide

An iSCSI LUN presented over two NICs and two switch paths looks redundant on a diagram, but until DM Multipath is configured and its timeouts are tuned, the host may still funnel all I/O down one path and take minutes to fail over when that path dies. Multipathing solves two problems at once: it presents multiple paths to the same device as a single /dev/mapper/mpathX device, and it decides which paths are usable and how quickly to abandon one. This article covers reliable configuration, the timeout hierarchy that governs failover speed, and how to test that failover actually works before it happens for real.

Install, enable, and stop multipathing local disks

dnf install -y device-mapper-multipath
mpathconf --enable --user_friendly_names n
systemctl enable --now multipathd
systemctl reload multipathd

multipath -v2 -l
multipathd show paths raw format "%d %w" | grep sda

--user_friendly_names n makes the device name the WWID rather than mpatha, which matters when several hosts connect to the same LUNs and you want the same identifier everywhere. Without a blacklist, multipath will happily build a multipath device over the host's internal boot disk — the single most common way to break a working server. Identify the WWID and exclude it, or set find_multipaths on so devices with only one path are left alone.

blacklist {
    wwid WDC_WD800JD-75MSA3_WD-WMAM9FU71040
}

Path grouping and policy

Path grouping determines whether the LUN is active/active or active/passive, and it must match what the array actually supports. Most iSCSI arrays behave active/passive per controller, so a policy that assumes active/active spreads I/O across paths that cannot serve it.

defaults {
    user_friendly_names no
    path_grouping_policy group_by_prio
    prio alua
    path_checker tur
    failback immediate
    no_path_retry 5
    rr_weight priorities
}

devices {
    device {
        vendor "NETAPP"
        product "LUN.*"
        path_grouping_policy group_by_prio
        prio ontap
        path_checker tur
    }
}

path_checker tur uses the SCSI Test Unit Ready command, which most devices support and which is fast enough to be useful. failback immediate returns I/O to a recovered preferred path as soon as it comes back; a delayed failback is sometimes preferred when a flapping path would otherwise cause repeated disruption.

The timeout hierarchy that decides failover speed

Setting Scope Effect
fast_io_fail_tmo multipathd global or per-protocol Time before I/O fails fast on a broken path; overrides iSCSI recovery_tmo
dev_loss_tmo multipathd or iSCSI layer How long the SCSI layer keeps a failed device before removing it
replacement_timeout multipathd Globally overrides iSCSI recovery_tmo; lower priority than fast_io_fail_tmo
no_path_retry multipathd Number of retries before queuing is disabled and I/O errors surface

The behaviour that surprises people: every reload of multipathd resets recovery_tmo to the value of fast_io_fail_tmo on managed devices. So a single global change silently rewrites the iSCSI-layer timeout, and per-protocol overrides are the only way to treat FC and iSCSI differently in the same host.

overrides {
    dev_loss_tmo 60
    fast_io_fail_tmo 8
    protocol {
        type "scsi:iscsi"
        dev_loss_tmo 60
        fast_io_fail_tmo 120
    }
}

Verification and failover testing

multipathd show paths format "%d %P"
multipathd show config
multipath -ll
systemctl status multipathd

multipath -ll prints the path groups, the policy, the priority of each path and the state of each member — active, failed or ghost. A LUN showing all paths in one group with equal priority on an array that is actually active/passive is a misconfiguration that will surface as performance complaints rather than errors.

Test failover deliberately: with a light write load running, disable one path (unplug the cable, or administratively bring the switch port down), then confirm within seconds that the surviving path is active, that the write completes, and that no SCSI errors reached the application log. Re-enabling the path should return it to the preferred group without a manual intervention if failback is immediate. Do this for every path, not just the first.

Beyond the host

Multipath configuration only decides how the host uses the paths you built, so verify the fabric side too: redundant NICs on separate switches, separate VLANs or separate physical fabric where the design calls for it, and Jumbo frames if the array expects them. Arrays with their own preferred-path semantics — ONTAP with ALUA, FlashArray with its own host management — reward reading the vendor's recommended multipath settings rather than using the distribution default; see Pure Storage FlashArray host management and ONTAP disk management commands. Multipathing protects against path failure only; disk media failure and controller loss still need RAID or array-level redundancy, and local software RAID is discussed in Linux mdadm RAID build and recovery.

原文链接:https://docs.redhat.com/en/documentation/red_hat_enterprise_linux/10/html/configuring_device_mapper_multipath/modifying-the-dm-multipath-configuration-file