Linux Software RAID with mdadm: Build and Recover - 夜莺博客

Linux Software RAID with mdadm: Build and Recover

Software RAID on Linux is more predictable than most hardware controllers, provided the configuration is persisted in the right two places. This guide walks the full lifecycle with mdadm: preparing GPT-partitioned devices, creating RAID 1/5/6/10 arrays, adding an internal write-intent bitmap so reboots do not trigger full resyncs, replacing a failed disk on a live array, and growing an array online. It also covers the two steps people skip - writing mdadm.conf and rebuilding the initramfs - which are the most common reason an array fails to assemble after a reboot.

Prepare the devices

apt install -y mdadm
parted /dev/sdb --script mklabel gpt
parted /dev/sdb --script mkpart primary 1MiB 100%
parted /dev/sdb --script set 1 raid on

Partition even when you could use the whole disk: it makes replacement easier, and it lets you trim a small amount of space so that a replacement drive from a different batch still fits. The raid partition flag is not strictly required for Linux md, but it stops other tooling from touching the device and is essential if a different OS ever sees the disk.

Create the array

mdadm --create /dev/md0 --level=6 --raid-devices=4 --bitmap=internal /dev/sd{b,c,d,e}1
cat /proc/mdstat

Useful level guidance: RAID 1 for OS and boot volumes, RAID 10 for latency-sensitive databases, RAID 6 for capacity-oriented data arrays where rebuild time is measured in hours. The internal bitmap is the highest-value flag in this entire article - it records which regions were being written, so after an unclean shutdown only those regions are resynced instead of the whole device.

Persist or lose it

mdadm --detail --scan | tee -a /etc/mdadm/mdadm.conf
update-initramfs -u        # Debian/Ubuntu
dracut --force             # RHEL/Fedora

Auto-detection exists but is fragile, especially when a disk has moved between controllers or the array UUID changed. Writing the scan output to mdadm.conf and regenerating the initramfs is what makes boot reliable. If the root filesystem sits on an md device and you skip this, a failed boot is a rescue-ISO event.

Replace a failed disk correctly

mdadm --detail /dev/md0            # look for faulty / removed
mdadm /dev/md0 --fail /dev/sde1
mdadm /dev/md0 --remove /dev/sde1
# swap the disk, then partition it identically
mdadm /dev/md0 --add /dev/sde1
watch cat /proc/mdstat

There is a better strategy for a disk that is showing errors but still in service: add the replacement first and use --replace ... --with .... The rebuild then happens while the old disk remains available, so you keep redundancy during the recovery window - which is exactly the period when a second failure is most likely.

Grow and monitor

mdadm /dev/md0 --add /dev/sdf1
mdadm --grow /dev/md0 --raid-devices=5
mdadm --grow /dev/md0 --size=max
echo check > /sys/block/md0/md/sync_action

Reshaping rewrites all data blocks and can run for a very long time; it is checkpointed, so a power loss is survivable, but it will resume. Schedule it deliberately, and cap rebuild throughput during business hours with echo 200000 > /proc/sys/dev/raid/speed_limit_max. Run monthly check scrubs and alert on any non-zero mismatch count.

Related: ZFS zpool and RAIDZ administration if you are choosing between md and ZFS, and Linux multipath configuration for the equivalent redundancy layer on SAN-attached disks.

原文链接:https://wiki.archlinux.org/title/Mdadm