MinIO Distributed Erasure-Coded Cluster Deployment - 夜莺博客

MinIO Distributed Erasure-Coded Cluster Deployment

A single-node MinIO is an S3 server on top of a filesystem: when the volume fills or the process dies, the bucket is unavailable. Distributed mode is structurally different — MinIO is launched with a list of hosts and drives, computes an erasure-coded layout (default EC:4 across a 16-drive set, meaning 12 data plus 4 parity), and spreads every object across nodes so the cluster tolerates drive and node failures. This guide covers a four-node cluster, which is the practical floor because three nodes give you EC:2 with no headroom for a second failure during recovery.

Prepare the Drives and Hosts

# on each of the four nodes: dedicate the drives, XFS is the recommended filesystem
mkfs.xfs -f -L minio-disk1 /dev/vdb
mkdir -p /mnt/disk1
echo "UUID=$(blkid -s UUID -o value /dev/vdb) /mnt/disk1 xfs defaults,noatime,nodiratime,inode64 0 2" >> /etc/fstab
mount -a

# resolution: every node must be able to resolve every other node
cat >> /etc/hosts <<'EOF'
10.20.0.11 node1
10.20.0.12 node2
10.20.0.13 node3
10.20.0.14 node4
EOF

Drives must be capable of serving on their own, and the directories must be fresh — MinIO refuses to start distributed mode on a directory that already contains foreign data. Keep the cluster nodes less than 15 minutes apart in time (run NTP) or erasure sets behave oddly after a clock jump.

Environment File and systemd Unit

### /etc/default/minio
MINIO_ROOT_USER="admin"
MINIO_ROOT_PASSWORD="CHANGE_ME_LONG_RANDOM_64CHAR"
MINIO_VOLUMES="https://node{1...4}.example.com/mnt/disk{1...4}"
MINIO_OPTS="--address :9000 --console-address :9001 --certs-dir /etc/minio/certs"
MINIO_SERVER_URL="https://s3.example.com"
MINIO_BROWSER_REDIRECT_URL="https://console.s3.example.com"
MINIO_API_REQUESTS_MAX="1600"
MINIO_API_REQUESTS_DEADLINE="10s"

### /etc/systemd/system/minio.service
[Unit]
Description=MinIO
After=network-online.target
Wants=network-online.target

[Service]
Type=notify
User=minio-user
Group=minio-user
EnvironmentFile=-/etc/default/minio
ExecStart=/usr/local/bin/minio server $MINIO_OPTS $MINIO_VOLUMES
Restart=always
LimitNOFILE=1048576
LimitNPROC=65535
TimeoutStopSec=120
SendSIGKILL=no
OOMScoreAdjust=-1000

[Install]
WantedBy=multi-user.target

The brace syntax matters: {1...4} with three dots is expanded by MinIO, while {1..4} with two dots is expanded by your shell — the latter changes the erasure-set ordering and therefore the fault tolerance. Start the service on all nodes; until every peer is up, the cluster reports itself as not quorate, which is expected.

Load Balancer in Front

frontend s3
  bind *:443 ssl crt /etc/haproxy/certs/s3.pem
  default_backend minio_servers

backend minio_servers
  balance leastconn
  option httpchk GET /minio/health/cluster
  http-check expect status 200
  default-server inter 5s fall 3 rise 2 ssl verify none
  server node1 10.20.0.11:9000 check
  server node2 10.20.0.12:9000 check
  server node3 10.20.0.13:9000 check
  server node4 10.20.0.14:9000 check

Health-check the /minio/health/cluster endpoint rather than /minio/health/live: the cluster endpoint returns 200 only when the node is part of a quorate cluster, which is the correct signal for routing client traffic.

Verify, Heal and Expand

mc alias set local https://s3.example.com $MINIO_ROOT_USER $MINIO_ROOT_PASSWORD
mc admin info local
mc mb local/app-uploads
mc version enable local/app-uploads
mc ilm rule add --noncurrent-expire-days 30 local/app-uploads

# replace a failed drive, then heal its data
mc admin heal start local/
mc admin heal status local/
  • Failover drill: stop MinIO on one node and confirm reads and writes continue; if the bucket goes offline, the endpoint list on the surviving nodes is wrong.
  • Expansion: add a new server pool to the same command line (... http://node{5...8}/mnt/disk{1...4}). Each new pool must use the same erasure-coding parity as the original, and new objects are placed in proportion to free space per pool.
  • Node replacement: reuse the same hostname and IP, mount fresh empty drives, and start MinIO with the same MINIO_VOLUMES value — the cluster recognises the gap and heals it.

Related reading: Ceph placement groups and autoscaler and ZFS zpool and RAIDZ vdev administration.

原文链接:https://github.com/minio/minio/blob/master/docs/distributed/README.md