GlusterFS Replicated Volumes: Create, Heal and Split-Brain - 夜莺博客

GlusterFS Replicated Volumes: Create, Heal and Split-Brain

GlusterFS builds a scale-out filesystem from bricks — directories on ordinary servers — and its replicate translator keeps copies of every file across those bricks. The design is simple to deploy and complex to operate, because the interesting work is not creating the volume but keeping replicas converged. This guide covers creating replicated and arbiter volumes, inspecting heal state, and resolving the split-brain condition that eventually appears in every long-lived replicated volume.

Create the trusted storage pool, then the volume

gluster peer probe host2
gluster peer probe host3
gluster peer status

gluster volume create volname replica 3 \
    host1:/data/brick1/brick \
    host2:/data/brick1/brick \
    host3:/data/brick1/brick

gluster volume start volname
gluster volume info volname
gluster volume status volname

For two-way replication with a lightweight tie-breaker, use an arbiter volume: the third brick stores metadata only, so it costs far less capacity while still preventing split-brain by holding the quorum vote.

gluster volume create volname replica 2 arbiter 1 \
    host1:/data/brick1/brick \
    host2:/data/brick1/brick \
    host3:/data/brick1/brick

Arbiter is the right default for capacity-sensitive deployments, and it can be combined with a distributed-replicate layout so each subvolume gets its own arbiter vote.

Growing and shrinking a volume

gluster volume add-brick volname host4:/data/brick1/brick

gluster volume remove-brick volname host4:/data/brick1/brick start
gluster volume remove-brick volname host4:/data/brick1/brick status
gluster volume remove-brick volname host4:/data/brick1/brick commit

gluster volume replace-brick volname host2:/data/brick1/brick \
        host2:/data/brick2/brick commit force

remove-brick is a three-phase operation for a reason: start migrates data off the brick, status shows progress, and commit only then rewrites the volume definition. Committing early loses whatever had not migrated. With replica volumes, fix-layout is usually required after an expand.

Watching heal state

gluster volume heal volname info
gluster volume heal volname info summary
gluster volume heal volname info split-brain
gluster volume heal volname full

The self-heal daemon walks the .glusterfs/indices trees on each brick and repairs entries that differ. Entries in heal pending that never reach zero is the signature of an inconsistent volume: either a brick is unreachable, or a file is in split-brain and the daemon refuses to guess which copy is correct.

Resolving split-brain

Split-brain means two bricks both believe their copy is authoritative. Gluster will not pick a winner without an explicit policy, and that is deliberate — the wrong automatic choice silently destroys data.

gluster volume heal volname split-brain bigger-file <FILE>
gluster volume heal volname split-brain latest-mtime <FILE>
gluster volume heal volname split-brain source-brick host1:/data/brick1/brick <FILE>
gluster volume heal volname split-brain source-brick host1:/data/brick1/brick

For GFID split-brain the same policies apply, but the file must be referenced by its absolute path as seen from the mount point, and with source-brick you must run the command per file rather than once for the volume. Files can also be inspected and resolved from the mount with extended attributes:

getfattr -n replica.split-brain-status /mnt/gluster/data/file.dat
setfattr -n replica.split-brain-choice -v "brick3" /mnt/gluster/data/file.dat
setfattr -n replica.split-brain-heal-finalize -v "brick3" /mnt/gluster/data/file.dat

To stop the problem recurring, enable an automatic policy so future conflicts resolve without an operator at 3 a.m.:

gluster volume set volname cluster.favorite-child-policy mtime

Accepted values are size, ctime, mtime and majority. majority picks a copy whose size and mtime agree on more than half the bricks, which is the most defensible automatic choice.

Operational checklist

  • Alert on heal pending counts, not just brick availability.
  • Never force a brick back into a replica without checking its heal state first.
  • Keep arbiter bricks on independent hardware — an arbiter sharing a failure domain with a data brick removes the protection you paid for.
  • Compare approaches when planning: the object-storage alternative is described in Ceph OSD failure troubleshooting, and for single-node file servers the snapshot replication workflow in ZFS send and receive replication is often simpler to operate.

原文链接:https://docs.gluster.org/en/main/Administrator-Guide/Setting-Up-Volumes/