Longhorn Distributed Storage for Kubernetes - 夜莺博客

Longhorn Distributed Storage for Kubernetes

Persistent storage is where most self-managed Kubernetes clusters get stuck, because the cloud-native answer (attach a cloud disk) is unavailable on bare metal. Longhorn fills that gap: it turns local disks on your nodes into replicated block storage for PersistentVolumes, with snapshots, backups to object storage, and a UI that makes recovery a few clicks. This guide covers the prerequisites that decide whether Longhorn will be reliable, StorageClass design, replica placement, and the failure behaviour you should test before trusting it.

What Longhorn is and is not

Is:     replicated block devices (iSCSI/NVMe-oF) for ReadWriteOnce volumes
        snapshots, backup to S3/NFS, offline rebuild, volume expansion
Is not: a filesystem, a POSIX shared-filesystem layer, or an HDFS replacement
ReadWriteMany  -> backed by NFS or share-manager, with its own trade-offs
Database-optimised performance: check latency, Longhorn adds a hop

For a shared-filesystem workload, or a database with strict latency requirements, evaluate alternatives (Ceph RBD, local NVMe with application-level replication) rather than forcing Longhorn into a shape it does not fit. Longhorn's strength is operational simplicity for RWO volumes.

Prerequisites

- Containerd or Docker with the open-iscsi client installed on EVERY node
  (apt install open-iscsi && systemctl enable --now iscsid)
- a dedicated, preferably SSD-backed filesystem at /var/lib/longhorn
- enough free space per node for replicas + snapshots
- nfs-common if you plan to use RWX volumes
- kernel modules: iscsi_tcp must be loadable
- MTU consistency across nodes for the storage network
# verify the prerequisites the installer checks anyway
kubectl create namespace longhorn-system
kubectl -n longhorn-system apply -f longhorn.yaml
kubectl -n longhorn-system get pods -w

Missing iscsid on one node is the number one installation failure: the pods run, volumes schedule, and then every volume that lands on that node fails to attach.

StorageClass design

kind: StorageClass
apiVersion: storage.k8s.io/v1
metadata: { name: longhorn-fast }
provisioner: driver.longhorn.io
allowVolumeExpansion: true
reclaimPolicy: Delete
parameters:
  numberOfReplicas: "3"
  staleReplicaTimeout: "30"
  dataLocality: "best-effort"      # keep a replica on the consuming node
  diskSelector: "ssd"
  nodeSelector: "storage"
  fsType: "ext4"
volumeBindingMode: WaitForFirstConsumer
StorageClass variants worth defining
  longhorn-fast  3 replicas, SSD nodes, dataLocality best-effort
  longhorn-bulk  2 replicas, HDD nodes, no locality (cheaper)
  longhorn-test  1 replica, no backup (never use for real data)

WaitForFirstConsumer is important: it lets the scheduler place the pod before choosing a node, otherwise a volume can be provisioned on a node that cannot run the pod.

Replica count and quorum

3 replicas  tolerates 1 node down (quorum maintained) -> default choice
2 replicas  tolerates 1 node down but no further failure during rebuild
1 replica   no redundancy; only for scratch
Rule: replicas should be an odd number and spread across failure domains.
Longhorn spreads replicas across nodes automatically; verify with kubectl.
kubectl -n longhorn-system get replicas.longhorn.io -o wide
kubectl -n longhorn-system get volumes.longhorn.io

Snapshots and backup

# snapshot = local, fast, same disk (survives accidental deletion, not disk loss)
# backup   = snapshot shipped to S3/NFS (survives node and cluster loss)

kubectl -n longhorn-system apply -f - <<'YAML'
apiVersion: longhorn.io/v1beta2
kind: BackupTarget
metadata: { name: default, namespace: longhorn-system }
spec:
  backupTargetURL: s3://my-bucket@ap-southeast-1/longhorn
  credentialSecret: longhorn-s3-creds
YAML

# recurring job: nightly backup, hourly snapshot, keep 7 backups
#   RecurringJob CR: task=snapshot / backup, cron, retain, labels

Without a BackupTarget, snapshots live on the same disk as the data — useful for rollback to a point in time, useless for hardware failure. Define the recurring jobs at install time; retrofitting backup after an incident is not a plan.

Failure behaviour worth testing

Scenario                      Expected outcome
One node reboots              volume stays available (3 replicas)
One node lost permanently     replica rebuilt onto a surviving node with space
Node out of space             replica fails; volume degrades but keeps serving
Two nodes lost at once        3-replica volumes become read-only/offline
Volume detached mid-write     fsck/ext4 journal replay on next attach
Upgrade of Longhorn           volumes stay attached; engine image rollout

The read-only transition when quorum is lost is intentional: it prevents split-brain writes. Size nodes so that rebuilding a replica after a failure does not fill the remaining disk — a full storage node turns one failure into two.

Performance notes

- dataLocality: best-effort removes a network hop for local reads
- do NOT enable dataLocality: strict with 3 replicas on a 3-node cluster
  (it forces the pod onto the replica node and hurts scheduling)
- test with fio, not with a database's own benchmarks
- storage network should be separate from the pod network if possible

Baseline the numbers before putting a database on it: fio benchmarking gives you the iodepth/latency picture, and the SAN comparison in iSCSI multipath tuning frames where Longhorn sits relative to traditional storage.

FAQ

Q: Can Longhorn back VMs? Yes — and it integrates with KubeVirt for VM disks, which is one of its most common uses.
Q: How does it differ from Ceph? Ceph is a general storage platform (block, file, object) with more tuning surface; Longhorn is deliberately narrower — RWO volumes managed from within Kubernetes. See Ceph placement groups for the Ceph side.
Q: What is the minimum cluster size? Three nodes for meaningful redundancy; single-node works but only for snapshot-based rollback.

原文链接:https://longhorn.io/docs/