etcd Maintenance: Compaction, Defrag and Snapshot Restore - 夜莺博客

etcd Maintenance: Compaction, Defrag and Snapshot Restore

etcd is the source of truth for every Kubernetes cluster, and it is one of the few components whose failure mode is unrecoverable without a backup. Everything it stores is versioned, so the keyspace grows with every write and every update — which is why compaction, defragmentation and snapshot backups are not optional hygiene but the difference between a cluster you can recover and one you cannot. This guide covers each maintenance operation with the commands that perform it and the symptoms that tell you it is overdue.

Why etcd needs maintenance at all

Every write to etcd keeps history. Deleting a key does not remove the data; it adds a tombstone at a higher revision while the previous revision remains in the backend database. Two mechanisms guard the keyspace: compaction removes old revisions, and defragmentation releases the file-system space that compaction merely marked as free inside the database file.

If etcd runs out of space, the space quota trips and raises a cluster-wide alarm that puts the cluster into limited-operation mode: only key reads and deletes are accepted. In a Kubernetes cluster that means the API server can still read, but no object can be created or updated — a visible, total outage for anything that writes.

Compaction

# keep 1000 revisions
etcd --auto-compaction-mode=revision --auto-compaction-retention=1000

# keep one hour of history
etcd --auto-compaction-retention=1

With a retention period greater than one hour, etcd compacts every hour while maintaining the retention window; at one hour or less, it compacts at the retention interval. Manual compaction is also available:

etcdctl compact <revision>

Compaction does not shrink the file. It creates free space inside the backend database, which is exactly why the next step exists.

Defragmentation

etcdctl defrag
# Finished defragmenting etcd member[127.0.0.1:2379]

etcdctl defrag --cluster
# Finished defragmenting etcd member[http://127.0.0.1:2379]
# Finished defragmenting etcd member[http://127.0.0.1:22379]
# Finished defragmenting etcd member[http://127.0.0.1:32379]

Defragmentation blocks reads and writes on the member being defragmented, so never defragment every member at once in production. Roll through them one at a time and verify health between members. When etcd is stopped, the data directory can be defragmented offline:

etcdctl defrag --data-dir /var/lib/etcd

Clearing a NOSPACE alarm

The standard recovery sequence after a quota alarm is compact, then defrag, then disarm the alarm.

rev=$(etcdctl --endpoints=:2379 endpoint status --write-out="json" \
      | egrep -o '"revision":[0-9]*' | egrep -o '[0-9].*')
etcdctl compact $rev
etcdctl defrag
etcdctl alarm disarm
etcdctl put newkey 123        # confirms writes work again

Skipping the compaction step and disarming the alarm only postpones the outage: the backend file is still oversized and the quota trips again on the next few writes.

Snapshot backup — the part that saves you

etcdctl snapshot save backup.db
etcdutl --write-out=table snapshot status backup.db

+----------+----------+------------+------------+
|   HASH   | REVISION | TOTAL KEYS | TOTAL SIZE |
+----------+----------+------------+------------+
| fe01cf57 |       10 |          7 | 2.1 MB     |
+----------+----------+------------+------------+

Take snapshots on a schedule and verify each one with snapshot status. A snapshot that has never been status-checked is an assumption. Restoring is an offline operation: stop the etcd member, clear its data directory, and restore the snapshot into the data directory before starting it again — then bring members back one at a time and confirm the cluster elects a leader with a healthy endpoint status.

Monitoring signals that matter

  • etcd_server_quota_backend_bytes versus the actual backend file size — this ratio is your early warning.
  • DB size and DB size in use: a widening gap means defragmentation is overdue.
  • fsync and backend commit latency — rising values cause missed heartbeats and leader elections long before anything visibly breaks.
  • Alarm status. Scrape it; do not wait for a human to notice.

Defragmentation and compaction are also worth scheduling around your control-plane activity, since they compete with the API server for the same disk. If the cluster spans hosts, plan the maintenance in the same window as the virtual networking work described in kube-proxy iptables versus IPVS, and pair the snapshot routine with a tested node-recovery plan such as the one in Kubernetes CNI comparison.

原文链接:https://etcd.io/docs/v3.5/op-guide/maintenance/