Proxmox Backup Server: Prune, Garbage Collection and Verify - 夜莺博客

Proxmox Backup Server: Prune, Garbage Collection and Verify

Proxmox Backup Server is easy to install and easy to leave on default maintenance, which is how datastores fill up with data that was "restored" but never actually verified. Three scheduled jobs do the real work: prune decides which snapshots survive, garbage collection reclaims the disk space that pruned snapshots still reference, and verification proves the surviving chunks are readable. This guide covers all three, plus the maintenance mode you should use before touching the underlying storage.

The three jobs and what each one actually does

  • Prune removes snapshot metadata according to a retention rule set. It does not free meaningful space, because the chunks those snapshots referenced may still be needed by remaining snapshots.
  • Garbage collection (GC) is what frees space: it walks the datastore and deletes chunks no longer referenced by any snapshot.
  • Verify re-reads chunk data and compares it against stored checksums, detecting bit rot and silent corruption.

Because prune only unlinks metadata, a datastore that prunes aggressively but never runs GC will still report low free space. And because GC only removes unused chunks, running GC without prune just re-scans the same set.

Retention settings that make sense

PBS uses keep-* options, each counting backwards from the newest snapshot:

  • keep-daily: 13 — combined with keep-last, this guarantees at least two weeks of daily recovery points.
  • keep-weekly: 8 — two full months of weekly points.
  • keep-monthly: 11 — nearly a year of monthly boundaries.

Set retention on the datastore or per backup group, then create a prune job so it happens without anyone remembering. Prune jobs can be scoped to a single namespace.

Scheduling prune and garbage collection

proxmox-backup-manager prune-job list
proxmox-backup-manager garbage-collection start <datastore>
proxmox-backup-manager garbage-collection status <datastore>

For most deployments a weekly GC is the right interval — frequent enough to keep space predictable, infrequent enough not to hammer the disk. The schedule can be cleared when a datastore is archived or during maintenance:

proxmox-backup-manager datastore update <datastore> --delete gc-schedule

The GUI equivalent lives on the Prune & GC tab of each datastore, including a Start Garbage Collection button and a GC status view showing chunks removed and space reclaimed.

Verification: the job people skip

proxmox-backup-manager verify <datastore> --read-threads 1 --verify-threads 4 --ignore-verified false

Thread counts range from 1 to 32, with 1 reader and 4 verify threads as the documented default. Verify jobs run on a calendar schedule and can either skip already-verified snapshots or re-verify everything after a set period.

The re-verification advice from the PBS documentation is worth stating plainly: re-verify all backups at least monthly, even if a previous verification succeeded. Physical media degrades over time, so a backup that verified clean last year is not proof of anything today. A good pattern is an hourly or daily job that checks new snapshots, plus a weekly or monthly job that re-checks everything.

Notifications and maintenance mode

PBS can email results for scheduled verification, garbage collection and synchronisation tasks. By default notifications go to the address configured for root@pam, and you can set a different recipient per datastore. Route them to a mailbox someone reads; a silent verify failure is worse than no verify job, because it creates false confidence.

# read-only mode: blocks writes, allows reads and restores
proxmox-backup-manager datastore update <datastore> --maintenance-mode read-only

# offline mode: blocks both reads and writes
proxmox-backup-manager datastore update <datastore> --maintenance-mode offline

Use read-only to keep restores available while you work on the storage layer, and offline when the backing device must not be touched at all. Both modes are per datastore.

A maintenance routine worth copying

  1. Prune job daily; GC weekly; verify new snapshots daily and re-verify everything monthly.
  2. Alert if GC reports an unexpectedly large amount of reclaimed space — that usually means a backup group stopped running.
  3. Test a restore quarterly, into an isolated network, not just a file-level restore.
  4. Watch datastore growth against capacity and plan the next disk before you need it.
  5. Correlate with the hypervisor: quorum and cluster issues are covered in Proxmox VE quorum loss recovery, and for file-level retention patterns the approach in Restic forget and prune retention is a useful comparison.

原文链接:https://pbs.proxmox.com/docs/maintenance.html