restic Backup Retention: forget, prune and S3 Repos - 夜莺博客

restic Backup Retention: forget, prune and S3 Repos

Deduplicating, snapshot-based backup tools make it trivial to store terabytes of history and surprisingly easy to store it forever, because deleting a snapshot does not delete data. restic separates the two operations: forget removes snapshot references, prune removes the data that only those snapshots referenced. Getting retention right is therefore both a policy design question and a repository maintenance question - and the difference between a 200 GB and a 4 TB bucket.

Repository and snapshot model

export AWS_ACCESS_KEY_ID=...
export AWS_SECRET_ACCESS_KEY=...
restic -r s3:s3.eu-west-1.amazonaws.com/my-backups init
restic -r s3:s3.eu-west-1.amazonaws.com/my-backups backup --tag daily /etc /var/www
restic -r s3:s3.eu-west-1.amazonaws.com/my-backups snapshots

A snapshot is a reference to a tree of deduplicated blobs. Two consequences matter operationally: snapshots are cheap, and deleting one snapshot reclaims nothing until prune runs. Tagging backups (--tag daily, --tag pre-upgrade) is what makes later policy decisions possible.

Retention as a policy, not a list of IDs

restic forget --keep-daily 7 --keep-weekly 5 --keep-monthly 12 --keep-yearly 75 --prune
restic forget --keep-within 7d --keep-within-weekly 1m --keep-within-monthly 1y
restic forget --keep-last 10 --keep-tag important --dry-run

The --keep-* options describe what to keep; everything else in the group is removed. --keep-within style policies are better when backups are irregular, because "the last 7 days of snapshots" beats "the 7 most recent snapshots" when a nightly job failed for two nights. --dry-run prints the decision without touching anything - run it before every policy change.

Grouping is a safety feature

When a policy runs, restic first groups snapshots by host and paths, then applies the policy to each group independently. This is deliberate: without grouping, a retention policy aimed at web servers could delete unrelated database snapshots that happen to be older. Change the grouping only when you understand the consequences:

restic forget --group-by paths,tags --keep-daily 3
restic forget --group-by '' --keep-last 1     # applies policy to everything

Set the same --group-by on backup and forget, otherwise the grouping used at retention time will not match the grouping you intended when the snapshots were created.

Reclaiming space with prune

restic forget --keep-daily 7 --prune
restic prune --max-repack-size 0
restic check --read-data-subset=5%

prune needs read, write and delete access to the repository, downloads metadata, and can be slow on large repositories or high-latency S3 buckets. --max-repack-size 0 tells it not to repack packs for optimisation, which minimises scratch disk usage on a small server at the cost of a slightly less compact repository. Run prune on a schedule with enough time budgeted that it never overlaps the next backup.

Append-only repositories need different rules

If the repository is protected with append-only mode (for example S3 object lock or a restricted credential that can only create objects), forget and prune cannot run from the backed-up host - which is the point, because a compromised host must not be able to destroy history. Use a separate, well-secured administrative client for maintenance, and prefer --keep-within policies there: they keep legitimate snapshots even when an attacker has injected their own, so an attacker cannot use retention to wipe the good backups silently. Watch for more snapshots than usual or odd timestamps as an intrusion signal.

Operational checklist

  • Alert on backup age, not just on failure: a job that has not run in 36 hours is an incident.
  • Tag by purpose and keep critical pre-change snapshots with --keep-tag regardless of age.
  • Verify restores, not backups: restore a directory to a scratch path weekly and check --read-data-subset monthly.
  • Remember that restic refuses to act on an "empty" policy (--keep-last 0), an intentional guard against wiping a repository with one typo.
  • Store credentials outside the backed-up host, and never in the same place as the repository secret.

Related: ZFS send/recv snapshot replication, ONTAP SnapMirror break and resync policy, and SMART disk health monitoring.

原文链接:https://restic.readthedocs.io/en/latest/060_forget.html