Skip to main content

Applying a retention policy

Delete old backup artifacts under a configured policy, having first confirmed exactly which artifacts that policy selects.

When to use this​

Use this when you are enforcing a policy by hand: after writing or narrowing one, when a storage backend is filling up, or when you want to see what the automatic sweep will do before it does it.

Do not use it to explore what a policy means. sentinel retention preview tells you what will be deleted, not why the rules combine the way they do; Retention explains that. If you want long-horizon calendar tiers rather than a simple "keep the last N", start at Configuring GFS retention tiers and come back here to apply it.

Deletion is permanent

sentinel retention apply removes backup artifacts from storage. There is no trash, no grace period, and no undo. Sentinel will not stop you deleting the only usable restore point for a database. Always run sentinel retention preview with the same --job first and read every line it prints.

Before you start​

  • A retention: block on the job, or under defaults:, with at least one of keep_last, keep_days, or a positive gfs: tier. A block containing none of those does nothing at all, and the job is skipped entirely by an all-jobs run.
  • Delete permission on the job's storage backend, not just read. Deletion is implemented for local, s3, gcs, and azure; see the failure notes below for google-drive.
  • A readable history database at history_db_path. Retention evaluates the recorded execution history, never a directory listing, so an artifact with no history row is invisible to it and will never be deleted.
  • An understanding of how the two flat rules combine, because it is the opposite of what most people expect. Read step 1 before you write a policy.

Steps​

1. Write the policy, knowing that the flat rules intersect​

version: "1.0"
history_db_path: ./.sentinel/history.db

defaults:
storage:
type: local
local_path: ./backups
retention:
keep_last: 7
keep_days: 30
caution
keep_last and keep_days narrow each other

Together these mean "at most 7, and none older than 30 days". They do not mean "whichever keeps more". Each rule contributes to a delete set, so a backup that either rule discards is deleted even when the other rule would have kept it. With keep_last: 7, the eighth most recent backup is deleted regardless of keep_days, even if it is an hour old.

This is issue #158. Sentinel's own README and operator runbooks describe these two rules as a union; they are wrong. Size keep_last for the retention you actually want, and treat keep_days as a further ceiling on age rather than as a floor on coverage.

Only the flat family and the GFS family combine as a union of keeps. Adding a gfs: block can only ever protect more; it can never make an existing keep_last delete more than it already did.

Two inheritance rules bite here. defaults.retention is inherited as a whole block, so a job declaring retention: { keep_last: 14 } gets exactly that and loses the keep_days: 30 from defaults. And a job whose only retention key is dry_run: true counts as declaring a policy, blocks the inheritance, and then fails validation with a message that says nothing about inheritance:

backup 'app-postgres': keep_last, keep_days, or gfs is required when retention is enabled

2. Preview, scoped to one job​

sentinel retention preview --config sentinel.yaml --job app-postgres

Each candidate is printed with its size and the rule that selected it:

retention preview for app-postgres
- ./backups/app-postgres_2026-06-11T02-00-00.sql (48210944 bytes) - exceeded keep_last
- ./backups/app-postgres_2026-05-28T02-00-00.sql (47993088 bytes) - exceeded keep_last, not retained by gfs

Preview reads history, computes candidates, prints them, and stops. It contacts no storage backend and modifies no database, so it is safe against production at any time.

This output goes to standard error

Retention prints its report on stderr, not stdout. sentinel retention preview ... > report.txt captures an empty file. Use 2>&1 > report.txt or ... > report.txt 2>&1 depending on which stream you want.

Anything absent from the list is being kept by something. Two exclusions are applied before printing: the single most recent successful backup is always retained, and so is the full baseline of the chain that the most recent backup belongs to.

Always pass --job. Without it the summary uses the word "deleted" in preview mode too, so sentinel retention preview --config sentinel.yaml prints total deleted: 3 backups while deleting nothing whatsoever. Only the single-job form labels its output retention preview for <job>.

3. Apply, scoped to the same job​

Once the preview matches your intent:

sentinel retention apply --config sentinel.yaml --job app-postgres

The same calculation runs, then each candidate's artifact is deleted from storage one at a time, and only if every artifact deletion succeeded are the matching history rows removed in a single transaction. That ordering is deliberate: a history row pointing at a missing file is recoverable, whereas a deleted row pointing at a surviving file leaves an artifact nothing will ever clean up again.

sentinel retention apply --dry-run is the same code path as preview, with the header reading retention apply for <job> instead.

note
retention.dry_run is honoured, since the fix for issue #157

A job carrying dry_run: true is never deleted from, by any path: not by sentinel retention apply, and not by the automatic sweep that runs after a scheduled backup.

The key and the --dry-run flag combine, and neither can cancel the other. Either one is enough to make a run report without deleting. Turning off a configured dry_run: true means editing the configuration file, which is deliberate: a safety switch that a flag could silently disable is not a safety switch.

On v1.4.0 and earlier the key was ignored and backups were deleted for real, including under the post-backup sweep. If you are running one of those versions, do not rely on this key; use sentinel retention preview instead. Fixed by issue #157.

4. Let the scheduler take over​

After each successful scheduled backup, Sentinel evaluates that one job's policy in real deleting mode, so a policy stays enforced without anyone running a command. A one-shot sentinel backup --config never triggers it. Retention failures there are warnings and do not turn a successful backup into a failed one, which also means a repeatedly failing sweep is easy to miss.

Verify​

Confirm the artifact count dropped and that the backend is still healthy:

sentinel storage status --config sentinel.yaml --output text

Confirm the history agrees with storage, since the two can diverge if a deletion failed partway:

sentinel monitor list --config sentinel.yaml --job app-postgres --last 90d

Every remaining row should name a file that still exists. Then re-run the preview: on a policy that has just been applied it should report nothing further to delete.

Finally, if the job produces incremental chains, confirm you did not cut through one. This is the failure that stays invisible until a restore. The check is named after a configured restore job rather than the backup job, and takes it as a positional argument:

sentinel restore validate-chain app-postgres-restore --config sentinel.yaml

If it goes wrong​

Google Drive retention. Deletion works for local, S3, GCS, Azure and Google Drive.

On v1.4.0 and earlier, Google Drive jobs failed at the deletion step with retention delete not supported for storage type 'google-drive', leaving both the artifacts and the history rows in place. The backend had always implemented deletion; only the retention path lacked a case for it. preview gave no hint either, since it returns before storage is consulted, so a policy looked configured, looked previewed, and enforced nothing. Fixed by issue #169.

preview now warns up front when a job's storage type cannot be deleted from, so this shape cannot recur quietly for a backend added later.

An all-jobs run reports errors and exits non-zero. Per-job failures are listed individually on stderr, under a line naming how many of the jobs with a policy failed, and the command exits 1. A cron entry wrapping it will alert.

On v1.4.0 and earlier this exited 0. Failures were summarised as a single retention completed with errors line with no detail and then discarded, so nothing noticed that retention had stopped working and the first symptom was a full disk. The --job path exited 1 on the same failure, so testing with --job showed correct behaviour and hid the difference. Fixed by issue #168.

A candidate is reported as refusing to report ... as deleted: no object at that path. Retention checks that an artifact exists before removing it, and reports the ones it could not find rather than counting them as deleted. Its history row is kept on purpose: either the object was removed out of band, in which case the row is stale and sentinel repair will say so, or the recorded path is wrong, in which case the row is the only evidence. Either way, nothing is lost by keeping it.

On v1.4.0 and earlier this was silent, and on GCS it was systematic. Deletion was reported from the delete call alone, which every backend implements idempotently, so an object that was never there counted as removed. The GCS backend compounded it by flattening prefixes when resolving the object, so backups/pg/dump.sql was addressed as dump.sql: retention deleted nothing, reported success, and dropped the history rows. Storage grew while every signal said it was being pruned. Fixed by issue #182.

Artifacts are gone but history still lists them. On v1.4.0 and earlier, one failed deletion skipped the history transaction for the entire job, including artifacts that had been removed. History rows are now cleared for exactly the artifacts whose deletion was confirmed, so a partial failure no longer leaves the rest inconsistent. Fix the underlying permission or path problem and re-run retention apply.

Manifests and sidecars are still there. Retention deletes the artifact path recorded in history and nothing else, so <artifact>.manifest.json survives, along with engine side artifacts such as an archived binary log tarball or an oplog archive. They are small but they accumulate, and a stranded manifest makes the storage listing misleading. This is issue #159. Remove them out of band.

A restore fails on a chain whose artifacts all appear to be present. Retention removed the baseline of a completed chain while its incrementals survived. Only the baseline of the chain currently being extended is protected. Keep keep_last comfortably larger than incremental_backup.max_chain_depth + 1 so a chain ages out as a unit.

A job is never swept. An all-jobs run skips any job whose policy would not act, which includes an absent, empty, or all-zero retention: block. Confirm the job's own block, remembering that any positive retention key on the job stops defaults.retention being inherited at all.

{/* sources: internal/domain/retention/policy.go, internal/domain/retention/types.go, internal/cli/retention.go, internal/cli/retention_helpers.go, internal/cli/retention_cleaner.go, internal/cli/backup.go, internal/adapters/monitor/retention.go, internal/config/types.go, internal/config/loader.go, internal/config/validator.go, docs/runbooks/apply-retention.md */}