Skip to content

Backup, point-in-time recovery and the restore drill

Back up Quire, restore it to a point in time, and prove it with the restore drill.

The design is docs/architecture/23-ops.md section 8. This is the runbook for the Docker Compose product. It is written to be followed by someone who did not write it; if a step is unclear, that is a defect in this document.

What is protected, and how

Asset How Where
The database WAL archived continuously, every 60 seconds at most, from the first boot pgwal volume
The database Base backups with pg_basebackup, daily by default (backup-scheduler) pgbackup volume
The database Encrypted copies of the base backups and the WAL, every five minutes (backup-offsite) A separate store you name
Files The files volume. Copy it with your host’s backup tool, or use versioned object storage files volume
Secrets docker/.env, above all QUIRE_MASTER_KEY (and any QUIRE_MASTER_KEY_RETIRED still in use), QUIRE_BACKUP_ENCRYPTION_KEY, and docker/secrets/audit-signing-key.pem Keep a copy off this host
Search indexes, caches, renditions Not backed up; rebuilt

Targets: a recovery point within 60 seconds of the failure, and a restore within 60 minutes for a 500 GB database.

Two mistakes are common. A database restored without its files renders broken pages. A database restored without QUIRE_MASTER_KEY cannot decrypt the SSO, webhook and integration credentials it holds; until a master key rotation finishes with nothing unresolved (key-rotation.md), that includes the retired keys. Both are part of the backup.

Taking backups

A base backup of the whole cluster:

docker compose -f docker/compose.yaml --profile backup run --rm backup

It keeps the newest QUIRE_BACKUP_KEEP base backups (default 5) and prunes WAL the oldest one no longer needs, so the archive cannot grow without bound. Schedule it daily with cron or a systemd timer on the host:

15 2 * * * cd /srv/quire && docker compose -f docker/compose.yaml --profile backup run --rm backup >> /var/log/quire-backup.log 2>&1

Or let the stack schedule it: the backup profile runs backup-scheduler, which takes a base backup every QUIRE_BACKUP_INTERVAL_HOURS (default 24), and backup-offsite, described next.

docker compose -f docker/compose.yaml --profile backup up -d

Encrypted off-host copies

Both volumes live on the same host as the database, and a backup on the machine that failed is not a backup. backup-offsite copies every base backup and every archived WAL segment to a separate store through the storage port, encrypted, and keeps them there under retention:

  • Encryption. AES-256-GCM with QUIRE_BACKUP_ENCRYPTION_KEY (or the file named by QUIRE_BACKUP_ENCRYPTION_KEY_FILE): 32 bytes, from openssl rand -hex 32. Each file has its own nonce and an authentication tag, so a copy is unreadable without the key and any change to it is detected. Keep the key with QUIRE_MASTER_KEY, away from this host and away from the backup store. Without the key there is no restore.
  • Where. QUIRE_BACKUP_STORAGE_DRIVER is s3, azure or local (a mounted remote disk at QUIRE_BACKUP_STORAGE_ROOT). The settings are the file storage ones with a QUIRE_BACKUP_ prefix: QUIRE_BACKUP_S3_ENDPOINT, QUIRE_BACKUP_S3_BUCKET, QUIRE_BACKUP_S3_ACCESS_KEY_ID, and so on. Use a different bucket and, ideally, a different account from the files, with credentials that can write but not delete if the provider allows it.
  • Retention. The newest QUIRE_BACKUP_OFFSITE_KEEP base backups (default QUIRE_BACKUP_KEEP, else 7) and the WAL the oldest of them needs; older sets and segments are deleted from the store.
  • When. Every QUIRE_BACKUP_SHIP_INTERVAL_SECONDS (default 300). Shipping is idempotent: what is already stored is skipped, and a base backup counts as stored only once its manifest is written, last.

The same command runs by hand:

docker compose -f docker/compose.yaml run --rm backup-offsite bun apps/worker/src/backups/main.ts ship
docker compose -f docker/compose.yaml run --rm backup-offsite bun apps/worker/src/backups/main.ts verify

To restore on a new host, bring a set back first, then follow the steps below with the fetched directory in place of the pgbackup volume and the fetched wal-archive in place of pgwal:

bun apps/worker/src/backups/main.ts fetch base-20260924T021500Z /srv/restore

Restoring to a point in time

Use this after data loss: a bad import, a deleted course, a contract migration you need to undo. It replaces the live database, so rehearse it with the drill below first.

  1. Choose the target time, in UTC, just before the damage: 2026-09-24 09:30:00+00. The audit log (/admin/audit) usually shows the moment.
  2. Stop everything that writes: docker compose -f docker/compose.yaml stop web content worker scheduler collab
  3. Keep the damaged cluster until the restore is verified:
    docker compose -f docker/compose.yaml stop postgres
    docker run --rm -v quire_postgres18-data:/from -v quire_postgres-damaged:/to alpine cp -a /from/. /to/
  4. Unpack the newest base backup older than the target into the data volume, and ask for targeted recovery:
    docker run --rm -v quire_pgbackup:/backups:ro -v quire_postgres18-data:/var/lib/postgresql postgres:18-alpine sh -euc '
      base="$(ls -1d /backups/base-* | sort | tail -n 1)"   # or the one before the target
      rm -rf /var/lib/postgresql/18/docker && mkdir -p /var/lib/postgresql/18/docker
      tar -xzf "$base/base.tar.gz" -C /var/lib/postgresql/18/docker
      touch /var/lib/postgresql/18/docker/recovery.signal
      chown -R postgres:postgres /var/lib/postgresql/18/docker && chmod 700 /var/lib/postgresql/18/docker'
  5. Recover: start Postgres once with the recovery settings, from a Compose override so the normal file is untouched:
    # docker/compose.recover.yaml
    services:
      postgres:
        command: [postgres, -c, "restore_command=cp /var/lib/postgresql/wal-archive/%f %p",
                  -c, "recovery_target_time=2026-09-24 09:30:00+00",
                  -c, recovery_target_action=promote, -c, archive_mode=off,
                  -c, max_connections=200, -c, hba_file=/etc/postgresql/pg_hba.conf]
    The override replaces the whole command, so it repeats the two settings recovery depends on: max_connections no lower than the primary’s (otherwise recovery aborts with “insufficient parameter settings”) and the mounted pg_hba.conf.
    docker compose -f docker/compose.yaml -f docker/compose.recover.yaml up -d postgres
    docker compose -f docker/compose.yaml logs -f postgres   # wait for "database system is ready"
  6. Check it before letting anyone in: the audit chain (docker compose -f docker/compose.yaml run --rm worker bun tooling/audit-verify/run.ts), and that the lost data is back.
  7. Return to normal: docker compose -f docker/compose.yaml up -d. This restarts Postgres with archiving on, and a new WAL timeline begins. Take a fresh base backup straight away.

The project name quire prefixes each volume; docker volume ls shows the exact names.

The scheduled verification drill

backup-offsite also runs a drill every QUIRE_BACKUP_DRILL_INTERVAL_HOURS (default 168, weekly), and again at the next pass after one fails. It fetches the newest off-host base backup and every WAL segment after it, decrypts each (which proves the key still opens them and nothing was altered), compares each file with its manifest, checks the archive is a Postgres data directory, and checks the WAL from the backup onwards has no gap. The report is written to the store as reports/drill-<time>.json and to the service’s log; a failed drill names the file or the first missing segment.

The restore drill

A backup that has never been restored is not a backup. The drill restores the newest complete base backup taken before the target, plus the WAL archive, into a scratch Postgres that shares nothing with the live one, and proves the result:

docker/scripts/restore-drill.sh                                  # to ninety minutes ago
docker/scripts/restore-drill.sh --target "2026-09-24 09:30:00+00"

The target is UTC in exactly that form. It needs a base backup older than it and archived WAL beyond it: on a new install, take a base backup and wait for the next archived segment (at most a minute with writes) before choosing a target after the backup. The drill needs Docker and bash on the host, nothing else.

Each step fails the drill:

  1. Recoverability: the scratch cluster replays to the target and opens.
  2. Completeness: every table’s row count against the live database (tooling/restore-drill). The live database has moved on since the target, so a table may differ by the larger of 500 rows and a tenth of its size, in either direction (writes make it lag, deletions make the restore hold more); a partition created after the target is not a lost table. Widen the allowance on a busier install with QUIRE_DRILL_MAX_BEHIND and QUIRE_DRILL_MAX_DRIFT_RATIO. A missing or emptied table fails.
  3. Integrity: the audit hash chain verifies on the restored copy.
  4. Usability: the application role reads through row-level security.
  5. Time: start to green, against QUIRE_DRILL_RTO_SECONDS (default 3600).

It never writes to the live database or its volumes: the backup and WAL volumes are mounted read only and the scratch cluster is removed at the end, pass or fail.

Set QUIRE_DRILL_REPORT to a path to have a JSON report written, pass or fail, and run it on a schedule from the Docker host:

30 3 1 * * cd /srv/quire && QUIRE_DRILL_REPORT=/var/log/quire-drill.json docker/scripts/restore-drill.sh >> /var/log/quire-drill.log 2>&1

Run it monthly and before every upgrade. A failed drill blocks the upgrade. Once a quarter, have someone who did not write this runbook perform a real point-in-time restore on a spare host, using only this document.

Files

Local files live in the files volume. Back it up with the database, at the same time, and restore both together:

docker run --rm -v quire_files:/files:ro -v "$PWD":/out alpine tar -czf /out/files-$(date -u +%Y%m%d).tar.gz -C /files .

With object storage, turn on bucket versioning and keep 35 days of noncurrent versions; point-in-time recovery for files is then the bucket’s own.

Navigation

Type to search…

↑↓ navigate↵ selectEsc close