The design is docs/architecture/23-ops.md section 8. This is the runbook
for the Docker Compose product. It is written to be followed by someone who
did not write it; if a step is unclear, that is a defect in this document.
What is protected, and how
Asset
How
Where
The database
WAL archived continuously, every 60 seconds at most, from the first boot
pgwal volume
The database
Base backups with pg_basebackup, daily by default (backup-scheduler)
pgbackup volume
The database
Encrypted copies of the base backups and the WAL, every five minutes (backup-offsite)
A separate store you name
Files
The files volume. Copy it with your host’s backup tool, or use versioned object storage
files volume
Secrets
docker/.env, above all QUIRE_MASTER_KEY (and any QUIRE_MASTER_KEY_RETIRED still in use), QUIRE_BACKUP_ENCRYPTION_KEY, and docker/secrets/audit-signing-key.pem
Keep a copy off this host
Search indexes, caches, renditions
Not backed up; rebuilt
Targets: a recovery point within 60 seconds of the failure, and a restore
within 60 minutes for a 500 GB database.
Two mistakes are common. A database restored without its files renders
broken pages. A database restored without QUIRE_MASTER_KEY cannot
decrypt the SSO, webhook and integration credentials it holds; until a master
key rotation finishes with nothing unresolved (key-rotation.md),
that includes the retired keys. Both are part of the backup.
Taking backups
A base backup of the whole cluster:
docker compose -f docker/compose.yaml --profile backup run --rm backup
It keeps the newest QUIRE_BACKUP_KEEP base backups (default 5) and prunes
WAL the oldest one no longer needs, so the archive cannot grow without bound.
Schedule it daily with cron or a systemd timer on the host:
15 2 * * * cd /srv/quire && docker compose -f docker/compose.yaml --profile backup run --rm backup >> /var/log/quire-backup.log 2>&1
Or let the stack schedule it: the backup profile runs backup-scheduler,
which takes a base backup every QUIRE_BACKUP_INTERVAL_HOURS (default 24),
and backup-offsite, described next.
docker compose -f docker/compose.yaml --profile backup up -d
Encrypted off-host copies
Both volumes live on the same host as the database, and a backup on the
machine that failed is not a backup. backup-offsite copies every base backup
and every archived WAL segment to a separate store through the storage port,
encrypted, and keeps them there under retention:
Encryption. AES-256-GCM with QUIRE_BACKUP_ENCRYPTION_KEY (or the file
named by QUIRE_BACKUP_ENCRYPTION_KEY_FILE): 32 bytes, from
openssl rand -hex 32. Each file has its own nonce and an authentication
tag, so a copy is unreadable without the key and any change to it is
detected. Keep the key with QUIRE_MASTER_KEY, away from this host and away
from the backup store. Without the key there is no restore.
Where.QUIRE_BACKUP_STORAGE_DRIVER is s3, azure or local (a
mounted remote disk at QUIRE_BACKUP_STORAGE_ROOT). The settings are the
file storage ones with a QUIRE_BACKUP_ prefix: QUIRE_BACKUP_S3_ENDPOINT,
QUIRE_BACKUP_S3_BUCKET, QUIRE_BACKUP_S3_ACCESS_KEY_ID, and so on. Use a
different bucket and, ideally, a different account from the files, with
credentials that can write but not delete if the provider allows it.
Retention. The newest QUIRE_BACKUP_OFFSITE_KEEP base backups (default
QUIRE_BACKUP_KEEP, else 7) and the WAL the oldest of them needs; older
sets and segments are deleted from the store.
When. Every QUIRE_BACKUP_SHIP_INTERVAL_SECONDS (default 300). Shipping
is idempotent: what is already stored is skipped, and a base backup counts
as stored only once its manifest is written, last.
The same command runs by hand:
docker compose -f docker/compose.yaml run --rm backup-offsite bun apps/worker/src/backups/main.ts shipdocker compose -f docker/compose.yaml run --rm backup-offsite bun apps/worker/src/backups/main.ts verify
To restore on a new host, bring a set back first, then follow the steps below
with the fetched directory in place of the pgbackup volume and the fetched
wal-archive in place of pgwal:
bun apps/worker/src/backups/main.ts fetch base-20260924T021500Z /srv/restore
Restoring to a point in time
Use this after data loss: a bad import, a deleted course, a contract
migration you need to undo. It replaces the live database, so rehearse it
with the drill below first.
Choose the target time, in UTC, just before the damage:
2026-09-24 09:30:00+00. The audit log (/admin/audit) usually shows the
moment.
Stop everything that writes:
docker compose -f docker/compose.yaml stop web content worker scheduler collab
Keep the damaged cluster until the restore is verified:
docker compose -f docker/compose.yaml stop postgresdocker run --rm -v quire_postgres18-data:/from -v quire_postgres-damaged:/to alpine cp -a /from/. /to/
Unpack the newest base backup older than the target into the data
volume, and ask for targeted recovery:
docker run --rm -v quire_pgbackup:/backups:ro -v quire_postgres18-data:/var/lib/postgresql postgres:18-alpine sh -euc ' base="$(ls -1d /backups/base-* | sort | tail -n 1)" # or the one before the target rm -rf /var/lib/postgresql/18/docker && mkdir -p /var/lib/postgresql/18/docker tar -xzf "$base/base.tar.gz" -C /var/lib/postgresql/18/docker touch /var/lib/postgresql/18/docker/recovery.signal chown -R postgres:postgres /var/lib/postgresql/18/docker && chmod 700 /var/lib/postgresql/18/docker'
Recover: start Postgres once with the recovery settings, from a
Compose override so the normal file is untouched:
The override replaces the whole command, so it repeats the two settings
recovery depends on: max_connections no lower than the primary’s
(otherwise recovery aborts with “insufficient parameter settings”) and
the mounted pg_hba.conf.
docker compose -f docker/compose.yaml -f docker/compose.recover.yaml up -d postgresdocker compose -f docker/compose.yaml logs -f postgres # wait for "database system is ready"
Check it before letting anyone in: the audit chain
(docker compose -f docker/compose.yaml run --rm worker bun tooling/audit-verify/run.ts),
and that the lost data is back.
Return to normal: docker compose -f docker/compose.yaml up -d. This
restarts Postgres with archiving on, and a new WAL timeline begins. Take a
fresh base backup straight away.
The project name quire prefixes each volume; docker volume ls shows the
exact names.
The scheduled verification drill
backup-offsite also runs a drill every QUIRE_BACKUP_DRILL_INTERVAL_HOURS
(default 168, weekly), and again at the next pass after one fails. It fetches
the newest off-host base backup and every WAL segment after it, decrypts each
(which proves the key still opens them and nothing was altered), compares each
file with its manifest, checks the archive is a Postgres data directory, and
checks the WAL from the backup onwards has no gap. The report is written to
the store as reports/drill-<time>.json and to the service’s log; a failed
drill names the file or the first missing segment.
The restore drill
A backup that has never been restored is not a backup. The drill restores the
newest complete base backup taken before the target, plus the WAL archive,
into a scratch Postgres that shares nothing with the live one, and proves the
result:
docker/scripts/restore-drill.sh # to ninety minutes agodocker/scripts/restore-drill.sh --target "2026-09-24 09:30:00+00"
The target is UTC in exactly that form. It needs a base backup older than it
and archived WAL beyond it: on a new install, take a base backup and wait for
the next archived segment (at most a minute with writes) before choosing a
target after the backup. The drill needs Docker and bash on the host, nothing
else.
Each step fails the drill:
Recoverability: the scratch cluster replays to the target and opens.
Completeness: every table’s row count against the live database
(tooling/restore-drill). The live database has moved on since the target,
so a table may differ by the larger of 500 rows and a tenth of its size, in
either direction (writes make it lag, deletions make the restore hold more);
a partition created after the target is not a lost table. Widen the
allowance on a busier install with QUIRE_DRILL_MAX_BEHIND and
QUIRE_DRILL_MAX_DRIFT_RATIO. A missing or emptied table fails.
Integrity: the audit hash chain verifies on the restored copy.
Usability: the application role reads through row-level security.
Time: start to green, against QUIRE_DRILL_RTO_SECONDS
(default 3600).
It never writes to the live database or its volumes: the backup and WAL
volumes are mounted read only and the scratch cluster is removed at the end,
pass or fail.
Set QUIRE_DRILL_REPORT to a path to have a JSON report written, pass or
fail, and run it on a schedule from the Docker host:
Run it monthly and before every upgrade. A failed drill blocks the upgrade.
Once a quarter, have someone who did not write this runbook perform a real
point-in-time restore on a spare host, using only this document.
Files
Local files live in the files volume. Back it up with the database, at the
same time, and restore both together:
docker run --rm -v quire_files:/files:ro -v "$PWD":/out alpine tar -czf /out/files-$(date -u +%Y%m%d).tar.gz -C /files .
With object storage, turn on bucket versioning and keep 35 days of
noncurrent versions; point-in-time recovery for files is then the bucket’s
own.