---
title: "Backup, point-in-time recovery and the restore drill"
description: "Back up Quire, restore it to a point in time, and prove it with the restore drill."
image: "https://docs.quirelms.com/og.png"
---

> Documentation Index
> Fetch the complete documentation index at: https://docs.quirelms.com/llms.txt
> Use this file to discover all available pages before exploring further.

# Backup, point-in-time recovery and the restore drill

<span id="backup-point-in-time-recovery-and-the-restore-drill"></span>

The design is `docs/architecture/23-ops.md` section 8. This is the runbook
for the Docker Compose product. It is written to be followed by someone who
did not write it; if a step is unclear, that is a defect in this document.

## What is protected, and how <!--quire:what-is-protected-and-how-->

| Asset | How | Where |
| --- | --- | --- |
| The database | WAL archived continuously, every 60 seconds at most, from the first boot | `pgwal` volume |
| The database | Base backups with `pg_basebackup`, daily by default (`backup-scheduler`) | `pgbackup` volume |
| The database | Encrypted copies of the base backups and the WAL, every five minutes (`backup-offsite`) | A separate store you name |
| Files | The `files` volume. Copy it with your host's backup tool, or use versioned object storage | `files` volume |
| Secrets | `docker/.env`, above all `QUIRE_MASTER_KEY` (and any `QUIRE_MASTER_KEY_RETIRED` still in use), `QUIRE_BACKUP_ENCRYPTION_KEY`, and `docker/secrets/audit-signing-key.pem` | Keep a copy off this host |
| Search indexes, caches, renditions | Not backed up; rebuilt | |

Targets: a recovery point within 60 seconds of the failure, and a restore
within 60 minutes for a 500 GB database.

Two mistakes are common. A database restored **without its files** renders
broken pages. A database restored **without `QUIRE_MASTER_KEY`** cannot
decrypt the SSO, webhook and integration credentials it holds; until a master
key rotation finishes with nothing unresolved ([key-rotation.md](/ops/key-rotation/)),
that includes the retired keys. Both are part of the backup.

## Taking backups <!--quire:taking-backups-->

A base backup of the whole cluster:

```sh
docker compose -f docker/compose.yaml --profile backup run --rm backup
```

It keeps the newest `QUIRE_BACKUP_KEEP` base backups (default 5) and prunes
WAL the oldest one no longer needs, so the archive cannot grow without bound.
Schedule it daily with cron or a systemd timer on the host:

```cron
15 2 * * * cd /srv/quire && docker compose -f docker/compose.yaml --profile backup run --rm backup >> /var/log/quire-backup.log 2>&1
```

Or let the stack schedule it: the `backup` profile runs `backup-scheduler`,
which takes a base backup every `QUIRE_BACKUP_INTERVAL_HOURS` (default 24),
and `backup-offsite`, described next.

```sh
docker compose -f docker/compose.yaml --profile backup up -d
```

## Encrypted off-host copies <!--quire:encrypted-off-host-copies-->

Both volumes live on the same host as the database, and a backup on the
machine that failed is not a backup. `backup-offsite` copies every base backup
and every archived WAL segment to a separate store through the storage port,
encrypted, and keeps them there under retention:

- **Encryption.** AES-256-GCM with `QUIRE_BACKUP_ENCRYPTION_KEY` (or the file
  named by `QUIRE_BACKUP_ENCRYPTION_KEY_FILE`): 32 bytes, from
  `openssl rand -hex 32`. Each file has its own nonce and an authentication
  tag, so a copy is unreadable without the key and any change to it is
  detected. Keep the key with `QUIRE_MASTER_KEY`, away from this host and away
  from the backup store. Without the key there is no restore.
- **Where.** `QUIRE_BACKUP_STORAGE_DRIVER` is `s3`, `azure` or `local` (a
  mounted remote disk at `QUIRE_BACKUP_STORAGE_ROOT`). The settings are the
  file storage ones with a `QUIRE_BACKUP_` prefix: `QUIRE_BACKUP_S3_ENDPOINT`,
  `QUIRE_BACKUP_S3_BUCKET`, `QUIRE_BACKUP_S3_ACCESS_KEY_ID`, and so on. Use a
  different bucket and, ideally, a different account from the files, with
  credentials that can write but not delete if the provider allows it.
- **Retention.** The newest `QUIRE_BACKUP_OFFSITE_KEEP` base backups (default
  `QUIRE_BACKUP_KEEP`, else 7) and the WAL the oldest of them needs; older
  sets and segments are deleted from the store.
- **When.** Every `QUIRE_BACKUP_SHIP_INTERVAL_SECONDS` (default 300). Shipping
  is idempotent: what is already stored is skipped, and a base backup counts
  as stored only once its manifest is written, last.

The same command runs by hand:

```sh
docker compose -f docker/compose.yaml run --rm backup-offsite bun apps/worker/src/backups/main.ts ship
docker compose -f docker/compose.yaml run --rm backup-offsite bun apps/worker/src/backups/main.ts verify
```

To restore on a new host, bring a set back first, then follow the steps below
with the fetched directory in place of the `pgbackup` volume and the fetched
`wal-archive` in place of `pgwal`:

```sh
bun apps/worker/src/backups/main.ts fetch base-20260924T021500Z /srv/restore
```

## Restoring to a point in time <!--quire:restoring-to-a-point-in-time-->

Use this after data loss: a bad import, a deleted course, a contract
migration you need to undo. It replaces the live database, so rehearse it
with the drill below first.

1. **Choose the target time**, in UTC, just before the damage:
   `2026-09-24 09:30:00+00`. The audit log (`/admin/audit`) usually shows the
   moment.
2. **Stop everything that writes**:
   `docker compose -f docker/compose.yaml stop web content worker scheduler collab`
3. **Keep the damaged cluster** until the restore is verified:
   ```sh
   docker compose -f docker/compose.yaml stop postgres
   docker run --rm -v quire_postgres18-data:/from -v quire_postgres-damaged:/to alpine cp -a /from/. /to/
   ```
4. **Unpack the newest base backup older than the target** into the data
   volume, and ask for targeted recovery:
   ```sh
   docker run --rm -v quire_pgbackup:/backups:ro -v quire_postgres18-data:/var/lib/postgresql postgres:18-alpine sh -euc '
     base="$(ls -1d /backups/base-* | sort | tail -n 1)"   # or the one before the target
     rm -rf /var/lib/postgresql/18/docker && mkdir -p /var/lib/postgresql/18/docker
     tar -xzf "$base/base.tar.gz" -C /var/lib/postgresql/18/docker
     touch /var/lib/postgresql/18/docker/recovery.signal
     chown -R postgres:postgres /var/lib/postgresql/18/docker && chmod 700 /var/lib/postgresql/18/docker'
   ```
5. **Recover**: start Postgres once with the recovery settings, from a
   Compose override so the normal file is untouched:
   ```yaml
   # docker/compose.recover.yaml
   services:
     postgres:
       command: [postgres, -c, "restore_command=cp /var/lib/postgresql/wal-archive/%f %p",
                 -c, "recovery_target_time=2026-09-24 09:30:00+00",
                 -c, recovery_target_action=promote, -c, archive_mode=off,
                 -c, max_connections=200, -c, hba_file=/etc/postgresql/pg_hba.conf]
   ```
   The override replaces the whole command, so it repeats the two settings
   recovery depends on: `max_connections` no lower than the primary's
   (otherwise recovery aborts with "insufficient parameter settings") and
   the mounted `pg_hba.conf`.
   ```sh
   docker compose -f docker/compose.yaml -f docker/compose.recover.yaml up -d postgres
   docker compose -f docker/compose.yaml logs -f postgres   # wait for "database system is ready"
   ```
6. **Check it** before letting anyone in: the audit chain
   (`docker compose -f docker/compose.yaml run --rm worker bun tooling/audit-verify/run.ts`),
   and that the lost data is back.
7. **Return to normal**: `docker compose -f docker/compose.yaml up -d`. This
   restarts Postgres with archiving on, and a new WAL timeline begins. Take a
   fresh base backup straight away.

The project name `quire` prefixes each volume; `docker volume ls` shows the
exact names.

## The scheduled verification drill <!--quire:the-scheduled-verification-drill-->

`backup-offsite` also runs a drill every `QUIRE_BACKUP_DRILL_INTERVAL_HOURS`
(default 168, weekly), and again at the next pass after one fails. It fetches
the newest off-host base backup and every WAL segment after it, decrypts each
(which proves the key still opens them and nothing was altered), compares each
file with its manifest, checks the archive is a Postgres data directory, and
checks the WAL from the backup onwards has no gap. The report is written to
the store as `reports/drill-<time>.json` and to the service's log; a failed
drill names the file or the first missing segment.

## The restore drill <!--quire:the-restore-drill-->

A backup that has never been restored is not a backup. The drill restores the
newest complete base backup taken before the target, plus the WAL archive,
into a scratch Postgres that shares nothing with the live one, and proves the
result:

```sh
docker/scripts/restore-drill.sh                                  # to ninety minutes ago
docker/scripts/restore-drill.sh --target "2026-09-24 09:30:00+00"
```

The target is UTC in exactly that form. It needs a base backup older than it
and archived WAL beyond it: on a new install, take a base backup and wait for
the next archived segment (at most a minute with writes) before choosing a
target after the backup. The drill needs Docker and bash on the host, nothing
else.

Each step fails the drill:

1. **Recoverability**: the scratch cluster replays to the target and opens.
2. **Completeness**: every table's row count against the live database
   (`tooling/restore-drill`). The live database has moved on since the target,
   so a table may differ by the larger of 500 rows and a tenth of its size, in
   either direction (writes make it lag, deletions make the restore hold more);
   a partition created after the target is not a lost table. Widen the
   allowance on a busier install with `QUIRE_DRILL_MAX_BEHIND` and
   `QUIRE_DRILL_MAX_DRIFT_RATIO`. A missing or emptied table fails.
3. **Integrity**: the audit hash chain verifies on the restored copy.
4. **Usability**: the application role reads through row-level security.
5. **Time**: start to green, against `QUIRE_DRILL_RTO_SECONDS`
   (default 3600).

It never writes to the live database or its volumes: the backup and WAL
volumes are mounted read only and the scratch cluster is removed at the end,
pass or fail.

Set `QUIRE_DRILL_REPORT` to a path to have a JSON report written, pass or
fail, and run it on a schedule from the Docker host:

```cron
30 3 1 * * cd /srv/quire && QUIRE_DRILL_REPORT=/var/log/quire-drill.json docker/scripts/restore-drill.sh >> /var/log/quire-drill.log 2>&1
```

Run it monthly and before every upgrade. A failed drill blocks the upgrade.
Once a quarter, have someone who did not write this runbook perform a real
point-in-time restore on a spare host, using only this document.

## Files <!--quire:files-->

Local files live in the `files` volume. Back it up with the database, at the
same time, and restore both together:

```sh
docker run --rm -v quire_files:/files:ro -v "$PWD":/out alpine tar -czf /out/files-$(date -u +%Y%m%d).tar.gz -C /files .
```

With object storage, turn on bucket versioning and keep 35 days of
noncurrent versions; point-in-time recovery for files is then the bucket's
own.

Source: https://docs.quirelms.com/ops/backup-restore/index.mdx
