Skip to content

Upgrading without downtime

Upgrade a self-hosted Quire without downtime.

The rules are in docs/architecture/23-ops.md section 7 and docs/architecture/07-data.md section 4.1. This is the procedure.

The guarantee that makes it safe

Release R runs correctly against schema R and schema R minus one. Every schema change is split into expand, transition and contract:

  1. Expand: add the new column, table or index. Old code ignores it.
  2. Transition, for at least one release: new code writes both shapes and reads the new one; a resumable job backfills old rows.
  3. Contract: drop the old shape, in a later release, alone.

So at every moment of a rolling upgrade, old and new processes can share one database. There are no down migrations: a migration that dropped a column an hour ago cannot give back the rows written in that hour.

The schema-compat CI job checks the guarantee on every release by running the previous release’s tests against the new schema.

Before you start

  1. Read the release notes. A release that needs a maintenance window says so, with the estimate; at most one per release.
  2. Run the restore drill, or confirm it ran green for this release (backup-restore.md). A failed drill blocks the upgrade.
  3. Take a base backup: docker compose -f docker/compose.yaml --profile backup run --rm backup.

Docker Compose, one host

export QUIRE_RELEASE=2026.10.0            # or set it in docker/.env
docker compose -f docker/compose.yaml pull   # or build
docker compose -f docker/compose.yaml run --rm migrate
docker compose -f docker/compose.yaml up -d --no-deps web content collab
docker compose -f docker/compose.yaml up -d --no-deps worker scheduler

The order is deliberate:

  1. Migrate first, while the old release serves traffic. Expand migrations are invisible to it.
  2. Web tier next. On SIGTERM each web process flips /readyz to draining, finishes in-flight requests within 30 seconds, closes streams with a reconnect hint, and exits. stop_grace_period is 40 seconds so Compose never cuts a healthy drain.
  3. Workers last, so the newest event shape is produced before the newest consumer expects it. Workers stop fetching at once and get 120 seconds; a job that cannot finish is fetched again elsewhere, which is safe because every job is idempotent. The scheduler hands leadership over on its next tick.

On one host, Compose replaces each container in turn, so there is a short gap per service. For no gap at all, run the web tier as two containers behind your own proxy (an override file adding a second web service without a published port), and recreate them one at a time, waiting for each to report healthy before the next.

Several hosts or an orchestrator

Use the same order: migrate once from a single job, then roll the web tier with surge one and unavailable zero, then the workers. Point readiness probes at /readyz and liveness at /healthz.

With dedicated tenant databases, the migrate step does both: it migrates the control database first, then every database listed in ops.tenant_database, one at a time, each under its own lock. A failure in one tenant database does not stop the others. When every database is done it compares the migration ledgers and exits non-zero unless every database has applied exactly the migrations the control database has, naming each one that is behind or ahead. The same command installs the queue tables in each database, because the worker consumes a pinned tenant’s jobs where they were written.

bun apps/worker/src/migrate.ts   # what the Compose step runs
bun run db:migrate:all                            # the same, from a checkout

Each dedicated database is reached through the name it is registered under. A database registered as env:QUIRE_DB_NORTHWIND_URL needs:

Variable Used for
QUIRE_DB_NORTHWIND_URL The application role, for the web tier and the worker
QUIRE_DB_NORTHWIND_URL_MIGRATOR The migrator role, for this command and for moves
QUIRE_DB_NORTHWIND_URL_SUPERUSER Optional: reapplies the bootstrap (roles, schemas, helpers) before migrating

A registered database with no _MIGRATOR connection is reported as a failure, never skipped. The web tier can roll once the control database is done. A tenant database behind by an hour warns; by a day it pages.

pgvector

From migration 0264 the grounding corpus uses a pgvector HNSW index where the server has the extension; the Compose postgres service is built with it (docker/postgres.Dockerfile). The first migrate after switching images creates the extension through the superuser bootstrap, and 0264 then adds a generated vector column and builds the index. Adding the column rewrites app.ai_chunk once under an exclusive lock, so grounding requests wait for it; nothing else touches that table.

On a server without pgvector, 0264 logs a notice and changes nothing, and retrieval stays exact. With a pgvector older than 0.8 the column and index are built but retrieval stays exact until the extension is upgraded (alter extension vector update), because filtered HNSW scans need 0.8’s iterative scans. To enable it later on a server without it, install the extension, run the bootstrap again (or create extension vector as a superuser), then as quire_migrator:

set maintenance_work_mem = '1GB';  -- the HNSW build is much faster in memory
select ops.ai_chunk_enable_vector_index();

It is idempotent and returns enabled or unavailable. Run it on each dedicated tenant database as well.

Rolling back

Rolling back code is always available: set QUIRE_RELEASE to the previous tag and up -d again. That works because the schema is compatible in both directions within a release.

Rolling back schema is not offered. What cannot be undone, and how to recover from it:

Not reversible Recovery
A contract migration that dropped a column Point-in-time restore to before the drop into a new database, extract, merge
An in-place data change The same, then reconcile writes since
Sent webhooks and events Compensating events, never deletion
Sent email A human writes the follow-up
The audit hash chain Never rewritten; append a correction entry

That is why a contract migration ships alone: a restore then has a clean boundary.

Checking the upgrade

docker compose -f docker/compose.yaml ps           # every service healthy
curl -fsS http://localhost:8080/readyz             # ready, and what is configured
docker compose -f docker/compose.yaml logs migrate # the migrations applied
Navigation

Type to search…

↑↓ navigate↵ selectEsc close