Skip to content

Master key and signing key rotation, and break-glass access

Rotate the master key that protects stored credentials, and use break-glass access.

The controls in 21-compliance.md section 14 that an auditor asks for by name. This page is the procedure; the records it produces are the evidence.

The keys

Every stored credential is sealed with a fresh data encryption key (DEK). The DEK is wrapped by the master key (the KEK), and the reference of the master key is stored beside it (key_ref, or the reference inside a packed value). Rotating the master key re-wraps the DEKs. It never decrypts or re-encrypts a credential.

Setting Meaning
QUIRE_MASTER_KEY The current master key: 32 bytes, base64. Every new secret is wrapped under it
QUIRE_MASTER_KEY_VERSION Its version label. v1 when unset. Raise it whenever you change the key
QUIRE_MASTER_KEY_RETIRED Earlier keys that stored secrets may still be under, as v1=<base64>,v0=<base64>. Read, never written

The web tier, the worker and the bun run kek:rotate command read the same three settings. They must all have the same values, or one of them cannot open what another sealed.

Without QUIRE_MASTER_KEY each subsystem keeps the key it derives from QUIRE_SECRET_KEY. That works, the System health page shows it as degraded, and it stays readable after you set a master key, which is how the first rotation moves everything off it. Anyone who can read the process environment can decrypt every stored credential, so a production installation should have a master key, kept in a secret store and not in the same backup as the database.

Rotating

The key age shows on Platform console, Security, Master key, and as the metric quire.secrets.master_key.age (days). The daily platform.key_age schedule (03:41 UTC) writes a reminder entry to the platform audit chain when the key reaches 365 days, and again every 30 days until it is rotated. Rotate on the reminder, and whenever a key may have been exposed.

  1. Generate the new key: openssl rand -base64 32.
  2. Set QUIRE_MASTER_KEY to it and QUIRE_MASTER_KEY_VERSION to the next label (v2). Move the old key to QUIRE_MASTER_KEY_RETIRED as v1=<old base64>. Keep a copy of both somewhere other than this host.
  3. Deploy the web tier and the worker with the new settings. New secrets are now wrapped under env:QUIRE_MASTER_KEY:v2; old ones still open through the retired key.
  4. Request the rotation, with the reason that will sit in the audit trail:
    • in the console: Security, Master key, Rotate the master key; or
    • on a shell with the same environment: bun run kek:rotate request --reason "Annual rotation, ticket SEC-114".
  5. The worker re-wraps a slice a minute (the scheduler’s platform.key_rotation schedule) and resumes after a restart. To finish it in one sitting: bun run kek:rotate run. Watch it with bun run kek:rotate status.
  6. When the record shows the rotation completed with zero unresolved and zero failed, remove the retired key from QUIRE_MASTER_KEY_RETIRED and redeploy. Until then keep it: a value it could not move is still wrapped under the old key.

What the job walks

Every store that holds a wrapped DEK: the ones in SEALED_STORES (apps/worker/src/key-rotation.ts). Control database stores are walked on the control database; organisation stores are walked one organisation at a time under row-level security, in whichever database holds the organisation, so a tenant pinned to a dedicated database is rotated in that database. A test fails when the schema gains a wrapped-key column that the list does not name, and another when the credential review classifies a sealed column the list misses.

The record

  • ops.key_rotation: one row per rotation, with its reason, who requested it, its state and its totals (re-wrapped, already current, unresolved, failed).
  • ops.key_rotation_progress: one row per store and scope once walked, with the key references it could not read and how many values sat under each. A resumed rotation skips these.
  • Platform audit chain: platform/key_rotation_request (with the reason), one platform/key_rotation_store per store with its counts, and platform/key_rotation_complete or platform/key_rotation_fail; platform/key_age_reminder for the reminder.
  • Metrics: quire.secrets.master_key.age and quire.secrets.rewrap.outstanding (values the last rotation could not move).

When values are unresolved

An unresolved value is wrapped under a key reference this installation does not hold, or is not in the shape its column promises. The progress record names the reference (for example env:QUIRE_MASTER_KEY:v0 (unreadable)). Restore that key into QUIRE_MASTER_KEY_RETIRED and run another rotation, or, if the key is gone for good, have the administrator of the organisation enter the credential again: it is then sealed under the current key. Failed rotations show their error on the record; fix the cause and request again.

Signing keys

Separate from the master key: each organisation signs its OpenID Connect tokens and LTI messages with its own RSA key, published at /.well-known/jwks.json. Nothing here needs an operator. The hourly platform.signing_keys schedule publishes a successor seven days before the current key’s ninety are up, a week later the successor starts signing and the old key becomes retiring, and ninety days after that the old key is deleted and leaves the key set. Every step is a platform/signing_key_advance entry on the platform audit chain.

To replace an organisation’s key early, for example after an exposure:

  • in the console: Security, Master key, Publish a new signing key (needs platform/keys_manage); or
  • on a shell with the worker’s environment: bun run kek:rotate signing-keys rotate --tenant <slug or id> --reason "Key exposed, INC-3310". bun run kek:rotate signing-keys status lists every organisation’s keys by stage.

The new key is published at once and starts signing after seven days, when the current key retires. The week is deliberate: relying parties cache the key set, and a shorter overlap fails every tool at once. The retiring key stays in the key set for ninety more days so tokens it already signed keep verifying; if the exposure means it must stop being trusted sooner, deleting its row is a change made with the operator’s own database access under a change record (break-glass access is read only), and tokens signed with it then fail verification. The forced rotation is platform/signing_key_rotate in the audit chain, with the reason. The worker needs the same QUIRE_MASTER_KEY settings as the web tier to wrap the new key; bun run kek:rotate for the master key re-wraps signing keys with everything else (oauth_signing_key is in SEALED_STORES).

Break-glass production access

No one holds standing access to production. When something cannot wait, an owner issues a break-glass grant: Platform console, Security, Break-glass access.

  • A grant has a scope (one organisation, or the platform registry), a reason of at least 20 characters that names the incident or ticket, and a window of 5 to 240 minutes. It expires by itself: it is checked against the clock on every statement.
  • It can be issued to the owner who issues it, or to another owner (the two person form). Only the person it was issued to can use it. Issuing needs platform/break_glass_issue and using needs platform/break_glass_use; both are owner only by default.
  • Statements run through the gateway, not on a database login: read only, one at a time, confined to the organisation or to the control registry, with a five second timeout and at most 500 rows. Binary values are shown by size.
  • The platform audit chain records the issue (with the reason), the revocation, every statement before it runs (platform/break_glass_statement, refused ones with outcome denied) and every result (platform/break_glass_result). ops.break_glass_statement holds the audit entry ids, so the issuance record joins to its audit entries.
  • Writes are not offered. A change that cannot wait for a release uses the operator’s own database access under its own change record, outside this product, and the record should cite the incident reference used here.

Why not issue database credentials: a Postgres login outlives the session that asked for it, bypasses the row-level security the application relies on, and cannot write to this product’s audit chain, so its statements would be audited only as far as someone shipped the server log. The gateway makes the audit trail a property of the access rather than a practice around it.

To answer an audit request: list the grants in the period (Break-glass access), open a grant’s history for its statements and the audit entry ids, and read those entries on the platform audit chain (bun run audit:verify --platform proves the chain is intact).

Navigation

Type to search…

↑↓ navigate↵ selectEsc close