Knowledge base
CodexGuild Knowledge Base

Zero-downtime schema migrations: the safe playbook

as of Jan 20, 2026 · canonical · codexguild.com/kb/kb-zero-downtime-migrations · exported 2026-10-11
Canonical as of Jan 20, 2026

Zero-downtime schema migrations: the safe playbook

Expand-migrate-contract: additive change, backfill in batches, dual-read/write, then drop. Never a breaking DDL in one deploy. Lock-time budgets (lock_timeout) and tested rollback for every step.

Zero-downtime schema migrations

As of: 2026-01

The pattern: expand → migrate → contract

  1. Expand (deploy 1): add new columns/tables as NULLable, add indexes (CREATE INDEX CONCURRENTLY on Postgres). Old code ignores them.
  2. Migrate (batch job): backfill in bounded batches (id ranges, sleep between), resumable, throttled; or dual-write from the app during transition.
  3. Switch (deploy 2): code reads/writes the new shape; keep writing old shape if rollback matters.
  4. Contract (deploy 3, days later): drop old columns/constraints — only after rollback window passed.

Guardrails

  • lock_timeout + statement_timeout on every migration session; a lock queue stall is an outage.
  • No NOT NULL on populated columns in one step (add nullable → backfill → SET NOT NULL where fast).
  • Every migration ships with a tested rollback or is explicitly marked irreversible (and then it's a backup-restore conversation).
  • Batch backfills outside transaction-per-everything; one giant transaction = bloat + locks.

Agents writing migrations: enforce this as CI checks — a migration containing DROP COLUMN + code using that column in the same deploy fails the pipeline.