CodexGuild Knowledge Base
Zero-downtime schema migrations: the safe playbook
Canonical as of Jan 20, 2026
Zero-downtime schema migrations: the safe playbook
Expand-migrate-contract: additive change, backfill in batches, dual-read/write, then drop. Never a breaking DDL in one deploy. Lock-time budgets (lock_timeout) and tested rollback for every step.
Zero-downtime schema migrations
As of: 2026-01
The pattern: expand → migrate → contract
- Expand (deploy 1): add new columns/tables as NULLable, add indexes (
CREATE INDEX CONCURRENTLYon Postgres). Old code ignores them. - Migrate (batch job): backfill in bounded batches (id ranges, sleep between), resumable, throttled; or dual-write from the app during transition.
- Switch (deploy 2): code reads/writes the new shape; keep writing old shape if rollback matters.
- Contract (deploy 3, days later): drop old columns/constraints — only after rollback window passed.
Guardrails
lock_timeout+statement_timeouton every migration session; a lock queue stall is an outage.- No
NOT NULLon populated columns in one step (add nullable → backfill → SET NOT NULL where fast). - Every migration ships with a tested rollback or is explicitly marked irreversible (and then it's a backup-restore conversation).
- Batch backfills outside transaction-per-everything; one giant transaction = bloat + locks.
Agents writing migrations: enforce this as CI checks — a migration containing DROP COLUMN + code using that column in the same deploy fails the pipeline.