Quick answer
Rolling back application code is unsafe when the older version cannot correctly read or write the database’s current schema or data. Keep database changes backward-compatible during the rollback window, migrate incrementally, and delay destructive cleanup. If compatibility has already been broken, consider a tested forward fix or feature bypass instead of automatically restoring the previous application version. (prisma.io)
Scope: This guide uses PostgreSQL 18 documentation and Kubernetes Deployments. The column-renaming example is illustrative, not a tested production runbook. Adapt it to your database version, migration tooling, application framework, and deployment controller.
1. Check the database before choosing a rollback target
A Kubernetes Deployment rollback restores an earlier Pod template. It does not reverse database migrations or restore database contents. Treat application rollback and database recovery as separate operations. (kubernetes.io)
Before approving rollback, check three compatibility boundaries:
- Schema: Does the target application still have the columns, tables, and constraints it expects?
- Data: Can it interpret records written by the newer release?
- Writes: Will its inserts and updates remain valid without discarding newer information?
A structurally compatible schema is not enough. For example, an older application may reject a newly introduced status value even though the column type has not changed. GitLab’s migration guidance similarly warns that changing data formats can break application code regardless of whether migration happens before or after deployment. (docs.gitlab.com)
Record the approved rollback target against the current database state, not merely against the last successful build. A release that previously worked is not necessarily usable now.
For incident triage, assemble:
- Application versions running across web processes and workers.
- Applied migrations and active backfills.
- Current feature-flag settings.
- The last compatibility-tested recovery target.
- Any destructive migration or irreversible data transformation.
Use that record to decide what can safely be reversed.
2. Preserve compatibility with expand-and-contract migrations
Expand-and-contract separates a database transition into stages:
- Expand: Add the new structure while retaining the old one.
- Migrate: Update clients and move existing data.
- Switch: Start using the new structure while preserving a recovery path.
- Contract: Remove the obsolete structure after its consumers are retired.
Prisma’s guidance describes introducing parallel structures, updating clients to write both representations, migrating existing data, validating the new interface, and only later removing the original structure. (prisma.io)
Backward compatibility must cover both reads and writes. Adding a nullable column may preserve existing application behavior; adding a required column without a usable default can invalidate older inserts. PostgreSQL documents these constraint and default behaviors. (postgresql.org)
Define a rollback floor: the oldest application release approved for the database’s present state. Move that floor deliberately as the migration progresses.
Do not bundle expansion, application cutover, and destructive cleanup into one recovery decision. Small, reversible changes and preparation for unsuccessful deployments are consistent with AWS operational-excellence guidance.
3. Worked example: replace a column without breaking older code
Assume an illustrative table:
customers (
id bigint PRIMARY KEY,
full_name text
)
The goal is to replace full_name with display_name, without changing its meaning.
The breaking shortcut is:
ALTER TABLE customers
RENAME COLUMN full_name TO display_name;
PostgreSQL renames the column without changing its stored data. However, application queries still using full_name no longer have that column available. GitLab specifically identifies direct column renames as a deployment compatibility hazard. (postgresql.org)
Use this staged plan instead:
| Stage | Application behavior | Recovery rule |
|---|---|---|
| Original | Reads and writes full_name |
Original release remains usable |
| Expansion | Both columns exist | Test the original release against the expanded schema |
| Bridge release | Reads full_name; writes both columns |
Use this as the later cutover’s recovery target |
| Read cutover | Reads display_name; still writes both |
Return reads to full_name if verified |
| New-only release | Reads and writes only display_name |
Retire older recovery targets |
| Cleanup | full_name removed |
Only new-only releases are eligible |
This is an application of expand-and-contract, not a universal migration recipe. (prisma.io)
First, expand the schema:
ALTER TABLE customers
ADD COLUMN display_name text;
Next, deploy the bridge release everywhere. For this example, require each name change to update both columns atomically in the same database transaction. Include API processes, scheduled tasks, importers, and queue workers—not just web Pods.
Then backfill existing rows. Only proceed once all relevant writers follow the bridge behavior. Otherwise, an old writer can update full_name while leaving an already-populated display_name stale.
An illustrative bounded update is:
UPDATE customers
SET display_name = full_name
WHERE id >= :lower_id
AND id < :upper_id
AND display_name IS DISTINCT FROM full_name;
The colon-prefixed values are parameters supplied by your migration tooling, not literal SQL to paste unchanged. Use short transactions and restartable progress tracking. GitLab recommends batched background migrations where copying data in a regular migration would take too long. (docs.gitlab.com)
Finally, validate and switch reads. Keep both writes active during the rollback window. Later, deploy a new-only release, retire incompatible recovery targets, and approve column removal separately.
If you temporarily return to old-only writers, treat the new column as potentially stale and reconcile it before another cutover.
4. Account for PostgreSQL locks—not just data rewrites
A migration can be logically compatible and still disrupt traffic.
PostgreSQL’s ALTER TABLE acquires an ACCESS EXCLUSIVE lock unless a particular operation documents a weaker lock. Adding an ordinary nullable column without a default does not require rewriting the table, but that does not make it lock-free. (postgresql.org)
An ACCESS EXCLUSIVE lock conflicts with every table-lock mode and blocks ordinary reads. Locks are normally held until the transaction ends. Keep schema-changing transactions short; do not place a long backfill in the same transaction as the column addition. (postgresql.org)
For a staging rehearsal, a migration transaction might use:
BEGIN;
SET LOCAL lock_timeout = '2s';
SET LOCAL statement_timeout = '10s';
ALTER TABLE customers
ADD COLUMN display_name text;
COMMIT;
These timeout values are illustrative, not recommended production thresholds. lock_timeout limits time spent waiting to acquire locks; statement_timeout limits statement duration. Set them according to your workload’s approved limits. (postgresql.org)
If the statement fails, roll back the failed transaction and inspect the database state before retrying. Avoid an uncontrolled retry loop.
Your rehearsal should include concurrent reads, writes, and a deliberately long transaction. Observe lock waits, request latency, database load, and migration failure handling.
5. Make release gates test recovery—not just deployment
A useful gate demonstrates that recovery works after the new release has written data.
For the column example, rehearse this sequence in staging:
- Apply expansion to a representative, safely handled dataset.
- Run the original application against the expanded schema.
- Exercise mixed original and bridge instances.
- Complete bridge deployment and backfill.
- Enable new reads and create or update records.
- Return to the bridge release or disable new reads.
- Confirm those records remain readable and editable.
Also test backfill interruption and restart, ORM-generated queries, delayed workers, and dependency handling.
Kubernetes rolling updates can run old and new ReplicaSets during the transition. A successful rollout is therefore not proof that every application/database combination was safe. (kubernetes.io)
Define explicit stop conditions before production deployment:
- A required compatibility test fails.
- Any writer still uses an unaccounted-for path.
- Backfill progress or data consistency is uncertain.
- Lock waits or user-facing errors exceed approved limits.
- The proposed recovery target needs a removed structure.
- Feature-flag behavior differs from the rehearsed behavior.
Attach the gate results to the release record. Record migration identifiers, recovery target, flag state, and the person authorized to approve progression.
6. Choose between rollback, feature bypass, and rolling forward
During an incident, prefer the least disruptive verified mitigation.
- Application rollback: Appropriate when the target release remains compatible with current schema and data.
- Feature bypass: Useful when a flag or runtime setting can avoid the failing behavior.
- Forward fix: Consider when the old application is incompatible but a narrowly scoped correction can restore service.
- Data recovery: Consider separately when corruption or destructive changes require restoration.
Microsoft’s incident-management guidance distinguishes rollback, fallback, feature bypass, and emergency deployment, and recommends predefined approval rules for mitigation. (learn.microsoft.com)
A feature flag changes application behavior; it does not recreate a dropped column or reconstruct transformed data. For this example, verify that disabling new reads leaves dual writes active and that every relevant process observes the intended setting.
A backup is not a routine deployment undo button. PostgreSQL point-in-time recovery reconstructs database state using a base backup and archived write-ahead logs. Its native physical recovery operates on the database cluster, not an individual migration or table. It also requires the necessary continuous WAL sequence. (postgresql.org)
Consequently, choosing a recovery point before a bad migration also excludes later legitimate transactions from that recovered state. Replaying farther forward may include the unwanted changes again.
Before restoration, document the proposed recovery point, acceptable data loss, restore-and-cutover procedure, write fencing, and reconciliation of external effects such as messages or payments. Provider-specific restoration behavior remains a separate check.
7. Verify recovery and postpone destructive cleanup
Do not close the incident because Pods are healthy. Verify the affected user operations and the database/application contract.
For the illustrative migration, check:
SELECT count(*) AS mismatched_rows
FROM customers
WHERE display_name IS DISTINCT FROM full_name;
Before new-only writes begin, this query checks the example’s intended equality invariant. It does not prove complete correctness. It scans the table, so budget its production cost or use a reviewed partitioned validation process.
Pair data checks with:
- Create, read, update, and delete tests for affected records.
- Database errors associated with the recovered application version.
- Lock waits, latency, and replication health.
- Queue processing and delayed-job behavior.
- Confirmation of applied migrations and active flags.
- Records written before, during, and after mitigation.
Use explicit closure criteria and document remaining work rather than treating deployment completion as incident resolution. (learn.microsoft.com)
Avoid three shortcuts:
- Automatically running “down” migrations: Reversing a schema operation does not necessarily reconstruct lost data.
- Checking only missing values: A populated replacement column can still contain stale values.
- Dropping the old column immediately: Cleanup removes the recovery path for clients that still depend on it. (prisma.io)
Keep expansion artifacts until the compatibility window ends. Remove obsolete flags, migration jobs, and columns through a separate approved change.
The recovery test is simple: Can this exact application version safely read and write the database as it exists now?