Database Migrations That Break Rollback as an Option
Some migrations destroy your ability to roll back the moment they execute.

Every engineer who has shipped a schema change carries the same assumption into the deploy: if something goes wrong, the team can always roll back. That assumption is structurally wrong for a specific and identifiable category of migrations, not merely risky. The mechanism has nothing to do with bugs in the migration script or a slow database engine. A migration can execute exactly as written, pass every test, and still destroy the rollback option the moment it mutates, deletes, or re-semanticizes data that the previous version of the application cannot interpret. At that point, redeploying the old binary does not restore safety. It connects a known-good application to a dataset it no longer understands.
The Medium postmortem illustrates how quietly this happens. Staging held a fraction of the rows production carried, and the migration that looked instant and harmless at that scale locked the production users table for more than four hours once it ran against the real dataset. Nobody decided to eliminate rollback. The option disappeared the moment the migration began executing, long before anyone in the incident channel understood what had happened.
What "rollback" actually means, and why the word covers four different operations
Before identifying which migrations break rollback, the term itself needs a precise definition, because "rollback" is routinely used to describe four distinct operations, each with its own risk profile and its own authorization requirements. OneUptime's framework separates traffic rollback, code rollback, schema rollback, and data restoration. Traffic rollback sends requests back to a previous application version. Code rollback redeploys a known application artifact. Schema rollback reverses a database definition change. Data restoration replaces current data with an earlier copy or a point-in-time snapshot.
These four operations do not compose safely by default, and treating them as interchangeable is where most recovery plans fail before they are even tested. Data restoration discards every legitimate write made after the restore point. A team invoking it is trading an outage for a smaller, quieter data-loss event. Reversing a schema does not reverse a data transformation that already ran against it, so a table can look structurally correct while the values inside it remain in a new and incompatible format. A traffic rollback can fail outright if the old code hits rows the new binary already wrote in a shape the old code cannot parse.
A team's incident runbook should specify, for each release, which of these four operations is permitted, what data loss that operation authorizes, and who has the standing to invoke it. Most runbooks specify none of this.
The operations that eliminate rollback: a taxonomy of irreversibility
Schema changes are not uniformly risky. Some leave the previous application version fully functional against the new database state, and some do not, and the ones that break compatibility share a common structure: they change what stored data means to the application, not merely what the schema definition looks like on paper.
Destructive constraint enforcement is the first category. Adding a NOT NULL check or a similar constraint to an existing column forces the database to scan and validate every row, and on a large table that operation can hold locks for hours rather than seconds. The Medium postmortem shows what happens when this goes wrong in the worst way: the lock did not release even after the session that started the migration was killed, and the migration continued running in a ghost state with no clean abort path. Mid-migration rollback simply is not available in this scenario. An ALTER TABLE running against tens of millions of rows cannot be interrupted with a Ctrl+C and leave the table in a usable condition.
Dropping columns or tables is the second category, and it is the most absolute. Any old code that still references a dropped column fails immediately, with no graceful degradation path available to it. OneUptime's safe precondition for executing a drop is proof that no old readers or writers remain active anywhere in production, including in whatever systems a rollback plan might still reach, and almost no team verifies that condition formally before running the drop. Nothing short of a full data restoration can reopen it.
In-place data transformations that rewrite stored values form the third category, and they are the hardest for teams to recognize as dangerous, because the schema itself may look unchanged. When existing rows are rewritten into a new representation, old code may no longer understand the new format even if the column names never changed. Vadim Kravcenko's account of splitting a single name column into first_name and last_name fields makes the mechanism concrete: once rows had already been backfilled into the new structure, any request that slipped through during an attempted rollback produced records with half-missing names, and there was no easy way to reconstruct the original values. A schema rollback in this situation restores the table definition. It cannot undo the data the new binary already wrote in the new shape, because the information needed to reverse that transformation was discarded the moment it ran.
Adding a required column without a safe default is the fourth category. Old writers have no way to supply a value for a column they don't know exists, so any write attempted by the old binary violates the new constraint and fails. OneUptime's compatibility guidance lays out the only safe shape for this kind of change: add the column as nullable or with a default, backfill the existing rows, and only then enforce the constraint as a separate step. Skipping any one of those three steps eliminates rollback for whichever binary version is still running the old code path.
Column renames and type changes make up the fifth category. During a rolling deploy, old and new binaries are querying the same database simultaneously, and a rename with no coexistence period breaks one version or the other the moment it takes effect. Type changes carry an added layer of unpredictability, because conversion semantics, index invalidation, and replication behavior all vary by database engine and by version, so the locking and data-loss risk of a given type change cannot be read off the DDL statement alone.
Why the rollback window closes faster than teams expect
Teams tend to reason about rollback as though it remains available right up until someone decides to abandon it. In practice, the new binary, along with every background worker and asynchronous consumer attached to the same database, begins writing new-format data from the moment the migration completes, and each of those writes narrows the window a little further.
The first compounding factor is the binary-schema compatibility window inherent to any rolling deploy. New code and old code are querying the same database at the same time while containers cycle through the rollout, and a migration that is perfectly forward-compatible with the new binary can break the old binary's writes immediately. Kravcenko frames this as three simultaneous worlds every migration script has to tolerate at once: the upgrade path, where new and old code hit the same database while containers roll; the downgrade path, which tends to happen at three in the morning with tired humans making the call; and the awkward middle, where dual writes, ghost tables, and backfills are all running in the background at the same time. Most bugs hide in that middle state.
The second compounding factor is that asynchronous consumers extend the exposure window well past what the request-serving fleet suggests. Queue consumers, scheduled jobs, exports, read replicas, and disaster recovery environments can all continue running older code long after the primary application tier has finished its rollout, so a schema change that looks completely safe from the perspective of the request-serving fleet can still break a background worker that nobody was watching. Running a contract step, such as dropping the old column, before those async consumers have fully drained eliminates rollback for those paths even while the primary tier looks healthy.
The third compounding factor is that shared infrastructure multiplies the blast radius once point-in-time restore becomes the only remaining option. The perrotta.dev case from September 2026 shows how far this can go: pinning the old container image back through GitOps was not sufficient, because the old binary refused to even start against the new schema, and the only recovery path left was an AWS RDS point-in-time restore. Several unrelated applications shared that same RDS instance, so an in-place restore would have rolled back every one of their databases, not just the one that had actually failed. Recovery instead required standing up a new instance, restoring to it, dumping only the affected database, renaming the live database out of the way, and restoring from that dump, a sequence that kept rollback two ALTER DATABASE statements away but cost everything written after the restore point.
The fourth compounding factor is that a single root cause generates many downstream alerts, which makes the clock harder to read at the exact moment it matters most. When a database locks up or a schema change breaks compatibility, the signal an on-call engineer sees is never one clean alert. It's connection timeouts across multiple services, server-error spikes across multiple endpoints, rising queue depth, failing health checks, and latency SLO breaches, all firing at once. Diagnosing a schema-caused outage from that noise, rather than from a direct read on what the migration actually did, costs time that the rollback window does not have to spare.
Inside a Real Migration Failure
The taxonomy above becomes concrete only against a real incident, where the decisions that seemed reasonable at the time, the exact moment the window closed, and the absence of any clean way back are all on the record.
The Medium postmortem from April 2026 is the clearest documented account available. The migration in question added a CHECK constraint, ALTER TABLE users ADD CONSTRAINT users_phone_number_check CHECK (phone_number IS NOT NULL), to a very large production users table. Every row had already been backfilled with a phone number; the team had verified zero nulls. Staging testing, run against a dataset a fraction of production's size, completed the migration in roughly 30 seconds, and the team reasonably concluded the change was safe to ship. Production held 84,000,000 rows against staging's 10,000, and the migration locked the users table for 4 hours and 12 minutes. Over that window, 12.4 million users could not log in, and the checkout API timed out on every single request. The lock did not release even after the session driving the migration was killed; the operation continued running in a ghost state with no clean point at which to abort it. No mid-migration rollback existed, because an ALTER TABLE of that size cannot be interrupted and left in a usable state. The rollback window had closed the instant the command was issued, four hours before the team recognized what was happening.
A GitLab incident from September 2025 corroborates the same structural failure in a different form. A migration script, 20250812153148_remove_fk_from_mrdc.rb, failed to acquire a lock on a busy table and failed outright. That single failed migration then blocked every subsequent deployment and every other pending migration until engineers marked it as executed, a workaround that deferred the underlying problem rather than resolving it. GitLab's actual fix was to write a replacement migration with improved lock-retry logic.
Both incidents share the same underlying shape. Each migration looked safe in pre-production, each one executed with no available abort path once it started, and each one blocked recovery for far longer than the responsible team's runbook had assumed it might.
The Expand-and-Contract Pattern
Expand-and-contract is the structural answer to everything the previous sections describe, because it takes a single large, irreversible migration and decomposes it into a sequence of smaller steps, each of which leaves the prior state intact until the very last one. Rollback stays available at every step except the final contract.
OneUptime's five-step structure lays this out concretely. The expand step makes an additive schema change that does not invalidate anything the current application relies on, such as adding a nullable column rather than one constrained by an immediate NOT NULL requirement. The next step deploys compatible code: the new application version tolerates both the old and new representations at once, transitional writes may populate both columns within a single transaction, and the old code continues running against the expanded schema without any change in behavior. The data migration step backfills existing rows in bounded batches with a stable checkpoint, while tracking replication lag, lock wait time, database CPU, and application latency throughout; completeness has to be confirmed by querying for missing or divergent values directly, not inferred from a job's exit status. The fourth step switches reads to the new representation, using a controlled release or a feature flag, and only stops writes to the legacy column after that switch has been observed to hold, an ordering that preserves the option to send traffic back to the old code for as long as possible. The contract step, finally, removes the old column, and it is treated as a fully destructive, separate release: it only happens once every old application version is confirmed absent from production and from the rollback inventory, async consumers have fully drained, replicas and change-data-capture consumers have been checked, and a fresh backup exists.
At every step before that final contract, reverting to the prior state is safe, because neither the schema nor the underlying data has yet been made incompatible with the previous binary. Zero downtime, understood this way, means something narrower than zero risk. It means the risk that a single large irreversible migration would otherwise carry all at once gets spread across a sequence of small steps, each one reversible on its own, until only the last one isn't.



