Point-in-Time Recovery Explained: PITR for Databases
Point-in-time recovery restores a database to any moment between backups by replaying transaction logs, undoing bad deploys and accidental deletes.
Point-in-time recovery (PITR) is the ability to restore a database to its exact state at any specific moment — not just to the moment of the last full backup, but to any second in between. It works by combining a periodic full backup with a continuous log of every change made since, so recovery replays exactly the transactions that happened up to (but not past) the moment you choose.
Why “restore the last backup” usually isn’t enough
A nightly full backup only protects the state of the database as of when it ran. If a bad deploy runs a destructive migration at 2 p.m. and the backup ran at midnight, restoring that backup throws away fourteen hours of legitimate writes along with the mistake. For any system handling real user data, that trade-off is rarely acceptable.
PITR closes that gap. Instead of a single all-or-nothing snapshot, you get a full backup plus a continuous, ordered record of changes — which means you can restore to 1:59 p.m., a minute before the bad migration ran, and lose nothing that mattered.
How it actually works
PITR relies on the same mechanism most relational databases already use for crash recovery: a write-ahead log. As covered in what is write-ahead logging, every change is recorded in a durable, sequential log before it’s applied to the actual data files. That log is the ingredient PITR needs.
The recovery process has two phases:
- Restore the base backup. The database is brought back to the state it was in at the time of the most recent full backup.
- Replay the log up to the target time. Every logged transaction between the backup and the chosen recovery point is reapplied in order, bringing the database forward to that exact moment — and no further.
Because the log records transactions in order with timestamps, “any point in time” really does mean any point, not just the intervals between scheduled backups. Some systems support recovering to a specific transaction ID instead of a timestamp, which is useful when you know exactly which write caused the problem but not precisely when it happened.
What PITR protects against — and what it doesn’t
PITR is built for a specific category of failure: logical errors introduced by the application or an operator — an unintended DELETE without a WHERE clause, a migration that drops the wrong column, a bug that corrupts records over the course of an hour. In all of these cases, the fix is the same: recover to the moment just before the damage started.
It does not, on its own, protect against everything:
- Storage failure. If the underlying disks holding both the backup and the log are lost together, PITR has nothing to replay from. Backups and logs need to live somewhere independent of the primary storage.
- Detection lag. PITR can only recover to before you know something went wrong. If a bug corrupts data silently for days before anyone notices, the log retention window has to be long enough to reach back that far — most systems don’t keep write-ahead logs indefinitely by default.
- Malicious or compromised credentials. If an attacker had write access and also deleted or tampered with the backups themselves, PITR depends on those backups being intact and stored somewhere the attacker couldn’t reach.
PITR and replication are complementary, not substitutes
It’s worth being explicit that database replication and PITR solve different problems. A replica protects against a server going down — it fails over quickly, but a bad write replicates to the replica just as fast as it replicated everywhere else, so replication alone does nothing against logical errors. PITR protects against exactly that case: recovering to before the bad write happened. Production systems generally run both — replication for availability, PITR-capable backups for recoverability.
Setting a recovery point objective
The practical question PITR forces you to answer up front is your recovery point objective (RPO): how much data loss is acceptable if disaster strikes right now? A system with continuous log shipping to durable storage can offer an RPO measured in seconds. A system that only ships logs once an hour accepts up to an hour of potential loss between the failure and the last shipped log segment. That number should be a deliberate choice, driven by what the data is worth, not an accident of default configuration.
The other number worth deciding deliberately is retention: how far back a recovery point can reach. A log retained for seven days lets you undo a mistake discovered a week later; one retained for a day does not. Since the log itself consumes storage that grows with retention window and write volume, this is a real trade-off between storage cost and how much time you’re giving yourself to notice a problem before the ability to fix it expires.
The takeaway
Point-in-time recovery restores a database to any specific moment, not just the time of the last backup, by combining a full backup with a replayable log of every subsequent transaction. It’s the tool for undoing logical mistakes — bad migrations, accidental deletes — that replication alone can’t fix, since a bad write replicates just as reliably as a good one. The two work together: replication keeps you available, PITR keeps you recoverable.
Tagged
Keep reading
Chisato · · 4 min read Postgres Logical Replication Explained
Logical replication streams row-level changes between Postgres databases instead of copying raw disk blocks — enabling selective sync, upgrades, and CDC.
The Lycoris Team · · 4 min read What Is Apache Kafka? Event Streaming, Explained
Apache Kafka is a distributed event-streaming platform built on a durable, append-only log. How topics, partitions, and consumers power real-time pipelines.
Chisato · · 4 min read A Hands-On Guide to AWS Spot Instances
AWS Spot Instances sell unused EC2 capacity at a steep discount. Here's how to launch one, handle interruptions, and pick workloads that fit.