In managed OLTP, restores have always been painfully slow and they get slower at scale. This often means that large production databases, where downtime costs the most, are left waiting the longest to recover.
The usual workarounds are difficult, expensive and risky. They involve extra replicas, extra environments and even a DBA on the restore which doesn’t fully guarantee you will be protected. Failover to a healthy replica helps when a machine dies, but it does not help when the bad write is already on the standby. That still means a restore, and a restore can mean hours of downtime.
This pain is caused by the architecture of traditional managed OLTP systems. Compute and storage ship as one machine, a restore starts with provisioning a new instance (you’re already waiting); snapshots live in object storage while Postgres runs on that volume, so the snapshot still has to be pulled onto disk (more waiting); WAL replay then has to close the gap from snapshot time to the exact recovery timestamp (even more waiting). As the database grows, this gets slower and more expensive.
The Lakebase Postgres architecture breaks the monolith and changes the restore mechanics. In Lakebase, compute and storage are decoupled, and database history is already kept in object storage in a way that’s instantly referenceable. In this architecture, a restore does not copy data into a new disk. Instead, it simply creates a branch at a timestamp which is a simple metadata operation, not a multi-hour copy-and-replay job.
In practical terms, time to restore drops to seconds, even if the database is 100 TB. And, it’s so simple that an agent can do it.

Traditional Postgres point-in-time recovery (PITR) is built from two ingredients: a base backup of the files and archived WAL for everything after that backup. In a managed Postgres environment, such as Amazon RDS, that backup is usually a snapshot sitting in object storage.
“Restoring from backup” to a particular moment in time T is actually a process that involves three parts:
This involves shipping compute and storage volumes (coupled) of at least the same as the primary. A small instance can come up in minutes, but a large instance with large EBS volumes usually takes longer causing you to wait before even starting the restore process.
RDS snapshots live in S3, so “restoring” means getting that snapshot out of object storage and onto Postgres’s disk.
That process is slow at large size, so RDS does not wait for it to be completed before making the restored instance available. For large volumes, this happens while most table and index pages are still in S3. But available does not mean the working set is on the Postgres volume. If a query touches a block that is not local yet, the volume fetches it from S3 on the spot while the rest keeps hydrating in the background.
These types of queries come with a latency that’s fine for an internal checkup, but not for production. The restore is only done when the data you actually need lives on the volume, and that is not quick for a large database. The larger the database, the more time (in hours) this will take.
The snapshot is only consistent as of snapshot time. To reach T, Postgres still has to replay the transaction logs archived after that snapshot. How long replay takes depends on how much happened between the snapshot and T. A snapshot from an hour ago will be much faster than a snapshot from last night. Plus if you had a heavy write day, that’ll be a lot of WAL to replay. Once again, this means a long wait (on top of the instance still hydrating from S3).
Unless the database is small, PITR is almost always a multi-hour operation. You have to provision the monolith, pull a snapshot out of S3, replay WAL and wait until enough of the volume is local to take traffic.
For that whole window, you may be suffering downtime. A healthy, high available (HA) replica can save you if the primary died and the replica still has good data. However, it may not save you from PITR so dropped tables and bad writes might already be on the standby.
In a survey, 50 developers running 1TB+ production Postgres were asked about their experience with restores:
This created potential adverse business impact:

In Lakebase Postgres, a modern architecture enables a different path for restores.
Compute and durable storage are split apart and connected by the WAL. Compute runs Postgres meaning it runs SQL, plans queries, applies MVCC, manages locks, and generates WAL (all the regular Postgres tasks). What it does not do is own the durable copy of your data.
Storage is what owns durability and history, and the job is divided in three parts executed by three distinct components:

When a write comes in:
In this design, the write path takes an interesting shape. The architecture above separates committing a transaction from materializing pages, which in simpler terms means: old page versions are never overwritten. Your database's history piles up as a timeline you can point at, not a single copy you mutate.
In the traditional path, a restore mostly means rebuilding that past state into a separate instance. But with Lakebase, there’s an immutable storage history to reference, so that step is unnecessary and is replaced by a different primitive: a branch.
Where traditional restores provision a new instance and copy data into it, a restore in Lakebase is a branch at a point in history. This new branch has its own independent compute, its own connection string and it can be queried completely independently of production. It is not a replica of the original instance, but it feels just like it.
Here’s how to use this primitive for a restore:
Since this is the concept that matters, let's reiterate:
This restore method completely eliminates data copies. The restored branch doesn’t need to copy data; it simply points at the image and delta layers that already exist up to that point in time.
The benefit is that scaling is not operationally scary anymore. If a restore is metadata work, how long it takes does not grow with the size of your data, and the mechanism stays the same.

Independent of how large your database is:
A human or an agent can always reach a queryable past state right away after an incident, even on a huge Postgres database.

Restores in Lakebase are a simple operation: create a branch at a timestamp. The loop is short enough that agents can treat it as an ordinary tool call, not only to resolve incidents.
That is the piece agent platforms like Replit or v0 actually productize to build versioning or undo features. A typical loop looks like this:
Traditional PITR is too slow and too heavy to support live workflows, but a branch at a timestamp is fast and cheap enough to be part of the product.
Traditional restores get slower and heavier as the database grows. Branch-based restores do not. History already lives outside compute, so a past point is something you can open as a branch, not something you rebuild. Regardless if you have 10 GB or 100 TB, restores look the same. Pick a timestamp, create the branch and attach compute. Failures at scale are no longer as scary.
Experience it for yourself: create a Lakebase Postgres database, load a good chunk of data, run a restore and wonder how you’ve been living without this for so long.