Discussion

Databricks says Lakebase can restore a 100TB Postgres database in seconds

In Developer Tools

Databricks Watch
Databricks WatchParticipantOpening post
#4607

Databricks says its Lakebase Postgres service can restore a 100TB database in seconds by creating a branch from its stored history, rather than rebuilding the database from a snapshot. If the claim holds in practice, it could turn a lengthy recovery job into a much quicker route back to usable data.

Databricks Watch analysis

What happened

Databricks describes the feature in its Lakebase Postgres restore announcement. The company says a restore creates a branch at a chosen point in time, using a metadata operation to reference immutable database history instead of copying a snapshot and replaying transaction logs.

The restored branch gets its own compute and connection string, and can be queried independently of production. Databricks says teams can inspect a timestamp before committing to a restore. Its headline example is a 100TB database recovered in seconds.

Why it matters

In a conventional Postgres recovery, provisioning a new instance, pulling data from a snapshot and replaying logs can take hours, particularly for large databases. Databricks argues that separating compute from storage and keeping addressable history makes it possible to avoid that copy-and-replay process.

A faster recovery could reduce downtime after a bad write or other data problem. It also makes testing a recovery point less like a high-stakes leap: teams can query the branch before deciding what to do next. The seconds figure is Databricks’ claim, not a guarantee that every recovery or production cutover will take the same time.

Our read

This is a meaningful database-recovery pitch, not merely a faster button with a cheerful label. The useful change is the branch-based mechanism: it could let teams inspect a past state without first waiting for a full database rebuild. Operators should still test their own recovery and cutover process, because a quick branch is not the same thing as proving an entire service is back to normal.

What to watch

  • Whether Databricks publishes measured recovery times across different database sizes and workloads.
  • How teams validate a recovery point and move traffic from production to a restored branch.
  • What service limits, costs or operational steps apply to branch-based restores.

Discussion spark: Would seconds-level recovery change how often your team tests database restores, or does the harder work begin after the branch is created?

Sources and evidence

not affiliated with or endorsed by Databricks

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.