Databricks says its Lakebase Postgres service can load 1TB of data in under five minutes by building database pages on distributed Spark compute instead of making the live Postgres server do the heavy lifting. The change could cut large data loads from hours to minutes without making application queries compete for the same resources.
Databricks Watch analysis
What happened
In a post published on 8 October, Databricks describes its Lake Transactional/Analytical Processing architecture, or LTAP. Spark builds valid Postgres pages and indexes in parallel, then writes them directly to storage. The Lakebase primary records the completed import with a compact WAL entry, rather than processing every imported row itself.
Databricks says an internal benchmark loaded 1TB in under five minutes. The company specifies that this measures building the data pages; parallelising index builds is still in progress. During beta, one customer reportedly loaded about one billion rows each day, with a job that had taken more than eight hours previously. Databricks says the new approach kept that customer’s operational workloads unaffected. Read Databricks’ technical post.
Why it matters
Bulk imports into conventional Postgres can compete with live queries for processing, memory and storage bandwidth. Databricks’ approach aims to move that work to separate Spark compute, so teams can refresh large datasets without scheduling around quieter periods or scaling their database primary to survive the import.
The practical appeal is not just a faster benchmark. If the approach works across real workloads, teams could keep operational data fresher while leaving the database serving customers with more room to do its day job.
Our read
This is a substantial change to the plumbing, with a clear potential payoff for teams moving large Lakehouse datasets into operational applications. The benchmark is promising, but it measures page construction, not every part of a finished import; index-building performance and results across other workloads will matter. Databases, like kitchens, are happier when one enormous delivery does not block everyone else.
What to watch
- Whether Databricks publishes results for parallel index building.
- How the approach performs across different datasets and import patterns.
- Whether customers can achieve similar gains without disrupting live workloads.
Discussion spark: Would you trust a database import built outside its primary server for production data, or would you wait for independent results across more workloads?
Sources and evidence
- Load terabytes of data in minutes into Lakebase Postgres (8 October 2026, 18:49 UTC)
not affiliated with or endorsed by Databricks