Databricks has introduced NEAREST BY, a SQL join for finding the nearest vectors across large batches. The important shift is that batch vector search can run inside Databricks Runtime, rather than sending each query row to a separate real-time search endpoint.
Databricks Watch analysis
What happened
In a Databricks engineering post, the company describes NEAREST BY as a top-k ranking join: for each query row, find the closest rows in another table. Queries can specify EXACT for exhaustive results or APPROX to allow an approximate method, such as an index.
Databricks also describes a fused Photon operator with a custom matrix-multiplication kernel, and an optional IVF vector index stored as a liquid-clustered Delta table. The post says embeddings can stay in Lakehouse tables, without a separate vector store to sync and operate. It also lists SQL functions for cosine similarity, inner product and L2 distance, plus vector normalisation and aggregation helpers.
Our top picks
- A join built for batches
NEAREST BY treats millions of searches as one large query, rather than a string of individual requests. - Approximation is explicit
EXACT promises exhaustive evaluation; APPROX is the signal that an approximate strategy is acceptable. - Photon gets a tailored engine path
A fused operator and custom kernel are designed to handle the heavy scoring work. - The index stays in Delta
The optional IVF index is an ordinary liquid-clustered Delta table, rather than another system to maintain. - Vector maths comes to SQL
Similarity, distance, normalisation and aggregation functions support querying and index construction.
Why it matters
Batch workloads such as deduplication, record enrichment and entity resolution care about finishing a large job efficiently, not shaving milliseconds off one lookup. Bringing that work into a distributed execution engine could make it easier to scale across a cluster while using its existing retry and spill-to-disk behaviour.
There is a useful portability wrinkle too: Databricks says creating or dropping an index cannot silently change a query’s results. Approximation has to be requested, not smuggled in as an optimisation.
Our read
This is a substantial infrastructure development, not merely a new SQL flourish. Keeping data, search and execution in one system could remove operational busywork for teams with large batch workloads. The real test will be how it performs on customers’ own data, and what the cost looks like beside specialised vector-search tools.
What to watch
- Which Databricks Runtime versions and configurations support NEAREST BY.
- How exact and approximate queries compare on representative workloads.
- Whether the Delta-based IVF index delivers useful pruning at large scale.
Discussion spark: For large batch searches, would you rather keep vectors inside your existing data platform, or use a dedicated vector database even if it means another system to operate?
Sources and evidence
- NEAREST BY Join: Scaling Vector Search in Databricks Runtime | Databricks Blog (5 October 2026, 20:00 UTC)
not affiliated with or endorsed by Databricks