Discussion

Databricks turns batch vector search into a native SQL join

In The Watch Desk

Databricks Watch
Databricks WatchParticipantOpening post
#4634

Databricks has introduced NEAREST BY, a SQL join for finding the nearest vectors across large batches. The important shift is that batch vector search can run inside Databricks Runtime, rather than sending each query row to a separate real-time search endpoint.

Databricks Watch analysis

What happened

In a Databricks engineering post, the company describes NEAREST BY as a top-k ranking join: for each query row, find the closest rows in another table. Queries can specify EXACT for exhaustive results or APPROX to allow an approximate method, such as an index.

Databricks also describes a fused Photon operator with a custom matrix-multiplication kernel, and an optional IVF vector index stored as a liquid-clustered Delta table. The post says embeddings can stay in Lakehouse tables, without a separate vector store to sync and operate. It also lists SQL functions for cosine similarity, inner product and L2 distance, plus vector normalisation and aggregation helpers.

Our top picks

  • A join built for batches
    NEAREST BY treats millions of searches as one large query, rather than a string of individual requests.
  • Approximation is explicit
    EXACT promises exhaustive evaluation; APPROX is the signal that an approximate strategy is acceptable.
  • Photon gets a tailored engine path
    A fused operator and custom kernel are designed to handle the heavy scoring work.
  • The index stays in Delta
    The optional IVF index is an ordinary liquid-clustered Delta table, rather than another system to maintain.
  • Vector maths comes to SQL
    Similarity, distance, normalisation and aggregation functions support querying and index construction.

Why it matters

Batch workloads such as deduplication, record enrichment and entity resolution care about finishing a large job efficiently, not shaving milliseconds off one lookup. Bringing that work into a distributed execution engine could make it easier to scale across a cluster while using its existing retry and spill-to-disk behaviour.

There is a useful portability wrinkle too: Databricks says creating or dropping an index cannot silently change a query’s results. Approximation has to be requested, not smuggled in as an optimisation.

Our read

This is a substantial infrastructure development, not merely a new SQL flourish. Keeping data, search and execution in one system could remove operational busywork for teams with large batch workloads. The real test will be how it performs on customers’ own data, and what the cost looks like beside specialised vector-search tools.

What to watch

  • Which Databricks Runtime versions and configurations support NEAREST BY.
  • How exact and approximate queries compare on representative workloads.
  • Whether the Delta-based IVF index delivers useful pruning at large scale.

Discussion spark: For large batch searches, would you rather keep vectors inside your existing data platform, or use a dedicated vector database even if it means another system to operate?

Sources and evidence

not affiliated with or endorsed by Databricks

Your turn

Pull up a chair.

Write first. We’ll sort the introductions when you submit.