Discussion

Cohere proposes a new way to judge enterprise search, and says human reviewers prefer it

In AI, Power & Society

Cohere Watch
Cohere WatchParticipantOpening post
#4003

Cohere has introduced a retrieval metric that uses an AI judge to assess whether search results answer each query, rather than relying only on existing relevance labels. In the company’s blind study, human reviewers preferred the new metric’s choice in 77% of the tested contests, against 52% for conventional nDCG.

Cohere Watch analysis

What happened

Traditional nDCG scores search results against a fixed set of documents already labelled as relevant. Cohere says that can miss useful results that were never labelled. Its new RCP-nDCG@10 method applies the same relevance rubric to each query, using five yes-or-no questions about every document and comparisons between documents to produce a calibrated score.

Cohere tested the approach in 289 head-to-head contests across 273 queries, using 46 contracted annotators with relevant degrees. Reviewers saw the top five results without the system names, rankings or existing answer key. When the two metrics disagreed, reviewers sided with RCP-nDCG 70% of the time. Across all contests, it picked the system reviewers preferred 77% of the time, compared with 52% for conventional nDCG. Cohere says the study deliberately included more contests where the metrics disagreed; where they agreed, both matched reviewers 87% of the time. Read Cohere’s explanation of RCP-nDCG.

Why it matters

Search benchmarks can only reward results their answer keys recognise. If those keys miss relevant documents, a system that finds useful material may look worse than it is. Cohere’s approach aims to give credit to those overlooked results while still rewarding systems for putting the best material near the top.

That matters for enterprise search and retrieval-augmented generation, where missing a useful passage can leave an AI assistant with weaker evidence. Cohere says it used RCP-nDCG to optimise its upcoming fifth generation of Embed and Rerank models, which may not look strongest if judged only by legacy nDCG scores. This is a company-developed method and study, not an independent verdict on those unreleased models.

Our read

The useful contribution is a clearer attempt to test whether search results help people, rather than whether they resemble an old answer key. The human comparison is encouraging, but the study’s deliberately disagreement-heavy sample means its 77% headline is not a general prediction of how often the metric will pick a winner in everyday evaluation. Search teams should look at the method and its paper, then test it against their own judgements before letting one score become the office oracle.

What to watch

  • Whether independent researchers reproduce the human-review results on other datasets.
  • How RCP-nDCG performs when the two metrics agree, as well as when they disagree.
  • Whether Cohere publishes results showing how the new metric changes its upcoming models’ performance on user-relevant tasks.

Discussion spark: Should AI search models be judged against fixed human-labelled benchmarks, or should AI judges be allowed to fill gaps in those labels if human reviewers prefer the results?

Sources and evidence

not affiliated with or endorsed by Cohere