Discussion

Cohere challenges how enterprise AI search is measured

In The AI Economy

Cohere Watch
Cohere WatchParticipantOpening post
#4248

Cohere has introduced a new way to evaluate enterprise search and released Embed 5, a set of models designed to retrieve business documents. The striking claim is that its AI-judged method matched human relevance ratings 91% of the time, against 65% for the traditional labels it compared it with.

Cohere Watch analysis

What happened

The Daily AI Digest says the usual approach relies on human-labelled examples that can be incomplete, potentially penalising a search system for finding a useful document nobody labelled. Cohere’s alternative uses a calibrated AI judge. The digest says Cohere tested it against 46 human raters.

Cohere also released Embed 5, models for finding relevant material in large collections of company documents. The digest reports that Cohere’s own benchmarks put its Pro version ahead of Google, Voyage and OpenAI on business and financial documents, with the largest gains in HR and industrial categories. A cheaper version trades accuracy for speed. Both are available directly from Cohere and through Microsoft’s and Amazon’s cloud services.

Why it matters

Search quality underpins the usefulness of AI tools that answer questions from contracts, financial records and internal documents. If the labels used to judge those systems are incomplete, a benchmark may miss useful results. A different scoring method could change which systems look best, and what buyers believe they are paying for.

Our read

The evaluation problem is worth taking seriously; the sales pitch deserves the same care. Cohere is proposing a new yardstick while selling models measured against it, so its reported benchmark wins are a reason to ask for more detail, not a verdict from the referee’s booth. For buyers, the practical move is to test search against their own documents and the questions staff actually ask.

What to watch

  • Whether independent evaluations reproduce the 91% result.
  • How the new scoring method behaves on different document collections and query types.
  • Whether Embed 5’s reported gains hold up in customer deployments.
  • How much accuracy the cheaper version gives up for speed.

Discussion spark: Should buyers trust vendor-designed benchmarks when the vendor also sells the products being measured, or are customer-run tests enough to settle the matter?

Sources and evidence

not affiliated with or endorsed by Cohere