Cohere has launched Embed 5, a family of embedding models for enterprise search, RAG and agent workflows. Its practical hook is that teams can index with the higher-quality Pro model, then use the faster, cheaper Fast model to search the same index, without rebuilding it.
Cohere Watch analysis
What happened
Embed 5 Pro and Embed 5 Fast are generally available through Cohere’s API and Model Vault, Microsoft Foundry and Amazon SageMaker. Cohere says both support text and image inputs, more than 100 languages and a 128K-token context window. Read Cohere’s announcement.
Our top picks
- One index, two models
Pro and Fast share an embedding space, so teams can index with one and query with the other without re-indexing. - Fast is priced for frequent searches
Text costs $0.08 per million tokens for Fast and $0.12 for Pro; image inputs cost $0.40 per million tokens for either tier. - Long, mixed-format documents are in scope
Both models accept text, images and fused text-image inputs, with a 128K-token context window and support for more than 100 languages. - Pro leads Cohere’s reported retrieval tests
Cohere says Pro achieved its highest average results across tests including financial documents, parsed PDFs and visually rich material. - Fast targets the live query path
Cohere says Fast averaged 2.4 times Pro’s document throughput across the context sizes it tested, a potential fit for high-volume search and agent loops.
Why it matters
Embedding models help search systems find relevant material before an AI model writes an answer. The shared space gives teams more room to balance retrieval quality against the latency and cost of repeated searches: Cohere recommends indexing with Pro and querying with Fast. That could be especially useful in agent workflows, where a task may trigger many searches.
The performance comparisons are Cohere’s own evaluations, not a guarantee that Embed 5 will win on every organisation’s data. The useful test is whether the new models improve retrieval on the documents and queries a team actually handles.
Our read
The shared embedding space is the most interesting part of this launch: it makes a quality-first indexing strategy compatible with a cheaper, faster query path. That is more actionable than another leaderboard victory lap. Teams should compare the tiers on their own corpus, paying particular attention to whether any retrieval-quality trade-off is worth the savings.
What to watch
- How Embed 5 performs on independent evaluations and customer datasets.
- Whether teams adopt Cohere’s suggested Pro-index, Fast-query setup.
- How the models’ image and fused text-image retrieval holds up on real enterprise documents.
Discussion spark: Would you index with the higher-quality Pro model and query with Fast, or keep one model for both stages to avoid even a small retrieval trade-off?
Sources and evidence
- Source update (30 September 2026, 09:51 UTC)
not affiliated with or endorsed by Cohere