Cohere has released Cohere Transcribe, a 2-billion-parameter automatic speech recognition model trained for 14 languages and licensed under Apache 2.0. It is a notable move beyond the text and retrieval tools for which the company is better known: audio can now enter the same enterprise toolbox without every recording first taking a trip to somebody else's hosted transcription service.

Cohere Watch analysis
What happened
The model takes audio and produces text through a Conformer-based encoder-decoder architecture. Cohere says it trained the system from scratch, with support spanning English, Arabic, Mandarin Chinese, Japanese, Korean, Vietnamese and eight other European languages. The weights and model card are available through Cohere Labs, alongside examples for offline use and online inference.
Cohere reports an average word error rate of 5.42 across the English tasks used by the Hugging Face Open ASR Leaderboard as it stood on 26 March 2026. That is the company's benchmark claim, not a universal promise: accuracy will still depend on language, accent, recording conditions and the sort of audio being shoved through the workshop hatch.
Why it matters
A permissively licensed model of this size gives teams another route for keeping recordings within infrastructure they control. That can matter for meeting transcripts, call archives, searchable media and regulated workflows where sending raw speech to an external service is awkward or prohibited. The compact scale also makes serious local evaluation more plausible than it would be with a much larger model, although practical hardware requirements still need testing in each deployment.
Our read
Cohere's own documentation says the model has no automatic language detection, can be inconsistent on code-switched audio, and does not provide timestamps or speaker diarisation. It may also transcribe non-speech sounds unless paired with a voice-activity or noise gate. In other words, the engine has arrived, but the dashboard, seatbelts and bloke pointing out who said what remain separate fittings.
What to watch
- Independent tests on noisy meetings, mixed-language speech and accents beyond the headline benchmark.
- Whether Cohere adds diarisation, timestamps and language detection without sacrificing local deployment.
- How the model's real throughput and memory use compare with hosted speech services on ordinary enterprise hardware.
Discussion spark: Would a strong local transcription model change where you process meetings and calls, or do diarisation and timestamps still make hosted speech services the practical choice?
Sources and evidence
- Introducing Cohere Transcribe: a new state-of-the-art in open-source speech recognition (26 March 2026)
- Cohere Transcribe model card (26 March 2026)
- Introducing Cohere-transcribe: state-of-the-art speech recognition (26 March 2026)
not affiliated with or endorsed by Cohere