ElevenLabs has launched its v4 and v4 Turbo speech models, with support for 90 languages, up from 70. The company also says the new architecture improves voice cloning speed, expressive control and context handling, with lower latency aimed at enterprise voice agents.
Watch Desk analysis
What happened
The new models combine a broader language range with changes to how generated speech handles expression and context. ElevenLabs says they are optimised for lower latency in enterprise voice-agent applications.
The company also says its annualised revenue run rate has passed $600 million and that it is planning an IPO in the coming years. Those are company claims and plans, not a confirmed listing timetable.
Why it matters
For developers building voice products, 20 additional supported languages and lower-latency generation could widen where these systems are usable. Better expression and context handling also matter when a voice agent needs to sound less like a sat-nav reading a script.
The announcement gives useful product claims, but no comparative measurements here establish how much faster or more expressive the models are in practice. The revenue figure and IPO ambition add a glimpse of the business behind the technology, though the product changes are the more immediate news for users.
Our read
This is a meaningful step for a fast-moving voice-AI market: more language coverage, plus a stronger pitch for conversational applications. The next question is whether the improvements hold up in real deployments, not just in a feature list. Developers should check language quality and latency against their own use cases before rebuilding anything around the new versions.
What to watch
- Whether ElevenLabs publishes comparative latency or voice-quality results.
- How well the added languages perform across accents and everyday conversational use.
- Whether enterprise voice-agent customers adopt the models at scale.
- Whether the IPO ambition develops into a concrete timetable or filing.
Discussion spark: For voice agents, which matters more: broader language coverage or reliably natural expression and context handling?
Sources and evidence
- ElevenLabs’ new v4 speech model supports more expression control and 90 languages (28 September 2026, 14:00 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.