On September 28, 2026, ElevenLabs launched two new text-to-speech models: Eleven v4 and Eleven v4 Turbo. The ElevenLabs v4 speech model raises language coverage to more than 90 languages, up from 70 in v3, adds finer control over expression, and can clone a voice from a 10-second sample. Splitting the release into two models also shows a tradeoff every voice product has to make between expressiveness and speed.
What ElevenLabs launched
According to TechCrunch, the biggest quality gains are reported in Japanese, Brazilian Portuguese, Mandarin and Cantonese. Both models can clone a voice from 10 seconds of audio. ElevenLabs says in its launch post that voices can speak any supported language while retaining the identity of the original, and that v4 adds support for Professional Voice Clones (PVC), its highest-fidelity cloning option. Both models are available in ElevenAgents, ElevenCreative and through the ElevenAPI. Pricing was not disclosed in the coverage we reviewed.
The company positions the two models differently. Standard v4 is aimed at narration, dubbing and expressive dialogue, while v4 Turbo targets conversational agents. TBreak notes that the roughly 150 ms median time to first speech quoted for Turbo comes from ElevenLabs' own testing, and that total response time in a real product also includes the language model, the network and the application layer.
Why two models: expressivity versus latency
A speech model can produce better delivery when it sees more text before it speaks, because tone, pacing and emphasis depend on what comes next. A live voice agent cannot wait for that. It has to start speaking almost as soon as the language model behind it starts generating words. TechCrunch reports that v4 Turbo can begin generating audio as soon as the underlying LLM starts producing an answer.
That is the basic reason for two models. The standard model can spend its budget on nuance across a paragraph or chapter. The Turbo model trades some of that headroom for speed, which matters because long pauses make conversations feel unnatural. ElevenLabs describes the Turbo latency as shorter than the average pause between two people talking. Other vendors, including OpenAI and Google, sell voice products into the same market, so the same tension shapes the whole category.
The model uses surrounding text context to shape delivery rather than treating each line as isolated.
TBreak, summarizing the v4 launch
Context-aware delivery for audiobooks and agents
Earlier text-to-speech systems often read each sentence in isolation, which can flatten a scene where a character gradually loses patience. v4 reads surrounding text and interprets tone, pacing, emotion, character and context, per ElevenLabs. It also keeps inline tags, such as a laugh or a spoken direction, and TechCrunch reports that tags can now be stacked so the model follows the sequence.
For long-form content such as audiobooks, the practical change is consistency: a narrator voice that stays recognizable across chapters while the emotion shifts with the story. For agents, the value is different. Customer-service calls involve escalations, holds and moments of frustration, and TechCrunch says v4 Turbo handles confrontations, escalations and holds differently. Independent evaluation will matter here, since the preference results ElevenLabs cites (about 75 percent of test listeners preferring v4 in blind comparisons) are the company's own figures.

Why multilingual speech is hard, especially tonal languages
Going from 70 to more than 90 languages is not only a vocabulary problem. Each language has its own rhythm, stress and intonation, and a voice has to keep its identity while those patterns change. Mandarin and Cantonese are tonal: the pitch contour of a syllable changes its meaning, so a model that gets the tone wrong can say a different word, not merely sound accented. Cantonese adds more tones than Mandarin, and Japanese relies on pitch accent. Brazilian Portuguese differs from European Portuguese in vowels and rhythm, so a generic Portuguese voice can sound off to native listeners. It is notable that these are the languages ElevenLabs singles out for the biggest gains.
Coverage claims also deserve a check. TBreak points out that the launch announcement does not specifically confirm Arabic support, and advises developers to verify language and accent availability before building production workflows.
Voice cloning from 10 seconds: the consent question
Cloning from a 10-second sample lowers the barrier for legitimate uses such as restoring a voice, localizing a creator's content or keeping a brand voice consistent. It also lowers the barrier for misuse, because a short public clip may now be enough to imitate someone. Professional Voice Clones sit at the other end, aimed at the highest fidelity from more source material.
Teams adopting any cloning feature should treat consent as a product requirement: written permission from the speaker, clear disclosure that audio is synthetic, controls on who can use a cloned voice, and a way to revoke it. The launch coverage we reviewed does not detail ElevenLabs' safeguards for v4, so buyers should read the current policies before deploying cloned voices.
What this means for builders
The main takeaway is that voice is becoming a set of models chosen per job, not a single engine. A long-form narration pipeline and a phone agent have different needs, and both are now served by separate models from one vendor. Metir uses text-to-speech and voice features in its own products and stays model-agnostic across voice engines, so teams can compare providers as releases like this arrive.
Sources:
- TechCrunch: ElevenLabs' new v4 speech model supports more expression control and 90 languages
- ElevenLabs: Introducing Eleven v4
- ElevenLabs: Eleven v4
- TBreak: ElevenLabs v4 brings 90 languages and an expressive voice model
Image credits
Header image: Neumann U87 condenser microphone, by Will Fisher via Wikimedia Commons, licensed under CC BY-SA 2.0. In-body image: Shure SM7, by The Midnite Wolf via Wikimedia Commons, released under CC0. Both images are illustrative and do not depict ElevenLabs or its models.
