ElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28. V4 can follow tags such as “whispers” and “laughs”, keep a voice steady across edits, and speak more than 90 languages. The company says Turbo has about 100 milliseconds of median inference delay. Both models support voice cloning with verified owner consent.
New to this? Read it in simple words
- ElevenLabs released two new models that turn written text into speech.
- Eleven v4 can follow directions about emotion, pauses and sounds in more than 90 languages.
- The Turbo version is designed to answer quickly during live calls.
- The company says every cloned voice needs the voice owner’s verified consent.
- Voice model
- An AI system that creates spoken audio from text.
- Latency
- The time between a request and the start of the answer.
- Voice cloning
- Making an AI voice that sounds like a real person.
What the two models do
ElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28. The company says both use a new architecture. V4 is aimed at produced audio, while Turbo is designed for real-time agents and phone calls.
Writers can add directions such as “laughs”, “whispers” or “long pause” inside a script. The model can also keep several speakers apart and add sound effects. It supports more than 90 languages.
The company says long recordings can be joined with a steady voice and pace. A user may regenerate one line without making the speaker sound like a different person.
Language and speed figures are ElevenLabs claims. Total call delay depends on the full system.
Speed for live agents
ElevenLabs reports about 100 milliseconds of median model delay for Turbo. It reports about 150 milliseconds before speech starts. A millisecond is one thousandth of a second.
The model can receive text while another AI is still writing it, then start returning audio before the sentence is complete. This streaming process can make a voice agent feel less slow.
These numbers come from the company’s own tests. Network delay, the language model and other software also affect the total wait in a real call. Users should test the full system, not only the speech model.
Cloning needs clear consent
Professional voice cloning returns after being unavailable in v3. ElevenLabs says every clone needs verified consent from the voice owner. It also says an AI speech classifier can detect generated audio.
The same features that improve films, games and support calls can also create convincing false speech. Clear labels, strong account security and careful storage of voice recordings remain important.
Independent tests from Artificial Analysis ranked Eleven v4 highly for voice quality and pronunciation. Benchmarks are useful, but they do not measure every accent, noisy call or misuse risk.
Sources
Every fact in this story comes from the sources below. Open them to check our work.
- 1
- 2Research · September 28, 2026ElevenLabs’ new v4 speech model supports more expression control and 90 languages Yahoo Tech
- 3
We checked ElevenLabs’ product page against an independent news report and the Artificial Analysis leaderboard. Speed, stability, consent and language support are company statements unless noted. We do not treat a benchmark score as proof of safe use.