In brief

ElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28. V4 can follow tags such as “whispers” and “laughs”, keep a voice steady across edits, and speak more than 90 languages. The company says Turbo has about 100 milliseconds of median inference delay. Both models support voice cloning with verified owner consent.

New to this? Read it in simple words
  • ElevenLabs released two new models that turn written text into speech.
  • Eleven v4 can follow directions about emotion, pauses and sounds in more than 90 languages.
  • The Turbo version is designed to answer quickly during live calls.
  • The company says every cloned voice needs the voice owner’s verified consent.
Words to know
Voice model
An AI system that creates spoken audio from text.
Latency
The time between a request and the start of the answer.
Voice cloning
Making an AI voice that sounds like a real person.

What the two models do

ElevenLabs released Eleven v4 and Eleven v4 Turbo on September 28. The company says both use a new architecture. V4 is aimed at produced audio, while Turbo is designed for real-time agents and phone calls.

Writers can add directions such as “laughs”, “whispers” or “long pause” inside a script. The model can also keep several speakers apart and add sound effects. It supports more than 90 languages.

The company says long recordings can be joined with a steady voice and pace. A user may regenerate one line without making the speaker sound like a different person.

Sources12

VOICE PATH 01
Expression for recordings, lower delay for live agents.

Language and speed figures are ElevenLabs claims. Total call delay depends on the full system.

Speed for live agents

ElevenLabs reports about 100 milliseconds of median model delay for Turbo. It reports about 150 milliseconds before speech starts. A millisecond is one thousandth of a second.

The model can receive text while another AI is still writing it, then start returning audio before the sentence is complete. This streaming process can make a voice agent feel less slow.

These numbers come from the company’s own tests. Network delay, the language model and other software also affect the total wait in a real call. Users should test the full system, not only the speech model.

Sources123

Professional voice cloning returns after being unavailable in v3. ElevenLabs says every clone needs verified consent from the voice owner. It also says an AI speech classifier can detect generated audio.

The same features that improve films, games and support calls can also create convincing false speech. Clear labels, strong account security and careful storage of voice recordings remain important.

Independent tests from Artificial Analysis ranked Eleven v4 highly for voice quality and pronunciation. Benchmarks are useful, but they do not measure every accent, noisy call or misuse risk.

Sources123

Sources

Every fact in this story comes from the sources below. Open them to check our work.

  1. 1
    Primary source · September 28, 2026Eleven v4 and Eleven v4 Turbo ElevenLabs
  2. 2
  3. 3
    Research · September 28, 2026Text to Speech Leaderboard Artificial Analysis
How we checked this story

We checked ElevenLabs’ product page against an independent news report and the Artificial Analysis leaderboard. Speed, stability, consent and language support are company statements unless noted. We do not treat a benchmark score as proof of safe use.