Added to our archive on 1 October 2026. This report is dated to the day of the event.

ElevenLabs opened its launch post with one line of dialogue: "I need you to stay calm." From a doctor speaking to a frightened patient it should sound gentle; from a game character shouting to his squad before a battle it should not. On 28 September 2026 the company released Eleven v4, which it called its most emotive text to speech model yet, and a faster variant, Eleven v4 Turbo, for voice agents and live conversation.

The release speaks to two groups. Producers of audiobooks, dubbing and game dialogue get a model they can direct in plain words, in more than 90 languages. Companies building phone and support agents get a Turbo model that starts speaking in about 150 milliseconds, by ElevenLabs' own measure. Until 12 October both cost 72 percent less than their list prices on the API.

  • What: Eleven v4 (eleven_v4) and Eleven v4 Turbo (eleven_v4_turbo), speech and dialogue models.
  • Price: USD 0.022 and USD 0.011 per 1,000 characters on the API until 12 October, then USD 0.08 and USD 0.04.
  • Why it matters: more than 90 languages, instant voice clones from 10 seconds of audio, faster replies for agents.

Directing a voice in plain words

ElevenLabs says Eleven v4 sits on an entirely new architecture that interprets tone, pacing, emotion, character and context from the text. Users can describe how a line should be delivered and add inline tags such as [laughs], [said angrily in French accent], [light rain] or [phone buzzing]. "Because the model understands the context of a whole scene, it generates natural dialogue where speakers respond to what's just been said, rather than stitching together isolated lines," the company wrote in its announcement.

The changelog sends eleven_v4 through the Text to Dialogue API, for content creation and long form audio, and eleven_v4_turbo through its WebSocket version, for agents and interactive apps. A single generation can run to 10,000 characters, according to the product page. Both models use two voice settings, Stability and Similarity. The Style and Speed sliders are not available, and SSML, the markup older speech engines used for pauses and pronunciation, is not supported, the documentation says.

Coverage grows to more than 90 languages, ElevenLabs says, from 70 in the previous version, TechCrunch reported, with the biggest quality gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese. When the output language differs from the language of the reference voice, Eleven v4 "generates fluent, natural-sounding speech in the target language", the documentation says, instead of carrying over the original accent.

Turbo for agents

ElevenLabs puts the median inference latency of Eleven v4 Turbo at about 100 milliseconds and its median time to first speech at about 150 milliseconds. On the company's own comparison, rival services took between 262 and 814 milliseconds, from Cartesia Sonic 3.6 to OpenAI's GPT-4o mini TTS. The figures are vendor measured.

In ElevenAgents, the company's platform for voice agents, the model is meant to adjust its tone during a call, "reassuring when a caller is frustrated, clearer when they are confused," ElevenLabs wrote. That market matters to the company: TechCrunch reported that more than 55 percent of its business comes from large companies and that its annualised revenue run rate has grown from about USD 330 million at the start of the year to more than USD 600 million. ElevenLabs also says Eleven v4 was ranked first by the benchmarking firm Artificial Analysis and preferred by about 75 percent of listeners in blind tests against rival services; both claims come from its launch materials.

Clones, designed voices and safeguards

Instant Voice Clones can now be made from 10 seconds of audio, ElevenLabs says, and v4 adds support for Professional Voice Clones, meant for the highest fidelity cloning. The documentation says this support "is currently rolling out to everyone and should be available within the next few days." Voices created with the company's Voice Design tool are a weaker fit: "Voice Design voices work with Eleven v4, but they may not be as performative or sound as good as with earlier models," the documentation warns.

The launch materials name no new safeguards for v4. ElevenLabs' product page says every v4 voice clone, instant or professional, "requires verified consent from the voice's owner," and that its AI Speech Classifier can detect the generated audio as AI generated. Professional clones have stricter rules. "For now, we only allow you to clone your own voice," the company's cloning guide says, and users must pass a verification step, reading lines aloud, before a clone is trained. "Even with their consent, you cannot clone someone else's voice."

Price and launch discount

Model (API)Until 12 OctoberList price
Eleven v4USD 0.022 per 1,000 charactersUSD 0.08
Eleven v4 TurboUSD 0.011 per 1,000 charactersUSD 0.04

The API price list keeps Eleven v3 at USD 0.08 per 1,000 characters. In ElevenLabs' own apps, subscribers on Creator and higher plans get three times the credits for v4 until the same date, and free accounts can use both models; the free plan includes 10,000 credits a month. Eleven v4 and v4 Turbo are live in ElevenAgents, ElevenCreative and the ElevenLabs API, and the discount ends on 12 October.