
Summary
- On September 28, ElevenLabs released Eleven v4, a voice model built on a new architecture, together with Eleven v4 Turbo for real-time use.
- Eleven v4 ranked first on the Artificial Analysis voice leaderboard and beat competing models 65% to 81% of the time in the company's blind preference tests.
- Turbo's median time to first speech is about 150 milliseconds, and both models support more than 90 languages and inline direction tags.
On September 28, ElevenLabs released a new voice model, Eleven v4, and its low-latency variant, Eleven v4 Turbo. The company introduced Eleven v4 as "our most emotive text-to-speech model yet," and both models support more than 90 languages. Eleven v4 took first place on the voice leaderboard of Artificial Analysis, a site that benchmarks AI models, and both models were available from launch day in ElevenAgents, ElevenCreative and ElevenAPI.
Eleven v4 is built on an entirely new architecture. ElevenLabs said the model reads a script the way a voice actor would, interpreting tone, pacing, emotion and character from context: who is speaking, what just happened, and how each line should land. In scenes with multiple speakers, the company said, it produces dialogue in which speakers respond to what was just said, rather than stitching together separately generated lines.
The company offered two numbers as evidence. In blind head-to-head preference tests ElevenLabs ran in September, listeners chose Eleven v4 about 75% of the time. By opponent, it won 81% against both Cartesia Sonic 3.6 and Inworld TTS-2, 72% against Google Gemini 3.8 Flash-Lite TTS and 65% against Gemini 3.8 Flash TTS. According to reports, Eleven v4 scored an Elo of 1,319 from 1,674 samples on the Artificial Analysis leaderboard.
Writing direction into the script is the centerpiece of this release. Users can describe in natural language how a line should be delivered, or insert inline tags such as [laughs], [said angrily in French accent], [light rain] or [phone buzzing] to specify emotion and even sound effects. When several tags are stacked, the model follows them in sequence. The official product page says v4 follows tag sequences more reliably than v3, and support for the International Phonetic Alphabet (IPA), which lets users set how names and acronyms are pronounced, has been significantly improved. The 2-minute-36-second official demo video METAL reviewed states that its scenes, a director and actor on set, a news anchor reporting a pigeon mystery and a pharmacy prior-authorization call, were generated from prompts alone with no edits.

Eleven v4 Turbo is the model for real-time conversation. Its median inference latency is about 100 milliseconds, and its median time from request to audible speech is about 150 milliseconds. In ElevenLabs' comparison using identical scripts and default settings, Cartesia Sonic 3.6 took 262 milliseconds, xAI TTS 362 milliseconds, Google Gemini 3.8 Flash-Lite TTS 685 milliseconds and OpenAI GPT-4o mini TTS 814 milliseconds. The company said this speed is faster than the average pause between two people talking. Turbo supports bidirectional streaming, starting audio while the large language model is still generating its answer, and was optimized together with ElevenAgents, ElevenLabs' agent platform, as a single system.
Customer reactions were included in the announcement. Ryan Peterson, SVP of Product for Agentforce Voice at Salesforce, said, "That's exactly where we're seeing Eleven v4 Turbo raise the bar, with faster, more natural responses." Oscar Daniels, Head of Credit Building Products at Spring Financial, said Eleven v4 Turbo "is the first one that does both," speed and quality, adding, "For automated sales workflows, this is the point where building stops feeling like an experiment."

Language coverage has also widened. According to reports, the previous generation supported 70 languages and v4 raises that to more than 90, with the largest quality gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese. When a voice recorded in one language speaks another, it keeps the original voice's identity while adopting a native speaker's accent. The company said Instant Voice Clones can copy a voice from just 10 seconds of audio, though according to reports the v4 documentation says a sample of one to two minutes is generally used. Professional Voice Clones, which were dropped in v3, are supported again, and stitching long scripts generated in several passes without the voice drifting has also been improved.
From an AI engineer's perspective, the core of this release is a design that splits expressiveness and speed out of one architecture. ElevenLabs noted that high-quality voice models have been slow, pushing businesses toward fast but monotone agents. Splitting the same technology into v4 for creation and Turbo for real time, and tuning the agent platform as one body with the model, is how the company separates itself from rivals that sell only models. That the more than 17,500 voices in its voice library carry over to both models also lowers switching costs.
The business is shifting in the same direction. According to reports, more than 55% of ElevenLabs' revenue comes from large companies, and its annualized revenue run rate has grown from about $330 million at the start of the year to more than $600 million. Earlier this year the company raised $500 million in a round led by Sequoia at an $11 billion valuation, and its headcount has passed 800. Co-founder and CEO Mati Staniszewski said the company is aiming for an IPO "in the next years" but did not commit to a timeline, according to reports. METAL has reported on ElevenLabs' general release of its real-time conversational voice model, v3 Conversational.
With startups such as Cartesia and Deepgram and large companies such as Google and OpenAI all shipping expressive voice models, ElevenLabs answered with two numbers: first place on the leaderboard and 150 milliseconds. The company's claim is that it has moved past imitating human voices to performing scripts, and that claim will now be tested in real settings such as customer service calls, audiobooks and game dialogue.
Sources
- ElevenLabs — Eleven v4: Our most expressive text-to-speech AI model yet →
- ElevenLabs — Eleven v4 and Eleven v4 Turbo Text to Speech models →
- TechCrunch — ElevenLabs' new v4 speech model supports more expression control and 90 languages →
- RuntimeWire — ElevenLabs ships v4 voice models, ranked No. 1 by Artificial Analysis →





Comments