METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

ElevenLabs Releases Eleven v4 Voice Model

ElevenLabs released Eleven v4, a voice model with stronger emotional expression, and its low-latency variant Eleven v4 Turbo. Turbo reaches first speech in about 150 milliseconds, and both models support more than 90 languages.

ElevenLabs Releases Eleven v4 Voice Model

Summary

  • On September 28, ElevenLabs released Eleven v4, a voice model built on a new architecture, together with Eleven v4 Turbo for real-time use.
  • Eleven v4 ranked first on the Artificial Analysis voice leaderboard and beat competing models 65% to 81% of the time in the company's blind preference tests.
  • Turbo's median time to first speech is about 150 milliseconds, and both models support more than 90 languages and inline direction tags.
Eleven v4: Our most expressive text-to-speech AI model yet

On September 28, ElevenLabs released a new voice model, Eleven v4, and its low-latency variant, Eleven v4 Turbo. The company introduced Eleven v4 as "our most emotive text-to-speech model yet," and both models support more than 90 languages. Eleven v4 took first place on the voice leaderboard of Artificial Analysis, a site that benchmarks AI models, and both models were available from launch day in ElevenAgents, ElevenCreative and ElevenAPI.

Eleven v4 is built on an entirely new architecture. ElevenLabs said the model reads a script the way a voice actor would, interpreting tone, pacing, emotion and character from context: who is speaking, what just happened, and how each line should land. In scenes with multiple speakers, the company said, it produces dialogue in which speakers respond to what was just said, rather than stitching together separately generated lines.

The company offered two numbers as evidence. In blind head-to-head preference tests ElevenLabs ran in September, listeners chose Eleven v4 about 75% of the time. By opponent, it won 81% against both Cartesia Sonic 3.6 and Inworld TTS-2, 72% against Google Gemini 3.8 Flash-Lite TTS and 65% against Gemini 3.8 Flash TTS. According to reports, Eleven v4 scored an Elo of 1,319 from 1,674 samples on the Artificial Analysis leaderboard.

Writing direction into the script is the centerpiece of this release. Users can describe in natural language how a line should be delivered, or insert inline tags such as [laughs], [said angrily in French accent], [light rain] or [phone buzzing] to specify emotion and even sound effects. When several tags are stacked, the model follows them in sequence. The official product page says v4 follows tag sequences more reliably than v3, and support for the International Phonetic Alphabet (IPA), which lets users set how names and acronyms are pronounced, has been significantly improved. The 2-minute-36-second official demo video METAL reviewed states that its scenes, a director and actor on set, a news anchor reporting a pigeon mystery and a pharmacy prior-authorization call, were generated from prompts alone with no edits.

Eleven v4가 경쟁 음성 모델 4종과의 블라인드 일대일 선호도 시험에서 이긴 비율(65~81%)을 비교한 막대 도표

Eleven v4 Turbo is the model for real-time conversation. Its median inference latency is about 100 milliseconds, and its median time from request to audible speech is about 150 milliseconds. In ElevenLabs' comparison using identical scripts and default settings, Cartesia Sonic 3.6 took 262 milliseconds, xAI TTS 362 milliseconds, Google Gemini 3.8 Flash-Lite TTS 685 milliseconds and OpenAI GPT-4o mini TTS 814 milliseconds. The company said this speed is faster than the average pause between two people talking. Turbo supports bidirectional streaming, starting audio while the large language model is still generating its answer, and was optimized together with ElevenAgents, ElevenLabs' agent platform, as a single system.

Customer reactions were included in the announcement. Ryan Peterson, SVP of Product for Agentforce Voice at Salesforce, said, "That's exactly where we're seeing Eleven v4 Turbo raise the bar, with faster, more natural responses." Oscar Daniels, Head of Credit Building Products at Spring Financial, said Eleven v4 Turbo "is the first one that does both," speed and quality, adding, "For automated sales workflows, this is the point where building stops feeling like an experiment."

Eleven v4 Turbo와 경쟁 음성 모델의 첫 음성까지 걸리는 시간 중앙값(150~814밀리초)을 비교한 막대 도표

Language coverage has also widened. According to reports, the previous generation supported 70 languages and v4 raises that to more than 90, with the largest quality gains in Japanese, Brazilian Portuguese, Mandarin and Cantonese. When a voice recorded in one language speaks another, it keeps the original voice's identity while adopting a native speaker's accent. The company said Instant Voice Clones can copy a voice from just 10 seconds of audio, though according to reports the v4 documentation says a sample of one to two minutes is generally used. Professional Voice Clones, which were dropped in v3, are supported again, and stitching long scripts generated in several passes without the voice drifting has also been improved.

From an AI engineer's perspective, the core of this release is a design that splits expressiveness and speed out of one architecture. ElevenLabs noted that high-quality voice models have been slow, pushing businesses toward fast but monotone agents. Splitting the same technology into v4 for creation and Turbo for real time, and tuning the agent platform as one body with the model, is how the company separates itself from rivals that sell only models. That the more than 17,500 voices in its voice library carry over to both models also lowers switching costs.

The business is shifting in the same direction. According to reports, more than 55% of ElevenLabs' revenue comes from large companies, and its annualized revenue run rate has grown from about $330 million at the start of the year to more than $600 million. Earlier this year the company raised $500 million in a round led by Sequoia at an $11 billion valuation, and its headcount has passed 800. Co-founder and CEO Mati Staniszewski said the company is aiming for an IPO "in the next years" but did not commit to a timeline, according to reports. METAL has reported on ElevenLabs' general release of its real-time conversational voice model, v3 Conversational.

With startups such as Cartesia and Deepgram and large companies such as Google and OpenAI all shipping expressive voice models, ElevenLabs answered with two numbers: first place on the leaderboard and 150 milliseconds. The company's claim is that it has moved past imitating human voices to performing scripts, and that claim will now be tested in real settings such as customer service calls, audiobooks and game dialogue.

Comments