
Summary
- Inworld AI has made its real-time speech synthesis model, Realtime TTS-2, generally available, and claims it now tops the Artificial Analysis leaderboard
- The model boasts under-150-millisecond latency, natural-language voice steering to adjust tone on the fly, and the ability to switch between more than 500 dialects while keeping the same speaker identity
- Just two weeks earlier, on August 19, Cartesia's Sonic-3.6 held the top spot on the same leaderboard, so the lead has already changed hands again
Inworld AI Declares Realtime TTS-2 Generally Available
Inworld AI announced in a video posted to its X account on September 2 (local time) that its real-time speech synthesis model, Realtime TTS-2, is now generally available (GA). The company describes the model as topping the Artificial Analysis leaderboard and as the fastest TTS system in its class. Text-to-speech, or TTS, is the underlying technology that turns written text into human-like speech, and it's used everywhere from call-center chatbots to audiobook narration.
To put that in context, Artificial Analysis is an independent benchmarking outfit that measures and ranks AI models on performance, speed, and price. Since the firm doesn't build models itself but simply scores them, a No. 1 ranking there is a metric the industry frequently cites as proof of performance.
Inworld AI said usage during the research preview phase far exceeded expectations, and it folded lessons from that period back into its research before shipping this GA model. The company called the launch "the biggest leap in Inworld's history."
The Leaderboard Lead Changes Hands Again After Two Weeks
This leaderboard has already changed hands once in the past month. On August 19, Cartesia's streaming TTS model, Sonic-3.6, topped both the provider-voice and controlled-voice categories. Now that Inworld AI is claiming the top spot on the same leaderboard, the lead has shifted again in under two weeks.

What's Different This Time
Inworld AI says it rebuilt both the serving stack and the model architecture from the ground up for this GA release. The result is "promptable" — meaning it can take natural-language instructions directly, much like a large language model, and understand conversational context. The "voice steering" feature lets users adjust tone and emotion on the fly using natural language, similar to directing an actor, and the model also supports switching between more than 500 dialects in real time while keeping the same speaker identity. The company says all of this happens in under 150 milliseconds, making it feel natural even in live conversation.
David, co-founder and CTO of LiveKit, which builds voice agent frameworks, tried out the model and said it "genuinely felt like talking to a person." He pointed to the model's ability to produce entirely different emotional ranges from the same text depending on the prompt as one of its strengths.
| Category | Inworld Realtime TTS-2 | Cartesia Sonic-3.6 |
|---|---|---|
| Announced | 2026-09-02 | 2026-08-19 |
| Leaderboard rank | No. 1 on Artificial Analysis (self-reported) | No. 1 in both provider and controlled voice categories |
| Latency | Under 150ms (self-reported) | Under 90ms (self-reported) |
| Key feature | Voice steering, 500+ dialect switching | State space model (SSM) architecture |

How to Try It
The video the company released is more of a demo than a detailed how-to guide. Inworld researchers took turns at the microphone, switching voice styles on the spot to demonstrate the voice steering feature, and pointed viewers to "realtime.ai" to try the model themselves. The announcement materials didn't include specifics on pricing, the API application process, or supported platforms.
Still, the potential applications within the model's stated capabilities are easy to picture. Hook it up to a customer-service chatbot, for instance, and it could switch between a calm tone and an urgent one using nothing but natural-language instructions, depending on the situation. In multilingual dubbing, it could re-record the same speaker's lines in different languages while preserving that speaker's vocal identity.
Editor's Take
It's no accident that the TTS leaderboard keeps changing hands every few weeks. It reflects just how tight the competition has become among companies like Inworld AI, Cartesia, and ElevenLabs for the real-time voice AI market.
Based on how teams have deployed real-time voice models in practice, clear and accurate reading used to be enough for a TTS system. That's now table stakes. Whether a model understands conversational context to adjust its own tone, and whether it can switch between languages while keeping the same speaker identity, are what actually drive purchasing decisions today.
For teams looking to deploy voice agents, it's safer to route through an agent framework like LiveKit and keep the TTS engine swappable, rather than locking your entire infrastructure to a single model. This case shows that a leaderboard's No. 1 title can have a shelf life of just two or three weeks.
There's a good chance Cartesia or ElevenLabs will respond within the next few weeks with their own announcement touting latency or Elo scores. This battle for the top spot looks set to continue for a while.





Comments