METAL LAB

Inworld AI Launches Realtime TTS-2, Reigniting Race for Top Voice Synthesis Spot

Inworld AI says it has reclaimed the No. 1 spot on Artificial Analysis, a ranking Cartesia had held for just two weeks

Summary

  • Inworld AI has made its real-time speech synthesis model, Realtime TTS-2, generally available, and claims it now tops the Artificial Analysis leaderboard
  • The model boasts under-150-millisecond latency, natural-language voice steering to adjust tone on the fly, and the ability to switch between more than 500 dialects while keeping the same speaker identity
  • Just two weeks earlier, on August 19, Cartesia's Sonic-3.6 held the top spot on the same leaderboard, so the lead has already changed hands again
Realtime TTS-2 is GA today, now the #1 model on Artificial Analysis and fastest TTS in the world of its class.

Inworld AI Declares Realtime TTS-2 Generally Available

Realtime official website

The image shows Cartesia, marked with a bold border, claiming the leaderboard's top spot (dotted circle) on August 19 via a dashed arrow, followed by Inworld, marked with a single dot, reclaiming that same spot on September 2 via another dashed arrow. Both arrows are dashed, signaling that this No. 1 position is provisional and could change again at any time.The image shows Cartesia, marked with a bold border, claiming the leaderboard's top spot (dotted circle) on August 19 via a dashed arrow, followed by Inworld, marked with a single dot, reclaiming that same spot on September 2 via another dashed arrow. Both arrows are dashed, signaling that this No. 1 position is provisional and could change again at any time.

Inworld AI announced in a video posted to its X account on September 2 (local time) that its real-time speech synthesis model, Realtime TTS-2, is now generally available (GA). The company describes the model as topping the Artificial Analysis leaderboard and as the fastest TTS system in its class. Text-to-speech, or TTS, is the underlying technology that turns written text into human-like speech, and it's used everywhere from call-center chatbots to audiobook narration.

To put that in context, Artificial Analysis is an independent benchmarking outfit that measures and ranks AI models on performance, speed, and price. Since the firm doesn't build models itself but simply scores them, a No. 1 ranking there is a metric the industry frequently cites as proof of performance.

Inworld AI said usage during the research preview phase far exceeded expectations, and it folded lessons from that period back into its research before shipping this GA model. The company called the launch "the biggest leap in Inworld's history."

The Leaderboard Lead Changes Hands Again After Two Weeks

This leaderboard has already changed hands once in the past month. On August 19, Cartesia's streaming TTS model, Sonic-3.6, topped both the provider-voice and controlled-voice categories. Now that Inworld AI is claiming the top spot on the same leaderboard, the lead has shifted again in under two weeks.

이미지: @inworld_ai (X)

What's Different This Time

Inworld AI says it rebuilt both the serving stack and the model architecture from the ground up for this GA release. The result is "promptable" — meaning it can take natural-language instructions directly, much like a large language model, and understand conversational context. The "voice steering" feature lets users adjust tone and emotion on the fly using natural language, similar to directing an actor, and the model also supports switching between more than 500 dialects in real time while keeping the same speaker identity. The company says all of this happens in under 150 milliseconds, making it feel natural even in live conversation.

David, co-founder and CTO of LiveKit, which builds voice agent frameworks, tried out the model and said it "genuinely felt like talking to a person." He pointed to the model's ability to produce entirely different emotional ranges from the same text depending on the prompt as one of its strengths.

CategoryInworld Realtime TTS-2Cartesia Sonic-3.6
Announced2026-09-022026-08-19
Leaderboard rankNo. 1 on Artificial Analysis (self-reported)No. 1 in both provider and controlled voice categories
LatencyUnder 150ms (self-reported)Under 90ms (self-reported)
Key featureVoice steering, 500+ dialect switchingState space model (SSM) architecture
opengraph image

How to Try It

The video the company released is more of a demo than a detailed how-to guide. Inworld researchers took turns at the microphone, switching voice styles on the spot to demonstrate the voice steering feature, and pointed viewers to "realtime.ai" to try the model themselves. The announcement materials didn't include specifics on pricing, the API application process, or supported platforms.

Still, the potential applications within the model's stated capabilities are easy to picture. Hook it up to a customer-service chatbot, for instance, and it could switch between a calm tone and an urgent one using nothing but natural-language instructions, depending on the situation. In multilingual dubbing, it could re-record the same speaker's lines in different languages while preserving that speaker's vocal identity.

Editor's Take

It's no accident that the TTS leaderboard keeps changing hands every few weeks. It reflects just how tight the competition has become among companies like Inworld AI, Cartesia, and ElevenLabs for the real-time voice AI market.

Based on how teams have deployed real-time voice models in practice, clear and accurate reading used to be enough for a TTS system. That's now table stakes. Whether a model understands conversational context to adjust its own tone, and whether it can switch between languages while keeping the same speaker identity, are what actually drive purchasing decisions today.

For teams looking to deploy voice agents, it's safer to route through an agent framework like LiveKit and keep the TTS engine swappable, rather than locking your entire infrastructure to a single model. This case shows that a leaderboard's No. 1 title can have a shelf life of just two or three weeks.

There's a good chance Cartesia or ElevenLabs will respond within the next few weeks with their own announcement touting latency or Elo scores. This battle for the top spot looks set to continue for a while.

Comments