METAL for iPhone

Read AI news in the METAL app.

Download METAL and discover fresh AI stories every day.

Download on the App Store

For iPhone · Free download

Search for METAL AI Magazine in the App Store on your iPhone.

METAL

Meta Unveils Muse Realtime Avatar

At Connect 2026, Meta showed a real-time video model that gives its personal agent Muse facial expressions and gestures. Voice and video come back together about 870 milliseconds after you finish speaking.

Meta Unveils Muse Realtime Avatar

Image: @AIatMeta (X) (video still)

Summary

  • On September 23 at Meta Connect 2026, Meta unveiled Muse Realtime Avatar, a real-time model that turns Muse's voice into an interactive avatar.
  • Meta distilled a 40-step teacher model into a two-step student, cutting compute by 60x, and the model streams 448x768 video at 25 frames per second with about 870 milliseconds of latency.
  • In Meta's own evaluation, overall preference was 78% against Runway Characters and 88% against HeyGen LiveAvatar.
@finkd just unveiled Muse Realtime Avatar

Meta unveiled Muse Realtime Avatar, a real-time video model that gives its personal AI agent Muse a face and body language, at its annual Meta Connect 2026 event on September 23 (US time). CEO Mark Zuckerberg introduced it himself from the keynote stage. Built by Meta Superintelligence Labs, the model turns Muse Realtime Voice, Muse's voice model, into an expressive, interactive avatar. According to Meta, its first use is live conversation in the Muse app.

The avatar's look comes from a reference image the user picks. According to Meta's research blog, a photographic portrait responds with subtle expressions, a full-body illustration gestures and shifts its posture as it speaks, and animals and everyday objects gain expressions without losing what makes them distinctive. The company said the avatar's appearance and mannerisms stay consistent frame by frame across many conversational turns.

On stage, Zuckerberg said, "We trained a whole new state-of-the-art real-time model that creates joyful animations of your Muse as you're talking to it," adding, "The details here really make the experience." In the 2-minute-2-second demo video Meta posted on X, an avatar named Agrippa, wearing a laurel wreath and a toga, greets the audience and then discusses with Zuckerberg the finish material and component layout for Project Charm, a small device meant to hang from a keychain. The same day, Meta previewed Muse Charm, a pocket-size device built for talking with Muse, and said it would share more later this year.

The core of the technology is a design in which voice and video share one stream. According to Meta, Muse Realtime Voice emits speech tokens that carry both what is said and how it is delivered. While an audio decoder turns those tokens into sound, Muse Realtime Avatar consumes the same tokens to generate video. The company says that sharing one stream keeps voice, lip motion, and expression in sync. The model is an audio-driven diffusion transformer that generates video in short chunks and passes only the newest output of each chunk forward as motion context for the next, so the amount of computation stays bounded even as a conversation gets longer.

Muse Realtime Voice의 음성 토큰을 Muse Realtime Avatar가 받아 2단계 확산으로 영상을 만들고 음성과 동기화해 내보내는 구조도

The speed comes from distillation. Meta said it shrank a large teacher model, which needs 40 diffusion steps with three-way classifier-free guidance, or 120 model evaluations, for each chunk, into a student model that finishes in two unguided steps. Compute fell by a factor of 60, while human raters' preference was close to an even split, at 45% for the student and 55% for the teacher. The company explained that self-forcing, a technique in which the student trains on context it generated itself, keeps errors from accumulating and the video from drifting over long conversations.

Meta also published results against commercial avatar services. Raters held two- to three-minute conversations with Runway's Runway Characters and HeyGen's HeyGen LiveAvatar in each product's own live-call interface, using matched avatar identities, and then compared the experiences. On overall preference, Muse Realtime Avatar scored 78% against Runway Characters and 88% against HeyGen LiveAvatar. By dimension, facial expressivity came in at 93% and 92% and lip sync at 78% and 94%, while the mannerism comparison with Runway Characters, at 56%, was not statistically distinguishable from parity, Meta said.

실시간 대화 평가에서 Muse Realtime Avatar와 Runway Characters, HeyGen LiveAvatar의 전체 선호도와 항목별 선호도를 비교한 막대그래프

The serving figures are specific as well. According to Meta, the model streams 448x768 portrait video at 25 frames per second, and about 870 milliseconds pass between the moment a user finishes speaking and the arrival of the first byte of the synchronized voice-and-video response. Running a single session on one NVIDIA GB200, each generation step produces eight frames, or 320 milliseconds of playback, in 20 milliseconds. The company said four-bit quantization-aware training and optimizations developed with NVIDIA raised serving capacity 8x over a two-step BF16 baseline, so one GB200 can handle 12 real-time sessions at once.

As a safeguard, Meta added a watermark. The company said Muse Realtime Avatar embeds a durable, invisible watermark throughout generated video using its own Meta Video Seal technology, without adding latency. Meta noted that the examples in its blog illustrate the model's capability and do not all reflect avatars available in the Muse app, and Muse is for users 18 and older. It did not say when the avatar will reach the app.

The avatar was one piece of a flood of Muse announcements that day. Meta also announced a voice mode that lets you design your Muse's voice by describing it, Muse on its AI glasses in the coming months activated by saying the agent's name, a dedicated email address for Muse, computer use that lets Muse operate apps on a Mac, and shopping connections with retailers including Walmart, Best Buy, and Sephora. Zuckerberg said in the keynote that "the centerpiece of our vision for what we're building is Muse," and that Meta is "making Muse free for a huge number of tokens, with the expectation that over time we will profit by taking a small fee from transactions," according to reports.

Muse is only a few weeks old. METAL previously reported that Meta launched its personal agent Muse in early September, and in this announcement Meta wrote that the response since launch had been "incredible." After the app, WhatsApp, the web, and the Mac desktop, Meta is widening the ways to reach Muse to glasses and a dedicated device.

Taken together, the demo video and research blog METAL reviewed show that what Meta is selling is presence more than model performance. A face that answers within 870 milliseconds with matching expressions and lip motion makes an agent feel less like a tool and more like a counterpart, and Meta has set that face on a business model of free tokens and transaction fees. The comparison was designed by Meta itself and no app launch date has been set, but the figure of 12 sessions per GB200 reads as a sign that the math for rolling this out at the scale of billions of people is already in place.

Comments