Image: METAL
Summary
- Google unveiled Gemini 3.8 Live with Live Avatar on September 24, and Google Cloud said on September 25 that it is generally available in Gemini Enterprise.
- It processes video and audio together to sync lip movements and expressions, calls tools in the background without pausing the conversation, and moves across 97 languages.
- Custom avatars can be built from one reference photo but are open only to allowlisted enterprises, and every audio and video output carries a SynthID watermark.
On September 24 Google unveiled Gemini 3.8 Live with Live Avatar, which puts a face on its real-time conversation model. It is a feature in which a voice AI that listened and answered now converses as an on-screen video avatar, syncing its lip movements and making facial expressions. It is available now in Gemini Enterprise, Google's workspace for business AI, and Google Cloud said the following day, September 25, that the feature is generally available (GA).
The announcement came about a week after the underlying model. According to reports, Google released Gemini 3.8 Live and Live Extended Thinking on September 15, and Live Avatar adds a visual layer on top of that foundation. Google Cloud said the technology was first previewed at Google Cloud Next 2026 and is now ready for enterprise production. METAL has previously reported on Google's release of the Gemini 3.8 Live voice models.
The core of Google's pitch is natural conversation. Research scientist Shuo-yiin Chang and software engineer CJ Zheng, who wrote the announcement, noted that conversation is inherently multimodal: we listen, look, speak, and use facial expressions. Live Avatar processes video and audio inputs at the same time and aims to feel like a video call with a person, through precise lip-syncing, natural expressions, and fluid turn-taking.
Work keeps running behind the avatar while it talks. Thanks to asynchronous tool calling, Live Avatar triggers tools and fetches data in the background without stopping the conversation. Google Cloud said that when a user interrupts, the model recovers without losing the conversation context or the backend transactions in progress. It can also understand live camera feeds and screen shares alongside audio in real time.

It covers 97 languages. Google said Live Avatar moves across 97 languages while adapting its lip-sync and expressions, without degrading video fidelity or introducing visual drift. Google Cloud added that the model detects language automatically. Deployment spans the web, mobile, and interactive kiosks in stores.
For brands, the notable part is how the face is chosen. Businesses can pick from a library of diverse preset avatars, or generate their own brand avatar from a single high-quality reference image. In a demo published by Google Cloud, an interactive custom avatar was built simply by adding system instructions and uploading one reference photo and an audio sample file. Custom avatar creation, however, is open only to enterprise customers on an allowlist and goes through a strict verification process.
Safeguards were announced alongside. Google said all output generated by its AI products is watermarked with SynthID. The imperceptible mark is woven directly into the audio and video output so AI-generated content remains detectable later. The service is offered with US and EU endpoints, provisioned throughput, enterprise compliance, and data governance, while Gemini 3.8 Live Extended Thinking, which reasons for longer, remains in private preview.
The Google Cloud post that METAL reviewed includes three demos. One is an insurance claims intake desk. As a customer talks and shows the damage on camera, the claim notebook fills itself in, while in the background an agent team built with ADK (Agent Development Kit) checks the policy, applies the intake rules, and builds the packet for the adjuster. Google has released the demo code as open source.

Customer numbers were shared too. Equal AI in India handles more than a million live calls a day across nine Indian languages. "Gemini 3.8 Live improved interruption handling, multilingual conversations, and tool-call reliability," said Akhilesh Damaraju, CEO of Equal AI, adding that "this AI doesn't just answer calls; it gets things done for you."
There is a car marketplace case as well. Cox Automotive built an AI shopping assistant for Autotrader that uses screen highlighting and tool calling to guide shoppers through vehicle search, comparison, and financing by conversation. "Shoppers increasingly expect to describe what they need in their own words rather than work through filters and menus," said Marianne Johnson, Chief Product Officer of Cox Automotive. Salesforce said it is bringing Gemini 3.8 Live to Agentforce through a collaboration between its AI Research group and Google.
Through a content marketer's lens, the announcement changes the question of who serves as a brand's face. Until now, competition among enterprise voice bots was about speed and cost, but Fabien Blanc-paques, Group Product Manager at Google Cloud, wrote that the priority has shifted to the quality of each interaction. If the face at a service desk or store kiosk is built from a single reference photo and speaks 97 languages, what a brand must manage is no longer an ad model but the tone and expressions of an avatar that holds millions of conversations a day. Embedding the fact that this face is AI through SynthID reflects Google's judgment that trust is the basic condition of this new channel.
Sources
- Google DeepMind — Introducing Gemini 3.8 Live with Live Avatar →
- Google — Introducing Gemini 3.8 Live with Live Avatar →
- Google Cloud — Power your agents: Gemini 3.8 Live with Live Avatar is now generally available →
- Digital Trends — Google has given Gemini 3.8 Live a face that can lip-sync, react, and use tools in the background →





Comments