What changed
Google introduced Gemini 3.8 Live with Live Avatar for Gemini Enterprise. It combines live dialogue with near-real-time generated video, so an agent can respond through a visual character instead of a disembodied chat box.
Google says the system processes visual and audio input together, produces synchronized speech and expressions, and can switch among 97 languages without losing lip sync. That makes multilingual service and guided walkthroughs the clearest early use cases.
What it can do
The agent can make asynchronous tool calls while the avatar continues the conversation. In Google’s example, a hotel check-in can keep moving while the system fetches data in the background.
Organizations can choose preset characters. Custom avatars can be generated from a reference image, though that option currently requires enterprise allowlisting.
Why it matters
All generated audio and video carries SynthID watermarking. The launch is an enterprise product, so it should be judged by latency, task completion and user trust in real deployments.
The larger shift is that an AI interface can maintain visual presence while it works. If the tool calling is reliable, customer support and instruction may feel closer to a video call than a chatbot.
Source published 2026-09-24. Coverage is based on the maker’s announcement and demonstration.
