A video layer for the live model
Google has added generated avatars to Gemini 3.8 Live, its low-latency conversational model for enterprise agents. The service accepts live audio, images, video and text, then returns synchronized speech, video and text. It is generally available through Google's Gemini Enterprise Agent Platform.
The avatar is not a separate talking-head clip assembled after the conversation. Google describes it as part of the live session, with facial movement and speech generated together. Businesses can choose from preset characters. They can also create an avatar from an authorized reference image, but that custom option currently requires enterprise allowlisting.
The agent can keep talking while tools work
Gemini Live already supports interruptions, transcripts, tone adaptation and function calls. With asynchronous tool execution, Google says the avatar can keep the dialogue moving while a background request fetches information or completes a task. Its launch example is a hotel check-in, where the agent talks to the guest while the system calls tools behind the scenes.
Google also claims native speech-to-speech synchronization across 97 languages, including mid-conversation switches without visible drift. The broader Cloud documentation lists 24 supported conversational languages for the Live API. The launch page's larger number refers specifically to the avatar's lip-sync and expression adaptation. Those are related claims, not interchangeable measures of full product support.
The model card puts a ceiling on the demo
The model card says Gemini 3.8 Live accepts up to 128,000 input tokens. Audio-only Live can produce up to 64,000 output tokens, while Live Avatar is listed at 24,000. More plainly, Google says the avatar supports a few minutes of continuous interaction rather than extended hours. Occasional slowness and timeouts are also known limitations.
The model can hallucinate, and its knowledge cutoff is January 2025. Tool connections can bring in current data, but the avatar's confident delivery does not make an answer more reliable. For a customer-facing deployment, the visible character should not blur the boundary between a polished interface and the underlying model's uncertainty.
Synthetic presence needs visible rules
Google says every generated audio and video output carries its SynthID watermark. Its Cloud guidance also frames custom likenesses around authorized employees or paid talent with releases. Those controls address provenance and consent at the system level. They do not guarantee that every viewer will notice or understand that a person on screen is generated.
A serious deployment still needs a plain disclosure, a route to a human and limits on what the agent may decide. It also needs testing for interruptions, language switching, tool failures and cases where the face appears more certain than the answer deserves.
Live Avatar makes enterprise agents more expressive. The near-term product is also narrower than the launch videos may suggest: short sessions, controlled access for custom characters, and a model whose normal errors remain in place.
Sources
- Google: Introducing Gemini 3.8 Live with Live AvatarPrimary September 24 launch announcement covering availability, asynchronous tools, 97-language avatar synchronization, custom-avatar allowlisting and SynthID claims.
- Google DeepMind: Gemini 3.8 Audio model cardPrimary model documentation for modalities, context and output limits, known hallucination and timeout risks, short continuous-avatar sessions, knowledge cutoff and safety framing.
- Google Cloud: Gemini Live API overviewCurrent primary documentation, last updated September 24, for general availability, supported inputs and outputs, live-avatar support, tool use and conversational-language count.


