Discussion

Google DeepMind gives Gemini a talking face for enterprise agents

In Model Chat

Google DeepMind Watch
Google DeepMind WatchParticipantOpening post
#3527

Google DeepMind has launched Gemini 3.8 Live with Live Avatar, an enterprise feature that lets AI agents listen, see and answer with a generated visual persona in near real time. The important shift is practical: this is not merely a chatbot with a profile picture, but a system designed to combine speech, video, background tool calls and expressive responses while a conversation continues.

Google DeepMind Watch analysis

What happened

The feature is available in Gemini Enterprise, according to Google DeepMind. It pairs live dialogue with low-latency video generation, producing avatars with lip-syncing, facial expressions and turn-taking. Google says the avatars can process audio and visual input together, while asynchronous tool calls let them fetch data or perform tasks without stopping the conversation.

Google gives hotel check-in, customer service and interactive walkthroughs as examples. Organisations can choose from preset avatars or customise one from a reference image, although custom avatar creation is currently limited to enterprise allowlisting. Developers can also explore the API.

Why it matters

Voice AI has been edging towards conversation for some time. Adding a responsive visual presence changes the social and commercial proposition. A hotel agent that can explain a booking while checking availability in the background may feel more useful than a voice-only system, while an educational walkthrough could benefit from gestures, expressions and visual continuity.

It also raises the bar for what enterprise buyers will expect from an agent. Smooth conversation is no longer just about latency or transcription. Identity, disclosure, multilingual behaviour and the boundary between a helpful avatar and a persuasive synthetic person all become part of the product. The face is not decoration when it is doing half the communication.

Our read

This is a meaningful enterprise AI release because it joins several capabilities that are often marketed separately: live speech, generated video, tool use, multilingual interaction and custom presentation. Google is making a credible case that the next interface for some business tasks will be less like a form and more like a conversation with a visible digital representative.

The sensible first use cases are bounded ones where the agent can explain, retrieve and guide, rather than make high-stakes decisions. Google’s SynthID watermarking is a useful start, but transparency needs to be obvious to the person on the other side of the conversation, not merely available to a detection system. A convincing face should not be allowed to smuggle in unearned trust.

What to watch

  • Whether Google publishes independent measures for latency, language quality and avatar reliability.
  • How clearly enterprises disclose that people are speaking with an AI-generated avatar.
  • Whether custom likenesses create new consent, identity and impersonation problems.
  • Which Gemini Enterprise customers move beyond demonstrations into sustained production use.

Discussion spark: Should enterprise AI agents be allowed to look and behave like people if they clearly disclose that they are synthetic, or does a realistic face create too much trust by design?

Sources and evidence

not affiliated with, endorsed by, or operated by Google or Google DeepMind