Talking AI Avatars Explained: How Meta, Google, and Synthesia’s Real-Time Faces Work

Cheerful woman on a video call with her laptop in a cozy living room

In the span of a single week, three of the biggest names in AI shipped the same futuristic thing: a face you can talk to. Meta announced a real-time avatar for its Muse assistant, Google launched Live Avatar inside Gemini 3.8 Live, and Synthesia let a TechCrunch reporter build a digital clone of himself that anyone can hold a conversation with.

It looks like the future arrived all at once. It didn’t — not exactly. One of these avatars you can try today, one ships in “the coming months,” and one is locked behind an enterprise gate. And none of them is a single piece of AI. Under the hood, every talking avatar is a four-part relay race, run in under a second.

How a talking AI avatar actually works

Forget the idea of one super-model doing everything. A real-time avatar is four separate AI models chained together, each handing off to the next fast enough that the result feels like a live conversation.

First, voice-to-text transcribes what you said. Then an agentic language model — the brain — figures out what to say back, sometimes pulling in tools or context along the way. Next, text-to-voice turns that reply into spoken audio in the avatar’s voice. Finally, a video model renders the face itself: the lips, the expressions, the subtle head movement, all synced to the audio.

The hard part isn’t any single step — it’s the handoff. Every millisecond of delay between “you stop talking” and “the face starts answering” breaks the illusion. If the lips drift out of sync with the audio, or the avatar freezes mid-sentence while the brain thinks, your brain instantly files it under “robot.” Keeping the whole chain under roughly a second is what separates a demo from a product.

Three launches, three very different realities

Synthesia is the one you can use right now. A TechCrunch reporter filmed a short session, and the company turned it into a personal avatar that visitors can actually talk to in real time. The guardrails are serious: personal avatars require photographed consent, so nobody can mint a convincing fake of you — or anyone else — without proof you agreed. That is the model shipping today.

Meta is the one you’re waiting for. The company announced its Muse Realtime Avatar on September 23, built around its Muse assistant (internal codename “Jolly”), but it ships in “the coming months,” as TechCrunch reported. Meta’s play is clearly consumer: alongside the avatar, it showed off a $150 Charm wearable that ships in time for the holidays — a signal that Meta wants its talking assistant on your body, not just in an app.

Google is the one you probably can’t touch. Gemini 3.8 Live with Live Avatar launched September 24 and is technically generally available — but only for enterprise customers. Want an avatar that looks like a specific person? That’s behind an allowlist. Google’s version of this race is being run inside boardrooms, not app stores.

Why this matters

The headlines make this look like a demo contest. It’s not — it’s a distribution fight, and the scoreboard already tells you who each company is really selling to.

Synthesia isn’t a toy company. It’s valued at $4 billion, pulls in over $100 million in annual recurring revenue, and its flagship enterprise product, Roleplay Sessions, sells exactly this technology to businesses for training, sales coaching, and onboarding. When a company at that scale ships real-time conversational avatars to the public, it’s not a party trick — it’s a product line coming downmarket.

That’s the real story of this three-way race: three companies, three go-to-market bets. Synthesia is productizing the tech for anyone with consent and a camera. Meta is betting consumers will want a face on their AI assistant — and a $150 wearable to carry it. Google is keeping its most powerful version behind enterprise contracts and an allowlist. The four-model stack is nearly identical across all three. What differs is who gets the keys, and when. Watch that, not the demos.

Meanwhile in Apps & Tools: Microsoft Edge Gets New AI “Copilot Mode” for Browsing — the browser itself is becoming the AI assistant.

Written by
Daniel covers apps, software, and productivity tools — from AI-powered apps and major platform updates to the tools teams actually use to get work done.