I Got My Own Interactive AI Avatar—Here’s How It Works

Synthesia, the AI video-generation company, has built a digital twin of a journalist—capable of listening and responding in real time. The avatar, made this summer, was trained to answer questions about a specific investigative story the journalist wrote on fraud in venture-backed startups. This marks the first time Synthesia has created such an interactive avatar for a journalist—or anyone outside its executive team.

How it was built

The avatar was created during a visit to Synthesia’s New York office. The team photographed the journalist in detail—with and without glasses—and recorded physical gestures and voice over two minutes. The company then built several versions: static ones that read scripted text, and interactive ones that can respond to questions based on the training material.

The training pipeline combines voice-to-text (converting spoken words into text), a language model that understands and responds to that text, text-to-voice (turning responses back into speech), and a video model that animates the avatar. While Synthesia supplies its own models, customers also have the option to plug in models from Cartesia, ElevenLabs, Google, or OpenAI. Deployment flexibility is built in, with avatars hosted either via Synthesia’s cloud or a customer’s preferred infrastructure.

Experience and limits

When interacting with the avatar, the journalist found that while the voice and likeness are close to her own, they weren’t perfect. The scripted avatars reproduced mannerisms convincingly, but the interactive version only answers questions directly about the prepared material. Queries outside that scope are redirected back to the topic at hand. Friends described the avatar as both fascinating and unsettling.

While the AI twin performed well in its envelope, it has clear boundaries. Its replies are deterministic—every answer is pre-trained. There’s no freestyle conversation. That means no spontaneous personality or responses beyond trained content. The author sees potential for both benefit and risk, warning that more open-ended avatars could lead to confusion about what’s real.

Broader context: avatars, trust, and AI’s role in communication

Synthesia launched earlier this year at a $4 billion valuation and crossed $100 million in annual recurring revenue. Apart from personal avatars, its products include “Roleplay Sessions,” where users role-play conversations with AI avatars, and APIs that let enterprises integrate avatar and voice-video models with their own tools. These offerings show how demand is rising for digital human tools in training, PR, communications, and beyond.

The journalist reflects on what this means for media: Could avatars replace human anchers or reporters? While some people are immediately resistant, others are open to the idea of a hybrid future where digital and human presences coexist. Central to this debate is trust—something built over time through human interaction. AI that tries to replace that could erode credibility, the author worries.

Life-like digital twins may offer new ways to communicate. They could handle routine tasks during absences, enhance training, or scale customer interactions. But with their rise come questions about identity, personhood, and ethical boundaries. In short: cool tech—yet complicated to live with.

My own takeaway? As avatars become more lifelike and interactive, they’ll force us to rethink what we expect from human connection, public communication, and authenticity. What gets gained when digital perfection replaces human nuance—and what might be lost—is something we need to watch closely.