When Alexandru Voica, head of corporate affairs at video-generation startup Synthesia, sent over a link this summer to the newest member of his PR team, the reaction was immediate surprise. The new hire was not a person at all, but an interactive virtual avatar of Voica himself, trained to answer common press questions about Synthesia โ what it does, how it works and why the company believes digital humans are becoming a serious business tool.
For a journalist who had just been on a panel discussing whether AI-generated text should be used in pitches, the avatar felt like a leap beyond the usual debate. It was, in effect, a talking digital stand-in for a corporate spokesperson โ a polished, responsive version of the kind of automated outreach that already fills inboxes. The experience framed a larger question now confronting the media, public relations and enterprise software worlds: when does AI-assisted communication stop being a convenience and start becoming a substitute for human interaction?
That question was no longer theoretical in September, when Synthesia invited the reporter to its new office space in New York. The company, originally based in the U.K., has become one of the best-known names in the digital avatar market alongside rivals such as D-ID, HeyGen and Colossyan. Earlier this year, it reached a $4 billion valuation, and last year it said it had crossed $100 million in annual recurring revenue, a milestone that signals real commercial traction rather than speculative hype.
Synthesia's core business is helping enterprises create training and communication videos using AI avatars. More recently, it launched Roleplay Sessions, a product designed to let employees practice scenarios such as sales pitches with an interactive AI avatar that responds in real time and scores their answers. The company is positioning these tools not as entertainment, but as infrastructure for corporate learning, internal communications and customer-facing workflows.
At the office opening, the company offered the reporter a chance to create a personal avatar. The answer was immediate: yes. The appeal was partly practical, partly personal. The outfit was good that day, the hair was in place, and the idea of a digital twin had a certain irresistible novelty. What might once have seemed like a gimmick now looked more like a preview of how people may increasingly present themselves online.
The result was something Synthesia said it had never done before: a digital avatar for a journalist, and the first one outside of Voica's own. The avatar was trained only on the reporter's story about why venture-backed startups commit more fraud than non-VC-backed startups, and it was designed to answer only questions related to that piece. Users were told they could ask it why the story was written, what the research paper was about and what the researchers found.
Creating the avatar required a visit to a small film studio inside Synthesia's office, where the company took numerous photos and recorded two minutes of the reporter's voice. Consent was required, underscoring one of the most important issues in the avatar economy: these systems depend on likeness, voice and identity, and their legitimacy rests on permission. From that material, Synthesia created a personal avatar that simply reads a script, as well as two interactive avatars that can listen and respond, each version produced with and without glasses.
The technical stack behind the interactive version combines voice-to-text, language, text-to-voice and video models. Synthesia uses its own video and voice models, but customers can also choose alternatives from other providers including Cartesia, ElevenLabs, Google and OpenAI. Enterprises can host avatars on their own cloud infrastructure or pay Synthesia to host them, a flexibility that reflects how the company is trying to fit into corporate IT environments rather than operate as a closed consumer app.
In practice, the avatar pipeline works in layers: speech is converted into text, a language model interprets the request and decides how to respond, the answer is turned back into audio, and a video model animates the avatar as it speaks. The result is a system that can appear conversational while remaining tightly controlled by the underlying training material and product design.
That control may be the key to Synthesia's pitch. The company is not promising a fully autonomous digital person, but a managed, branded, enterprise-safe version of one. Still, the broader implications are hard to miss. If a startup can build a convincing avatar of a journalist to explain a single article, it is not difficult to imagine the same technology being used by executives, sales teams, educators, recruiters and public-facing brands at scale.
For now, the digital twin remains a novelty with a narrow remit. But it also serves as a demonstration of where the market is heading: toward a world in which people may increasingly delegate parts of their communication to synthetic versions of themselves. What once sounded like science fiction is now being packaged as a product, sold to businesses and tested in public view โ one avatar at a time.
