When Alexandru Voica, head of corporate affairs at the video-generation startup Synthesia, sent a link this summer introducing the newest addition to the company’s public relations team, the result was immediately arresting. It was an interactive virtual avatar of Voica himself, meticulously trained to field common press inquiries about Synthesia, its underlying technology, and its core operational framework. Just a day prior, during an industry panel discussion, public relations professionals had debated the acceptable boundaries of AI-generated text in media pitches. Yet Voica’s interactive avatar bypassed that entire debate, presenting what felt like the absolute final boss of integrating artificial intelligence into media relations and corporate communications.

By September, Synthesia invited visitors to tour its freshly established office space in New York. Originally founded in the United Kingdom, Synthesia has rapidly emerged as a dominant force in the digital avatar ecosystem, alongside competing startups like D-ID, HeyGen, and Colossyan. The company achieved a staggering $4 valuation milestone earlier this year, according to recent financial disclosures, and reported crossing the $100 million threshold in annual recurring revenue last year.

Synthesia primarily enables enterprise organizations to construct interactive training videos using high-fidelity AI avatars. More recently, the company expanded its product suite by launching an agentic platform called Roleplay Sessions. This tool allows corporate employees to practice complex professional scenarios—such as high-stakes sales pitches—with an interactive AI avatar that dynamically responds to inputs, evaluates performance, and provides detailed scoring.

When visiting the new New York office, the opportunity arose to test the technology firsthand. Asked if the experiment should extend to creating a custom AI avatar, the response was immediate. Naturally, the prospect of generating a digital twin held immediate appeal, especially on a day when the outfit was well-chosen and the hairstyle was completely in place.

Until encountering this personal digital twin firsthand, personal sentiment toward avatars had leaned largely toward indifference. Yet it has become increasingly clear that these digital replicas are destined to weave themselves into the fabric of everyday online existence. Already, social media users on platforms like Instagram have begun generating digital likenesses to streamline content creation. This intersection of technology and identity remains deeply compelling, which ultimately dispelled any hesitation about presenting a personal digital twin to the public.

This marks the first time Synthesia has ever constructed a digital avatar for a journalist, or indeed for anyone outside of internal personnel like Voica. The interactive version was specifically trained on a previously published investigative story detailing why venture-backed startups statistically commit more fraud than their non-VC-backed counterparts. True to its programming, the interactive model will strictly answer questions concerning the contents of that specific report.

The creation process began inside a compact, professional film studio nestled directly within Synthesia’s office. Production staff captured numerous high-resolution photographs alongside a two-minute voice recording. After formal consent was granted, digital generation commenced, successfully yielding what can only be described as digital clones. The team produced both personal avatars—unidirectional replicas designed to recite arbitrary text supplied by the user—and interactive avatars capable of real-time listening and verbal response. Versions were generated both with and without glasses to provide versatile visual options.

Once the source material—the venture fraud article—was selected and fed into the system, Synthesia’s engineering team assembled the interactive agent. The underlying architecture relies on a sophisticated tech stack combining voice-to-text, agentic language models, text-to-voice, and advanced video synthesis. While Synthesia’s proprietary video and voice models power the core infrastructure, the platform also permits enterprise customers to integrate third-party alternatives from specialized labs such as Cartesia, ElevenLabs, Google, or OpenAI. Furthermore, corporations retain the flexibility to host these digital avatars on their preferred cloud infrastructure or rely on Synthesia’s managed hosting services.

The operational workflow relies on a coordinated sequence of artificial intelligence models. The voice-to-text engine converts spoken queries into written text. Next, the agentic language model interprets the text and determines the appropriate logical response or action. A text-to-voice model subsequently translates that response back into natural-sounding audio. Finally, Synthesia’s proprietary video model animates the visual avatar, synchronizing facial movements and expressions with the generated speech.

Broadly speaking, Synthesia’s commercial offerings fall into three distinct categories. The first is a traditional video-creation and distribution platform featuring classic avatars, where users input a written script for the digital figure to recite. The second is the agentic platform encompassing tools like Sessions, which facilitates interactive surveys and roleplaying exercises. The third is an API-driven ecosystem enabling developers to combine Synthesia’s video and voice primitives with external services to build custom interactive applications.

Constructing the customized avatars required just a couple of days. Initial testing began with the personal, unidirectional avatars by inputting a generic script to evaluate the fidelity of the synthetic voice. The text described the arrival of autumn in New York, a personal favorite season. The resulting vocal simulation proved remarkably accurate, successfully avoiding the slight hoarseness present during the initial live recording session.

Displaying the personal avatar to non-tech-savvy friends elicited a mixture of fascination and unease. Reactions shifted when introducing the interactive avatar, which operates deterministically—meaning it strictly adheres to its training parameters and refuses to deviate from the venture fraud article. Attempts to probe the model with unrelated personal inquiries, such as prior career history before joining TechCrunch or residential neighborhoods in New York, were met with polite redirection back to the core subject matter.

While friends observed that the voice on the interactive model lacked the exact cadence of a natural human delivery and the visual likeness lagged slightly behind the personal avatar, the overall execution remained uncomfortably close to reality. Family members proved equally intrigued. A parental review yielded high praise, accompanied by repeated, unsuccessful attempts to trick the model with obscure personal trivia known only to family members. Every deviation was deftly parried by the system, which systematically steered the conversation back to venture-backed startup fraud. A lighthearted joke followed from family testing: the realization that parenting responsibilities somehow expanded without warning into the digital realm.

Encounters with this technology naturally prompt broader reflections on the future trajectory of journalism. Would media consumers accept daily news broadcasts presented entirely by artificial avatars? Initial reactions from financial investors contacted for perspective were swift and absolute in their rejection. Such skepticism aligns with a broader public pushback against the rising tide of unvetted AI-generated content flooding social media and information-sharing networks.

Yet other industry observers remain less certain. Could digital avatars eventually augment—or even fully replace—working journalists? Would corporate chief executive officers prefer engaging with an AI avatar of a reporter rather than a living human counterpart?

The core appeal of a journalistic career lies deeply in human connection, investigative reporting, and the exploration of complex topics. At its foundation, journalism relies upon public trust—an intangible quality that appears exceedingly difficult to outsource safely to an algorithm.

Beyond media and publishing, however, the commercial appeal of self-cloning is glaringly obvious in corporate environments. The prospect of bypassing the daunting backlog of emails and catch-up work following a holiday or extended vacation—delegating routine queries to an always-on digital proxy—carries undeniable utility.

How avatar usage will ultimately permeate corporate America remains to be seen. Owning a digital twin leaves one with genuinely mixed emotions. After the initial novelty wears off, observing the avatar sit in silence as it waits for input triggers an eerie introspection. The lingering question remains whether the digital double might eventually blink, offer an unscripted remark, flash an independent smile, or signal a genuine awareness of its own existence.

Because these deterministic digital twins are strictly bounded by their programming, such autonomous self-awareness will never occur. Nevertheless, it is easy to imagine how vulnerable users might tumble into a form of digital psychosis when interacting with non-deterministic avatars powered by unconstrained, conversational chatbots.

Reflecting on these developments with venture capital investors brings a distinct generational perspective to light. Generation Z may struggle to embrace digital twins comfortably, given their distinctly science-fiction origins that have abruptly transitioned from cinematic fantasy into tangible reality. Yet, compared to humanoid robotics, digital avatars feel decidedly less jarring; when interactions cross into unsettling territory, the simple solution remains logging off.

Until the broader technological landscape settles, however, the personal avatar stands ready to deliver a concise rundown of the week’s top headlines across the publication.

By Basiran

Leave a Reply

Your email address will not be published. Required fields are marked *