The transition from static video generation to interactive artificial intelligence took a significant step forward recently as Synthesia, the unicorn startup valued at 4 billion dollars, began deploying digital twins capable of real time conversation. While the company has long been known for its ability to turn text into video using realistic avatars, its latest developments focus on interactivity and agentic behavior.
The potential of this technology was recently highlighted when Synthesia created an interactive digital avatar for a technology journalist, marking one of the first times the company has extended its digital twin capabilities beyond its own internal staff. Unlike standard video avatars that simply read a script, these interactive versions are designed to listen, process information, and respond to questions in a specific context.
The Shift from Video to Interaction
Synthesia is part of a competitive landscape of digital avatar startups, alongside players like HeyGen and D-ID. However, the company is pivoting its strategy toward what it calls an agentic platform. Earlier this year, Synthesia hit a 4 billion dollar valuation and surpassed 100 million dollars in annual recurring revenue (ARR), fueled by a shift into the enterprise training and corporate communications sectors.
The core of this evolution is a product called Roleplay Sessions. This platform allows employees to engage in live coaching with an AI avatar to practice high stakes conversations, such as sales pitches, customer support scenarios, or performance reviews. The avatar does not just speak; it listens to the user and provides a score based on their performance, offering a scalable way for large organizations to conduct soft skills training without requiring human roleplay partners.
How the Digital Twin is Created
Building a high fidelity digital twin involves a multi stage technical process. To create the avatar for the recent demonstration, Synthesia used a dedicated mini film studio to capture numerous high resolution photos and a two minute recording of the subject's voice.
The resulting digital twin relies on a sophisticated tech stack that integrates several different AI models:
- Voice to Text: A model that converts the user's spoken words into written text.
- Agentic Language Model: A large language model (LLM) that interprets the text and determines the appropriate response or action.
- Text to Voice: A model that converts the generated response back into audio. While Synthesia has its own voice models, it allows enterprise customers to integrate third party options from labs like ElevenLabs, OpenAI, or Google.
- Video Animation: Synthesia's proprietary video models, including its fourth generation expressive avatars, animate the digital face and body in real time to match the audio.
The company offers three distinct paths for using this technology: a video creation platform for standard script based videos, an agentic platform for live sessions, and an API that allows developers to integrate Synthesia's models into their own applications.
Deterministic vs. Open Ended AI
One of the most important distinctions in the current state of interactive avatars is the difference between deterministic and non-deterministic behavior. For the recent journalist demonstration, the avatar was programmed to be deterministic. It was trained specifically on a single research article regarding venture capital fraud and was instructed only to answer questions related to that topic.
When tested with personal questions, such as where the subject lived or their work history, the model consistently redirected the conversation back to the training data. This level of control is vital for enterprise clients who want to ensure that their AI brand ambassadors do not hallucinate or share unauthorized information.
In contrast, a non-deterministic avatar powered by a standard chatbot could potentially discuss any topic, but this poses risks for corporate reputation and accuracy. The deterministic approach ensures that the "digital twin" remains a faithful representative of specific knowledge rather than a general purpose conversationalist.
The Impact on Industry and Trust
The rise of digital twins raises profound questions about the future of professional identity and the nature of trust. In the world of journalism and public relations, the idea of an avatar acting as a press officer or a news anchor is already becoming a reality. Alexandru Voica, head of corporate affairs at Synthesia, uses an interactive version of himself to handle common press inquiries, effectively acting as the first point of contact for the company.
However, the human response to these clones remains complex. While some see the technology as a way to augment human productivity (allowing a version of oneself to answer questions while on vacation) others find the experience unsettling. During testing, observers noted that while the likeness is impressive, it can still feel "creepy" or "jarring" when compared to real human interaction.
For fields like journalism, the primary obstacle may not be the technology itself but the concept of trust. While an AI can report facts or answer questions about a story, the relationship between a reporter and their audience is built on human accountability, something that cannot yet be outsourced to an algorithm.
What Happens Next
Synthesia's expansion into the United States, including a new office in New York and a presence in Seattle, suggests a focused push into the American enterprise market. As inference costs continue to fall, the company plans to move beyond large corporate clients to offer interactive roleplay tools to smaller businesses and educational institutions.
The long term success of digital twins will likely depend on how well they can navigate the "uncanny valley" and whether they can provide enough utility to overcome the initial skepticism of users. For now, the technology is finding its strongest footing in controlled environments like corporate training, where the goal is consistent, measurable improvement in human skills.
Filed under: AI, TechNews, Startups, ProductLaunches, Software, DigitalTwins