Gil Perry and Sella Blondheim returned from a long trip through Central and South America with wonderful photographs and an odd problem: they did not want to post them. Both understood enough about facial recognition to know that a holiday picture was also biometric data. A password can be replaced. A face is stubbornly non-refundable.
So in 2017, Perry, Blondheim and computer-vision engineer Eliran Kuta founded D-ID - short for “de-identification” - to make faces unreadable to recognition systems while leaving the photos apparently unchanged to people. The first pitch was privacy. The first product sold to companies storing lots of images, with prices ranging from about $3 to a few cents per processed picture, depending on volume.
Today D-ID does almost the photographic negative of that job. Give its software a portrait and a script, and it makes the face speak. Give it a knowledge base and a language model, and that presenter can answer questions in real time. The company has become a useful case study in how to pivot without throwing away the hard part.
01 / The reversal
The product changed. The capability did not.
D-ID's original software was not publicly declared a technical failure. It won pilots, paying customers and a $13.5 million Series A during the first wave of Covid-19. But enterprise privacy was a narrower wedge than the opportunity forming around synthetic media. The team discovered that its ability to model faces could be used not only to conceal identity, but to animate it. MyHeritage supplied the decisive public proof.
In 2021, MyHeritage launched Deep Nostalgia, built with D-ID's Live Portrait technology. Old family photographs blinked, tilted their heads and seemed to notice the descendants staring back. Nearly 100 million faces were animated in eight months. The effect was sometimes moving, sometimes uncanny, and extremely shareable. It showed that facial synthesis could create an emotional consumer experience at enormous scale.
Same engine, opposite outcome
What changed the founders' minds was convergence. Face animation improved. Voice cloning made presenters multilingual. Then large language models could produce answers quickly enough to make a digital person more than a pre-rendered puppet. Perry has said the team decided around 2019 and 2020 to move from facial privacy toward entertainment-quality digital humans. LLM progress later turned those humans into a proposed interface for AI.
“We saw that voice cloning and digital humans would soon be able to move, speak, and interact.”Gil Perry, co-founder and CEO
02 / What it sells
A studio for creators, pipes for developers, faces for software
The easiest doorway is D-ID Studio. A marketer, trainer or teacher selects a stock avatar or creates a custom one, types a script, chooses a voice and generates video without a camera, actor or editing suite. Video Translate takes an existing recording, translates the speech, clones the speaker's voice and changes the lip movement to match. Video Campaigns personalizes messages and streams each variation only when the recipient clicks - a small but sensible economic detail, because unopened emails do not consume rendering credits.
For developers, the API and client SDK are the more consequential products. They embed streaming presenters and visual agents in websites or applications. A company can connect an agent to approved documents, an LLM and workflow hooks. The avatar listens, retrieves an answer, speaks it and can collect a lead or hand the user to a person. Microsoft says companies had created more than 150,000 D-ID visual agents by May 2025, producing 1.8 million messages and 340,000 minutes of interaction.
V4 Expressive Avatars add selectable sentiments, tighter lip sync and claimed end-to-end conversational latency below half a second. Agentic Videos apply the same system to recorded content: pause a training clip, ask the presenter a question, get a contextual answer, then continue. It is a more interesting bet than merely making another talking-head generator. Video becomes queryable.
The offer ladder - access expands with usage
*Monthly equivalent on the annual API plans observed during research. Credits, minutes, licensing and watermark rules vary; enterprise pricing is custom.
That pricing reveals the business model: free trial, usage credits, recurring SaaS tiers and negotiated enterprise contracts. Higher plans add commercial rights, more minutes, custom avatars, voice clones, branding, faster processing and security or integration support. The API creates distribution through other software, while Studio captures users who never want to read documentation.
03 / The market
Not every message needs a film crew. Not every message needs an avatar, either.
D-ID sits in a crowded AI-video market beside Synthesia, HeyGen, Tavus, Hour One, Colossyan and DeepBrain AI. Its alternatives also include ordinary production crews, localization agencies, screen recordings, a chat bubble and, frequently, a well-written web page. D-ID's distinction is the combination: animate from a single photo, generate through an API, stream in real time, translate the speaker and let the presenter converse.
The 2025 acquisition of Berlin-based simpleshow, for an undisclosed amount, makes the positioning clearer. simpleshow had spent years helping enterprises turn complicated material into short explainer videos. D-ID supplies the digital presenter and live conversation layer. Together they can help a company explain a policy, then allow an employee to question it. The combined company says it serves more than 1,500 large enterprises, including hundreds of Fortune companies.
Named customers span PepsiCo, Fidelity, J.P. Morgan, SoftBank, NTT, Deutsche Telekom, PwC, Deloitte, AXA and Gameloft. Use cases are less science fiction than the “digital human” label suggests: multilingual onboarding, product explanation, lead qualification, employee training and customer support. PepsiCo's Gatorade Sports Science Institute, for example, used a visual agent as a hydration consultant. Southern Illinois University School of Medicine built a virtual patient for clinical practice.
The sensible question is not “Can we make a digital person?” It is “Which repeated conversation is expensive, inconsistent or inaccessible today?”
04 / The copyable playbook
Steal the sequence, not the surface effect
A talking face is memorable in a demo. In production, the unglamorous pieces matter more: rights to the face and voice, current source material, response latency, an escalation route and a metric tied to an actual job. D-ID's own path suggests a practical sequence for companies experimenting with visual agents.
The five-step theft
- Choose one repeated, high-volume conversation - onboarding questions beat a vague “AI ambassador.”
- Start with approved scripts and a narrow knowledge base before opening the agent to broader material.
- Use one consented presenter, then localize language and format instead of rebuilding every video.
- Connect an action: book the meeting, retrieve the policy, capture the lead or route to a human.
- Measure completion, comprehension, escalation and cost per resolved task - not avatar views.
This approach copies D-ID's strongest strategic move. Keep the expensive capability constant and change the outcome around it. The founders had computer-vision expertise, patents and a team that understood facial data. They did not restart as a generic chatbot company. They turned the same technical asset from a privacy filter into a media engine, then connected that engine to language models.
05 / Where the face falls flat
The conditions that make the trick stop working
A cloned face or voice without explicit rights is a liability, not a productivity tool.
A photorealistic hallucination may feel more authoritative than text and therefore do more damage.
If the agent cannot solve, route or retrieve, it adds theatre to a chat box.
Even half a second matters in turn-taking; poor networks can make a “human” interface feel less natural.
Users should know when a presenter is synthetic and what will happen to their conversation data.
Grief, crisis, medicine and consequential advice need accountable people, not just warmer pixels.
D-ID's roots give it a credible reason to care about those constraints. The company publishes an ethics pledge around consent, transparency, privacy and anti-deception, and its Azure architecture uses moderation and security tooling. Yet the central risk cannot be engineered away completely: a human face lends emotional weight to whatever comes out of the model behind it.
That is also why the company is worth watching. Its first act asked how to keep machines from recognizing us. Its current act asks whether we will understand machines better when they resemble us. Between those ideas sits a business with usage-based software, enterprise distribution, a large developer community and a peculiar historical advantage: D-ID has already spent years thinking about what a face means to a computer.