The founding idea arrived in the negative. Gil Perry and Sella Blondheim had met during sensitive military service, and the ordinary theatre of social media - holiday snaps, profile pictures, proof that one had indeed visited South America - was not available to them in the ordinary way. Facial recognition could turn a casual photograph into a trail. They were social people living in the opening act of the selfie age, obliged to behave like fugitives from the group album.
Their annoyance became a technical question: could a photograph remain recognizable to friends while becoming useless to a recognition system? Perry, a computer science graduate of Tel Aviv University, built a prototype. With Blondheim and Eliran Kuta, an engineer with deep computer-vision experience, he formed D-ID. The name stood for de-identification, and the proposition was admirably literal. Alter the numbers beneath a face without making the face look altered.
In 2017, the company joined Y Combinator. The founders were early enough to sound alarmist. Perry later recalled that investors dismissed the threat because facial recognition did not yet seem capable enough to matter. The team kept pitching. Regulation, ubiquitous cameras and rapidly improving models soon made their concern look less eccentric. Perry wanted D-ID to become the standard layer of image protection.
The division of labor was deliberate. Perry became CEO and took responsibility for sales, business and marketing. Blondheim, the operator in the trio, led operations, product, legal and finance. Kuta built the research and development organization and directed the technology roadmap. Around them, D-ID assembled advisers whose backgrounds made the privacy claim harder to shrug off, including cryptographer Adi Shamir, privacy scholar Ann Cavoukian and former Microsoft chief privacy officer Richard Purcell.
That network supplied more than distinguished names. It connected the founders' personal concern to the institutional problem faced by companies holding huge collections of photographs. D-ID could speak to both sides of the same image: the person who did not want to be tracked and the organization obliged to protect biometric data. Perry's role was to translate between the laboratory and the buyer, a pattern that would become even more important once the company changed direction.
“We run fast and always need to stay ahead.”Gil Perry on D-ID's operating rhythm
The pivot hiding in the pixels
Startup pivots are often narrated as conversions: the old belief is discarded and a better one appears. D-ID's turn was stranger and more economical. The company changed what its face technology was for. A team trained to modify identity had also learned how to synthesize expressions, preserve likeness and manipulate visual detail. The difficult asset was not the privacy product. It was fluency in faces.
Around 2019 and 2020, Perry steered the company toward photo animation and digital humans. The direction seemed opposite to the original mission. D-ID had begun by hiding people from machine vision; now it would help machines perform a likeness. Yet the route through the technology was almost straight. The early privacy work had forced the team to understand the small signals by which algorithms read a face. Those lessons could also be used to make a still image move.
The public proof arrived through MyHeritage. Its Deep Nostalgia feature used D-ID technology to animate old family photographs. Ancestors blinked, tilted their heads and briefly seemed available to the present. More than 100 million photographs were animated. The response was not merely the usual internet fascination with a new filter. People brought memory, grief, curiosity and family mythology to the output. A technical demo had found an emotional job.
At TechCrunch Disrupt in 2021, Perry made the idea personal. He demonstrated Speaking Portrait using a childhood photograph of himself, mapping a performer's expressions onto the still image to stage an exchange between the adult founder and his younger face. It was a neat piece of founder theatre, and also a compact summary of the business: one picture, some speech and enough machine learning to trouble the boundary between archive and performance.
A face for the language model
Generative video soon collided with the rise of large language models. A scripted portrait could deliver a line; a model could invent the next line. Join speech recognition, language generation, voice synthesis and real-time facial animation, and the result begins to resemble a person on a video call. Perry calls the intended experience a natural user interface, or NUI. The acronym is his challenge to the GUI, the collection of windows, icons, menus and pointers that has trained several generations of humans to behave in ways computers understand.
His ambition reverses the lesson. The machine should adapt to the person's most practiced interface: conversation. No hunt through settings, no special syntax, no tiny help icon tucked into a corner. The user asks, interrupts, clarifies. The software replies with a voice and a face. D-ID has aimed these visual agents at training, sales, customer support and internal communication, places where explanation matters and a static answer often creates another question.
The bet carries a peculiar design burden. A button need only work. A face appears to make promises about attention, understanding and intent. The closer an avatar comes to human expression, the more carefully a company must handle consent, disclosure and the possibility of deception. Perry's public language has retained traces of D-ID's privacy origins. When the company launched Speaking Portrait, he argued for transparency and consent across the synthetic-media industry. The team that worried about machines recognizing faces now has to decide how machines should wear them.
Perry's style in interviews mixes the engineer's pipeline with the founder's slogan. He can reduce a visual agent to its sequence - speech becomes text, a language model supplies a response, the response becomes audio and video - then jump to the claim that software is moving from GUI to NUI. The compression is useful. It lets an audience see both the plumbing and the destination.
From product to platform
By 2024, D-ID was selling the shift as infrastructure rather than spectacle. The company offered a self-service studio, APIs and enterprise tools. Its position was no longer simply that anyone could make a photograph talk. Developers could build a visual agent into another product, customize it and operate it at scale. The viral magic trick had become a software layer.
In March 2025, D-ID announced a partnership with Microsoft that brought its avatars to Azure and pointed toward integrations with Teams and other business software. Microsoft later described a migration that helped D-ID expand its global reach and accelerate growth. The relationship also gave Perry's NUI thesis a practical venue: not a distant metaverse, but the applications where people already meet, learn and ask for help.
Six months later came a more traditional sign of company-building: D-ID acquired Berlin-based simpleshow. The target had spent years turning complicated corporate material into explainer videos and brought more than 1,500 enterprise customers. D-ID brought real-time avatars. Put them together and a training video no longer has to end when its script does. A viewer could interrupt, ask a question, try a role-play or take a quiz. Video, in this conception, becomes less like a broadcast and more like office hours.
The acquisition also reveals Perry's current priority. D-ID is competing for enterprise workflows, not simply for the most startling talking head. Simpleshow contributes customer relationships, libraries and a familiar way to make business explanations. D-ID contributes interactivity. Perry described the combination as more human, scalable and efficient, three adjectives that do not always coexist comfortably. The company's task is to prove they can.
Early in 2026, D-ID began folding simpleshow's workflow into its own offering. It then announced more expressive V4 visual agents and, in April, Agentic Videos: presentations designed to accept questions and respond within the experience. The vocabulary keeps moving - portrait, avatar, agent, agentic video - but the direction is consistent. Perry wants video to stop behaving like a sealed container.
The face remains the argument
There is a pleasing contradiction at the center of Perry's career. He became a founder because a face revealed too much, then built a larger company around the belief that software without a face reveals too little. One problem concerned identity. The other concerns presence. Both assume that the face is unusually powerful data.
This is why D-ID's story is more instructive than a simple tale of catching an AI wave. The company did not abandon its past when the market changed. It repurposed a hard technical education, found a consumer proof point, and kept moving toward a platform. Perry's useful instinct was to preserve the difficult knowledge while replacing the commercial frame around it.
Whether visual agents become a default interface will depend on prosaic matters: latency, accuracy, price, disclosure and whether anyone truly wants eye contact with the expense-report system. Human beings have spent decades learning to click. They may not surrender the mouse merely because an avatar has excellent lip sync.
Still, Perry's wager has the charm of an idea that can be explained without a diagram. Talk to the machine as you would talk to a person. Let it answer in kind. The rest is engineering, judgment and the occasional childhood photograph asking what became of you.