THE VOICE FILE
03 SEP 2026 / Deepdub launches Phantom Z 3.4 Conversational16 APR 2026 / Agentic Dubbing Co-Worker announced

COMPANY / AI & MEDIA / DEEPDUB

Deepdub wants your voice to travel without you

An actor’s voice can cross a border. An invoice can trip up a robot. Deepdub is building a business around the awkward details of making machines speak for people.

Consider an invoice for $1,240. To a person, it is a sum of money. To a speech model, it is a small examination: recognize the currency, expand the number, put the words in the right order. A pleasant voice that fails this examination can make a customer-service call unpleasant very quickly. Deepdub’s September 2026 model announcement dwells on exactly this sort of detail. It is a curious destination for a company associated with dubbing films. Yet the distance between a movie scene and an invoice is shorter than it looks. Both require a machine to understand what a voice is doing.

The story in four beats
  • The work: AI dubbing, multilingual speech and voices for conversational agents.
  • The buyer: media owners, localization teams and enterprise developers.
  • The offer: a virtual studio, managed production and voice APIs.
  • The useful wrinkle: human direction and licensed voices remain part of the production.

01 / THE ORIGINAL PROBLEMTwo brothers, an awkward translation problem

Brothers Ofir and Nir Krakowski founded Deepdub in 2019. Ofir became CEO; Nir became CTO. Their backgrounds were technical rather than theatrical: Ofir worked in the Israeli Air Force’s technology organization, while Nir helped establish the Shin Bet’s cyber unit. They brought that experience to a business in which a technically correct result can still sound wrong.

A film travels easily as a file. Its dialogue travels with considerable luggage. Translation must fit the scene; phrasing must fit the available time; a performance must survive both. Deepdub’s proposition was to use AI within localization, helping owners sell their work to audiences who speak another language.

The commercial opportunity was already visible when Insight Partners led a $20 million Series A in February 2022. Deepdub had a multi-series deal with Topic to bring foreign television into English. The financing was earmarked for sales, delivery and research in Tel Aviv. For a distributor, localization turns the question from “Can we release this?” into “Which audience can justify the next production bill?”

Deepdub’s founders standing between Captain America and Pink Panther figures
Two founders. Two conspicuously silent colleagues. Deepdub’s own founders’ photograph gives the voice business a suitably theatrical supporting cast.

02 / THE PRODUCTION FLOORA studio with more than one door

Deepdub GO, launched in July 2023, widened the invitation. Advertising agencies, educators and creators could use a virtual studio to transcribe, translate, generate voices and mix audio. They could edit the intermediate work, guide a delivery with a recording, or prompt its emotional expression. At launch, the company advertised 65 languages. That historical figure should not be confused with the coverage of every subsequent model.

The company also offers managed, “White Glove” production, with adapters and producers overseeing the work. Developers can integrate speech through APIs. These are different purchases: a workspace, an outcome, or a component inside another product. Published terms describe subscription packages; managed work depends on scope. Even the free API has a commercial boundary: its outputs are for internal testing.

The combination is useful positioning, although competitors have similar ambitions. ElevenLabs also offers dubbing and human-edited production. Deepdub’s case therefore rests on the fit of its voices, controls and production service to a particular buyer’s work, rather than on voice cloning alone.

03 / THE DIFFICULT SCENEThe scene that needed a person

Deepdub’s account of Q Studios Berlin’s English version of Freier Fall supplies a useful test. Human actors recorded directed English performances. Voice-to-voice modeling applied the original actors’ vocal identities; manual lip sync and sound work followed. The division of labor mattered.

Freier Fall / production sequence
  1. Cast target-language actors
  2. Direct and record performances
  3. Apply Deepdub voice modeling
  4. Complete manual lip sync
  5. Mix sound and review
The yellow step is the AI voice-modeling stage. This project retained human performance and manual finishing.
The machine gets a speaking part. It does not get the director’s chair.

The troublesome material was emotionally demanding: sarcasm, dark humor and crying. Around sixty lines needed a three-hour recording session. The case study says the result satisfied the original actors and reports dubbing costs at roughly one-third of a traditional approach. Those are project-specific claims, with savings redirected into finishing. Original unmixed audio and carefully selected references helped; this was a hands-on production.

The lesson is practical: budget for exceptions, test difficult scenes early, and keep a route back to recording. A team without suitable source material, performance direction or permission to use the voices cannot simply assume the same result.

04 / THE RIGHTS BEHIND THE SOUNDThe voice has an owner

Deepdub’s royalty program makes a less glamorous part of the business explicit. Artists submit recordings, producers assess them, and approved voices become available for selection. The stated arrangement pays artists when their voices are used in projects, with rights governed by the agreement. A voice bank is a collection of permissions as well as sounds.

Security belongs in the same conversation. Deepdub received Trusted Partner Network content-security accreditation in 2023. Its published ethics policy says training data is collected legally; its localization offering says customer data is isolated and excluded from model training. For an unreleased film, those provisions can matter before anyone presses play. They also explain why enterprise buyers evaluate a vendor’s handling of assets alongside the output.

05 / A DIFFERENT AUDITIONNow the invoice gets a speaking part

That enterprise pitch now extends to agents. Deepdub supplies speech to Notch’s contact-center platform and announced a Wonderful partnership in 2026. The API can adjust delivery, tempo and accent. The application still needs to manage the conversation; speech generation is one part of the system.

“Deployments don’t stall on the 95% a model gets right”

OFIR KRAKOWSKI / PHANTOM Z 3.4 LAUNCH, SEPTEMBER 2026

Phantom Z 3.4 emphasizes spoken numbers, dates and context-sensitive Hebrew pronunciation. Deepdub reports a 150-millisecond p95 time-to-first-audio in real-time mode at 48 kHz. That measures the speech component under its stated conditions, rather than the time needed to resolve a customer’s request. An April 2026 studio co-worker announcement tackles another task: organizing segments, refining dialogue and checking exports under creative supervision.

Phantom Z 3.4 / company-reported150 ms

P95 time-to-first-audio in real-time mode at 48 kHz. A speech-delivery measure, not a guarantee for an entire agent stack.

For a prospective user, the sensible audition follows the actual job. Try the surname, the hurried interruption, the joke that depends on emphasis. Ask a native speaker to listen. Count corrections and human handoffs along with generated minutes. Deepdub’s most interesting promise is that voices can travel. Its production work shows how much attention a convincing journey still requires.