David AI · Audio data for voice models2024 · Founded in San Francisco2025 · $50M Series B2026 · Human preference voice leaderboardDavid AI · Audio data for voice models2024 · Founded in San Francisco2025 · $50M Series B2026 · Human preference voice leaderboard

Company profile / Voice AI

The $1,000 Conversation That Became an Audio Business

David AI began with a small robotics request and a stubborn recording problem: two people talking are much harder to capture cleanly than they sound. Today it licenses the resulting kind of data to voice AI labs and tests whether their machines are any good at talking back.

The first customer wanted a robot to listen. Tomer Cohen and Ben Wiley had barely started David AI, and their initial tool was a phone-calling app built over a weekend. It let people talk while the founders collected the kind of conversational audio a robot might need. The pilot brought in about $1,000. In startup folklore, such a sum usually appears as comic relief. Here it was a clue: a customer would pay for properly recorded conversation because the plentiful audio already online was surprisingly poor raw material.

The short version

  • David AI collects and licenses speech datasets for labs building voice models, rather than selling a consumer voice assistant.
  • Its early technical lesson: overlapping speakers need to be recorded on separate channels at the source.
  • Its catalog spans two-person talk, 15-plus languages, group conversations and expert dialogue.
  • In 2026 it added public voice-model evaluation, including about 153,000 human preference ratings.

Cohen and Wiley met at Scale AI, a company that helped make training data a commercial category. They entered Y Combinator’s summer 2024 class with an idea about voice. During customer conversations, a humanoid robotics team made the opportunity concrete. The founders’ account of that first pilot, told in a YC interview, shows how their target market widened: a robot might need to understand a request, but so might a wearable, a game character, or a phone agent. The common need was speech that resembled life rather than a voice actor reading one sentence in a silent booth.

YC interviewer Diana Hu with David AI founders Tomer Cohen and Ben Wiley in an interview thumbnail
Two founders, one interviewer, and a headline about the data beneath the voice. The robotics customer is the one you cannot see.

01 / The recording problem

Everyone talks over everyone

The web is full of talk. Podcasts, videos and meetings make it tempting to imagine that voice AI has all the training material it could want. But a transcript misses timing, tone and the awkward little overlap when one person starts before another has finished. A single mixed recording also makes it hard to tell which sound belongs to which speaker. Wiley put the supply problem bluntly in the YC interview: “There’s no real, like, common crawl for audio.”

Their first technical obstacle was channel bleed. Software can try to pull voices apart after a recording, but the founders found that route inadequate for the models they wanted to serve. So they collected conversation with speakers separated at capture. It is a small operational detail with large consequences: the model can learn turn-taking and interruption from the conversation while still knowing whose voice made each sound. The rest of the job is hardly simple. Accents, dialects, microphones, rooms, emotions and language all change what a sentence sounds like.

Why capture matters / schematic
Speaker A / clean channel
Speaker B / clean channel
Two streams, one conversation. Separating speakers during recording preserves the messy timing without forcing a model to guess who spoke.

This is the company’s real trade. It produces audio data and metadata for the people making speech recognition, translation, synthesis and conversational AI systems. Its public catalog presents four different shapes of speech rather than one generic pile of sound. Converse is natural English dialogue between two people on separate channels. Atlas follows that format across more than 15 languages and records accent and dialect information. Chorus brings in three or more speakers to help train separation and diarization. Dialog collects conversations among experts across domains.

Converse

Two speakers, separate channels, natural English dialogue.

Atlas

15-plus languages with accent and dialect metadata.

Chorus

Three or more voices for separation and diarization.

Dialog

Expert conversations across subjects.

02 / The business

A library with a sales desk

David AI turned that collection work into licensed products. A prospective buyer describes a use case, receives sample data, and negotiates access to a dataset. The company says an existing dataset can be delivered within one or two days after a license is agreed. It also works with research teams on new formats. Prices are not public, so the $1,000 first pilot is a founding detail, not a current rate card.

The model matters because a dataset made for one narrow request can become a repeatable asset if other teams need the same shape of speech. David AI describes its process as a sequence: hypothesize which capability a model lacks, design the data, run a small collection, evaluate the effect, then scale. That turns the expensive business of gathering voices into a product decision. Customers buy access to the resulting work instead of hiring a team to reinvent every recording rig and quality check.

01Hypothesize a capability
02Design the data
03Collect a small set
04Evaluate and revise
05Scale the collection
06License and improve

By May 2025, the company said its library held roughly 100,000 hours in more than 15 languages. In its Series A announcement that month, it said it had passed an eight-figure annual revenue run rate and served most of the seven largest US technology companies commonly called the “Magnificent Seven,” plus nearly every leading audio AI lab. The customer names and exact sales figures are private, so those claims remain company-reported. What is public is the funding sequence: $5 million in seed financing in January 2025, $25 million in May, and a $50 million Series B that October. Meritech led the last round, with NVIDIA participating.

$1kEarly robotics pilot, per founders
100kApproximate audio hours reported in 2025
$80mPublicly announced funding rounds

03 / A new test

A voice can be correct and still be wrong

In 2026, David AI began publishing research on a related problem: how to tell whether a speaking machine is good. An answer can contain the right facts and still sound wooden, rude or strangely delighted by bad news. In a July study, its researchers compared audio-model judges with human listeners. The automated judges tracked people more closely on the words spoken than on acoustic qualities such as naturalness and emotional fit. That finding gives David AI a second kind of product to sell: evaluation shaped by what listeners actually hear.

Its September single-turn speech-to-speech leaderboard used 819 human-recorded prompts, seven model families, 283 participants and about 153,000 comparative ratings. The researchers tested content and manner, from instruction following and pronunciation to conversational register and engagement. The result was a reminder that “good voice” is not one number. Their published ranking put GPT-Realtime 2.1 first on this prompt distribution, but the team also flagged methodological limits, including the effect of chosen voices and the fact that single-turn tests cannot describe a long conversation. Sensible evaluation starts by knowing what its score leaves out.

“Voice AI apps are only as good as the models underneath them, and the models are only as good as the data underneath them.”Tomer Cohen, YC Root Access interview

David AI sits below the visible application layer. The labs and tech companies using its data may put their own brands on the assistant, call center agent or robot. Its closest alternatives are collecting audio in-house, licensing public corpora, or hiring a general data supplier. The company’s argument is that audio has enough peculiar failure modes to reward a specialist: know the use case, record the right kind of speech, inspect the result in a model, and repeat. The founders’ Scale AI experience gave them a view of data operations; their bet was that speech deserved a shop of its own.

The approach has boundaries. A lab that already owns recordings matching its target speakers, devices and acoustic conditions may have little reason to license another library. A buyer needing exclusive rights to a peculiar voice or setting would need a custom collection, with different economics from a shared catalog. Even David AI’s 2026 leaderboard is deliberately single-turn; a team building a companion that must remember a conversation would need longer tests. These are useful constraints, because they keep the product question specific: which data is missing, and will adding it improve the model?

There is a useful lesson in the $1,000 pilot for anyone trying to build an infrastructure company. The check confirmed a buyer; the botched assumptions about easy audio confirmed a problem. David AI did not begin by needing a grand theory of all voice interaction. It needed to learn why a robot’s conversation was hard to record, and then make that answer useful to more than one robot. A voice model may eventually sound effortless. The effort, as this company demonstrates, is often filed under “data.”