The speech-AI company that turns spoken words into text - and then into meaning - for the developers building the next wave of voice software.
Software is learning to listen. AssemblyAI builds the models that let it - and hands them to developers as an API.
When Dylan Fox left a research-engineer role at Cisco in 2017, the conventional wisdom held that speech-to-text was a finished problem. The big cloud providers had transcription APIs; the accuracy was "good enough." Fox, who had watched the first wave of voice assistants like Amazon Alexa arrive, took a different reading: good enough was not the same as solved, and almost nobody was building speech tools with developers in mind.
That gap became AssemblyAI. The company builds deep-learning models that convert audio and video into text, then layers on what it calls audio intelligence - the ability to tell you who spoke, what a conversation was about, how people felt, and which words should be redacted before the data goes anywhere. All of it is delivered through a single API that a developer can wire up in an afternoon.
The result is infrastructure most people never see. When a meeting notetaker produces a summary, when a contact center scores a call, when a medical scribe drafts a clinical note, there is a reasonable chance AssemblyAI's models did the listening. The company reports processing more than a million hours of audio a day and powering billions of end-user experiences.
Started in Y Combinator and backed early by Daniel Gross, AssemblyAI grew from a side project into a Series C business without ever pivoting away from its first customer: the engineer who wants human-level speech understanding and would rather not train a model to get it.
AssemblyAI sells models, not dashboards. Each product is an endpoint developers call and pay for by usage.
The flagship batch model, tuned for high-accuracy transcription of meetings, media and calls with punctuation, formatting and speaker labels.
Released 2024Low-latency streaming transcription near 300ms - fast enough for live captions and conversational voice agents. Multilingual support added in 2025.
Released 2025A customizable model that blends LLM-style reasoning with audio processing to understand context and vocabulary, not just recognize sounds.
Beta 2025Runs Claude models over a transcript to return summaries, chapters, sentiment, Q&A and action items in a single API call.
Since 2023Speaker diarization, sentiment analysis, topic detection, auto chapters, entity detection and content moderation as composable features.
Since 2021Turn detection and streaming for conversational agents, plus PII redaction, profanity filtering and moderation before data leaves the pipe.
2022-2025Raw transcription has become close to a commodity - plenty of providers can turn clean audio into a wall of text. The hard, valuable part starts after that: separating speakers on a noisy call, catching a product name the model has never heard, judging sentiment, and stripping out a customer's credit-card number before it is stored.
AssemblyAI's bet is that developers will pay for the second layer. Rather than shipping a transcript and walking away, its stack keeps going - into diarization, summaries, redaction and reasoning over what was said. For a call-analytics startup or a healthcare scribe, that difference is the entire product.
AssemblyAI leans on documentation, sample code and a widely followed YouTube channel to win engineers - then backs it with the guarantees enterprises require:
Company-reported figures for Universal-Streaming vs. a competing model. Lower is faster.
AssemblyAI's users are the companies building voice into their products: meeting notetakers, contact-center and conversation-analytics tools, medical scribes, media and podcast platforms, and a fast-growing set of real-time voice agents. Because it sits underneath, its logo rarely appears on the surface - but the products do.
Revenue is usage-based: customers pay per hour or minute of audio processed, plus a charge for each intelligence feature they switch on. A self-serve tier lets individual developers start with a credit card, while enterprise contracts add volume discounts, uptime SLAs, security certifications, EU data residency and hands-on support.
It is a classic developer-infrastructure motion - land with a single engineer running a test, expand as their product scales and its audio volume climbs. AssemblyAI has said usage grew sharply through 2025, with streaming alone reaching well over a million hours a week.
Dylan Fox starts the company to build a developer-friendly speech-to-text API.
Goes through YC with early backing from Daniel Gross.
Accel leads; the platform expands beyond transcription into audio intelligence.
Insight Partners leads a round to scale models and infrastructure.
Funds "superhuman" speech models; launches the LeMUR LLM layer.
Ships new flagship models, real-time streaming, multilingual support and EU residency.
Led by Accel. Insight Partners, Y Combinator, Keith Block, Nat Friedman, Daniel Gross, Smith Point Capital.
Led by Insight Partners.
Led by Accel, with Daniel Gross and Nat Friedman.
Voice AI in 2026 is a busy neighbourhood. Deepgram competes hard on telephony accuracy and on-premises deployment; Google Cloud Speech-to-Text wins buyers already committed to its ecosystem; OpenAI's Whisper set a strong open-source baseline; ElevenLabs comes at voice from the synthesis side; and challengers like Speechmatics and Gladia press on price and languages.
AssemblyAI's chosen lane is the developer who values accuracy, documentation and a single vendor that handles both transcription and understanding in the cloud. Where rivals emphasise bare-metal control or the lowest error rate on noisy phone lines, AssemblyAI competes on breadth of intelligence features, ease of integration and the enterprise guarantees that let a prototype grow into production. It is not trying to be the only speech API - it is trying to be the default one for teams building on top of speech.
Fox has framed the longer arc as building "superhuman" speech models: systems that eventually understand audio better than a human transcriptionist could, then expose that capability simply enough that any product can use it. Whether or not the industry consolidates around a handful of providers, AssemblyAI has positioned itself as one of the layers the rest of the stack quietly depends on.
It provides AI APIs that transcribe audio and video into text and add higher-level understanding - summaries, sentiment, speaker labels, topic detection and PII redaction - so developers can build speech features without training their own models.
Dylan Fox founded the company in 2017. He previously worked as a research engineer at Cisco and serves as CEO.
More than $113M in total across seed to Series C, including a $50M Series C announced in December 2023 led by Accel, with investors such as Insight Partners, Y Combinator, Nat Friedman and Daniel Gross.
Other speech and voice-AI providers including Deepgram, Google Cloud Speech-to-Text, OpenAI Whisper, ElevenLabs, Speechmatics, Amazon Transcribe and Microsoft Azure Speech.
Meeting notetakers, call-center and conversation analytics, medical scribes, podcast and media tools, live captions, and real-time voice agents - anywhere software needs to turn speech into text and insight.