LIVE AssemblyAI processes 1M+ hours of audio every day FUNDING $50M Series C led by Accel to build superhuman speech models SHIP Universal-Streaming delivers real-time transcription near 300ms SCALE 600M+ inference calls per month, 2B+ end-user experiences REACH 99 languages supported across the platform FOUNDED 2017 in San Francisco by Dylan Fox, ex-Cisco LIVE AssemblyAI processes 1M+ hours of audio every day FUNDING $50M Series C led by Accel to build superhuman speech models SHIP Universal-Streaming delivers real-time transcription near 300ms SCALE 600M+ inference calls per month, 2B+ end-user experiences REACH 99 languages supported across the platform FOUNDED 2017 in San Francisco by Dylan Fox, ex-Cisco
Company Profile Speech & Voice AI San Francisco • Founded 2017

AssemblyAI

The speech-AI company that turns spoken words into text - and then into meaning - for the developers building the next wave of voice software.

Founder / CEODylan Fox
Total Funding$113M+
HeadquartersSan Francisco
Team~92 people
AssemblyAI brand image and logo
AssemblyAI
San Francisco, California - the voice-AI infrastructure behind meeting tools, call centers and medical scribes you have probably used without noticing.
1M+Hours of audio / day
99Languages supported
~300msStreaming latency
$113M+Total raised
The Dispatch

A quiet layer under a loud idea

Software is learning to listen. AssemblyAI builds the models that let it - and hands them to developers as an API.

When Dylan Fox left a research-engineer role at Cisco in 2017, the conventional wisdom held that speech-to-text was a finished problem. The big cloud providers had transcription APIs; the accuracy was "good enough." Fox, who had watched the first wave of voice assistants like Amazon Alexa arrive, took a different reading: good enough was not the same as solved, and almost nobody was building speech tools with developers in mind.

That gap became AssemblyAI. The company builds deep-learning models that convert audio and video into text, then layers on what it calls audio intelligence - the ability to tell you who spoke, what a conversation was about, how people felt, and which words should be redacted before the data goes anywhere. All of it is delivered through a single API that a developer can wire up in an afternoon.

The result is infrastructure most people never see. When a meeting notetaker produces a summary, when a contact center scores a call, when a medical scribe drafts a clinical note, there is a reasonable chance AssemblyAI's models did the listening. The company reports processing more than a million hours of audio a day and powering billions of end-user experiences.

Started in Y Combinator and backed early by Daniel Gross, AssemblyAI grew from a side project into a Series C business without ever pivoting away from its first customer: the engineer who wants human-level speech understanding and would rather not train a model to get it.

“Voice AI infrastructure for developers building products that transcribe, understand, and act on speech.” How AssemblyAI describes itself
Products & Services

One API, several ways to listen

AssemblyAI sells models, not dashboards. Each product is an endpoint developers call and pay for by usage.

Speech-to-Text

Universal-2

The flagship batch model, tuned for high-accuracy transcription of meetings, media and calls with punctuation, formatting and speaker labels.

Released 2024
Real-time

Universal-Streaming

Low-latency streaming transcription near 300ms - fast enough for live captions and conversational voice agents. Multilingual support added in 2025.

Released 2025
Speech Language Model

Slam-1

A customizable model that blends LLM-style reasoning with audio processing to understand context and vocabulary, not just recognize sounds.

Beta 2025
LLM Layer

LeMUR

Runs Claude models over a transcript to return summaries, chapters, sentiment, Q&A and action items in a single API call.

Since 2023
Audio Intelligence

Understanding Suite

Speaker diarization, sentiment analysis, topic detection, auto chapters, entity detection and content moderation as composable features.

Since 2021
Voice Agents & Safety

Voice Agent API + Guardrails

Turn detection and streaming for conversational agents, plus PII redaction, profanity filtering and moderation before data leaves the pipe.

2022-2025
The Problem It Solves

Hearing is easy. Understanding is the work.

Raw transcription has become close to a commodity - plenty of providers can turn clean audio into a wall of text. The hard, valuable part starts after that: separating speakers on a noisy call, catching a product name the model has never heard, judging sentiment, and stripping out a customer's credit-card number before it is stored.

AssemblyAI's bet is that developers will pay for the second layer. Rather than shipping a transcript and walking away, its stack keeps going - into diarization, summaries, redaction and reasoning over what was said. For a call-analytics startup or a healthcare scribe, that difference is the entire product.

How It Stands Apart

Developer-first, then enterprise-ready

AssemblyAI leans on documentation, sample code and a widely followed YouTube channel to win engineers - then backs it with the guarantees enterprises require:

99.99% uptime SLA SOC 2 Type 2 Healthcare BAA EU data residency 99 languages
By The Numbers - Streaming Latency

Real-time transcription, median latency

Company-reported figures for Universal-Streaming vs. a competing model. Lower is faster.

AssemblyAI Universal-Streaming~307 ms
Deepgram Nova-3 (reference)~516 ms
Who Uses It & How It Earns

Sold by the hour, spread across an industry

The customers

AssemblyAI's users are the companies building voice into their products: meeting notetakers, contact-center and conversation-analytics tools, medical scribes, media and podcast platforms, and a fast-growing set of real-time voice agents. Because it sits underneath, its logo rarely appears on the surface - but the products do.

ZoomClickUpFireflies GranolaHeyGenCallRail CalabrioLiveKitRetell ApolloMetaviewCommure

The business model

Revenue is usage-based: customers pay per hour or minute of audio processed, plus a charge for each intelligence feature they switch on. A self-serve tier lets individual developers start with a credit card, while enterprise contracts add volume discounts, uptime SLAs, security certifications, EU data residency and hands-on support.

It is a classic developer-infrastructure motion - land with a single engineer running a test, expand as their product scales and its audio volume climbs. AssemblyAI has said usage grew sharply through 2025, with streaming alone reaching well over a million hours a week.

The Story So Far

From a YC batch to infrastructure

2017

AssemblyAI is founded

Dylan Fox starts the company to build a developer-friendly speech-to-text API.

2019

Y Combinator

Goes through YC with early backing from Daniel Gross.

2021

$28M Series A

Accel leads; the platform expands beyond transcription into audio intelligence.

2022

$30M Series B

Insight Partners leads a round to scale models and infrastructure.

2023

$50M Series C + LeMUR

Funds "superhuman" speech models; launches the LeMUR LLM layer.

2024-25

Universal-2, Universal-Streaming, Slam-1

Ships new flagship models, real-time streaming, multilingual support and EU residency.

The Cap Table

$113M+ raised, four rounds

Series C • Dec 2023

$50M

Led by Accel. Insight Partners, Y Combinator, Keith Block, Nat Friedman, Daniel Gross, Smith Point Capital.

Series B • 2022

$30M

Led by Insight Partners.

Series A • 2021

$28M

Led by Accel, with Daniel Gross and Nat Friedman.

Where It Fits

A crowded market, a clear lane

Voice AI in 2026 is a busy neighbourhood. Deepgram competes hard on telephony accuracy and on-premises deployment; Google Cloud Speech-to-Text wins buyers already committed to its ecosystem; OpenAI's Whisper set a strong open-source baseline; ElevenLabs comes at voice from the synthesis side; and challengers like Speechmatics and Gladia press on price and languages.

AssemblyAI's chosen lane is the developer who values accuracy, documentation and a single vendor that handles both transcription and understanding in the cloud. Where rivals emphasise bare-metal control or the lowest error rate on noisy phone lines, AssemblyAI competes on breadth of intelligence features, ease of integration and the enterprise guarantees that let a prototype grow into production. It is not trying to be the only speech API - it is trying to be the default one for teams building on top of speech.

Fox has framed the longer arc as building "superhuman" speech models: systems that eventually understand audio better than a human transcriptionist could, then expose that capability simply enough that any product can use it. Whether or not the industry consolidates around a handful of providers, AssemblyAI has positioned itself as one of the layers the rest of the stack quietly depends on.

Speech-to-text is close to a commodity. Understanding speech is not. That distinction is the whole company. The thesis, in one line
Watch & Explore

Interviews, demos and docs

The Questions

Frequently asked

What does AssemblyAI do?

It provides AI APIs that transcribe audio and video into text and add higher-level understanding - summaries, sentiment, speaker labels, topic detection and PII redaction - so developers can build speech features without training their own models.

Who founded AssemblyAI and when?

Dylan Fox founded the company in 2017. He previously worked as a research engineer at Cisco and serves as CEO.

How much funding has AssemblyAI raised?

More than $113M in total across seed to Series C, including a $50M Series C announced in December 2023 led by Accel, with investors such as Insight Partners, Y Combinator, Nat Friedman and Daniel Gross.

Who are AssemblyAI's main competitors?

Other speech and voice-AI providers including Deepgram, Google Cloud Speech-to-Text, OpenAI Whisper, ElevenLabs, Speechmatics, Amazon Transcribe and Microsoft Azure Speech.

What can you build with AssemblyAI?

Meeting notetakers, call-center and conversation analytics, medical scribes, podcast and media tools, live captions, and real-time voice agents - anywhere software needs to turn speech into text and insight.

Find AssemblyAI

Links & social