AssemblyAI is a San Francisco speech-AI company that builds and serves models turning audio and video into accurate text, plus higher-level 'audio intelligence' like summaries, sentiment, speaker labels, and PII redaction. Founded in 2017 by Dylan Fox, it sells a developer-first API used to add transcription, real-time streaming, and voice-agent capabilities to software. The company has raised more than $113M across seed to Series C and reports processing over a million hours of audio a day for customers ranging from startups to large enterprises.
Corti is a Copenhagen-based healthcare AI company that builds clinical-grade models and APIs helping developers and health systems automate medical conversations. Its Symphony model powers speech-to-text, medical coding, clinical note generation, and agentic workflows, serving roughly 100 million patient interactions a year across providers, insurers, EHR vendors, and public-safety dispatch centers - including the NHS. Corti positions itself as privacy-compliant, healthcare-specialized infrastructure that outperforms general-purpose models on medical tasks.
SpeakIn Technologies is a voiceprint-recognition and identity-security company founded in Silicon Valley in 2015 and headquartered in Shenzhen. Using deep neural networks and its iVector engine, SpeakIn turns the unique physical signature of a human voice into a biometric key, offering both 1:1 verification (is this the right speaker?) and 1:N identification (who is speaking, out of many?). Its SDK and cloud APIs are used across banking, public security, and smart devices, with reported accuracy near 98% and customers including Tencent, Lenovo, ASUS, and China Merchants Bank.
Jason Chicola is the founder and CEO of Rev, the Austin-based speech-to-text company that pairs 50,000 freelance transcriptionists with best-in-class AI to turn spoken words into searchable text. An MIT-trained engineer who became the third employee at oDesk (now Upwork), he built Rev on a simple bet: people will pay for curated quality, and people everywhere want to work from home. Today Rev serves over 100,000 clients including 60% of the Fortune 500, and Chicola is steering the company toward AI tools for the legal system, where accuracy is not optional.
Lawrence Chen (Chen Haoliang) is the founder and CEO of SpeakIn Technologies, a Shenzhen-based voiceprint recognition company he started in 2015. An MIT dropout who previously worked on Google Glass's human-machine interaction program, Chen built SpeakIn to verify identities through voice rather than passwords, claiming 99.5% accuracy and selling into Chinese banks and public security bureaus. He was named to Forbes 30 Under 30 Asia in 2018.
Thoth AI is a Singapore-headquartered global AI data solutions company that builds the high-quality datasets, evaluation frameworks, and human-in-the-loop operations that production AI systems depend on. Spanning 200+ languages and 100+ countries, it handles data collection and annotation across images, video, 3D point cloud, text, and audio, along with RLHF, model evaluation, content moderation, and multilingual customer experience for teams building LLMs, VLMs, multimodal systems, and applied AI like robotics.
Rev is an American speech-to-text company that pairs the world's most accurate AI speech recognition with a global network of human transcriptionists to deliver transcription, captions, and subtitles at up to 99% accuracy. Founded in 2010 by six MIT-connected entrepreneurs, Rev serves over 100,000 customers and more than a million users across legal, media, education, and enterprise, and has increasingly focused its AI on the legal market with tools for depositions, evidence, and case prep.