HeyGen wants to make the camera optional. Its digital twins, translation tools and video agents turn one human performance into a production system that can speak to almost anyone.
Kalpa Labs is a San Francisco audio research lab building a generalist speech model - one system that handles speech-to-text, text-to-speech, voice cloning, and speech-in/speech-out reasoning with the kind of instruction-following and in-context learning that large language models brought to text. Founded in 2025 by ex-Google Assistant ML lead Prashant Shishodia and ex-high-frequency-trading engineer Gautam Jha, the company is a Y Combinator Fall 2025 startup aiming to collapse today's fragmented speech stack into a single, steerable model for developers and voice-agent builders.
Murf AI is a voice technology company that turns text into natural-sounding speech. Founded in 2020 by three IIT Kharagpur friends, it started as a self-serve AI voiceover studio for creators and marketers and has since expanded into dubbing, voice cloning, a low-latency text-to-speech API (Falcon), and conversational voice agents. Murf says its platform is used by more than 10 million people and 300+ Forbes 2000 companies across 190+ countries, offering 200+ voices in 35+ languages.
Speechify is the world's largest text-to-speech platform, turning any text - PDFs, books, emails, Google Docs, web pages - into natural-sounding audio. Founded by dyslexic Brown University student Cliff Weitzman and his brother Tyler, the company grew from a personal reading aid into an AI voice empire spanning consumer apps, a voice cloning and dubbing studio for creators, and a text-to-speech API for developers. It serves 50+ million users across mobile, Chrome extension, and desktop, offers hundreds of voices in 50+ languages, and won the 2025 Apple Design Award for Inclusivity.
Kapwing is a browser-based, collaborative video editing platform that uses AI to turn raw clips, transcripts, and ideas into shareable videos. Founded in 2017 by ex-Googlers Julia Enthoven and Eric Lu, it serves millions of creators, marketers, and educators who want to skip the desktop software and just make a video.
Personal AI (legally Human AI Labs, Inc.) builds an identity-based memory layer and Personal Language Models (PLMs) - compact, individually trained AI models that capture a person's or organization's knowledge, voice, and decision-making. Instead of a single generic chatbot, the platform lets people and enterprises create AI personas trained on their own data, keeping ownership and privacy with the user. Founded in 2020 and based in San Francisco, the company positions its smaller, memory-grounded models as a more accurate, lower-cost, privacy-first alternative to centralized cloud LLMs for sales, healthcare, legal, HR and knowledge work.
Zyphra is a superintelligence research and product company building efficient, open foundation models and a full-stack AI cloud. Best known for its SSM-hybrid Zamba models, the ZAYA1 reasoning family trained entirely on AMD hardware, and the Zonos text-to-speech system, Zyphra argues that frontier intelligence should run cheaply, locally, and outside the walled gardens of a handful of labs. In June 2025 it raised a $100M Series A at a $1B valuation led by Jaan Tallinn.

ElevenLabs is an AI research and product company that builds human-like voice technology. Founded in 2022 by two Polish engineers frustrated by badly dubbed movies, it grew from zero to $200M ARR in three years and reached an $11 billion valuation by February 2026. Its platform covers text-to-speech, voice cloning, AI dubbing, conversational agents, and speech-to-text across 70+ languages, used by everyone from independent creators to over 60% of Fortune 500 companies.