# Kalpa Labs

> Kalpa Labs is a San Francisco audio research lab building a generalist speech model - one system that handles speech-to-text, text-to-speech, voice cloning, and speech-in/speech-out reasoning with the kind of instruction-following and in-context learning that large language models brought to text. Founded in 2025 by ex-Google Assistant ML lead Prashant Shishodia and ex-high-frequency-trading engineer Gautam Jha, the company is a Y Combinator Fall 2025 startup aiming to collapse today's fragmented speech stack into a single, steerable model for developers and voice-agent builders.

- **Founded:** 2025
- **Headquarters:** San Francisco, California, United States
- **Founders:** Prashant Shishodia (Co-founder & CEO), Gautam Jha (Co-founder & CTO)
- **Team size:** 2 (founders); hiring research and audio engineers
- **Products:** Generalist Speech Model, Studio, Realtime, API
- **Notable:** Accepted into Y Combinator's Fall 2025 (F25) batch., Pretrained speech models from 800M to 4.8B parameters on roughly 2 million hours of mixed audio., Trained an 800M-parameter speech model for under $1,000 by removing bottlenecks in audio tokenization.

## Products & services

- **Generalist Speech Model** — A single foundation model spanning speech-to-text, text-to-speech, voice cloning, and speech-in/speech-out reasoning, with instruction-following and in-context learning. Models pretrained from 800M to 4.8B parameters on roughly 2 million hours of mixed audio.
- **Studio** — A scene-direction tool for writing dialogue, casting voices, setting mood, and generating multi-speaker scenes with emotion and consistent character voices.
- **Realtime** — A browser-based conversational interface for live, low-latency dialogue with the speech model.
- **API** — REST API access to the same models, with documentation designed to be readable by both humans and AI agents.

## Achievements

- Accepted into Y Combinator's Fall 2025 (F25) batch.
- Pretrained speech models from 800M to 4.8B parameters on roughly 2 million hours of mixed audio.
- Trained an 800M-parameter speech model for under $1,000 by removing bottlenecks in audio tokenization.
- Reported 59.3% win rate against ElevenLabs eleven-flash and 54.0% against eleven-turbo in blind preference tests.
- Shipped three products - Studio, Realtime, and an API - into open beta.
- Demonstrated emergent capabilities such as contextual accent adaptation, prosodic awareness, and disfluency handling without explicit tags.

## Latest updates

- **2025-09** — Reported seed round backed by Y Combinator and Nexus Venture Partners.
- **2025** — Joined Y Combinator's Fall 2025 batch as a two-person audio research lab.
- **2025** — Launched Studio, Realtime, and API products in open beta.

## Links

- Website: https://kalpalabs.ai
- LinkedIn: https://www.linkedin.com/company/kalpalabs
- Twitter/X: https://x.com/KalpaLabsAI

---

Profile page: https://yespress.io/kalpa-labs-yc-f25
Published by YesPress — https://yespress.io
Last updated: 2026-07-30
