YC F25  Koyal turns audio into cinematic video 5 music videos crossed 1.5M+ views each Paid pilots with Universal Music & T-Series 22 YC startups used Koyal for launch videos Public beta live at beta.koyal.ai Founded by siblings Mehul & Gauri Agarwal YC F25  Koyal turns audio into cinematic video 5 music videos crossed 1.5M+ views each Paid pilots with Universal Music & T-Series 22 YC startups used Koyal for launch videos Public beta live at beta.koyal.ai Founded by siblings Mehul & Gauri Agarwal
COMPANY AI · Filmmaking · San Francisco

The studio where you never pick up a camera

Koyal, a Y Combinator F25 startup, turns a song, a podcast, or a plain voice memo into a finished film. Its founders borrowed the trick from Pixar: record the sound first, then build everything around it.

There is a small production secret hiding inside every Pixar film. Before a single frame is drawn, the actors go into a booth and record their lines. The voice comes first. The pictures get built to fit the sound, not the other way around. Most people assume it happens in reverse. It doesn't - and that inversion is the entire idea behind Koyal.

Koyal, a company in Y Combinator's Fall 2025 batch, is an agentic AI filmmaking platform. In plain terms: you give it audio - a song, a podcast clip, a recorded story, a script read aloud - and it returns a finished, personalized video. Not a five-second novelty loop, but a sequence with consistent characters, coherent settings, changing moods, and camera angles that move the way a director would move them.

The company is run by two siblings, Mehul and Gauri Agarwal. Mehul is the CEO; Gauri is the CTO. Both spent time on Meta's Instagram video team before leaving to build this, and their backgrounds run through Carnegie Mellon and the MIT Media Lab. The technical seed of the company was a research paper at NeurIPS 2024. What started as an academic result became, within a year, a tool that music labels were paying to use.

01 / The pitchWhat Koyal actually does

Point most AI video tools at a blank screen and they hand you a prompt box and a coin flip. Type a sentence, get a clip, hope the result resembles what was in your head. Koyal works from a different starting material. The audio is the brief. If you feed it a song, the melody, the pacing, and the emotional tone of the track become the instructions for what appears on screen.

1

Audio In

Song, podcast, voiceover or script read aloud.

2

Read Emotion

The system captures tone and pacing from the sound.

3

Direct

Cast characters, build worlds, edit scenes in plain language.

4

Film Out

A finished video with consistent characters and shots.

FIG. 1 - The audio-first pipeline, start to finish.

The word the team keeps using is "agentic," and here it means something specific. Instead of you writing careful prompts and stitching clips together, the platform handles the tedious parts - blocking a scene, keeping a character's face the same from shot to shot, matching lips to dialogue - and leaves you the parts a director actually cares about. You storyboard. You give notes. You say make it warmer, move the camera back, put her in the red coat. The system carries out the instruction.

"Koyal is inspired by how Pixar builds stories: audio first, visuals second." Koyal, on its own approach

02 / The hard partThe problem everyone else keeps breaking

Anyone who has played with AI video knows the failure that ruins it: continuity. A character appears in the first shot, then shows up in the third looking like a cousin who resembles them but isn't quite them. Backgrounds drift. The red coat turns maroon, then brown. For a single viral clip this doesn't matter. For anything longer than a few seconds - a music video, a short film, a launch reel - it matters completely.

Koyal made that boring, unglamorous problem the center of the product. The company has built what it describes as state-of-the-art work in three areas: multilingual script-to-audio-to-video, agentic lip-sync with dynamic movement, and multi-character consistency. That last one is the whole game. Keeping several characters looking like themselves across an entire piece is the difference between a demo and a deliverable.

1.5M+Views per video, x5
22YC startups used it
2Global labels piloted
F25YC batch

03 / The customersWho is actually paying

Koyal's early traction reads in two directions at once. On one side are the big names. The company ran paid pilots with Universal Music and T-Series, working alongside Bollywood production houses including Maddock Entertainment. Out of those pilots came five music videos for artists who own Grammy and Oscar hardware - A.R. Rahman, Ricky Kej, and Shankar Mahadevan. Each of the five crossed 1.5 million views.

Why the label math works

A major label sits on thousands of tracks and a video budget that can only cover a handful of them. Traditional production means crews, sets, and weeks per song. Koyal's proposition flips the constraint: the song is the brief, so a back catalog becomes a pipeline of videos instead of a pile of audio. That is the reason a pilot at this scale even happens.

OBSERVATION - The unglamorous economics behind the glamorous clients.

On the other side is a quieter, scrappier crowd: startups. Twenty-two Y Combinator companies used Koyal to make their launch videos - the kind of polished thirty-second reel a founder needs for a demo day or a landing page, and normally cannot afford to shoot. The public beta, at beta.koyal.ai, opened that same capability to anyone. It hit #4 Product of the Day when it launched on Product Hunt in November 2025.

04 / The wrinkleA captcha for faces

Generating video that stars a real person raises an obvious question - whose face is that, and did they agree to it? Koyal's answer is a protocol it calls CHARCHA, a "character captcha." Before an AI character can wear someone's likeness, the system requires verified consent. It is a small feature with a large implication: consent built into the pipeline as a gate, not bolted on afterward as a disclaimer. In a category where the ethics are usually an afterthought, making it a step is a deliberate choice.

The camera, the crew, the set, the render farm - Koyal is betting all of it is optional, and that the only irreplaceable input is a human voice and a human decision.

05 / The fieldWhere it sits on the map

The AI video space is crowded and loud. Runway, Pika, Luma, Kling, Google's Veo, OpenAI's Sora - most of them are general-purpose engines optimized for the single best clip you can coax from a prompt. Koyal is not trying to win that race. Its bet is on the workflow around the clip: audio-first input, agentic direction, and consistency across a whole piece rather than brilliance in one shot. Different starting point, different finish line.

Where the effort goes - illustrative positioning, not a benchmark
Prompt-to-clipgeneral tools
Character
consistency
Koyal focus
Audio-first
input
Koyal focus
Consent
built-in
CHARCHA

06 / The peopleThe team behind it

Koyal is small and technical by design - a founding team drawn from CMU, MIT, and Meta, with research credentials to match. The sibling structure is unusual for a startup, and it shows up in how the work is split: one running the company, one running the technology.

Mehul Agarwal
Founder & CEO

CMU computer science and ML. Formerly on Meta's Instagram video team.

Gauri Agarwal
Co-Founder & CTO

MIT Media Lab and CMU. Also came from Meta's Instagram video work.

THE FOUNDERS - A brother-sister team, which makes Koyal both a company and a family project.

07 / The takeawayWhat it adds up to

Koyal raised a seed round through Y Combinator and is still early - a public beta, a handful of marquee pilots, a claim about consistency that longer projects will keep testing. The interesting part is not the hype, which the company mostly avoids. It is the framing. If the camera really is optional, the bottleneck in making a film stops being equipment and access and becomes something closer to taste. What you can imagine, and how well you can direct it. Koyal is a bet that this is where video is heading, and that the door in is a voice memo.