Signal / 09.28.26
NEW Multi-Speaker 2.0 separates overlapping voicesLIVE Dialogue isolation reported at 11msSCALE 100M+ minutes processedNEW Multi-Speaker 2.0 separates overlapping voicesLIVE Dialogue isolation reported at 11msSCALE 100M+ minutes processed

Company profile / AI audio infrastructure

The Song Was Finished. AudioShake Made It Editable Again.

Most recordings arrive as a single, stubborn object. AudioShake built a business by asking what happens when music, speech and noise can be taken apart again.

A finished song is a small miracle of commitment. Dozens of decisions - the singer here, the snare there, the bass tucked underneath - are pressed into a stereo file. For a listener, this is convenient. For anybody who wants the singer without the snare, it is a locked room. The key was the original multitrack session, and that key may be sitting on a damaged drive, hiding in a mislabeled archive, or may never have existed at all.

AudioShake makes replacement keys. Its software takes a mixed recording and estimates the parts inside it: vocals, drums, bass, guitars, dialogue, music, effects, even separate people speaking over one another. The result is not time travel and not the original studio session. It is a set of newly made stems useful enough to remix a catalog song, dub a film, clean a sports broadcast, align lyrics, train an audio model or rescue a voice from a noisy home video.

The quick cut

  • What it sells: a web studio, enterprise API, real-time SDK, lyric tools and custom audio-data services.
  • Who buys: labels, studios, broadcasters, sports organizations, localization shops and AI developers.
  • How it charges: annual enterprise fees plus usage; public contract prices are not disclosed.
  • The proof: 40-plus enterprise contracts and more than 100 million processed minutes reported in 2025.
  • The catch: poor sources and crowded arrangements can still leave artifacts; exacting restoration remains human work.

The karaoke question

The origin is pleasingly unserious. Jessica Powell and Luke Miner were living in Tokyo and doing a great deal of karaoke. They wondered whether any song could be turned into a karaoke song. If a machine could remove the lead vocal, why stop there? Why not take any sound and pull it into parts?

Powell had been a Google communications executive. Miner had led data science at Plaid. They founded AudioShake in 2020 and launched it the following year. The early attraction was easy to hear: upload a song, get its ingredients. Yet the valuable discovery was not that people enjoyed the trick. It was that companies owned enormous libraries of finished audio they could no longer properly edit.

Disney Music Group, one of the company's early customers, had recordings whose constituent stems were missing. Without them, an old master is less useful for sync licensing, spatial mixes or new fan experiences. AudioShake's collaboration added instrument separation and lyric transcription to the catalog workflow. EMPIRE uses the technology for stems and lyrics. Singa uses it to turn licensed original masters into karaoke tracks across more than 2,000 venues and for two million registered home users.

“Until now, most of the world’s audio has been effectively read-only.”AudioShake, announcing its Series A

The expensive part was hiding after the demo

This is the change that matters. AudioShake stopped looking merely like a music utility and started behaving like infrastructure. A flattened soundtrack can become dialogue, music and effects for localization. A noisy recording can yield cleaner speech for captions. Incidental music in a stadium feed can be removed before it triggers a copyright claim. Old audio can become structured data for search or model training.

The customers changed the scope because their expensive problems rhymed. They did not all need karaoke. They all needed one sound that was trapped inside several others. By October 2025 the company said it had more than 40 enterprise contracts, nearly 400 percent year-over-year revenue growth, and more than 100 million minutes of audio processed. Its named customers included Universal Music, Disney Music Group, Warner Music Group, Warner Bros. Discovery, BET and NFL Films.

40+enterprise contracts reported in 2025
100M+minutes of audio processed
~400%year-over-year revenue growth reported

The commercial arrangement reflects that seriousness. AudioShake sells annual contracts with usage charges, rather than pretending every job fits a cheerful $9.99 subscription. The interface matters, but the durable product is integration: an API for a catalog, an SDK inside an application, a model on an edge device, a data service built around a rights holder's own material. The 2025 Series A brought in $14 million, led by Shine Capital, and took publicly reported funding to $19 million.

Ghazi Shami, founder and CEO of EMPIRE, standing beside a framed platinum record
Ghazi Shami of EMPIRE, beside the traditional proof that a song has travelled. AudioShake works on what happens when the finished object needs to travel somewhere new.

The machine fails usefully

The instructive failure arrived in a 1936 Popeye cartoon. Audionamix, an audio-restoration specialist, used AudioShake to pull out the vocals. That first pass worked - and revealed another defect: the sound of a microphone apparently muffled by a hand. A general separation model had exposed the peculiar problem underneath. A human engineer still had to repair it.

That is a better account of professional AI than the usual vanishing-worker fantasy. Separation can remove the long, coarse part of the job. It cannot know the intention behind every damaged syllable or decide how much artifact a remaster should tolerate. AudioShake says its models work best with high-resolution WAV or AIFF files, though customers have used lesser sources. A crowded arrangement, lossy MP3, rare instrument or historically odd recording can produce leakage and warbling. Sometimes an original session, a local tool or a restoration suite such as iZotope RX is the better answer.

This also explains its position in a crowded market. Moises and LALAL.AI make separation approachable for individuals. Demucs and Spleeter are open-source. Logic, RipX, SpectraLayers and RX keep the work near the editor. AudioShake is aiming at the organization that values fidelity, many target types, predictable processing, deployment choices and somebody responsible for the integration. The moat, if it holds, is workflow rather than spectacle.

Eleven milliseconds changes the venue

Post-production offers a forgiving luxury: time. Live television does not. In April 2026 AudioShake introduced Dialogue RT, reporting 11 milliseconds of end-to-end latency on NVIDIA DGX Spark hardware. That number is small because the tolerance is small. Push audio much beyond the roughly 10-to-15-millisecond envelope and lips begin to betray the system.

11ms

Reported end-to-end latency for isolating dialogue from a live mixed feed.AudioShake Dialogue RT on NVIDIA DGX Spark

Now the separation can happen inside the broadcast chain. Commentary and crowd become independent controls. A captioning engine receives cleaner speech. A rights problem can be removed while the event is still occurring. In September 2026, Multi-Speaker 2.0 pushed in another direction, splitting overlapping voices into discrete tracks. The humble karaoke premise has wandered into stadiums, archives, dubbing rooms and voice systems.

The part worth copying

The obvious imitation would be to build another separator. The more useful lesson is about product expansion. AudioShake began with a delightful demonstration, then followed the expensive obstruction it revealed. It did not need to persuade customers to create new audio. The audio already existed, often in enormous quantities. It needed to make a fixed asset useful one more time.

  1. Begin with the locked asset. Find material customers already own but cannot adapt, search or redeploy.
  2. Sell the recovered control. The value is not a stem file; it is a sync placement, faster dub, clean caption or usable training set.
  3. Let serious users widen the market. Adjacent problems in music, film, sport and speech shared the same technical bottleneck.
  4. Keep the specialist in the loop. Automation earns trust when it hands experts better material, not when it claims their judgment.

The conditions are equally plain. This approach loses its advantage when original stems are readily available, when privacy requires fully local processing that the chosen deployment cannot support, when the source is too degraded for the required standard, or when a cheap consumer tool is already good enough. Enterprise infrastructure only makes sense when quality, scale, rights and integration cost more than the contract.

There is something charming about a company that set out to silence a karaoke singer and ended up making machines better listeners. But the business is not really about silence. It is about optionality. A recording once had to remain what it was. AudioShake's wager is that, given the right separation, finished is merely a format.