Breaking signal Modulate takes voice AI beyond transcription  •  500 million conversations analyzed  •  From game chat to fraud detection

Company profile / Voice intelligence

Modulate Wants AI to Hear What the Transcript Missed

The startup that learned to police Call of Duty voice chat is making a larger bet: the valuable part of a conversation is often the part a transcript throws away.

A transcript is a ruthless editor. It keeps the words and fires the hesitation, the laugh, the tightening voice, the sarcastic stress on one syllable. That bargain is fine when the job is searchable meeting notes. It is less fine when a bank wants to know whether a caller is being coached by a fraudster, a game studio needs to distinguish trash talk from a threat, or a voice agent has failed to notice that its customer is quietly furious. Modulate has built a company around the material that gets cut.

The Somerville, Massachusetts startup calls its field voice intelligence. Its first durable product, ToxMod, listens to live speech in games and social spaces, identifies moments that may violate a platform's rules, and gives moderators clips and context to review. Its newer platform, Velma, applies the same audio-first premise to transcription, synthetic-voice detection, fraud, customer experience, compliance and supervision of AI agents. Modulate says its systems have improved more than 500 million conversations and protected over 40 million consumers. Those are company figures, but the named deployments are substantial: Activision uses ToxMod in Call of Duty, while customers publicly identified by Modulate include Riot Games, Rockstar Games and Rec Room.

Abstract Swiss-style composition of waveforms passing through a listening circle, with one risky signal isolated in orange
The orange waveform has been asked to step aside for a brief conversation. The teal waveforms are pretending not to stare.

The pivot was hiding in the sound

Modulate did not begin as moderation software. Co-founders Mike Pappas and Carter Huffman met as MIT physics students, reportedly after Pappas solved a problem Huffman was working through on a hallway whiteboard. Huffman later built machine-learning models for spacecraft at NASA's Jet Propulsion Laboratory; Pappas worked on cloud security at Bridgewater Associates. On Christmas Day in 2015, Huffman began wondering whether the style-transfer techniques then remaking images could be applied to audio. The pair incorporated Modulate in 2017.

The original product was a real-time voice skin: software that could change how a player sounded while keeping the emotion and cadence of the original performance. It was technically demanding. Games leave little computing power or patience for a voice effect that lags. By late 2019, the team could play together using different synthetic voices without wrecking the experience. A $4 million seed round followed in March 2020, led by 2Enable Partners and joined by Sierra Ventures, Hyperplane, Third Kind, Everblue and others.

Then the useful side effect became the business. To change a voice convincingly, Modulate's models had to notice timbre, prosody, volume and emotion in real time. Those signals could also help tell whether a voice chat contained friendly banter, escalating harassment, grooming or a credible threat. ToxMod launched in 2020. The company had not abandoned its difficult primitive; it had found a problem with a clearer budget and a sharper cost of failure.

“You don't just want a transcript - you want to really understand what's going on.”Carter Huffman, co-founder and CTO

A filter for the impossible listening queue

Voice moderation has an ugly arithmetic. Among Us VR developer Schell Games faced nearly 50,000 hours of audio each month, according to a Modulate case study. Humans cannot listen to all of it, and player reports cover only what people notice and bother to submit. Keyword systems lose tone and context. Transcribe-everything pipelines add cost and still struggle with overlapping speakers, game noise, slang, accents and deliberately evasive users.

ToxMod works as triage. Its client software connects to a game's live voice system, performs an initial check, and escalates relevant speech for deeper analysis in Modulate's cloud. A dashboard can surface clips, timestamps, behavioral context and confidence for human reviewers. The game publisher still writes the code of conduct and decides what happens next. In Activision's case, voice chat may be monitored and recorded to investigate violations, but Activision operates the program and handles enforcement. ToxMod is the detection layer, not the judge.

500M+conversations improved, company reported
40M+consumers protected, company reported
18languages supported by ToxMod

That distinction makes measured interventions possible. Rec Room used ToxMod data to experiment with warnings and short mutes instead of treating every offense as grounds for an immediate ban. Its case study reports roughly a 75 percent reduction in repeat toxic offenses after warnings and about a 90 percent reduction after mutes and bans. The lesson is less “AI fixed toxicity” than “better visibility gave the platform more choices.” A warning can correct a player who lost their temper. A persistent predator demands something else.

The game lobby becomes a laboratory

Gaming gave Modulate an unusually hostile test environment. The audio is noisy. People interrupt one another. Teenagers invent language faster than policy teams can document it. Bad actors adapt when they learn how detection works. Decisions need to arrive quickly enough to matter, at a cost low enough to cover enormous volumes. If a voice model only performs in a clean podcast recording, it has learned the easy half of listening.

This is Modulate's clearest difference from transcript-first competitors. The standard stack converts speech to text and sends that text to a large language model. Modulate's models also examine the source audio: pauses, stress, emotion, speaker changes, background sound, interaction patterns and synthetic-voice markers. The company argues that sarcasm, vulnerability or escalation can change the meaning of identical words. Its current system, called the Ensemble Listening Model, coordinates more than 100 component models rather than asking one general model to do every job.

Specialization has practical appeal. A lightweight block can handle speaker timing; another can look for an acoustic cue; another can interpret the conversational pattern. Modulate says this design improves cost, speed and explainability, though many of its benchmark comparisons are self-published and the broader market will test those claims in production. What is already credible is the engineering pressure created by years of low-latency, high-volume deployments. AWS has described how ToxMod plugs its SDK into a customer's voice infrastructure and scales audio ingestion in the cloud.

Audio processing
Acoustic signals
Perceived intent
Behavior models
Conversation context

Beyond the ban button

In 2025, Modulate carried the stack outside its gaming home. VoiceVault looks for social engineering, impersonation, urgency scams and signs of synthetic speech while a call is happening. The intended buyers include financial services, insurers, logistics companies and contact centers. A Twilio integration lets customers connect ToxMod or VoiceVault to existing call flows through Media Streams. This is a sensible expansion: the angry delivery call and the abusive game lobby are different products, but both contain behavioral information a transcript may flatten.

Velma 2.0, introduced in January 2026, makes the platform argument explicit. Velma Transcribe turns batch or streaming audio into structured text with speaker and timing information. Velma Deepfake Detect watches for synthetic speech across a call. The broader conversation-intelligence API can return structured JSON about what happened and how it happened. In June, Modulate opened API access without requiring an enterprise contract, giving developers a way to test their own audio in a playground.

That produces more concrete possibilities than the phrase “conversation intelligence” suggests. A contact center could route a distressed caller to a human sooner. An insurer could flag a suspicious claims call for review. A voice-agent company could detect that its bot is repeating itself while the customer grows frustrated. A compliance team could find risky statements without relying on a magic keyword. A social app could spot a younger voice in an adult space. None of these systems should make irreversible judgments without governance, but each can shrink the haystack.

Where Modulate fits: part trust-and-safety platform, part speech API, part fraud infrastructure. Its nearest competitor changes with the buyer - Unitary or ActiveFence in moderation, Deepgram or AssemblyAI in transcription, Pindrop or Reality Defender in voice fraud and synthetic-media detection. The common alternative is an in-house chain of transcription, a general model and human review.

The business inside the listening layer

Modulate is a business-to-business software company. Large game deployments and moderation workflows imply enterprise contracts, integration support and pricing tied to usage and requirements. The newer API creates a usage-based, self-serve route for developers. Public pricing remains limited, so the important commercial question is whether one shared audio stack can support several products without turning the company into a custom-services shop.

The balance sheet gives it room to try. Public reports put total funding at $36 million: an approximately $2 million seed round in 2019, the $4 million round in 2020, and a $30 million Series A led by Lakestar in August 2022. Modulate had more than 40 employees by the end of 2023; the company record supplied for this profile lists 52. Its culture materials describe a mix of machine-learning researchers, audio engineers, gamers and policy work. Employee-led guilds formed around groups including LGBTQ+ staff, parents and people new to Boston.

Culture matters here because voice intelligence is never merely acoustic. A model that cannot separate reclaimed language from abuse, or exuberance from danger, will punish the communities it claims to protect. Modulate has publicly discussed bringing in outside expertise and preserving human review. Privacy creates a parallel obligation. Proactive listening catches what reports miss, but it also creates surveillance risk. On-device triage, limited escalation, audit trails, consent, retention controls and clear platform rules are product features, not paperwork around the product.

The bet hiding between the words

Modulate's expansion arrives as two trends collide. Synthetic voices make fraud cheaper, while voice agents put machines into more consequential conversations. At the same time, speech-to-text is becoming a commodity feature. If transcription prices keep falling, the margin moves upward into interpretation: Who is speaking? What changed? Is this voice real? Did the customer become confused before the agent noticed?

The risk is breadth. Game moderation, insurance fraud, medical transcription and AI-agent supervision have different data, regulations, buyers and definitions of success. Strong Call of Duty credentials do not automatically produce a banking business. Large cloud vendors and specialist startups can bundle adjacent features. Modulate must show that audio-native understanding is not just richer than a transcript, but valuable enough to change an operational decision.

Still, the company's path contains a useful principle. Its founders started with a playful technical feat, watched what the machinery learned along the way, and followed that capability toward a more urgent problem. ToxMod turned listening into a workflow. Velma is the attempt to turn that workflow into infrastructure. The pitch is not that every conversation needs a machine critic. It is that when a conversation truly matters, keeping only the words may be the most expensive form of selective hearing.

Keep listening

Explore Modulate's product material, developer documentation and public channels, or watch the team explain the technology in its own demos and interviews.