Modulate raises $25M to build the understanding layer for voice AI — 100+ models, 98.9% deepfake detection
The Boston audio-AI company announced $25 million led by Future Ventures, bringing total funding to $60M. Its Velma platform listens for emotion, intent, and synthetic voices — and its models analyze over 10 million hours of audio a month.

Modulate announced $25 million in new funding on Monday, led by Future Ventures with participation from Hyperplane and Lakestar, bringing the Boston-based company's total funding to $60 million. The money is going toward AI research, product and engineering, developer relations, and partnerships — and toward getting more developers to build on a premise most of the voice-AI world still ignores: that understanding a conversation takes more than a transcript.
'Voice is becoming a primary interface for AI, and that creates a whole new set of problems that can't be solved from a transcript,' said Carter Huffman, Modulate's CEO and co-founder, in the announcement. Transcripts capture words. They drop tone, emotion, emphasis — and they can't tell you whether the voice on the line is a person at all.
Audio-native, not transcript-native#
Modulate's models analyze raw audio directly for signals like emotion, tone, intent, emphasis, synthetic speech, and conversational behavior. Its flagship Velma platform composes those signals into higher-level events: suspected fraud, a failing AI agent, harassment, customer dissatisfaction, policy violations. The company claims Velma detects true positives with twice the accuracy of traditional large language models while producing seven times fewer false positives.
The differentiator is speed. Velma operates in real time, so applications can respond to — or intervene in — a conversation while it is still happening. That makes it a supervision layer for voice AI agents, not just a post-call analytics tool.

100 models instead of one giant#
Under the hood, Velma runs on Modulate's Ensemble Listening Model (ELM) architecture — more than 100 specialized audio models orchestrated together, each selected and combined for the task at hand. The company says this approach delivers up to 1,000x greater inference efficiency than a single large model, cutting the cost, energy, and memory needed to analyze audio at scale.
The scale numbers give that claim something to stand on. Modulate's models now analyze more than 10 million hours of audio per month and have processed over 600 million hours in total. Its transcription technology recently ranked number one on Hugging Face's Open ASR Leaderboard, and its deepfake detection model tops Hugging Face's deepfake speech benchmark — with 98.9% accuracy on public benchmark data. The transcription API is priced at $0.03 per hour for batch processing.
What it is actually used for#
The use cases read like a map of where voice AI is breaking down in the real world. Modulate's technology is used to protect healthcare institutions from deepfake hackers, detect child grooming in voice conversations, reduce extremism and harassment on social platforms, monitor whether voice agents are actually performing as intended, and — a striking one — mask the voices of agents operating in high-risk scenarios so they can't be identified.
That last category is the honest edge of the business: as voice AI agents become cheaper to deploy, the systems that watch them become infrastructure. Modulate wants that layer to be its models rather than every developer's homegrown attempt. 'Developers shouldn't have to rebuild the audio intelligence layer every time they create a new voice experience,' Huffman said.

Why this round matters#
The funding lands at an inflection point. Money is pouring into AI systems that can speak naturally — ElevenLabs launched its Eleven v4 voice model today on the same thesis of more expressive generation. Modulate is betting on the other side of the interaction: that the harder, stickier problem is understanding what is happening in a voice conversation, including whether it is happening at all or is synthetic.
Steve Jurvetson, co-founder of Future Ventures, framed it as a technical lead being extended: 'Modulate has gained a significant technical lead in audio-native AI, and the market opportunity is expanding quickly.' The bet is that audio intelligence becomes a foundational layer of the AI stack — the kind of thing you rent from an API instead of training yourself. Sixty million dollars says enough customers are about to agree.
Sources
- Modulate press release — 'Modulate Raises $25M to Scale Its Lead in Frontier Audio-Native AI' (via ACCESS Newswire, September 28, 2026)
- SecurityWeek — 'Modulate Raises $25 Million to Advance Deepfake Detection' (September 28, 2026)
- PYMNTS — 'Modulate Raises $25 Million to Expand Audio AI Efforts' (September 28, 2026)
- GamesBeat — 'Modulate raises $25M to scale its audio-native AI' (September 28, 2026)