Modulate, a company specializing in audio-native artificial intelligence, has raised $25 million in a fresh funding round aimed at strengthening defenses against deepfake voice fraud, impersonation, and other audio-based threats. This investment, which brings the company’s total funding to $60 million, was led by Future Ventures with contributions from Hyperplane and Lakestar.
At the heart of Modulate’s offering is Velma, a real-time platform engineered to catch events ranging from fraud attempts and harassment to customer dissatisfaction, policy violations, and even lapses in performance by AI-driven voice agents. Rather than relying solely on transcripts, Velma delves into how speech is delivered—examining tone, emotion, emphasis, intent, and even signs of synthetic generation.
How the Technology Works
Modulate’s system is built on its Ensemble Listening Model (ELM), which orchestrates over 100 specialized audio models to handle different aspects of audio understanding. This architecture aims to be far more efficient than using a single foundation model for every task—it can reduce the computational, memory, and energy costs of audio analysis by substantial margins.
In terms of accuracy, the company reports an impressive detection performance on deepfake voice: an average equal error rate (EER) of just 1.104% across 14 Speech DF Arena datasets as of August 19, 2026. That translates into nearly 98.9% accuracy. The model—featuring around 316 million parameters—can deliver streaming verdicts after just 2.5 seconds of speech.
Performance, Pricing, and Use Cases
Modulate’s tools are already processing more than 10 million hours of audio each month, and have amassed over 600 million hours overall. It has also recently placed first in transcription accuracy on Hugging Face’s Open ASR Leaderboard among 88 models evaluated as of July. Pricing begins at $0.03 per hour for batch transcription, while deepfake detection comes in at $0.25 per hour.
The new funding will be channeled into expanding APIs and SDKs, developing models tailored for specific industries, integrating partnerships, and supporting varied deployment settings. Target industries include fraud prevention, healthcare security, contact center management, social platform moderation, child safety, and oversight of AI voice agents.
Limits & Considerations
While Modulate’s results are strong, the company advises caution when deploying in production settings. Telephone audio can be compressed, languages beyond English and Mandarin may present challenges, background noise or replay attacks may reduce effectiveness, and high detection thresholds could lead to increased false positives. The benchmarks currently used don’t fully capture streaming behavior, end-to-end narrowband telephony, or many non-English languages.
According to Modulate’s CEO, the growth of voice as a primary AI interface necessitates moving beyond what transcripts reveal. The company sees audio understanding becoming foundational infrastructure for identity checks, monitoring of automated agents, and detecting manipulation. As synthetic voices grow more convincing, combining this technology with multifactor verification, transaction controls, and human review remains essential to prevent detection scores from becoming single points of failure.
Analysis: Modulate’s raise reflects rising urgency around audio-based threats, especially vishing and executive impersonation attacks. Its focus on how speech is expressed—not just what is said—is strategically important, offering more nuanced detection. Real-world performance, especially in noisier or more diverse linguistic contexts, will be the next test. Industry watchers should check how Modulate balances detection accuracy with usability—for example, what happens when false positives increase—and whether its model scales across global deployments, varying network qualities, and evolving synthetic voice techniques.