Meta has rolled out a new real-time voice dictation tool for Mac named Muse Voice Transcribe. Developed by Meta Superintelligence Labs, this streaming voice model arrives with a suite of features aimed at making audio input smoother and more intelligent—especially for those who write by voice or work across multiple languages.
What Muse Voice Transcribe Brings to the Table
Muse isn’t just a simple speech‐to‐text tool. It combines automatic speech recognition with speaker diarization and endpointing. That means it transcribes speech as it happens, can distinguish among more than 20 speakers in the same recording, and knows when one person has stopped talking before the next starts—all without any extra, post-processing work.
Training covered over 70 languages, with 25 of them validated for launch. Muse supports audio sessions longer than an hour and even handles fluid switching between languages within or between sentences. Users can also adjust language, keywords, and context to boost recognition accuracy.
A standout feature is Muse’s “adaptive delay” system: instead of a fixed latency, the model dynamically decides how much audio to buffer before committing each word. Easier speech can be processed quickly; complex or less clear words get more context to improve accuracy.
Where and How It’s Available
Muse Voice Transcribe is now live through Meta’s Model API at a cost of $3 per 1,000 audio-minutes (roughly $0.18/hour). Developers can integrate it in their own apps. On the Mac side, it powers system-wide dictation via Meta AI for Mac and is built into Muse Code. To dictate, Mac users simply hold the Fn key in any app to start using the feature.
According to benchmarks, Muse ranks first on the Artificial Analysis streaming speech-to-text leaderboard as of September 1. That gives it a measurable edge in speed, accuracy, or both compared to alternatives.
Why this matters: Voice input continues to grow across AI assistants, content creation, accessibility, and productivity tools. The ability to transcribe across dozens of languages, switch between them in real time, and handle multiple speakers without lag or manual editing sets a new standard. For developers, Muse raises expectations—and the bar—for what streaming voice models need to deliver.