Meta unveils real-time AI transcription with multi-speaker and multi-language support

Meta's Muse Voice Transcribe is a real-time audio model that can transcribe speech from over 20 speakers and handle multiple languages, including code-switching. It uses adaptive delay to improve accuracy and is trained on 70+ languages. The model is available via Meta's API and powers dictation in the Meta AI Mac app.
Muse Voice Transcribe represents Meta Superintelligence Lab's entry into streaming speech recognition, joining a wave of recent releases from the division including a coding agent and an open-weight model. The system employs adaptive delay, pausing longer on difficult words while committing quickly to simpler ones, which Meta says improves accuracy on imperfect real-world audio.
The model is trained across more than 70 languages, with 25 validated at launch, and can sustain hour-long sessions tracking over 20 distinct speakers. Pricing sits at $3 per 1,000 audio minutes through Meta's Model API. The technology already powers dictation in the Meta AI Mac desktop app, which can extend voice features to other applications. A demo version is available on Meta's research blog.
This release could reshape how people interact with voice technology across languages and group settings. Real-time transcription that distinguishes speakers and handles code-switching may benefit journalists, interpreters, and multilingual workplaces, while the Mac app integration could bring these capabilities to everyday users. However, pricing and platform availability may limit adoption, and competition with Google's similar offering could accelerate innovation or fragment the market. Privacy considerations around continuous audio processing remain a factor for users and developers alike.