Meta launches low-cost real-time transcription API with multi-speaker identification
Meta Superintelligence Labs has introduced Muse Voice Transcribe, a speech-to-text model that processes audio live and includes endpoint detection and speaker diarization for over 20 participants. The public API costs $0.18 per hour of audio, supporting sessions longer than an hour and seamless switching between languages. This pricing undercuts many existing transcription services while adding real-time capabilities.
Meta Superintelligence Labs has released Muse Voice Transcribe, a speech-to-text API built for live audio processing. The model detects when a speaker finishes talking and can distinguish between multiple voices, identifying over 20 participants in a single session. It handles continuous audio beyond one hour and permits mid-session language changes without interruption.
The service is priced at $0.18 per hour of audio, a rate that sits below many established transcription offerings. By combining real-time processing with multi-speaker separation at this cost, the tool targets use cases such as live meetings, customer support calls, and event captioning where both speed and speaker clarity matter.
This pricing could pressure existing transcription providers to adjust their rates, potentially benefiting small businesses and independent creators who rely on affordable captioning or meeting notes. Real-time speaker identification may improve accessibility for deaf and hard-of-hearing users in group settings, but it also raises consent questions, as conversations could be transcribed and attributed without participants' explicit knowledge. Organizations adopting this tool may need to weigh operational efficiency against data privacy responsibilities.