MobbleOpen in Mobble ⇢
Technology · Artificial intelligence · published 2026-09-01 · via Engadget

Meta unveils real-time AI transcription with multi-speaker and multi-language support

Image via Engadget
Image via Engadget

Meta's Muse Voice Transcribe is a real-time audio model that can transcribe speech from over 20 speakers and handle multiple languages, including code-switching. It uses adaptive delay to improve accuracy and is trained on 70+ languages. The model is available via Meta's API and powers dictation in the Meta AI Mac app.

Expanded Detail

Muse Voice Transcribe represents Meta Superintelligence Lab's entry into streaming speech recognition, joining a wave of recent releases from the division including a coding agent and an open-weight model. The system employs adaptive delay, pausing longer on difficult words while committing quickly to simpler ones, which Meta says improves accuracy on imperfect real-world audio.

The model is trained across more than 70 languages, with 25 validated at launch, and can sustain hour-long sessions tracking over 20 distinct speakers. Pricing sits at $3 per 1,000 audio minutes through Meta's Model API. The technology already powers dictation in the Meta AI Mac desktop app, which can extend voice features to other applications. A demo version is available on Meta's research blog.

Context

This release could reshape how people interact with voice technology across languages and group settings. Real-time transcription that distinguishes speakers and handles code-switching may benefit journalists, interpreters, and multilingual workplaces, while the Mac app integration could bring these capabilities to everyday users. However, pricing and platform availability may limit adoption, and competition with Google's similar offering could accelerate innovation or fragment the market. Privacy considerations around continuous audio processing remain a factor for users and developers alike.

Expanded detail and Context are AI-generated analysis; the linked article remains the authoritative source.
Read the full article at Engadget →
Related stories
Sony's ULT Tower speakers gain Auracast wireless linking and live EQ controls · Consumer gadgets
This summary is AI-generated and original to Mobble; the linked article is the authoritative source. Original headline: “Meta's new AI transcription model can distinguish between multiple speakers and languages in real-time.” Browse more stories.