Google's Gemini 3.5 Transcribe improves speech-to-text with formatting and filler removal
Google introduced Gemini 3.5 Transcribe, an AI model for speech recognition and transcription that supports over 85 languages. It can remove filler words, format unstructured speech into structured text, and attribute speech to up to three speakers. The model will soon enable speech-to-text in any web field in Chrome and is available in several Google products.
The model's speaker-attribution capability, which distinguishes up to three voices with word-level timestamps, positions it as a practical tool for podcasters, journalists, and meeting note-takers working with multi-person recordings. Its handling of alphanumeric sequences like order numbers and postal codes suggests attention to real-world business use cases beyond casual dictation.
Google is embedding the technology across its ecosystem, including Android's Rambler feature, the macOS Gemini app, and upcoming Chrome integration that would enable voice input in any web form field. Availability extends to Antigravity, Search Live, Gemini Live, Docs, Keep, and Gmail, with developer access through APIs.
This technology could meaningfully reshape how people interact with digital devices, reducing reliance on typing for everyday tasks like composing emails, searching, or filling out forms. Professionals who transcribe interviews or meetings may save considerable time, while accessibility gains could benefit users with mobility or vision impairments. However, reliance on cloud-based transcription raises questions about audio data privacy and the accuracy of speaker attribution in sensitive contexts, which may warrant careful consideration as adoption spreads.