MobbleOpen in Mobble ⇢
Technology · Artificial intelligence · published 2026-09-25 · via Techgenyz

Sarvam AI Launches Saaras V4 Speech Model for 22 Indian Languages

Image via Techgenyz
Image via Techgenyz

Sarvam AI has introduced Saaras V4, an automatic speech recognition model covering 22 Indian languages plus English, with general availability on September 2. The model uses an audio encoder and a 3-billion-parameter hybrid state-space decoder trained in-house, targeting code-switching, dialect variation, and noisy audio. It offers five output modes and supports REST, Batch, and WebSocket APIs, including real-time use.

Expanded Detail

Sarvam AI unveiled Saaras V4 on August 24, with general availability following on September 2. The system pairs an audio encoder with a three-billion-parameter hybrid state-space decoder that Sarvam trained internally. It is intended for 22 Indian languages and English, especially speech affected by language mixing, regional pronunciation differences, background noise, and varied textual representations.

Developers can choose five output styles: standard transcription, English translation from supported Indic speech, verbatim capture including disfluencies, Roman-script transliteration, and mixed-language output. The model is accessible through REST, Batch, and WebSocket interfaces, enabling real-time applications. Sarvam has also integrated it into its Realtime API and broadened keyterm prompting across speech endpoints.

Context

Saaras V4 could make voice interfaces more usable for speakers of many Indian languages, including those who mix languages or speak in noisy settings. Developers and businesses may build real-time transcription, translation, and Roman-script tools more easily. This may improve access to digital services for users whose speech patterns were previously poorly served, though actual benefits depend on deployment, accuracy, and affordability.

Expanded detail and Context are AI-generated analysis; the linked article remains the authoritative source.
Read the full article at Techgenyz →
This summary is Al-enhanced to contain extended analysis and broader social context. The original is {NAME); the linked article is the authoritative source. Original headline: “Saaras V4: Sarvam AI’s 22-Language Speech Recognition Model.” Browse more stories.