Sarvam AI Releases Saaras V4: A Speech-to-Text Model for All 22 Indian Languages and Global English
What it does
Sarvam AI launched Saaras V4, a speech-to-text model supporting all 22 Indian languages plus global English. It combines an audio encoder with a 3-billion parameter hybrid state-space decoder. The model includes keyterm prompting that handles up to 50 terms for more accurate recognition. It also offers five distinct output modes through one model and can stream transcriptions with first-token latency under 150 milliseconds. Saaras V4 is accessible via Sarvam’s API at a rate of ₹30 per hour.
Why it matters
Handling all official Indian languages in one speech-to-text framework is a significant step for localization and accessibility. Supporting global English along with these regional languages makes Saaras V4 practical for pan-Indian and international applications. The low latency streaming capability can improve real-time transcription in call centers, media, education, and government services. Keyterm prompting reduces errors for domain-specific or frequently used vocabulary, which can cut down manual correction time and improve automated workflows. The pricing at ₹30 per hour sets a clear benchmark for enterprises looking to build multilingual voice applications with predictable costs.
Who it is for
The model aims at developers and businesses building voice-enabled services targeting the Indian market and global English speakers. This includes call centers, media companies, educational platforms, and government agencies that require robust transcription across many languages. Builders focusing on real-time communication apps will find the low latency streaming valuable. The keyterm prompting feature is useful for sectors needing domain-specific accuracy, such as finance, healthcare, or legal. Startups and SMEs looking for cost-effective multilingual speech recognition APIs gain a viable option with Saaras V4.
The catch
While the model claims extensive language support, accuracy may still vary widely across 22 vastly different languages and dialects. Streaming latency under 150 milliseconds is promising but needs validation under real-world conditions with noisy or accented speech. The pricing is competitive, but costs can accumulate with heavy usage compared to open-source alternatives, which might pressure budgets for smaller players. The hybrid state-space decoder architecture sounds advanced but is less common in open research, raising questions about ease of integration and customization for atypical use cases.
What to watch next
Check how well Saaras V4 performs on actual multilingual transcription tasks, especially for underrepresented Indian languages with complex phonetics. Monitor its adoption among Indian enterprises requiring scalable voice interfaces. Look for case studies demonstrating keyterm prompting’s impact on transcription accuracy and productivity. Watch Sarvam’s roadmap for API extensions or enhancements beneficial to developers, such as fine-tuning capabilities or offline deployment. Also, observe how competitors respond in pricing or language coverage to this broad, low-latency offering.
AI Quick Briefs Editorial Desk