Models & Research

Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Lang…

· August 28, 2026
Google AI Releases Gemini 3.5 Transcribe: A Speech-to-Text Model Reporting 2.6% Average WER Across 85+ Lang…

What it does

Google AI has launched Gemini 3.5 Transcribe, a new speech-to-text model available as two separate API endpoints: streaming and batch. The streaming endpoint offers near real-time transcription with sub-second latency but skips speaker diarization and word timestamps. The batch endpoint processes audio slower but includes both speaker labels and word-level timing, and costs half as much. On accuracy, the streaming option hits a 4.0% word error rate across more than 85 languages, while the batch endpoint improves to 2.6%.

Why it matters

Splitting transcription into two endpoints signals a clear trade-off between speed, detail, and cost, which directly affects how voice-driven products get built. Real-time applications like live captioning or voice assistants benefit from the streaming endpoint’s low latency despite losing detailed speaker info and word timing. Meanwhile, transcription workflows that emphasize accuracy and analysis, such as legal transcripts or media indexing, gain from the batch endpoint’s richer output at lower cost. This separation makes it easier for operators to optimize performance and cost depending on the use case.

Who it is for

Builders and operators focused on voice AI products now have more control over balancing speed, detail, and expense. Streaming suits live agents, customer support calls, and interactive voice response systems where speed trumps granularity. Batch fits transcription services, content search, and archiving scenarios where in-depth speaker context and precise timing are critical. Enterprises handling multilingual audio can also leverage Gemini 3.5’s broad language coverage to expand transcription capabilities globally.

The catch

Dropping speaker diarization and word timestamps in streaming mode limits its usefulness for applications that rely on identifying who said what and when. Users must choose between real-time speed or detailed transcripts, increasing integration complexity. Also, the model’s accuracy gains come with an implied need for more compute power and infrastructure to match, despite batch pricing cuts. Transitioning from previous Google models may require retooling pipelines since Gemini 3.5 does not unify endpoints.

What to watch next

Tracking Gemini 3.5 Transcribe’s adoption will reveal if splitting endpoints becomes a new standard or complication for voice AI builders. Watch for how Google expands language support, lowers cost further, or integrates speaker diarization back into streaming. Competitors like OpenAI and Microsoft might respond with similar multi-endpoint approaches or aggressive feature bundles. Finally, notice if real-time transcription accuracy continues to close the gap with batch solutions, possibly reshaping industry norms.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.