Models & Research

Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis

· October 3, 2026
Microsoft AI Releases MAI-Transcribe-2-Streaming: #1 Real-Time Speech-to-Text Model on Artificial Analysis

What it does

Microsoft AI launched MAI-Transcribe-2-Streaming, its first real-time speech-to-text model. It tops a benchmark of 38 streaming models on Artificial Analysis, scoring 2.5% word error rate at around 0.13 seconds for final transcripts and 0.12 seconds for initial partial results. The model supports continuous language detection across 60 languages. It is available in public preview via Microsoft Foundry with a promotional price of $0.54 per audio hour.

Why it matters

Real-time transcription has been constrained by a trade-off between speed, accuracy, and language coverage. MAI-Transcribe-2-Streaming pushes this balance substantially, delivering near-instant transcripts with low error across many languages. This performance puts pressure on competitors to improve latency without sacrificing accuracy. Its price point also signals tougher competition in cost-efficient live transcription services. For businesses relying on multilingual voice data capture, this model offers a more scalable and affordable option to integrate real-time transcription at global scale.

Who it is for

Developers building voice-enabled applications now get a high-performing, multi-language streaming transcription engine. Enterprises aiming to automate customer support, transcription, or voice analytics can apply this model to improve response times and transcript quality. Content creators and meeting platforms targeting an international audience gain built-in continuous language detection without separate pre-configuration. Microsoft Foundry access eases integration for cloud-first operators ready to deploy advanced speech-to-text features quickly.

The catch

Though the introductory pricing is competitive, costs may rise after the initial period. Continuous language detection across 60 languages is promising but could introduce complexity if users require customization or handling of dialects and accents outside the model’s default scope. Public preview means some features or performance aspects could still evolve before a full release, so early adopters should test thoroughly.

What to watch next

Monitor whether Microsoft expands MAI-Transcribe-2-Streaming’s language and accent robustness, and how it handles domain-specific speech such as medical or legal jargon. Watch if competitors respond with faster, cheaper real-time models to maintain market share. Cost adjustments post-introductory offer will be critical for sustained adoption. Finally, follow Microsoft Foundry’s rollout as it becomes the main distribution channel for this and similar AI services.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.