Cartesia Ships Sonic-3.6: A Streaming TTS Model That Now Leads Both Artificial Analysis Speech Arenas
What it does
Cartesia has launched Sonic-3.6, a new streaming text-to-speech (TTS) model that uses state space models instead of the common transformer architecture. It streams audio with a sub-90 millisecond time-to-first-audio performance. Sonic-3.6 currently ranks first on both Artificial Analysis speech leaderboards: 1,283 Elo on the Provider Voice leaderboard and 1,123 on the Controlled Voice leaderboard, which levels the playing field by cloning models onto the same eight reference voices. The model is available in beta through Cartesia’s API for developers and businesses.
Why it matters
Sonic-3.6 changes the TTS playing field by leading in quality benchmarks while delivering fast streaming response times. Using state space models rather than transformers offers lower latency for real-time applications like voice assistants, live broadcasts, or accessibility tools. Ranking top on both leaderboards means Sonic-3.6 reliably outperforms competitors in both natural voice delivery and controlled benchmarking tests that isolate the speech synthesis engine itself. This builds pressure on transformer-based TTS providers and raises the bar for streaming speed, a critical factor in voice interfaces.
Who it is for
Developers, business operators, and startups seeking high-quality, low-latency streaming text-to-speech have a new option to evaluate. Sonic-3.6’s API access allows practical integration into products that need rapid, natural-sounding audio synthesis. This is particularly useful for real-time communication platforms, interactive voice response systems, or any application where speed and synthesis clarity directly impact user experience.
The catch
While Sonic-3.6 is impressive in benchmarks and latency, it remains in beta, meaning stability and enterprise-level support details are still emerging. Using a newer architecture like state space models instead of the more established transformer models could involve integration quirks or compatibility considerations with existing infrastructure. The controlled voice ranking highlights strengths in synthesis engine quality but may not directly translate to every use case’s voice customization needs.
What to watch next
Tracking Sonic-3.6’s adoption and Cartesia’s roadmap for moving beyond beta will show if this technology can reshape streaming TTS expectations. Observing if competitors respond by optimizing transformers for latency or pivoting architectures will shed light on the future of real-time speech synthesis. Watch how Sonic-3.6 performs in real-world deployments and if Cartesia expands their voice options or integrates more deeply into communication platforms.
AI Quick Briefs Editorial Desk