Models & Research

PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Func…

· July 31, 2026
PolyAI Releases Dialog-RSN-1: An Audio-Native Dialog Model That Fuses Turn-Taking, Speech Recognition, Func…

What it does

PolyAI has launched Dialog-RSN-1, a dialog model designed to process caller audio directly rather than relying on automatic speech recognition (ASR) transcripts. The model integrates turn-taking, speech recognition, function calling, and response generation into one unified audio-native system. Text-to-speech (TTS) is kept separate, preserving control over the spoken output’s voice characteristics. Unlike always-on streaming models, Dialog-RSN-1 operates on a request basis and delivers responses in under 300 milliseconds during live deployments.

Why it matters

Dialog-RSN-1 changes operational dynamics for voice-driven customer service and conversational AI. By avoiding dependence on ASR transcripts, it reduces intermediate processing errors and latency. Fusing the dialog tasks into a single model streamlines the tech stack and can cut development and maintenance complexity. Separating TTS from dialog generation also keeps voice customization flexible, which matters for brand consistency and user experience. The sub-300ms response time can improve real-time interactions, keeping conversations more natural and less robotic, which supports better customer satisfaction and efficiency.

Who it is for

This model targets conversational AI builders, contact center operators, and enterprises invested in voice automation who need faster, more integrated dialog handling. PolyAI’s approach suits use cases requiring robust turn-taking alongside dynamic function calls like API integrations or backend commands without adding processing overhead. Businesses focused on reducing friction in voice UX, lowering latency, and simplifying system architecture stand to benefit most.

The catch

Dialog-RSN-1 offloads TTS to maintain voice control, meaning it still requires a reliable external speech synthesis system to complete the voice interaction pipeline. Operating as a request-based large language model (LLM) may limit continuous stream-based scenarios or highly dynamic interactions in some contexts. Adoption may also hinge on integration complexity and how well it pairs with existing telephony and backend systems.

What to watch next

Monitor how PolyAI’s Dialog-RSN-1 performs at scale and whether its performance gains sustain across diverse languages and noisy environments. Watch for additional function calling capabilities or tighter integrations with popular contact center platforms. Competitors may respond by bundling similar end-to-end audio-native dialog systems. Pay attention to how separating TTS affects voice branding adoption and whether this architectural choice becomes a new industry norm.

AI Quick Briefs Editorial Desk

Stay ahead of AI Get the most important AI news delivered to your inbox — free.