Frontier models
Dialog-RSN-1
Introducing the most advanced frontier dialog model, purpose-built for mission-critical, enterprise conversations.
HIGHLIGHTS
Introducing Dialog-RSN-1
PolyAI’s audio aware conversational model that listens to the audio itself, and makes decisions about your conversation while you talk.
-
A transcript can't tell a model whether someone has finished their sentence or is still thinking, which is why most voice agents talk over people; Dialog-RSN-1 hears the call itself, so it knows the difference.
-
Quick to answer. Never talks over you. And you shape its behavior with plain instructions.
-
Under 300ms in production, and consistent enough to write an SLA against.
-
The model that answers also transcribes. The most accurate we have ever tested.
-
It hears every part of the call: tone, timing, hesitation. It speaks through the TTS you control, so your brand never gets mispronounced.
BUILT DIFFERENT
A third dialog architecture
Dialog-RSN-1 hears a call the way a person does, catching hesitation, picking up frustration before it escalates, and knowing exactly when to speak.
PERFORMANCE
The fastest model that fits inside a live call
Frontier voice AI, benchmarked against the best and proven in live calls.
Dialog-RSN-1 is first to respond, last to interrupt
Most voice agents run as a cascade: speech converted to text, handed to a separate model to reason over, then converted back to speech, with delay added at every handoff. Dialog-RSN-1 skips that translation and reasons over the call directly. In production, half its responses land in 280 milliseconds, against 860 milliseconds for GPT-realtime-2.1, and even its slowest one in ten replies, at 500 milliseconds, is still faster than GPT-realtime's typical response. For a caller, that's the difference between a pause that feels like a natural beat in conversation and a pause that feels like the line went quiet.
Dialog-RSN-1 tops real-time capable models across benchmarks
Because Dialog-RSN-1 is a single model, its latency figures include recognition and turn-taking right out of the box. Cascaded systems, on the other hand, exclude separate VAD and ASR steps, which typically add an entire second of delay end-to-end. When compared like for like, the real-world performance gap is much wider than the table suggests.
The safety metrics follow a similar rule. The table reflects only the model’s baseline guardrails. Once deployed, a PolyAI system adds independent safety layers, pushing live-system protection significantly higher.
Turn-taking & barge-in
Nothing breaks a caller's trust faster than being talked over. To stay stable around background noise, most automated systems simply turn off interruption handling altogether. This keeps the agent talking and leaves the caller completely unable to break in. Dialog-RSN-1 never forces you to make that compromise. It handles interruptions smoothly and yields naturally to the caller.
Dialog-RSN-1 outperforms on transcription quality.
Dialog-RSN-1 redefines accuracy by rethinking the architecture. Instead of relying on a separate speech recognition model, the exact same intelligence that responds to you is the one transcribing the audio. It knows the context, the history, and the prompt. By answering first and transcribing second, it delivers industry-leading accuracy that outpaces even the strongest APIs on the market.
The comparison that matters is the one you run yourself
During onboarding we run this same harness against a sample of your own traffic and share the full per-dimension results with your technical team under NDA, including the methodology and the per-turn scores. If a dimension matters more in your operation than it does in our benchmark, we weight it your way.
Measured on real conversations
Median first response in live production — 3x faster than GPT Realtime at 860ms.
Highest Dialog-Eval quality of any real-time capable model (next best: 77.0).
Turn-taking accuracy, vs. 61.3 for the next best real-time baseline.
Word error rate, where dedicated ASR APIs given the same context return 5.7–13.2%.
Built for mission-critical conversations
Complex service calls
Warranty intake end to end, from identity to service request.
Regulated & high-stakes
Banking verification, payments, healthcare intake, outage lines.
Turn-taking-sensitive flows
Reference numbers, spellings and card digits, heard in full.
Outbound & back office
Phone menus and holds, machine detection, CRM read and write.
Real-world outcomes
Relative increase in containment at a national restaurant group — calls resolved without a human handoff.
Response latency cut at a large insurance provider after migrating to Dialog-RSN-1.
Choose your model
Safe, independent, flexible. Continuously trained on hundreds of millions of real business interactions, and scored in the open on Dialog-Eval, our independently judged benchmark.
Dialog-RSN-1
Our flagship. Reasons over raw call audio for turn-taking, recognition and response in one audio-native model. First to respond, last to interrupt.
Latency in production
Turn-taking accuracy
Word error rate
Best for: Phone calls
Dialog-RAVEN-3.5
Text-based channels and multilingual. The fastest of the cascades, carrying the same agent across chat, SMS and messaging.
Cascade latency
Dialog-Eval quality
Languages available
Best for: Text-based channels and non-English deployments
Secure by design. Compliant by default.
Independent guardrails at every layer of the conversation, strict data residency and retention controls, and the certifications enterprise security teams expect — ISO/IEC 27001:2022, SOC 2+ Type 2, HIPAA (BAA available), PCI DSS, Cyber Essentials Plus, and GDPR-ready EU residency.
Nothing is locked in
Dialog-RSN-1 is one option within a model-agnostic, API-first platform, not a one-way door.
Model choice stays open
- Same agent runs on Dialog-RAVEN-3.5 or third-party models
- Guardrails, tools, and knowledge carry over unchanged
- A/B test or fall back, no rebuild required
- Self-hosted on PolyAI-managed GPUs
- No dependency on third-party inference APIs
Adopt it in pieces
- Use the full stack, or just part of it
- Turn-taking prediction on its own
- Recognition plus response, without the rest
- Composable, not an all-or-nothing swap
Your data, your systems
- Transactions are deterministic calls to your APIs
- Business logic and system of record stay yours
- Data retained for the contract term only
- Then deleted or anonymised via PII redaction
Frequently asked questions
-
For voice agents on the phone, wherever audio understanding, turn-taking, and latency matter. For everything else, see the model comparison above.
-
Speech-to-speech bakes the voice into the model weights and streams always-on audio, which is costly and hard to control. Dialog-RSN-1 is audio-native on the input only: it hears the call, keeps speech generation in a separate TTS layer you control, and runs on demand like a normal request-based model. In our testing, GPT Realtime scores on par with cascades on audio-aware evaluations.
-
The cascade's bottleneck is structural: the LLM only sees the ASR's best guess. Dialog-RSN-1 removes the bottleneck instead of patching around it: recognition, turn-taking, function calling, and response generation all share one context.
-
Nothing changes: Dialog-RAVEN-3.5 stays fully supported. Switch projects to Dialog-RSN-1 in Dialog Studio whenever you're ready.
-
Pricing is tailored to your deployment. Contact sales for a quote.