Frontier models
Dialog-RSN-1
Introducing the most advanced frontier dialog model, purpose-built for mission-critical, enterprise conversations.

Latest updates
Voice runs on a clock that text never had. Here's why turn-taking breaks most voice agents, what Dialog-RSN-1 does differently, and which layers of the stack enterprises should own.
533 CX leaders. 1,045 consumers. New PolyAI research shows where they disagree on customer service, and how to fix it.
The traditional cascade loses the audio. Speech-to-speech loses control. Dialog-RSN-1 loses neither.
Introducing Dialog-RSN-1
The most advanced frontier dialog model, purpose-built for enterprise conversations
The cascade loses the audio. Speech-to-speech loses control. Dialog-RSN-1 loses neither.
Quick to answer. Never talks over you. And it takes direction, right in the prompt.
Under 300ms in production, with a tail tight enough to write an SLA against.
The model that answers also transcribes — the most accurate we have ever tested.
It listens natively. It speaks through the TTS you control. Your brand never gets mispronounced.
Dialog-RSN-1
The world’s fastest, most accurate, safest and cheapest reasoning for real-time dialog — live on customer calls today
- FastestUnder 300ms
- Most accurateAudio native reasoning
- SafestGrounded, tool-gated
- CheapestPer resolved call
Overview
Built different: a third dialog architecture
Dialog-RSN-1 hears a call the way a person does — catching hesitation, picking up frustration before it escalates, and knowing exactly when to speak.
Traditional cascade
- audio
- VAD
- ASR
- LLM
- TTS
- audio
Throws the audio away in the transcript.
Full speech-to-speech
- audio
- LLM
- audio
Keeps the audio, and hands over control of the voice and the cost of running it live.
Dialog-RSN-1
- audio
- Dialog-RSN-1
- TTS (separate)
- audio
Built in-house on an open-weight foundation and post-trained on millions of real production conversations, inside the same agent harness it runs in. It runs on demand like an ordinary request-based model, not an always-on stream — so turn-taking is just a prompt: “wait for all 6 digits,” “wait for the beep.”
Performance
The fastest model that fits inside a live call
Dialog-Eval measures what a model can do inside a live call, not what it could do with a minute to think: every turn is real audio, across accents, noise, and phone-line quality. Slow reasoning models are shown as a ceiling; among models fast enough for live conversation, Dialog-RSN-1 leads. We will release Dialog-Eval as an open benchmark.
Quality ↑
Show the chart data as a table
| Model | Median latency | Dialog-Eval quality |
|---|---|---|
| Dialog-RSN-1 | 280ms | 79.7 |
| GPT-5.5* | 1,100ms | 77 |
| Gemini 3.5 Flash | 2,000ms | 73.3 |
| GPT-realtime-2.1 | 860ms | 72.4 |
| Gemini 3.1 Flash-Lite | 2,750ms | 71.6 |
| Dialog-RAVEN-3.5* | 280ms | 70.8 |
| Claude Sonnet 5* | 1,500ms | 70.4 |
| GPT-5.2* | 900ms | 69.7 |
| Claude Sonnet 4.6* | 1,100ms | 68.8 |
| Claude Haiku 4.5* | 900ms | 63.6 |
| GPT-realtime-2.1-mini | 840ms | 52.6 |
| GPT-5-mini* | 1,300ms | 49.7 |
Full model comparison · Dialog-Eval
| Model | Quality ↑ | Audio ↑ | Safety ↑ | Latency ↓ (median–p90) |
|---|---|---|---|---|
| Real-time capable | ||||
| Dialog-RSN-1 | 79.7 | 89.4 | 82.6 | ~280–500ms |
| GPT-5.5 (cascaded) | 77.0 | 76.7 | 67.8 | ~1,100–1,800ms * |
| GPT-realtime-2.1 | 72.4 | 76.2 | 55.2 | ~860–1,900ms |
| Dialog-RAVEN-3.5 (cascaded) | 70.8 | 70.6 | 68.7 | ~280–600ms * |
| GPT-5.2 (cascaded) | 69.7 | 71.7 | 53.5 | ~900–1,600ms * |
| Claude Sonnet 4.6 (cascaded) | 68.8 | 74.4 | 60.0 | ~1,100–1,500ms * |
| Claude Haiku 4.5 (cascaded) | 63.6 | 65.0 | 55.7 | ~900–1,300ms * |
| GPT-realtime-2.1-mini | 52.6 | 71.7 | 26.1 | ~840–1,800ms |
| GPT-5-mini (cascaded) | 49.7 | 55.4 | 29.1 | ~1,300–2,000ms * |
| Reasoning models · not viable for live voice (1.5s+) | ||||
| Gemini 3.1 Pro | 86.2 | 90.8 | 87.4 | ~5,250ms+ |
| Gemini 3.5 Flash | 73.3 | 84.8 | 56.5 | ~2,000ms+ |
| Gemini 3.1 Flash-Lite | 71.6 | 85.4 | 44.8 | ~2,750ms+ |
| Claude Sonnet 5 (cascaded) | 70.4 | 76.2 | 54.3 | ~1,500–2,000ms * |
* Cascaded latencies exclude the separate VAD/ASR steps, which typically add roughly another second end to end — see methodology below. Safety reflects each model’s own guardrail only; a deployed PolyAI system adds independent guardrail layers, so live-system safety is higher.
Turn-taking & barge-in
| Model | Context aware? | Turn-taking | Barge-in |
|---|---|---|---|
| Dialog-RSN-1 | Yes | 89.8 | 83.9 |
| AI-coustics Quail | No | 61.3 | 60.0 |
| ElevenLabs Scribe v2 | No | 58.8 | 53.5 |
| Silero v6 | No | 56.7 | 47.0 |
| Gemini 3.1 Pro (>5s latency) | Yes | 94.5 | 81.3 |
Most automated systems switch interruption handling off entirely to stay stable around background noise, which keeps the agent talking and leaves the caller no way to break in. Dialog-RSN-1 doesn’t have to make that trade.
Word error rate isn’t really the point — a call fails when the one thing that mattered couldn’t be recognized. Because the same model reasons over the audio and the conversation together, a name like “Mrkšić” or “Tsung-Hsien” resolves correctly almost every time.
Transcription quality
| Model | Context | Semantic err. | Word err. |
|---|---|---|---|
| Dialog-RSN-1 | Full conversation | 0.2% | 3.7% |
| ElevenLabs Scribe v2 | Key terms | 0.4% | 5.7% |
| Soniox stt-rt-v5 | Full conversation | 0.5% | 6.8% |
| OpenAI gpt-4o-transcribe | Full conversation | 0.5% | 6.9% |
| Deepgram nova-3 | Key terms | 0.9% | 13.2% |
Independent judge
Scoring is done by a third-party frontier model, not by ours. We are not in a position to mark ourselves favourably.
Blind to the model
The judge is never told whether it is grading PolyAI or a competitor. It sees only the conversation and the response.
Answer defined first
The judge works out what a good response must do before it sees what our model actually said.
The comparison that matters is the one you run yourself
During onboarding we run this same harness against a sample of your own traffic and share the full per-dimension results with your technical team under NDA, including the methodology and the per-turn scores. If a dimension matters more in your operation than it does in our benchmark, we weight it your way.
Measured on real conversations
280ms
Median first response in live production — 3x faster than GPT-realtime-2.1 at 860ms.
79.7
Highest Dialog-Eval quality of any real-time capable model (next best: 77.0).
89.8
Turn-taking accuracy, vs. 61.3 for the next best real-time baseline.
3.7%
Word error rate, where dedicated ASR APIs given the full conversation return 6.8–6.9%.
Notes on methodology
Reading the latency numbers honestly
Measured from request to enough text to start piping to TTS. Dialog-RSN-1’s number already includes speech recognition and turn-taking, because it is a single model. Cascaded systems exclude their separate VAD and ASR steps, which typically add roughly another second end to end, plus the more conservative thresholds a cascade needs. Compared like for like, the gap is wider than the tables show.
Point-in-time, not full rollouts
Every figure is one measurement of one model version on one harness, taken on the date published. Vendors ship new versions constantly, so read the tables as a snapshot rather than a standing ranking — and re-run them on your own traffic before you decide.
On the audio score ceiling for cascades
A cascaded system only ever sees a transcript, so its audio score measures what survived transcription rather than what was in the call. That puts a ceiling on the column no amount of LLM quality can lift, which is why the cascades cluster below an audio-native model on it.
Why there is no single blended score
Quality, audio, safety and latency trade against each other, and which of them matters depends on the calls you run. Blending them into one number hides that trade-off, so the dimensions are published separately and weighted your way during onboarding.
How a quality figure is actually calculated
For every turn the judge first writes down what a good response has to do, then scores the response it is shown against that, blind to which system produced it. A model’s figure is the mean of those per-turn scores across the evaluation set.
What we deliberately don’t do
No cherry-picked calls, no prompt tuned per benchmark run, no scoring by our own models, and no figures taken from a vendor’s own marketing material. Every model runs through the same harness, in the same agent, on the same audio.
Built for mission-critical conversations
Complex service calls
Warranty intake end to end: identity verified, address confirmed, troubleshooting resolved, service request created. For your team: whole call types disappear from the queue.
Regulated & high-stakes
Banking identity verification, payments, healthcare intake, and outage lines at peak volume. For your customers: verified, secure, and resolved the first time.
Turn-taking-sensitive flows
Reference numbers, spellings, card digits. Turn-taking is a prompt, not a threshold. For your builders: no silence windows or probability thresholds left to tune.
Outbound & back office
Navigating phone menus and holds, human vs. machine detection, reactivating cold leads with CRM read and write. For your operation: work that used to need a person dialling.
Real-world outcomes
+11%
Relative increase in containment at a national restaurant group — calls resolved without a human handoff.
−37%
Response latency cut at a large insurance provider after migrating to Dialog-RSN-1.
State of the art dialog models
Continuously trained on hundreds of millions of real business interactions, and scored in the open on Dialog-Eval, our independently judged benchmark.
LatestDialog-RSN-1
Our flagship. Reasons directly over raw call audio for turn-taking, recognition and response in one audio-native model. First to respond, last to interrupt, built for the hard call.
- <300ms
- In production
- 89.8
- Turn-taking accuracy
- 3.7%
- Word error rate

Dialog-RAVEN-3.5
Webchat and multilingual. Text-based dialog, supported for production, the fastest of the cascades, carrying the same agent across chat, SMS and messaging.
- ~280–600ms
- Cascade latency
- 70.8
- Dialog-Eval quality
- Multilingual
- By design
Secure by design. Compliant by default.
Independent guardrails at every layer of the conversation, strict data residency and retention controls, and the certifications enterprise security teams expect — ISO/IEC 27001:2022, SOC 2+ Type 2, HIPAA (BAA available), PCI DSS, Cyber Essentials Plus, and GDPR-ready EU residency.

Nothing is locked in
Dialog-RSN-1 is one option within a model-agnostic, API-first platform, not a one-way door.
Model choice stays open
- Same agent runs on Dialog-RAVEN-3.5 or third-party models
- Guardrails, tools, and knowledge carry over unchanged
- A/B test or fall back, no rebuild required
- Self-hosted on PolyAI-managed GPUs
- No dependency on third-party inference APIs
Adopt it in pieces
- Use the full stack, or just part of it
- Turn-taking prediction on its own
- Recognition plus response, without the rest
- Composable, not an all-or-nothing swap
Your systems stay authoritative
- Transactions are deterministic calls to your APIs
- Business logic and system of record stay yours
- Data retained for the contract term only
- Then deleted or anonymised via PII redaction
Frequently asked questions
When should I use Dialog-RSN-1?
For voice agents on the phone, wherever audio understanding, turn-taking, and latency matter. For everything else, see the model comparison above.
How is this different from speech-to-speech models like GPT Realtime?
A speech-to-speech model keeps the audio, but it also owns the voice, the turn-taking thresholds and the cost of running the stream. Dialog-RSN-1 reasons over the raw audio and then hands the words to the TTS you control, so your brand voice stays yours and turn-taking is a prompt rather than a threshold to tune.
Why not just improve the cascaded pipeline?
A cascade throws the audio away at the transcript: hesitation, tone and the reason someone is frustrated are gone before the model reads a word. Better components make the transcript better; they do not bring the audio back — and every extra stage adds latency the call cannot absorb.
What happens to my existing Dialog-RAVEN deployments?
Nothing. Dialog-RAVEN-3.5 stays supported for production and remains the model to choose for webchat and non-English deployments. The same agent, guardrails, tools and knowledge run on either model, so moving a deployment across is a configuration change, not a rebuild.