Frontier models

Dialog-RSN-1

Introducing the most advanced frontier dialog model, purpose-built for mission-critical, enterprise conversations.

Frontier model Dialog-RSN-1 — fastest for real-time dialog, under 300ms.

Latest updates

  1. Voice runs on a clock that text never had. Here's why turn-taking breaks most voice agents, what Dialog-RSN-1 does differently, and which layers of the stack enterprises should own.

  2. 533 CX leaders. 1,045 consumers. New PolyAI research shows where they disagree on customer service, and how to fix it.

  3. The traditional cascade loses the audio. Speech-to-speech loses control. Dialog-RSN-1 loses neither.

See all updates

Introducing Dialog-RSN-1

The most advanced frontier dialog model, purpose-built for enterprise conversations

  1. The cascade loses the audio. Speech-to-speech loses control. Dialog-RSN-1 loses neither.

  2. Quick to answer. Never talks over you. And it takes direction, right in the prompt.

  3. Under 300ms in production, with a tail tight enough to write an SLA against.

  4. The model that answers also transcribes — the most accurate we have ever tested.

  5. It listens natively. It speaks through the TTS you control. Your brand never gets mispronounced.

Dialog-RSN-1

The world’s fastest, most accurate, safest and cheapest reasoning for real-time dialog — live on customer calls today

  • FastestUnder 300ms
  • Most accurateAudio native reasoning
  • SafestGrounded, tool-gated
  • CheapestPer resolved call

Overview

Built different: a third dialog architecture

Dialog-RSN-1 hears a call the way a person does — catching hesitation, picking up frustration before it escalates, and knowing exactly when to speak.

Dialog-RSN-1

  1. audio
  2. Dialog-RSN-1
  3. TTS (separate)
  4. audio

Built in-house on an open-weight foundation and post-trained on millions of real production conversations, inside the same agent harness it runs in. It runs on demand like an ordinary request-based model, not an always-on stream — so turn-taking is just a prompt: “wait for all 6 digits,” “wait for the beep.”

Performance

The fastest model that fits inside a live call

Dialog-Eval measures what a model can do inside a live call, not what it could do with a minute to think: every turn is real audio, across accents, noise, and phone-line quality. Slow reasoning models are shown as a ceiling; among models fast enough for live conversation, Dialog-RSN-1 leads. We will release Dialog-Eval as an open benchmark.

Quality ↑

01s2s3s4s5sDialog-RSN-1GPT-5.5*Gemini 3.5 FlashGPT-realtime-2.1Gemini 3.1 Flash-LiteDialog-RAVEN-3.5*Claude Sonnet 5*GPT-5.2*Claude Sonnet 4.6*Claude Haiku 4.5*GPT-realtime-2.1-miniGPT-5-mini*
Median latency (faster is better) · * run as a cascade. In production, Dialog-RSN-1’s entire response-time distribution is faster than GPT-realtime-2.1’s median — p50 280ms vs 860ms, p90 500ms vs 1,900ms.
Show the chart data as a table
Quality vs latency: first to respond, last to interrupt
ModelMedian latencyDialog-Eval quality
Dialog-RSN-1280ms79.7
GPT-5.5*1,100ms77
Gemini 3.5 Flash2,000ms73.3
GPT-realtime-2.1860ms72.4
Gemini 3.1 Flash-Lite2,750ms71.6
Dialog-RAVEN-3.5*280ms70.8
Claude Sonnet 5*1,500ms70.4
GPT-5.2*900ms69.7
Claude Sonnet 4.6*1,100ms68.8
Claude Haiku 4.5*900ms63.6
GPT-realtime-2.1-mini840ms52.6
GPT-5-mini*1,300ms49.7
  • Independent judge

    Scoring is done by a third-party frontier model, not by ours. We are not in a position to mark ourselves favourably.

  • Blind to the model

    The judge is never told whether it is grading PolyAI or a competitor. It sees only the conversation and the response.

  • Answer defined first

    The judge works out what a good response must do before it sees what our model actually said.

The comparison that matters is the one you run yourself

During onboarding we run this same harness against a sample of your own traffic and share the full per-dimension results with your technical team under NDA, including the methodology and the per-turn scores. If a dimension matters more in your operation than it does in our benchmark, we weight it your way.

Measured on real conversations

  • 280ms

    Median first response in live production — 3x faster than GPT-realtime-2.1 at 860ms.

  • 79.7

    Highest Dialog-Eval quality of any real-time capable model (next best: 77.0).

  • 89.8

    Turn-taking accuracy, vs. 61.3 for the next best real-time baseline.

  • 3.7%

    Word error rate, where dedicated ASR APIs given the full conversation return 6.8–6.9%.

Notes on methodology

Reading the latency numbers honestly

Measured from request to enough text to start piping to TTS. Dialog-RSN-1’s number already includes speech recognition and turn-taking, because it is a single model. Cascaded systems exclude their separate VAD and ASR steps, which typically add roughly another second end to end, plus the more conservative thresholds a cascade needs. Compared like for like, the gap is wider than the tables show.

Point-in-time, not full rollouts

Every figure is one measurement of one model version on one harness, taken on the date published. Vendors ship new versions constantly, so read the tables as a snapshot rather than a standing ranking — and re-run them on your own traffic before you decide.

On the audio score ceiling for cascades

A cascaded system only ever sees a transcript, so its audio score measures what survived transcription rather than what was in the call. That puts a ceiling on the column no amount of LLM quality can lift, which is why the cascades cluster below an audio-native model on it.

Why there is no single blended score

Quality, audio, safety and latency trade against each other, and which of them matters depends on the calls you run. Blending them into one number hides that trade-off, so the dimensions are published separately and weighted your way during onboarding.

How a quality figure is actually calculated

For every turn the judge first writes down what a good response has to do, then scores the response it is shown against that, blind to which system produced it. A model’s figure is the mean of those per-turn scores across the evaluation set.

What we deliberately don’t do

No cherry-picked calls, no prompt tuned per benchmark run, no scoring by our own models, and no figures taken from a vendor’s own marketing material. Every model runs through the same harness, in the same agent, on the same audio.

Built for mission-critical conversations

  • Complex service calls

    Warranty intake end to end: identity verified, address confirmed, troubleshooting resolved, service request created. For your team: whole call types disappear from the queue.

  • Regulated & high-stakes

    Banking identity verification, payments, healthcare intake, and outage lines at peak volume. For your customers: verified, secure, and resolved the first time.

  • Turn-taking-sensitive flows

    Reference numbers, spellings, card digits. Turn-taking is a prompt, not a threshold. For your builders: no silence windows or probability thresholds left to tune.

  • Outbound & back office

    Navigating phone menus and holds, human vs. machine detection, reactivating cold leads with CRM read and write. For your operation: work that used to need a person dialling.

Real-world outcomes

  • +11%

    Relative increase in containment at a national restaurant group — calls resolved without a human handoff.

  • −37%

    Response latency cut at a large insurance provider after migrating to Dialog-RSN-1.

State of the art dialog models

Continuously trained on hundreds of millions of real business interactions, and scored in the open on Dialog-Eval, our independently judged benchmark.

  • A sphere of white particles on black.
    Latest

    Dialog-RSN-1

    Our flagship. Reasons directly over raw call audio for turn-taking, recognition and response in one audio-native model. First to respond, last to interrupt, built for the hard call.

    <300ms
    In production
    89.8
    Turn-taking accuracy
    3.7%
    Word error rate
  • A wave of white particles on black.

    Dialog-RAVEN-3.5

    Webchat and multilingual. Text-based dialog, supported for production, the fastest of the cascades, carrying the same agent across chat, SMS and messaging.

    ~280–600ms
    Cascade latency
    70.8
    Dialog-Eval quality
    Multilingual
    By design

Secure by design. Compliant by default.

Independent guardrails at every layer of the conversation, strict data residency and retention controls, and the certifications enterprise security teams expect — ISO/IEC 27001:2022, SOC 2+ Type 2, HIPAA (BAA available), PCI DSS, Cyber Essentials Plus, and GDPR-ready EU residency.

Read the trust centre →

HIPAA compliant, AICPA SOC 2, ISO 27001, PCI DSS compliant, GDPR and Cyber Essentials.

Nothing is locked in

Dialog-RSN-1 is one option within a model-agnostic, API-first platform, not a one-way door.

  • Model choice stays open

    • Same agent runs on Dialog-RAVEN-3.5 or third-party models
    • Guardrails, tools, and knowledge carry over unchanged
    • A/B test or fall back, no rebuild required
    • Self-hosted on PolyAI-managed GPUs
    • No dependency on third-party inference APIs
  • Adopt it in pieces

    • Use the full stack, or just part of it
    • Turn-taking prediction on its own
    • Recognition plus response, without the rest
    • Composable, not an all-or-nothing swap
  • Your systems stay authoritative

    • Transactions are deterministic calls to your APIs
    • Business logic and system of record stay yours
    • Data retained for the contract term only
    • Then deleted or anonymised via PII redaction

Frequently asked questions

When should I use Dialog-RSN-1?

For voice agents on the phone, wherever audio understanding, turn-taking, and latency matter. For everything else, see the model comparison above.

How is this different from speech-to-speech models like GPT Realtime?

A speech-to-speech model keeps the audio, but it also owns the voice, the turn-taking thresholds and the cost of running the stream. Dialog-RSN-1 reasons over the raw audio and then hands the words to the TTS you control, so your brand voice stays yours and turn-taking is a prompt rather than a threshold to tune.

Why not just improve the cascaded pipeline?

A cascade throws the audio away at the transcript: hesitation, tone and the reason someone is frustrated are gone before the model reads a word. Better components make the transcript better; they do not bring the audio back — and every extra stage adds latency the call cannot absorb.

What happens to my existing Dialog-RAVEN deployments?

Nothing. Dialog-RAVEN-3.5 stays supported for production and remains the model to choose for webchat and non-English deployments. The same agent, guardrails, tools and knowledge run on either model, so moving a deployment across is a configuration change, not a rebuild.

Build at the frontier of dialog