Ask an enterprise buyer what they want from a dialog agent, and you’ll hear two demands that don’t sit comfortably together: an agent that performs like a skilled human and one that behaves with the predictability of a machine.

That tension explains why voice AI remains far harder to deploy well than most people assume, and the difficulty starts before governance even enters the conversation.

Our CTO and Co-Founder Shawn Wen sat down with Machine Learning Street Talk to discuss this exact problem.

PolyAI is the agentic dialog platform for the enterprise. We build dialog agents that finish what customers start, and we built Dialog-RSN-1 , our audio-native dialog model, because voice is where that job is hardest. Here's why, and what it takes to get it right.


Text solved a different problem than voice

Most of the recent progress in AI reasoning happened in text, a world with excellent documentation and no timer against it. A model can take as long as it needs before it answers. Voice removes that luxury entirely. When a caller pauses mid-sentence, the system has a fraction of a second to decide whether they’re finished speaking or still gathering their thought. If someone starts talking while the agent is still mid-response, then the system has to work out in real time whether that’s an interruption to yield to or background noise to ignore.

Handling this well is a timing problem more than a reasoning problem. Timing is one of the hardest things to teach a model, because voice sits one step closer to the physical world than most AI applications do, and machines have never had access to the millions of everyday human conversations that would let them learn how real turn-taking works.

The ChatGPT moment that wasn’t

When ChatGPT first demonstrated how capable language models had become, the obvious next step looked like voice: pipe speech recognition and text-to-speech around the same model and the problem would resolve itself. A wave of wrapper companies formed around exactly that idea.

The reality proved harder. Speech recognition and text-to-speech had each individually gotten very good, but stitching them together with a language model in the middle didn’t produce a system that could hold a real conversation. Understanding pauses, handling interruptions, and reading when someone has finished speaking are dynamics that a cascade of separately-tuned components struggles to reproduce, no matter how good each individual piece is on its own.

MIT found that 95% of organizations saw no measurable return from their generative AI pilots , and in our part of the market the reason is usually because of a general-purpose model with a prompt in front of it, sold as an agent.

When the behavior lives in a prompt, it drifts every time the model underneath changes, and testing happens mostly on text rather than on the audio a live call contains. A prompt is an instruction the model may or may not follow on any given turn. A trained behavior is what the model does by default, even when the caller interrupts, trails off, or has a baby crying in the background. That's why we trained the dialog behavior into the Dialog-RSN-1 itself.

What enterprises actually want

Enterprises want two things from a dialog agent:

  1. Identity, and a branded voice that says the company’s name correctly, every time, in a tone that matches the brand.
  2. An agent that performs as well as their best human team member, while remaining fully governable, with no surprises and a complete audit trail for every decision it makes.

That combination is a difficult bar to clear, and it’s the real reason enterprise adoption moves slower than the underlying technology would suggest. Trust builds once a team can point to one specific workflow, show a real efficiency gain there, and watch the agent perform better than expected. That’s usually where broader adoption starts.

Why demo-quality models break in production

The newest speech-to-speech models sound close to indistinguishable from a real conversation. But if you deploy the same model in a live contact center, the cracks in capabilities are clear.

The most common failure is interruption handling tuned for a confident, steady speaker rather than a first-time caller who hesitates and pauses often. A model tuned on typical conversational pacing reads those pauses as “done speaking” and talks over them. That single moment breaks a customer’s trust in the system.

Fixing this means training on the conversations that happen in production: real calls across banking, utilities, logistics, hospitality, and retail, with the noise, crosstalk, and hesitation those calls contain. Counterintuitively, cleaning that data up first, stripping background noise, smoothing out crosstalk, makes the resulting model worse, because it never learns to handle the messy conditions it will face on a live call. That's why we train our own dialog models rather than rent someone else's.

A new & better way to build voice AI

For years, voice agents have been built in one of two ways. The first was a cascade that chains three separate tools: one transcribes the caller, one decides what to say, and one decides when to say it. Each step strips something out, so by the time the AI responds it's working from a best-guess transcript with the tone, hesitation and background noise removed.

The second way, speech-to-speech models, fix that by hearing and speaking in one model. The trade-off is that the voice is baked in, so you have little control over how it sounds or how it says your brand name, and it's expensive to run at contact center scale.

Dialog-RSN-1 takes a third path. A single model hears the raw audio of the call and handles turn-taking, understanding and calling your systems. A separate voice engine then speaks the reply, so you keep full control of your brand's voice. Before anything else, the model decides whether the caller is still talking or has finished, so a pause mid-sentence reads to the model as "the customer is still talking," not "done." It responds in under 300 milliseconds on average, which means customer conversations feel far more natural and seamless.

What enterprises actually need to own

Enterprises want to own AI the same way they own their brand, their product, and their customer relationships. That ambition ran into a wall at the model layer, as tuning a foundation model takes expertise most enterprises don’t have on staff, the GPU cost runs well past what most budgets would approve. The model layer got crossed off the list of things worth owning, and the API relationship took its place.

What stayed on the list was the harness: the layer of prompts, tools, and orchestration logic that decides how the underlying model behaves for a given business. It’s the one part of the stack an enterprise can genuinely build in-house without a research team attached, which explains why so many are building one, even where it means solving problems a specialized platform has already solved.

Voice sits outside that build list for a different reason. It carries a governance burden most of the stack doesn’t: turn-taking has to work call after call, in real time, before any decision-making even starts. A fast, natural front door to your brand that answers the phone in your company’s own voice, while the actual decision-making stays inside your enterprise’s own system, lets each side own what it’s equipped to own. The platform handles the talking, and the enterprise keeps the deciding.

Most companies never sat down and decided which of the model, the harness, and voice were worth owning. They built all three because nobody told them a choice was available.

How that ownership actually works

That choice is what we built PolyAI around. We own the dialog layer, and the rest is yours to decide:

  • Business users build in Dialog Studio
  • Engineers build in the ADK and the CLI
  • Anyone can build over the API and MCP. If you ever leave, your flows, evals, and memory leave with you.

Ownership of the model means deployment ownership: Dialog-RSN-1 is licensed to you and runs in your own environment, so the core intelligence of your customer operation is something you hold rather than something you rent from a lab whose roadmap you don't control.

The division between talking and deciding only holds if the handoff stays auditable, which is why we build PolyAI as a glass box, not a black box. Every agent is governed on every turn: every tool call is visible, every action is attributed, and you can roll back anything the agent does. Every response ties to a citation, and every process is visible as a flow you can inspect. Ownership, in practice, means visibility into what the agent did and why, more than possession of every line of code that produced it.

The same control extends to how the agent improves. PolyAI’s Wren reviews every conversation, drafts a fix, proves it with a test, and waits for approval. The dial extends to full autonomy as your trust builds.

Where this goes next

Voice is moving toward systems that are audio-native by default, with turn-taking built into the model rather than bolted on as a separate module. Consumer voice and enterprise voice look set to keep diverging: one optimized for personality and entertainment, the other for auditability, governance, and getting the customer’s problem solved as fast as possible.

The underlying lesson holds regardless of where the architecture goes next. Voice is a fundamentally different problem from text, shaped by real-time constraints that text never had to deal with. Enterprises that treat it that way, and choose systems trained on the messy reality of real calls rather than the clean conditions of a demo, are the ones seeing dialog agents hold up once they go live.

Make every customer feel heard. Instantly. Speak to our team today.