The half second that shapes every call: voice AI latency and customer satisfaction

  • Daniel Lavalette

    Senior Coordinator, Content Writer and Social

  • 30 Sept 2026
  • 7 min
The half second that shapes every call: voice AI latency and customer satisfaction

Share

Your customer calls about a refund. The agent answers straight away, confirms her name and date of birth, and asks how it can help, and so far it feels like any other call. Then she asks, "Has my refund cleared?" and the line goes quiet for a second, long enough for her to wonder whether anyone is still there.

By every measure on your team’s dashboard, the call is a success: answered, resolved, no human needed. What she remembers is the pause, and the small doubt it cast over the answer that followed, enough that she checks her banking app later to be sure. When 86% of customers say fast, accurate resolution affects whether they buy from a brand, how quickly and naturally that answer arrives is part of the resolution she judges you on.

Voice AI latency affects customer satisfaction because callers interpret silence in specific ways. A reply within the 200 to 500 milliseconds people expect keeps the caller's trust. After about one second, callers start to think the agent doesn't know or the line has dropped, and they may rate the call lower even when it resolved the request.

Why silence on a call hurts more than a slow web page

A late reply hurts most on a call because your caller has nothing to look at. A web page can show a spinner. On a call, silence is the only signal the caller gets. So she fills the quiet herself. Our design principles put it plainly: "Silence on a voice call feels like a dropped connection." That's when callers start saying "hello?" to check someone is still there.

Once callers read silence as an agent that doesn't know or a dropped line, the question becomes how much they will accept. People reply to each other in roughly 200 to 500 milliseconds, says Shawn Wen, our co-founder and Chief Technology Officer. Under 200 milliseconds feels like an interruption, and over a second the caller thinks something has gone wrong. Response time runs from the end of the caller's speech to the first sound of the reply, and the phone network's delay adds to your agent's reply time.

We recommend an operating target of a median under 500 milliseconds and a 95th-percentile (P95) response under one second, measured at the caller's handset. The median keeps a typical call feeling natural, and the P95 limit keeps your slowest calls under the one-second point where callers start to think something has gone wrong.

Most of the delay happens before the agent says a word

During any call, an agent first has to know the caller has finished speaking, a step called turn detection that Shawn Wen calls "the largest cost in most turns, and the least discussed." Then speech-to-text turns their words into text, and the language model figures out what the caller needs: "Has my refund cleared?" means checking the status of one specific payment.

Next, the agent looks up that refund in the billing system and waits for a reply. Only then can text-to-speech start speaking. Shawn's post on voice latency walks through how to shorten each of these steps.

In a cascaded voice agent, each step is a separate system handing off to the next, and the caller's voice is flattened into a transcript along the way, losing the pauses and tone that signal a finished sentence. Dialog-RSN-1, PolyAI's audio-native dialog model, listens to the caller's audio directly, along with the agent's instructions and the conversation so far, so it can tell a pause or a dropped line from a finished sentence. That tackles turn detection, the highest cost Shawn describes. When your caller pauses to find a refund reference, it waits, then answers when she finishes. It reliably responds in under 300 milliseconds, with very few slow outliers, so the pauses callers notice are rare.

Write the slow calls into your SLA

Hearing the agent over a real phone line shows you how it performs today, and agreeing targets with your vendor keeps it that way as call volumes grow and your setup changes. An average hides the calls that hurt you, because a single latency figure without its distribution describes only the calls nobody would have complained about.

When you set latency targets with your vendor, agree on each of these:

  • Name the percentile, with P95 at minimum.
  • Fix the start and stop points of end-to-end latency: end of caller speech to first audio reaching the caller, the delay your customers feel.
  • Set the test conditions as real telephony and your own call mix during peak hour.
  • Ask for monthly P95 reports per call route, with a plan to fix any route that misses its target.
  • Ask for latency reported separately for turns that involve a system lookup, like the refund check.

Then read latency next to customer satisfaction (CSAT) scores, sentiment, first-call resolution, and in-conversation abandonment on the same calls, so you can see whether the slow calls are the unhappy ones.

Resolution still matters most, and a call counts as resolved when the action lands in your system of record and the caller confirms it. Latency shapes how that resolution feels, so each month, take your slowest 5% of calls, the ones above your P95, and check them against PolyScore, our automated 1 to 5 quality score for engaged conversations. If those calls score lower, you'll see exactly where speed is costing you.

When a lookup takes time, tell the caller what the agent is doing

No matter how much you shorten everything else, some waits are out of your hands. When the agent checks a refund, it has to wait for your billing system to respond, and that can take a few seconds. The right acknowledgment covers that wait. In a 2025 study, natural, contextual acknowledgments improved how participants rated response time at 4-second and 6.5-second delays, while artificial wait indicators had no significant effect. So name the task. "Checking your refund now" tells the caller the agent heard them and knows what to do next.

Our delay-control tooling lets a designer schedule acknowledgments against a slow lookup, timed from when the lookup starts. For a refund check that takes several seconds, the agent might say "Let me pull up that refund" straight away, "Still checking that refund for you..." half a second later, and "Thanks for waiting!" two seconds after that. Lines can include details from the call, so they stay specific to what the caller asked. When a lookup comes back in under a second, skip the fillers, because a line that plays for no reason sounds worse than a short pause.

A right answer still has to arrive on time

Think back to that refund call. The agent had the answer, but a second of quiet made her sound unsure, and while your dashboard logged a resolved call, she left remembering the pause.

A reply within the half second people expect from each other would have kept the call feeling like a conversation, and she'd have heard an agent that knew its answer. Hold your agent to that, measured at the caller's handset and including the slowest calls. Dialog-RSN-1 responds in under 300 milliseconds, built for exactly that rhythm.

Hear Dialog-RSN-1 on your own calls: request a demo.

Frequently asked questions

What is an acceptable latency for a voice AI agent?

Set an operating target of a median under 500 milliseconds from the end of the caller's speech to the first sound of the reply, with P95 under one second over real phone lines. PolyAI co-founder and CTO Shawn Wen says people reply to each other in roughly 200 to 500 milliseconds, and past one second the caller starts to think something has gone wrong.

Is 20ms audio latency good?

Twenty milliseconds is the default size of one phone audio packet and says nothing about how fast the agent answers. Judge the agent on end-to-end response time at the caller's handset, since a call can move 20-millisecond packets perfectly and still leave the caller waiting 1.5 seconds.

What is acceptable VoIP latency?

The International Telecommunication Union's G.114 standard (2003, still in force) treats one-way delay under 150 milliseconds as transparent and sets 400 milliseconds as the limit for network planning. Those figures cover transmission between two people and add to the agent's response time on a voice AI call.

How do you reduce voice AI latency?

Start where the time goes: turn detection and slow system lookups. Stream speech-to-text and text-to-speech, speed up or pre-fetch the lookups the agent waits on, and use a model that handles turn-taking itself instead of a silence timer. Then measure end-to-end latency at the caller's handset again.

Hear a voice agent on your own calls

Book a 30-minute demo and we’ll run a conversation from your world through a live agent.