How voice AI transforms call centers: technology, ROI, and deployment in 2026
Voice AI for call centers combines speech recognition and dialog reasoning with function calling to answer customers' questions, understand natural speech, and take actions inside your customer relationship management (CRM), billing, scheduling, or booking systems. An agent completes the caller's request without a live agent, handing the call to a person with full context when it can't.
What voice AI for call centers is, and how it differs from IVR
Touch-tone interactive voice response (IVR) completes only a small fraction of calls on its own. It sends the majority of callers into menus and queues or a zero-out to a live agent before it does anything useful. ContactBabel's Inner Circle Guide 2024 found that self-service handled only 16% of calls entirely.
Legacy IVR and first-generation voice automation talked to callers and directed them to menus and queues. Some also offered callbacks. A voice agent replaces the menu with a conversation. The caller says what they need in their own words, and the agent acts in the system behind it before confirming the outcome.
The difference is easiest to see side by side.
Why voice AI matters for contact centers now
Executive pressure to adopt AI in customer service is at a historic high, yet production deployment remains limited. That gap between mandate and implementation is where the real business case for voice AI lives.
The technology changed between Gartner's December 2024 deployment survey and its 2026 executive-pressure survey. Dialog models now call functions in enterprise systems mid-conversation and ground answers in retrieved knowledge instead of scripted responses. In PolyAI's case, the model, Dialog-RSN-1, also reasons directly over the caller's audio. The agent can tell a genuine pause from a dropped line and read hesitation before frustration builds, because it reasons over the caller's audio rather than reducing the call to a transcript first. A transcript strips out vocal tone and timing, including hesitation, the cues a human agent relies on to tell a frustrated caller from a satisfied one.
How voice AI works: the core technology stack
A caller on a mobile phone in a parking lot says, "I need to move my appointment to Thursday afternoon." In the classic stitched architecture, four separate systems handle that sentence in sequence:
Every time a voice AI system passes a caller's words from one component to the next, it loses a little time and a little context. The researchers behind VoiceBench, 2026 and SD-Eval, 2024 found that keeping speech recognition and language reasoning in separate stages tends to get the words right, while processing raw audio all in one go tends to pick up on tone and what is happening in the background around the words. Neither approach is free of trade-offs.
In plain terms, separating recognition and reasoning tends to get the words right; combining them tends to capture tone and context better.
PolyAI's answer is to fuse the input side while keeping the output side separate. PolyAI’s Dialog-RSN-1 combines turn-taking and speech recognition with function calling and response generation in one model. It reasons directly over raw audio, so it can use timing and line conditions to understand what the caller means. Hesitation also helps it decide when they've finished speaking. Text-to-speech stays in a separate layer that your team controls, so brand voice and delivery remain a configuration choice, and the model still produces a transcript for records.
Latency and real-time conversational performance
Conversations feel natural at 300 milliseconds or under. Push past 800 milliseconds and callers start to hear the hesitation, the beat before the agent responds that gives away it's not a person.
Dialog-RSN-1 responds in under 300ms because it listens and thinks in one step. Most voice AI passes a caller's words through separate systems before it can reply, and each handoff adds a delay that compounds by the time the caller hears an answer. Dialog-RSN-1 removes those handoffs, so your agent keeps pace with the conversation instead of making callers wait for it to catch up.
The results show up where they matter most. One insurance provider using PolyAI cut response latency by 37%. A restaurant group using PolyAI saw an 11% lift in call containment, the metric that mattered most to their team.
Ask any voice AI vendor what their latency number actually measures, and ask them to prove it holds up on a real call, not just in a lab.
Use cases and automation workflows
Inbound self-service pays back fastest. It takes over the calls your agents already handle by rote, so your team's time goes where it actually matters. The workflows PolyAI runs in production include account management, authentication and identity verification, call routing, billing and payments, booking and reservations, order management, troubleshooting, appointment setting, and FAQ handling. Deployments show the range:
- Utilities: PG&E's 2025 case study reports that its voice agent reduced payment transactions from roughly 7 minutes to 3.5 minutes.
- Healthcare and hospitality: Audibel cut abandonment from 46% to 2% and reduced wait times by 87%, from 10–15 minutes to under two minutes.
- Golden Nugget's agent handles 34% of central reservation calls and completes 87% without human intervention.
Post-call analysis runs on the same conversations. PolyScore grades every call automatically, while Smart Analyst answers plain-language questions about conversation data. Together, these give you and your team the answer to "which conversational moments cause dissatisfaction" without a manual QA sample.
Authentication in voice deployments typically pairs knowledge-based verification against account data with secure payment capture. For card payments, PolyAI integrates with PCI Pal so the caller enters digits by DTMF and the card number never enters the transcript.
Human escalation and agent handoff
A caller who has to repeat their account number and problem to a human after talking to the agent has lost every second the agent saved. PolyAI passes context on transfer through SIP headers and structured JSON payloads. The payload gives the receiving agent the caller's identity and conversation context.
The handoff API includes telephony failover, and a documented emergency and crisis guardrail routes distressed callers to a person immediately. For webchat, the same live-agent handoff works into NICE, Salesforce, Zendesk, Webex, Amazon Connect, and Genesys.
The escalation rate you design for matters as much as the mechanism. According to PolyAI research, 86% of customers say fast and accurate resolutions influence whether they buy from a brand ( PolyAI, State of Customer Conversations 2026 ). Resolution speed affects both service performance and revenue. An agent that escalates early and cleanly protects both satisfaction and revenue better than one that holds the caller too long.
Your agents, in turn, receive the complex, emotionally sensitive conversations where human judgment creates the most value after the agent completes the routine identity and lookup work.
Design the handoff first, then decide what the agent should resolve on its own.
ROI and performance metrics to expect
PolyAI commissioned and paid for the Forrester TEI study published in June 2025. Forrester interviewed four decision-makers across energy, healthcare, hospitality, and insurance with 480,000 to 16 million annual calls and 65–850 agents.
Forrester then built a composite US organization handling 4 million calls a year with 200 agents, with PolyAI handling 25% of calls in Year 1, 35% in Year 2, and 40% in Year 3. The risk-adjusted three-year results:
Forrester modeled about $2.6 million USD in PolyAI usage costs and about $119,000 in professional services over the three years, plus 200 internal labor hours over four weeks for implementation .
Named deployments show what the composite looks like in a single organization:
- PG&E serves 5.2 million households and handles 16 million calls a year, with demand spiking during outages. PolyAI's own measurement puts the agent's end-to-end containment rate at 67%. PG&E has also saved 35,000 labor hours and lifted customer satisfaction on outage calls by 22%. Customer effort fell by 25%.
- UniCredit's Zagrebačka banka started with 25% call abandonment and a 2-4 minute agent wait after the IVR. Together, we automated 27% of calls and routed calls 83% faster. We also cut abandonment by 10 percentage points and raised Net Promoter Score by 14 points. The team went live in three months against a projected year.
- Peppermill Resort Spa Casino handles 10,000-15,000 calls a month across five properties. It documented 187% ROI from labor savings alone and a 92% reduction in contact center call volume.
- Golden Nugget generated 3,000 reservations and $600,000 USD in revenue at one property in a single month from calls that previously went unanswered.
- Fogo de Chão answers 100% of guest calls and completes 88% of bookings. It expects more than $7 million USD in incremental revenue.
Enterprise and telephony integration
PolyAI runs on top of the CCaaS platform you already have, instead of replacing it. The enterprise systems the agent acts in are the ones that determine resolution rates:
- CCaaS and telephony: Five9, NICE CXone, Twilio Flex, Amazon Connect, Genesys, Dialpad, Talkdesk, and custom SIP. A CXone customer, for example, can add an agent in front of existing queues.
- CRM: Salesforce, Zendesk, HubSpot, and Microsoft Dynamics 365 for account lookup, case creation, and CRM-triggered outbound.
- Healthcare records: Epic, ModMed, athenaOne, Cerner (Oracle Health), and Raintree; PolyAI's direct Epic integration launched in 2026 with PDS Health among the first to deploy at scale ( PR Newswire, 2026 ).
- Payments and knowledge: PCI Pal for secure payment, and Gladly, Zendesk, and ServiceNow for grounded answers.
Compliance and security standards
PolyAI holds SOC 2 Type II and ISO/IEC 27001 certifications, along with Cyber Essentials and Cyber Essentials Plus under the UK's NCSC framework. Its systems meet HIPAA requirements, comply with GDPR, and are built toward PCI DSS for payment card data, with data residency across the US, UK, Canada, and EU.
Deployment options
Your deployment path comes down to who builds the agent and how.
Dialog Studio is PolyAI's no-code builder. Describe what you want in plain language, and Wren, PolyAI's natural-language assistant, generates the agent, conversation flows, and guardrails from your prompt, then applies changes as you approve them. Your contact center team can build, edit, and ship changes without writing a line of code.
The Agent Development Kit (ADK) gives engineering teams the same platform through git, their own IDE, and CI/CD. Branch, build, test, and merge through the Agents API, with the same governance and approval gates Dialog Studio runs on. Both paths write to the same agent, so a contact center manager building an FAQ flow in Dialog Studio and a developer adding a Function step in ADK are working on one system.
PolyAI runs open at every layer. Most agents run on PolyAI's proprietary dialog model, and other models are supported too, so you're not locked into one. Dialog Studio connects to over 130 systems out of the box through a self-serve Connect Portal, covering telephony, CRM, knowledge, and productivity tools. Anything with an API, ADK connects to directly.
Start with Dialog Studio if your contact center team owns the agent day-to- day. Start with the ADK if your engineering org needs code review and CI/CD on anything customer-facing. Most enterprise teams use both.
Common mistakes when evaluating voice AI
Buyers repeat the same errors across vendors. The ones we see most often:
-
Containment tells you whether the AI kept the call in-channel. It says nothing about whether the customer's problem actually got solved. As our co-founder Nikola Mrksic put it: "We deflected 70% of all contact. That number is meaningless without another question: what happened next?" We report resolution for that reason. Ask any vendor for the same, and score a pilot on completed requests, not calls that simply left the queue.
-
Most voice AI still runs on a cascade: audio goes into a speech recognizer, the recognizer's best guess gets turned into text, and only the text reaches the model that decides what to say. The accent, the tone, the hesitation, all of it gets thrown away at that first step, which is why cascaded systems hit a hard ceiling on anything that depends on hearing the call rather than reading it. It's the reason we built Dialog-RSN-1 to reason directly over raw audio instead of a transcript. Whichever vendor you're evaluating, the test is the same: bring recordings of your own callers, with your own accents and your own line noise, and listen for where the confidence in the transcript doesn't match what was actually said.
-
On a real call, deciding when the caller has actually finished speaking costs more time than anything the model does, and a slow lookup to your own systems, a booking check, an account lookup, delays every architecture equally, model included. Telephony adds its own penalty on top: call routing is chosen per phone number, so a badly routed number can pay a latency tax on every single call until someone notices. None of this shows up when a vendor quotes you one clean number. Ask instead: how is turn-end detected, is transcription streaming or batched, what does the worst call in twenty look like rather than the average, and what does a turn cost when it has to call one of your systems. Any vendor who can't break the number down that way, us included, can't actually fix it.
-
A voice agent is only as useful as the systems it can act in, and every connector needs an owner on your side. Before you sign, ask who owns each connector on your team, and what happens the day your CRM schema changes under it.
-
The handoff to a human deserves the same design attention as the automated path, not less. That means deciding upfront what triggers a handoff and what context travels with it, not improvising once a caller is already frustrated. Build that path, including the crisis path, before you build the happy path, then test it with angry callers.
How to get started
Typical enterprise deployments with PolyAI go live in 3–8 weeks, and some launch in two. UniCredit's Zagrebačka banka went live in three months against a one-year projection, in Croatian, with banking-grade controls, which is a reasonable upper bound for a regulated first deployment. PolyAI partners with your team through four phases:
- Discovery and pre-build: Together we collect call recordings and contact center metrics . We then review agent training material and agree the IVR integration strategy. Analyzing real calls at this stage surfaces the intents that drive volume and the exceptions that break scripts.
- Build: PolyAI dialogue designers co-create the conversation design with your team in Dialog Studio or the ADK. They configure knowledge grounding and guardrails, then select voice and language settings for each market.
- Integration: Connect the SIP/PSTN trunk first, then connect the CCaaS platform. Add CRM and knowledge sources next, then supervisor tooling. Configure the SIP header and JSON handoff payload during integration.
- Test and launch: Validate behavior in sandbox and pre-release environments, then run a limited-traffic go-live. Expand by intent or by percentage of calls as PolyScore confirms quality.
Set the first-month target conservatively and use production results to revise it. PG&E set an initial target of 3.5% of calls handled by the agent; it reached 15% on the first day and 67% over time. Forrester's composite modeled 200 internal labor hours over four weeks for implementation, which is a useful planning figure for your own team's time.
After launch, the work continues on both sides. PolyAI's per-minute usage covers proactive performance improvements and maintenance. It also covers 24/7 support. The platform's continuous optimization cycle uses PolyScore and Smart Analyst to find failing turns. Analyst Agents help edit prompts and flows, then promote fixes through the same reviewed environments.
Your team gains time for the conversations that need a person, and the agent's resolution share grows month over month instead of plateauing at the pilot number.
FAQs
-
PolyAI supports more than 75 languages at the platform level across more than 25 countries, and Raven 3.5 maintains 100% language adherence across its 23 supported languages, meaning it does not drift into English mid-call. If you serve more than one variety of the same language, ask for live calls in each one specifically. Québécois and Parisian French, for example, aren't interchangeable just because both are French.
-
Dialog-RSN-1 reasons over the caller's audio instead of relying only on a transcript, and PolyAI post-trained it on millions of production and synthetic conversations. As a result, hesitation and crosstalk inform its decisions about what was said and whether the caller is finished. Line conditions inform those decisions as well. Test this with your own recordings, including mobile calls from noisy locations, and score how often the agent asks the caller to repeat themselves.
-
The agent needs a fast, context-preserving path to a human, and PolyAI documents an emergency and crisis escalation guardrail alongside a handoff API with telephony failover. Distressed callers reach a person quickly, with full context passed along so they don't have to repeat themselves. Disclosing that the caller is speaking with AI, and escalating early rather than holding the caller too long, both protect the experience.
-
PG&E's own case study reports its customer transaction score rose from 7.0 to 7.3 overall and from 5.8 to 7.1 on outage calls year over year, and UniCredit's Zagrebačka banka added 14 NPS points in six months.
Fogo de Chão records 95% guest satisfaction through its agent, and Robert Half's chat agent holds a 4.17 average customer satisfaction score while resolving more than 90% of candidate conversations end to end.
You can find these case studies, along with the Forrester TEI study, at See all resources .
Thank you!
More resources
Ready to hear it for yourself?
Get a personalized demo to learn how PolyAI can help you drive measurable business value.