Why do specialized AI models win in CX?
About the show
Hosted by Nikola Mrkšić, Co-founder and CEO of PolyAI, the Deep Learning with PolyAI podcast is the window into AI for CX leaders. We cut through hype in customer experience, support, and contact center AI — helping decision-makers understand what really matters.
Summary
General-purpose AI models are powerful, but when it comes to customer experience, power alone isn't enough.
In this episode of Deep Learning with PolyAI, Nikola Mrkšić sits down with Matt Henderson, VP of Research at PolyAI, to unpack why specialized AI models consistently outperform generalist models in real-world CX environments.
The conversation centers around the launch of Raven 3.5 and the broader philosophy behind building AI systems specifically designed for voice and customer interactions.
Together, Nikola and Matt explore:
- Why optimizing for CX requires balancing latency, cost, accuracy, and reasoning — all at once
- Why voice AI behaves differently from text-based AI, and what most models get wrong
- How training on real conversational data leads to more natural and reliable interactions
- How specialized models embed behavior directly into the model instead of relying on prompts
- What “auto-reasoning” is and how models learn when to think versus respond instantly
Watch the full episode to hear why specialized models are becoming essential for CX.
Key takeaways
- Specialization wins when you're optimizing everything at once: Raven targets sub-300-millisecond response times while also balancing cost, accuracy, instruction following, and multilingual consistency. A specialized model can make those trade-offs deliberately; a generalist model optimized for everything ends up optimized for nothing in particular.
- Constraints belong in the model weights, not the prompt: Raven is trained to ground its answers in the documents it has access to, cite its sources, flag out-of-domain requests, and stay consistent in the caller's language. With general models, teams layer prompt instructions on top of each other — what Nikola calls “layers of paint” — until the system becomes contradictory and impossible to interpret.
- Auto-reasoning teaches the model when to think: Reasoning improves quality but adds latency, so Raven 3.5 learns to reason only when it helps — warmed up with curated examples, then trained with reward signals that prefer the shortest reasoning that reaches the same quality. The hard part is teaching a model to know what it doesn't know.
- Real conversational data is the moat: Training on anonymized production phone conversations teaches the model how people actually talk — interruptions, recoveries, pauses — which general models aren't specialized for yet. Large reasoning models could top the benchmarks given unlimited time, so PolyAI uses them offline and distills their capability into smaller models that fit the latency budget.
-
[00:00:00] Matt Henderson: As a race, we are adapting towards how to speak to AI, but we're trying to meet people where they are right now. That's the great thing about working here for me is that we have so much data that we can train on. And it's all, you know, anonymized and everything, but it how people talk over the phone. And the model knows that this is speech, and, uh, it might be interrupted or it might need to recover from bad interruptions and stuff like that. We just train that in. The general models are not so specialized in voice speech
[00:00:42] Nikola Mrkšić: yet. Hi, everyone, and welcome to another episode of deep learning with PolyAI. Today with me is our VP of Research, Matt Henderson. Um, and, uh, we're here to talk about our, uh, new model, our new LLM, Raven 3.5. Matt, welcome to the podcast. Before we start, everyone, please like, share, subscribe. Um, and, uh, Matt, over to you.
[00:01:02] Matt Henderson: Yeah. Hey. Thanks for having me on again. Enjoyed the, uh, my first appearance. This is my second time on. Yeah. We're talking all things three point five today. Right?
[00:01:11] Nikola Mrkšić: Yeah. Absolutely. I mean, like, you know, I think we are very proud of, you know, our continued, uh, investment in r and d and just, like, how much you know, we are a research led company. Um, I think your PhD was the first of all, uh, of the all the people of the company still. Uh, and I think that, you know, we've had a thread from, well, 2012 to today of just working on this stuff. Right? Um, and, um, I think it'd be great to just tell the world, like, why is this special? Why is it different? Why is this not easy? What's special about it? I think there's more confusion than ever, and I think a lot of people have just decided that they're gonna either pick up the model from, like, OpenAI, Areanthropic or, you know, mister Oliver Quinn. Um, why are we doing this?
[00:01:57] Matt Henderson: Yeah. So I guess this is a question. Why train your own model? Right? Like, why go why specialize? And as you mentioned, competing models would be the generalist, the big, um, models on the cloud. Um, of course, a big reason for us is speed. We target and achieve a sort of sub three hundred milliseconds, uh, response time, and, uh, that just puts us, like, on the edge of optimizing. And I think at least for now, specialized model is always gonna be a generalist, especially when you're optimizing against so many things. Yeah. So we're optimizing for for latency cost. Those are easy. Right? But you also have to optimize for, uh, instruction following, then it becomes hard. And
[00:02:44] Nikola Mrkšić: can we maybe double click on that for, like, the you know, I think we've got, like, two clusters of audience, one being technical, and they'll know what you mean, and the others who, um, might not know what that means. I think there's, like, the whole, like, how do I make a system not hallucinate, right, uh, piece and, you know, like, what what are we doing to that end? Right?
[00:03:03] Matt Henderson: Yeah. Well, we train Raven to do things that are quite specific and specialized for the use case of, you know, customer support over the phone or over a web chat. One of those things is being grounded in the prompt or being grounded in the documents that it has access to. So it's specifically trained to respond to you and then to cite its the source of information for, uh, its, uh, what it said And optionally say, well, actually, this was an out of domain request. So you get that built in LLM powered, uh, flagging of what topics are being cited, what pieces of content and documents are being cited, and what are people asking that the system just can't answer yet. Just super useful. Not just for keeping the model from hallucinating. That helps as a sort of additional training signal, but also for your logging and just tracking all that stuff. Um, you wouldn't get that in a general model without sort of extra harnessing and framework around, like, prompting it. You just got something that's built into the model. You don't need prompt for it especially, and you serve you you get it for free.
[00:04:17] Nikola Mrkšić: Yeah. And, I mean, I I think when you look at, like, you know, people increasingly are playing with these models and with agentic systems. And, you know, it does feel like, you know, in the matrix, when you see the machine just building a more and more complicated thing. Right? It if you're using a general model and prompting your way around both stylistic things and then, like, the right tool calling, it ends up being quite, well, frankly, impossible to read what you've done or to know what you've done. And then it's kinda like just another layer of pan of paint just kinda like saying, hey. Please don't do this. In this case, do this. In this case, do that, and it starts getting contradictory. So think like a model that just inherently doesn't force you to that kind of, like, layers of paint gets you to, um, point where you build a system this much better and less complex. And less complex is always good because it means that it's more interpretable. Right?
[00:05:02] Matt Henderson: Yeah. So we basically put all of these kind of constraints that we know ahead of time into the model weights rather than into this into a dynamic prompt. So, you know, another example would be sort of consistency with what language it speaks. So we want Raven to be instructable in whatever language you're comfortable with. Yeah. That's typically in English. And then with the flick of a flip of a switch, then it it speaks in Spanish or Portuguese or Japanese or whatever. That's super confusing for the general models
[00:05:33] Nikola Mrkšić: because, you
[00:05:34] Matt Henderson: know, they you ask them, uh, the user asks them a question in Spanish, and it retrieves a document that's in English. All the all the prompting is in English, and it says, you know, there's maybe a reminder response using this document or something, and it's forgotten it's supposed to speak Spanish. And, uh, it differs part in English. It's a super bad error to make because one of those things that we can just put into our reward signal and and just Yeah. Be tighter the model.
[00:06:00] Nikola Mrkšić: Yeah. Yeah. No. I mean, I think and I think, like, when when you look at specialization, you kinda like you are what you eat. Right? Where like, if you're working on problems that require this level of precision and, like, multiminguality, Like, you kinda have to, like, get way better at it. And tell us a bit more about, like, the 3.5 aspect. Like, what's special relative to to Raven three? Like, auto reasoning would be one of one of the things. Right?
[00:06:21] Matt Henderson: Yeah. Auto reasoning is is the is the coolest new feature, I think, um, apart from well, maybe we'll dive into that in a bit. But, I mean, across the board, it's just sort of improvements on everything. We have a we start from a better pretrained model that's that has, uh, better multilingual capabilities. So we're just a big lift in in non English languages as well as English. We have these new features, like, out of domain detection. We work in web chat very well now. A few improvements in things like custom style following. So one of the other things we sort of built into Raven is a good output style for conversations that are happening via voice or, you know, optionally, web chat. Right? Um, there's a lot of kind of LLMEs style outputs that the models like GPT will you'll start to recognize and get really frustrated with if you're a caller, you know, being transferred three times and you get through to this. Like, I'd be really happy to drill down and get to the bottom of this for you.
[00:07:25] Nikola Mrkšić: And then I'll ask question it. Yeah. Right? I mean, like, I think I think what like, one thing that I've been, like, uh, trying to explain to a lot of people is latency is an obvious kinda, like, property of the system that makes it better. Right? If it, like, spends less time thinking and responds quickly without interrupting you when you've just made a pause, that's obviously a great thing. Right? I think, like, the barge in feature is one that is quite interesting where, you know, obviously, it's a great thing to show off. And in some places, it's incredible. In others, it can be quite disruptive if people accidentally, you know, interrupt the thing or another person speaking in the background does something. It is essential for a lot of the rappers to have good margin because their systems built on GPT models just won't shut up. And they're building this, like, UX assumption of like, to to anyone who's used the voice mode in ChargeGPT, it becomes quite obvious that you have to be very authoritative and, frankly, in human terms, quite rude when you speak to it and that you just have to cut it off and direct it where you want it to go. And if you do that, it's a bit of a different modality of conversation, but it is very powerful. But we what we see on the phone often is that people don't know how to do that. And there's this, like, discrepancy between, you know, how the power users of voice mode and child GPT, which are probably no more than point 1% of the planet, right, uh, behave with it versus everyone else who's just kinda, like, following the, like, rules of behavior from a human conversation. A lot of it is just, like, you don't want it to, like, speak forever and then ask you a question at the end of the whole thing.
[00:08:57] Matt Henderson: As a race, we are adapting towards how to speak to AI, but we're trying to meet people where they are right now. And, um, I guess we kind of get that a bunch because of all our training is on, you know, the data that we collect. That's the great thing about working here for me is that we have so much data that we can train on, and it's all, um, you know, anonymized and, um, uh, everything, but it how people talk over the phone. And the model knows that this is a speech, and, uh, it might be interrupted or it might need to recover from bad interruptions and stuff like that. We just train that in. Yeah. General the general models are not so specialized in voice speech yet. I mean yeah. So at this GPT real time is, like, a very good model that we compete with and and benchmark against.
[00:09:48] Nikola Mrkšić: I mean, like, I've heard a lot of customers. In fact, it's generating a lot of demand where people speak to it and they're like, I want that. Right? And then it's like, well, you know, that's a certain kind of car. Right? It's it might be a hypercar. Like, it's equally difficult. You know? You can't really drive it around London because you'll hit a curb wherever you go or indicate a model. It will hallucinate things and, you know, won't do, like, tool calling. But it's definitely showing us what's possible in terms of speed and the naturalness of dialogue. Right? I mean, I find it very impressive. We don't wanna talk about Raven four here, but I think so excited about that one. Right?
[00:10:23] Matt Henderson: Yeah. Raven four Raven Omni, we're starting to do audio. Yeah. So we're we, um, don't need speech recognition anymore, and we can sort of hear the user of native understanding of the speech or the or the audio. But, yeah, I guess that's the next podcast I'll appear on.
[00:10:40] Nikola Mrkšić: Maybe we stop on that thread. Tell me a bit about, you know, just, like, for the audience. Right? I mean, you know, I remember inheriting your code base and, you know, it was the wonderful library of Teano, and you couldn't even do, like, you know, TensorFlow with its optimizers was revolutionary in terms of how easy it was to, like, run an experiment. Like, why is this so hard when people, you know, live in a world where, like, a com you know, open clock can set up, like, an entire operating system and create, like, development pipelines that you needed a DevOps team for before? My patient. But, like, um, why is this still hard?
[00:11:14] Matt Henderson: I think yeah. We're we're training this generative model, right, that that can that you can speak to. And then I we can train and and and get a loss and think, okay. This looks fine. Even run it on a benchmark and and get some average number. But it really takes us to start talking to the model, getting in front of our, um, agent deployment teams and and in front of people to see what are the issues. You know? Is it calling tools robustly? Does it have some sort of weird style? When you start talking to a multi time conversation, what shows up? Maybe it's repetitive. It says great after every call, uh, after, you know, everything the user says. Um, we it's just the kinds of things that you only come across if you've been building conversational systems for, like, the length of time that that you and I have. We have a list of behaviors we wanna put into the model, and then we test for, and everyone had every one of those behaviors might have a sort of individual creative solution for, like, more more data, more preferences, a different reward signal, something like that.
[00:12:17] Nikola Mrkšić: When you talk about those kind of different metrics that you're optimizing for, right, like, I guess, we talked about style. We talked about kinda, like, construction following. Then there's the whole, like, balancing reasoning and latency and how you decide to do that. And then the naturalness of voice versus, like, the precision of answer and, like, the length of the answer. How do you, like, optimize for all that in training? And, like, like, post training setup, like, how does it go and work? Right?
[00:12:42] Matt Henderson: I know. That's your list as well. So there's also, like, uh, tool calling, and then there's all of those, but in all the different languages. And Pardon? Right. Yeah. And different modalities and what's on their web chat. Um, we also wanna check that those things, like, about hallucination and citing your topics and stuff work well. Uh, languages is particularly interesting because we would have a whole different style guide for each language. Now the you might wanna have different rules about politeness when it comes to an English conversation versus a Japanese conversation might be more formal. And in some languages, you can sort of avoid, but gender the person you're talking to to the pronouns you're using. And so we have certain kind of, uh, rules there. So the by the way, what the the question was, how do we balance all those things? We track all of them in all of our evaluations. Uh, and I guess making them measurable is the is the first important thing we do on the research team. So we have this internal benchmark, and you'll see that in our, uh, blog post announcing Raven 2.5 that we we beat, um, the big, you know, public nodules, d five, uh, the latest Claude sonnets on those benchmarks. And then each has an individual, um, uh, solution for making it work. And I guess I wanna start painting the picture of this isn't just launching a, you know, run DPO script, getting a model at the end and saying there, we we train 2.5. It has more data, blah blah blah. If you were to start plot the lineage of the final model checkpoint, right, it's it's the average of a bunch of runs of the reasoning tuning stage, which comes from a base model of an average of all these different one off DPO and GRPO and combined g r GPO, TR DPO runs. Um, it's become this kind of, like, messy law like, art of figuring out how to post train for all of those things at the same time.
[00:14:38] Nikola Mrkšić: Are we the more we talk about the data being like a moat and it absolutely is a moat. Like, you know, one way that Sean, our, um, our CTO described it and that I thought was really interesting was, like, you're basically kinda, like, just applying generations of different it's like you go to different grades as a model. Right? And, like, how you learn chemistry in grade five could very well inform whether you go for, like, physics or chemistry in, like, grade seven. Right? And then, like, your whole learning path takes a different, like, evolution where, really, like, what you did at what part of the training and how, like, you know, you got to rewrite, like, the the rewards and stuff. Kind of like it's not really a linear, like, training phase in the way that people used to think about it. Right? But it's really, like, a multistage, like, building of a system where it's more of a dark art than ever in a way. Right?
[00:15:36] Matt Henderson: It's pretty sort of surgical and targeted. And that like, yeah. Before you and I are doing our PhDs, we usually, you know, launch experiments, like, trains on this data set and it finishes and that. You know? It was randomly initialized usually. I think you did work on bringing in pretrained word embeddings and stuff, which is first steps to, you know, where we are now. But now we have these massive models, lots of data. We wanna sometimes we wanna reuse the sold model. It's it's really good. It's just not in this certain situation, it doesn't use the right tool. So then we we we fix it surgically.
[00:16:10] Nikola Mrkšić: You know, I think about this a lot. When people are like, what are we gonna be doing in the future? And, you know, if we were explaining what we're doing even to, like, a mathematician in, like, the thirties or forties, they'd probably think we lost our minds, or they were doing something very trivial and that it makes no sense. Right? Uh, they I mean, they'd understand everything. It's just like you know, you think of, like, people building, like, handcrafted products, you know, like leather goods or furniture or whatever, and you kinda see them, like, you know, polishing this one thing, shaping this other thing, leaving for a few days. And, you know, I kinda help but, like, draw the comparison where it's really, like, it's really not that different to this, where your intuition is, like, the meta level kinda, like, sequencer, and, like, it informs, like, how much even money and resources and time you want to invest in this, like, whole, like, sausage making thing. And
[00:17:00] Matt Henderson: Yeah. We we're sort of artisanal sort of model developers, um, and we we become, you know, reward function engineers and, um, you know, GRPO loss for mixing engineers. But at least we're not prompt engineers, so stuck with a generalist model, and all we have at our disposal is playing around with a with a prompt. We we we can back propagate and, uh, adjust, uh, model to exactly what we want. The auto reasoning stuff was, like, surprisingly difficult to train, I guess. And when it comes to reward engineering, you know, the model's always gonna try to, um, reward hack. And, uh, well, so with alternate reasoning, what that means is you're you want to get the benefits of reasoning, which means that the model takes some time before responding. Right? So it does some little generation and thinks mix to itself. But in our case, we don't that that would that adds latency necessarily to the to the call because it's doing it. So it couldn't, in theory, respond immediately. So we only want to do that in cases where it's gonna help. Um, so how do we teach the model when it should reason, when it shouldn't? So the naive thing to do is that you kind of just generate with and without reasoning. And if the if the one where it reasons is better, then but that then you train like that.
[00:18:19] Nikola Mrkšić: How do you fact, like, how long it reasons for into the balancing?
[00:18:24] Matt Henderson: Yeah. That's the that's the other tricky. That's one of our GRPO losses is making reasoning short because all of these, um, well, the typical sort of base model you would get doesn't have any kind of constraint like that. And so the original, we loop back to when reasoning first came in with DeepSeek. So we have some crazy long repetitive traces, and you know, really benchmark performs go way up, but, sure, you have to wait, like, five minutes to for it to actually reply to you. So we do that we do that in this artisanal way. Right? So we warm it up in a in SFT stage with some traces that we think are good. And then when it comes to later stages, we've got specific loss function to say, if you can get to the same quality response with shorter reasoning, that's you know, prefer the shorter reasoning reward. But I think one of the core difficulties you're super trying to teach the model to know what it doesn't know. Right? So the these models are sort of no famous for being confident bullshitters sometimes. Right? Yep. Yep. They're they're not well calibrated inside. So how can we teach it to know when it needs to stop and think? Uh, so that was particularly hard. With this, like, shortening the reasoning thing, those degenerate reward hacking cases, Well, I seem to be doing better and better every time I'll reason less, so I'm just never gonna reason. So there's these all these things are bouncing. And, you know, the trick with when to reason, well, we just warmed up with some SFT, basically. Some Yep. You should reason here. When it comes to this, we know this is a very difficult, you know, date time request rather than asking the model to figure out that it it can be unreliable without reasoning in those cases. We just want
[00:20:07] Nikola Mrkšić: like a providing a bit of structure that helps it, like, make that decision or learn to make that decision rather than letting it just infer from, like, data.
[00:20:16] Matt Henderson: That's right. So the beginning of training, it's it has something to latch onto, and then it and then it can kind of, like, figure out the other things that where reasoning helps. It's not obvious to why should it be obvious to the model that needs to reason to figure out how many r's are in the word raspberry? Yeah. That doesn't sound
[00:20:33] Nikola Mrkšić: Yeah. Because in the end of the day, the the, like, the the overall, like, um, signal whether it was good or not is one thing, but then, like, predicting the next word versus predicting it faster is quite a and you know how much faster it becomes quite a it really is some kind of, like, racing. Right?
[00:20:50] Matt Henderson: Yeah. There's a there's a, you know, there's a nice feature. So it's kind of if you want, just switch on auto reasoning. Raven will think when it, uh, when it thinks it should help, and it wouldn't add, uh, licensing.
[00:21:02] Nikola Mrkšić: How much better would would, like, one do in these benchmarks if it reasons if time wasn't an issue and, you know, like, we weren't optimizing? Like, how much better are reasoning models, you know, when we compare to, like, you know, sonic and stuff? If we took OPUS, for instance, because I can't detach myself from my, like, you know, uh, OpenClaw and what it does with the Opus model. I see a very distinct difference in, like, coding ability with one versus the other. How well would it do, like, relative to these models? It will not make much of a difference.
[00:21:33] Matt Henderson: Those are going to do very well in our benchmarks. Uh, large reasoning models where latency is not doesn't matter, then you can you can certainly top all of our our our benchmarks. So we wanna guard a bit against the the future where they do become, uh, instantaneous and fast. And, you know, when when you talk about using your your claw or clawed spot or whatever, open claw, There's a whole lot of, like, agent harnessing and stuff that sort of makes it respond very well that adds on latency. Yep.
[00:22:07] Nikola Mrkšić: Yep. The
[00:22:08] Matt Henderson: yeah. They you know, those I guess our our approach is to benefit from these types of models. I open models where we can run reasoning and do stuff that's inefficient. We used to do the offline enduring training and sort of distill it down into these, uh, smaller models.
[00:22:25] Nikola Mrkšić: Yeah. Because you could I think, Curtis, just like we're getting it ready for a race. It has to run a race. And at the end of the day, like, are we able to get most of the benefits in a faster model then of, you know, how much better a bigger model would do?
[00:22:40] Matt Henderson: Yeah. I we basic we we can, uh, get we we benefit from from all that all that teaching and achieve something which you just you you wouldn't be able to do with the latency budget and, uh, obviously, much much most.
[00:22:55] Nikola Mrkšić: Okay. Okay. Cool. Well, um, I think, like, with that, this was really just a teaser for what's gonna come in a few weeks' time with the release of the next model, but I think we've probably already said a bit too much about that. Thank you so much for joining. To everyone in total listening, like, check out the link. Check out Raven 3.5. Um, there are very exciting releases coming about the platform and all the models included with it over the coming few weeks. So, you know, subscribe, like, share, and we'll see you in the next one. Matt, thank you for for joining me today.
[00:23:27] Matt Henderson: Great. Thanks for