Is word error rate just a vanity metric?
About the show
Hosted by Nikola Mrkšić, Co-founder and CEO of PolyAI, the Deep Learning with PolyAI podcast is the window into AI for CX leaders. We cut through hype in customer experience, support, and contact center AI — helping decision-makers understand what really matters.
Summary
Word error rate is the number most of the voice AI industry can't stop quoting. It might also be the one buyers should trust least.
In this episode, Nikola Mrkšić and PolyAI agent design and engineering lead Oliver Shoulson dig into what a month of analysis across 100,000+ production call turns actually showed about how well voice agents understand people.
Their main finding: word error rate is buried in noise. Most mistranscriptions are harmless. A dropped filler word, a missed article, an inflection that's wrong in a language with heavy morphology — none of it changes what the customer meant. A smaller set are critical entity errors, like a single wrong digit in a phone number, and those are the ones that quietly break a call. When Oliver's team pulled the harmless errors out on their own, they explained less than 1% of the variation in conversation quality. Nikola calls them empty calories.
The bigger surprise: task success stayed almost perfectly flat as meaning-changing errors piled up. LLMs recover from mishearings far better than expected — callers pay in a little friction, not failure. The conversation also covers where context-aware ASR still falls short, and a blunt take on full-duplex interfaces like GPT live that are built to be talked over.
For anyone evaluating voice AI, here's the takeaway: stop buying on benchmark word error rate. Start measuring task success and conversation quality on your own calls.
Key takeaways
- Not all words are created equal: Classifying mistranscriptions by business impact — benign, semantic, and critical — shows that the benign ones (filler words, articles, formatting) explain under 1% of the variation in conversation quality. Judging an ASR system on raw word error rate means buying on noise.
- LLMs recover from errors better than you think: Task success stayed almost perfectly flat even as meaning-changing errors accumulated — callers pay in extra clarification turns, not in failed calls. The real cliff in conversation quality only appears at three or more meaning-changing errors, a cohort of just tens of calls in the whole dataset.
- Production reality beats vendor benchmarks: The headline production word error rate was 6.4% for English and 8.8% for Spanish (6.7% weighted) — nowhere near the 2–3% vendors quote from clean datasets, once phone low-pass filters, VoIP packet loss, and callers driving down the highway enter the picture.
- Semantic and critical errors behave in opposite ways: Semantic errors get punished on quality (−0.55) because they force clarification, but barely move handoffs. Critical entity errors — a wrong digit in a phone number — hit quality half as hard (−0.27) because they escalate to a human quickly and appropriately. The future is context-aware ASR that sees the conversation, not blind transcription.
-
[00:00:00] Oliver Shoulson: It's clear to me now more than ever that, like, not all words are created equal. Right? In context, you know, missing a yes or no could be everything or it could be nothing now that we're have we're starting to deal with audio models that, like, take prompting in and situational context in addition to the audio itself. Like, just going on, like, raw, how many words did we get right, like, is a pretty bad proxy for all the stuff that we actually care about.
[00:00:35] Nikola Mrkšić: Hello, everyone, and welcome to another episode of deep learning with PolyAI. Today, I've got Oliver back on the show with me. Good to see you again.
[00:00:43] Oliver Shoulson: Great to be here.
[00:00:44] Nikola Mrkšić: Yeah. Oliver is one of our elite Asian designers and a real kind of force behind a lot of work, whether it's, you know, pioneering OpenClaw or just, you know, bringing the good spirit of linguistics in. And today, after we spoke about some very interesting data that he showed me, I thought it would be really interesting to just kinda, like, have a whole episode that really looked into the magical notion of a word error rate. Right? So we build voice agents and how well these systems work feeds downstream from whether you're able to recognize what someone has said or not. Right? And I think most companies, especially those that focus predominantly on that component, really just talk about improving the word error rate and getting to a better and better one in their public benchmarks. And it's kinda treated as a holy grail by a lot of people, buying software, deciding what component to use, deciding whether it's good or not. And we'll talk a lot about whether it's good or not. I think, like, one anecdote from my peers, the supervisor was really good where I think it was the early days of deep learning, and we were all we just thought it would all be done in two years. I think that's the prevalent attitude that's continued ever since. And I remember Steve telling me, like, whoever thinks this future recognition will be solved in two years has worked on the problem for less than two years. And, you know, I've now spent more than two years working on the problem. I I feel like it was right. But, Oliver, like, maybe over to you to talk a bit about, like, the word error rates and the findings that you've had over the past, like, few months.
[00:02:14] Oliver Shoulson: Yeah. Totally. I so, I mean, I have much less experience, like, working on this, like, word error rate in academic context than you, but, you know, what I can tell you is that
[00:02:23] Nikola Mrkšić: But more than two years.
[00:02:25] Oliver Shoulson: More than two yeah. I guess so. So I first of all, I know that ASR is not solved. And second of all, I know now that word error rate is a pretty terrible proxy for conversation quality. Like, if you're using that as a way to evaluate both how well your ASR is performing in context and just as a way to understand, like like, how well are my customers being understood by the virtual agent, like, it is just utterly confounded with noise. And, And, you know, the thing I've been saying as I've been going through a lot of this data is, like, it's clear to me now more than ever that, like, not all words are created equal. Right? Like, in context, you know, missing a yes or no could be everything or it could be nothing. And depending on the turn length, and that could be 50% word error rate or it could be 1% word error rate. And when now that we have these models that are able to reason around ASR mis transcriptions and reason in context, and especially now that we're have we're starting to deal with audio models that, like, take prompting in and situational context in addition to the audio itself. Like, just going on, like, raw, how many words did we get right, like, is is a pretty bad proxy for all the stuff that we actually care about.
[00:03:41] Nikola Mrkšić: 100%. I think that I remember you know, and datasets evolve and change and gradually become more difficult. But, you know, we went from kinda, like, talking about a 5%, you know, state of the art word error rate to now you know? Recently, I was speaking with a CIO of a company, and he's like and you were like a 2% like everyone else. And I'm like, no one is at 2% for anything that really objectively matters in a complex conversation. And I think kinda kinda like, you know, my my favorite, like, peasant logic thing to break this is, like, to anyone who has Alexa or Google Assistant. It's like, does it work 98% of the time? And it's like, I think it's really far from that. But it somehow managed to permeate the consciousness of people where it's like, oh, yeah. It's all but done. Voice is solved. And I don't know. I guess we'll dig into your data and take a look at why it's not. But yeah.
[00:04:31] Oliver Shoulson: Totally. Okay. Well so I'm happy to pull up just some basic stats of, like, what we what we've done here. You know, I ran this analysis over the course of about a a month on on some of our really high volume production calls, analyzed over a 100,000 individual user turns. What we use kind of as the gold standard for word error rate here was a consensus between two different speech recognition vendors, um, or I should say audio audio LLM vendors, um, which were able to take the entire conversation and all of the context, and that definitely perform in a vacuum with the best word error rate compared to any ASR. They're, of course, a bit slower, and they don't produce, like, continuous transcription, so they're a bit harder to use in production. Um, but they they perform really well. And so when when we retro retroactively do analysis like this, we can use that kind of consensus between these two vendors as, like, a really good confident gold standard to compare against. And so this was over the course of around 25,000 calls. And what we did was when we found turns that were that differed in production from the gold standard that the consensus produced, we had an LLM judge basically classify those mistranscriptions, you might call them, as either benign. So these are, like, little formatting errors, kind of filler words, stuff that doesn't ultimately change the meaning. Semantic, which is changes that alter the meaning of the user input, but that could be reasoned around in context or that you could ask clarification about. And then critical, and these are specifically, like, what we would call entity flips. So things that change the value of a slot that the user is providing in such a way that obfuscates or changes the value in a critical way. So this is something like if you're collecting a phone number and you mishear a digit. That's a critical error because you need to get that a 100% right, and it's not necessarily obvious in context to that you misheard something. The the agent has no reason to believe it misheard something if it heard, you know, for some reason, three instead of eight. But so that would be something critical. And I can pull up for you just like a little to show you kind of the proportions that we're dealing with of these errors of the of the missed turns. So up to forty two percent so so around twenty twenty eight percent of missed transcriptions were deemed benign. Up to potentially forty two, some of them weren't judged for a variety of reasons. We have a relatively small proportion of just of these kind of meaning changing semantic errors. So this is when, like, content words get flipped, but in context, they can often be reasoned around or it's clear the what the clarification question should be. And then about half of the errors we saw are these entity flipping errors. And what this tells me actually is that the the proportion that are benign is a huge amount of obfuscating and confounding noise, basically, in the word error rate signal. I'll stop there for a second to, like, hear like, I'm sure you have thoughts about this, but, like, that's my first takeaway from this. No.
[00:07:21] Nikola Mrkšić: No. I mean, it's like I think it's, like, the empty calories in this half where, you know, like, anything people need to understand, it's not like a word is missed. Like, yeah, not all words are equal, but in particular, you know, articles, like whether you pick up and or the or whether someone might have said an or not and whether that's in or not is completely material. It it's, like, semantically empty. I think, you know, beyond that, like, plural or singular might matter for the meaning, but, again, like, a based system is probably gonna make it through that just fine. Right? Similarly, you know, like, a not omitted can be really, really devastating and creating, like, an antonym and opposite meaning of of a thing. So between all that, like, it it can be really hard to, like, decipher, like, you know, which ones really matter. And it's very easy to have even an you know, something with a much higher word error rate performing much worse. And it's plagued evals on dialogue systems forever, and that you improve something, you measure on this, like, intermediate task, you see a higher number, you roll it into production, and God forbid, you made changes downstream as well. At that point, you're completely lost for, like, where did you make progress? Where did you degrade? And it's it's tough. Right? And the more we integrate these components, the the harder it gets still. Right?
[00:08:36] Oliver Shoulson: Yeah. And it's not even like we can identify, like, particular classes of words that are semantically like like, in you know, we have this basic, like, distinction in linguistics between content words and function words where you have, like, your words that sort of mark syntactic information, like, which are your function words, so things like articles and and modal verbs and stuff like that. And then you have your content words, which actually sort of introduce the mental content into the sentence. But it's not even like you can necessarily just classify sort of an inventory of content and function words and know which ones are matter and which ones aren't. Like, it's obviously all contextual. So if the question is, you know, are you calling to check the status of an order, and the user just says yes, and you mishear the yes, like, that's a disastrous mishearing. Like, you you missed the whole content of what they said. But if they say, yes. I am, and you miss the yes, but you hear the I am, that's basically completely benign. Like, the like, you'll proceed and and interpret that exactly the same as if they had just said yes. Totally fine. It's like Yeah.
[00:09:37] Nikola Mrkšić: And I mean, you know, like, the entire rabbit hole, but, like, in different languages, it also, like, behaves very, very differently. Right? I mean, like, the morphology of Slavic languages in the cases, for instance, means that word error rates tend to be quite a bit higher.
[00:09:49] Oliver Shoulson: For all
[00:09:50] Nikola Mrkšić: intents and purposes and understanding, it doesn't really matter. In fact, there are entire groups of, say, I don't know, Serbian population that get cases wrong or use fewer than, like, what their language does. And, you know, like, it doesn't really matter. Like, not for this.
[00:10:04] Oliver Shoulson: That's so interesting because, like, the stem is it gets, like, the stem right, but it, like, it's yeah. It's, like, inflected wrong, and so that's counted as a word error. Yeah. Oh, that's so interesting.
[00:10:12] Nikola Mrkšić: And, you know, I think it's almost like you would then look at character level error rates, but I feel like it would probably have the exact same empty calories and maybe more actually than than the word error rate.
[00:10:22] Oliver Shoulson: Yeah. You know, one of the things that we were looking to sort of correlate or not correlate word error rate with is our proprietary evaluation, which is what we call our PolyScore, which gets run on all of our production calls. And so I thought I'd take a second to just sort of talk a little bit about PolyScore and what it consists of. So, actually, maybe I feel like you're maybe better suited to speak to, like, a little of the history of PolyScore and sort of what how that came about and how we Yeah. I'm I'm happy to.
[00:10:49] Nikola Mrkšić: I mean, look. I mean, I think that stuff like predict predictive, like, MPS scores and CSATs and just like a measure of, like, is it a good call or not has been something that's always needed. And, you know, whether you trust the vendor in showing you whether this is good or not is one piece. The other piece where it's just objectively very useful, 21 launching, the system is you wanna see the great calls and you wanna see the terrible calls. Right? So I think that, you know, as a measure of that, it's something that's been honed for years and, you know, continues to to be honed. It really composes it's composed of a measure of the quality of the conversation. Like, you know, is the fluidity, is the UX, is it a natural human sounding conversation? Are we interrupting each other? Or if they're interrupting, are we stopping? And very, very many things that have to do with everything from, you know, the quality of, like, the audio piece to kind of just, like, the fluidity of the back and forth. And then maybe most importantly, like, task success. Like, did the AI agent actually do what the AI agent is supposed to do in that particular, uh, conversation? And, you know, sometimes the caller is not happy with that. If they want a refund and you're not able to give them a refund, you know, like, the score may end up being high because the system behaved and did the right thing. But, yeah, it's pretty nuanced. And, you know, again, kind of like the LLM judges from the other one, it's it's not perfect, but it's pretty good.
[00:12:11] Oliver Shoulson: Yeah. And it's really nice that it it returns this kind of granular, like, constituent broken down score because we actually get to see some, like, interesting ways in which the number of meaning changing errors that occur over the course of the call affect conversation quality. So, obviously, like, the need for clarification and repeating of the same questions, but actually don't affect task success nearly as much as you might expect, which is kind of interesting. So, you know, the customer tends to pay in friction, but in terms of, like, the actual effect it has on getting where they're trying to go or even call outcomes like escalations and handoffs, you're actually it's kind of amazing how effective the model is at, like, uh, retaining that containment even at the expense of a little bit of friction of reasking those questions. Um, and just to sort of hammer in something I was saying before about this empty calories. I like that metaphor you're using. You know, benign only word error rate, like, as you might expect, varies widely and, like, predicts nothing on the poly score side. Like, we can we can account for maybe up to point 5% of variance in Poly score via the benign word error rate, which, first of all, goes to kind of validate the the LLM judge that was determining the severity of errors, but also tells you exactly what you'd expect, which is that, like, word error rate is confounded by all of this noise. And if you isolate the benign stuff, like, it's just it's literally just noise.
[00:13:35] Nikola Mrkšić: Yeah. Yeah. I mean, like, the the zero correlation is almost like it looks it almost looks rigged, Oliver.
[00:13:43] Oliver Shoulson: Yeah. The task success, zero core it does it does almost look rigged. I was kind of amazed by that. Like, you can actually see that even more. It's amazing how how the task success stays super flat even as we see the compounding effect of meaning changing errors. And this is sort of the flip side to word error rate, which is where we see these actually, I think, really valuable correlations, which is how many semantic or critical errors happen over the course of the call. And you see this really clear monotonic relationship between the number of these meaning changing errors and conversation quality on the poly score, for instance, which is this left graph over here, or handoff rate, which you can see in this right right graph, which crawls up with each meaning changing error that happens over the course of the call. That little dotted line you see on the top of that left graph over there is the task success. So, again, like, amazingly flat the entire time. And, really, it's just the user is paying in additional turns and clarification. But I like, my sort of rose colored glasses take on that is, like, that's an amazing vindication of the capacity of LLMs to recover from these mishearings. Because even if it requires a question to be asked more than once or a follow-up question to be asked, the user is still getting the task accomplished that they want to. They're just paying in a little bit of friction. And I think, like, that's something that, you know, back in the deterministic days, that, like, was just a complete nonstarter. You could not ask targeted follow ups. You could not repeat questions in a way to actually glean the clarification that you wanted.
[00:15:14] Nikola Mrkšić: No. Totally. I think, like, you know, maybe just to give, like, some contextual examples, but, you know, if you're being asked for, I don't know, your British, like, car license plates and you said, LO23EWN, and then, like, it repeated it and, you know, got the last character on. And then it would be like, no. No. No. It's like, NForNovember. Right? At that point, you're like, okay. Like, the last one? Yeah. Good. Or, like, repeat the last three. System repeats. It's that, not that. Right? And that's, like, very much how humans would talk to each other. And it's like Exactly. And it's just really you're right. Like, how much and, again, this has to be implemented. Right? It doesn't always work by default. Right? It can still be frustrating. And it is frustrating. We see it from the core core quality score. But I think, like, it's interesting in diagram how you have, like, the steepness of the core quality falling is higher. Right? Because I guess the quality is decreasing, but the functional capability of the system is there as long as the human stays in the call, which again is something that I guess we would have to kinda, like, derive from the data. Right?
[00:16:13] Oliver Shoulson: Yeah. Though, I mean, I I guess another thing that I was sort of pleasantly surprised by is to see how while we do have absolutely that slope downward over those first two errors, like, it's not as steep as you might think it would be. Like, you might think that people, when they get one critical or semantic error and it's clear that they've been misunderstood, they're like, screw this. Let me talk to a human. And so and on on neither round on neither graph, do you see this incredibly steep either degradation in call quality, meaning that, like, we're not able to recover from those errors, or an extremely steep increase in handoffs, which you might expect to see, which also points to this capacity for the model to retain engagement even after, you know, PolyScore, which detracts for repeated or clarified questions, like, is claiming that the call quality has degraded. And so only until up until that, like, third semantic or critical error, which is where we see that real cliff in conversation quality, You know? The first couple are not free, but, like, not as disastrous as you might expect them to be.
[00:17:20] Nikola Mrkšić: Yeah. So so wait. Like, the three plus on the handoffs going to 68, what is that one about?
[00:17:25] Oliver Shoulson: That one's you know, our Wilson band there, our 95% confidence interval is quite wide, so I don't wanna make any claims about that. You can see those little whiskers, but I think we just don't have enough data there. Because and and, again, this is also, like I was happy to see our cohort of three plus critical or semantic errors in in in our entire database of over a 100,000 turns is in the, like, tens of calls. So, you know, we have very few calls have three plus meaning changing turns. And the other thing to emphasize here, again, is, like, you know, I wouldn't be I wouldn't be very scientific if I didn't, like, whip out the correlation is not causation because other thing, like, you need to keep in mind is that things that affect conversation quality and handoff rate also affect things like word error rate and, like, critical error. So, like, if, like, the user is in a very noisy environment, like, that's this external factor that's going to affect all of this stuff. And so you can't necessarily say, oh, the ASR was bad, so it led to this compounding effect of errors. You can say, you know, the user is driving 70 miles an hour down the highway blasting music, and that's not gonna lead to a good experience.
[00:18:39] Nikola Mrkšić: Yeah. And, I mean, also, like, it's it's very complex if. Right? Because, like, to those conversations that get to have I guess, like, the caveat, just knowing what I know about us would be, like, we don't tend to proactively not hand off often if we get three, like, really bad turns. Right? So in places where they get to accumulate, especially if there's very few of them, those are gonna be those very, very long calls. And I think just I mean, totally, you know, you kinda, like the longer you stay in gameplay, maybe if you start the game with three lives, you know, like, it refills. Like, for every minute that you've stayed in the conversation, the sunken cost of rage quitting and demanding to be handed off. You don't know if the human can seamlessly continue the conversation. We have many clients where the flow has been implemented to kinda, like, persist and fill out the CRM partially, and it's great when it does. But for the most part and industry wide, it's not usually implemented well. So you're kinda used to just, like, you get out of that process, and then it's like, cool. So what was your name? What was your, like, Social Security number? And you're, oh, my god. Right? Like, that's the dot that CSAT is zero. Right? So I think, like, the yeah. Like, if you if you manage to last that long, then people are pretty committed. So it makes sense that the handoff would would fall because, you know, you're about to achieve something. You will persist even if your last name is Marchesic, and you really have to, you know, somehow convey that to an automated thing. What was the word error rate in that aggregate analysis?
[00:20:01] Oliver Shoulson: Um, our headline word error rate for English was 6.4.
[00:20:05] Nikola Mrkšić: Yeah.
[00:20:05] Oliver Shoulson: And for Spanish, which we had a smaller cohort of, was 8.8. Yeah. Which gives us a total of 6.7 weighted.
[00:20:12] Nikola Mrkšić: Yeah. I mean, you know, when you look at, like, how people benchmark against, you know, like, datasets, they're talking about 2%, 3%. And I think, you know, it's completely amidst the fact that, like, there's a different low pass filter on the phone and with VoIP that, you know, you might get several of them hitting the data and then just, like, lossy packets and stuff. So I really almost think that we should start producing a dataset of those. And there are some that are. I remember, like, evaluating we we used to collaborate with, uh, a bunch of kinda, like, carmakers for their voice agents back back at Cambridge. And those in car word error rates would literally go around Cambridge in a car and record samples there. It would be between 2040%. Now that's, like, car music, everything, and, like, an era of different models.
[00:21:03] Oliver Shoulson: But that's where people are calling customer service from. So you know? Sure.
[00:21:08] Nikola Mrkšić: It's very, very
[00:21:09] Oliver Shoulson: It's not a
[00:21:10] Nikola Mrkšić: yeah. Anything else that you kinda, like, think in this one is worth kinda, like, maybe sharing with the audience?
[00:21:14] Oliver Shoulson: Yeah. I mean, I guess the the last thing I wanted to share, which, again, I think is kind of intuitive if you think about it, is the difference between, like, in terms of handoffs versus quality, whether like, the different effects that semantic errors have versus critical errors. And, you know, what's kind of interesting is that, like, we punish the semantic errors on quality a lot more, but they don't move the handoff needle very much. And that makes a lot of sense when you consider that, like, semantic errors are easy to sort of disambiguate around. And so if we're punishing them on the quality score because we're reasking a question or we're asking a clarifying question, but they're not leading but they have basically no effect on handoff rate. Whereas, interestingly, you know, the critical errors have half of the quality degradation that semantic errors do, and that's because critical errors are getting escalated quickly, which is what you'd want to see. Like, when the model, like, misses an entity and is and fails to look up an account or fails to retrieve an account or validate the user or something, and something has gone really wrong to, like, obstruct the path of the flow, like, we want to escalate to a an agent quickly and in a way that does not degrade the quality. So you can see this, like, minus point two seven effect on quality that critical errors have versus the minus point five five that semantic errors have. And, again, that has a lot to do with how PolyScore scores quality and how it punishes Yeah. Clarification and disambiguation. But, like, again, I was happy to see that there's not this huge quality degradation around critical errors, which shows that we are escalating appropriately at the right times.
[00:22:42] Nikola Mrkšić: Yeah. Yeah. Yeah. Yeah. Yeah. This is a bit of a chicken and egg, but, like, I don't know if we can claim good signs, but it's it's really interesting. I mean, it's also like the the whole, like, you know, confirming stuff and, like, you know, how you're shipping out the quality if you're, you know, just kinda like the best call is, what's your name? What's your date of birth? Ideally, I'm an AI oracle. I get everything right, and we don't go through the awkward, like, s h o u l s o n. Right? Whatever. I mean, that's just, like, really not there not there yet in terms of good design. But if you correct that, you know, you're doing it with a human just the same as you are with a machine.
[00:23:20] Oliver Shoulson: Right. I mean, that's why that's why I kept trying to emphasize that, like, a lot of this, not an artifact, but a a result of the way that PolyScore works, which is that it tends to punish that kind of, like, sort of the opposite of what you're saying, which is where we can't we don't just, like, ask and answer to every single question. So, you know, we should interpret conversation quality through that lens and understand that a lot of these cases where we're seeing reduced quality, they're still navigating the conversation in a very natural and human like way. And so that's why it's helpful to know sort of what that score consists of.
[00:23:53] Nikola Mrkšić: Yeah. Just so maybe, like, if we break out a bit from just, you know, like, the analysis itself. You know, if you could have, like, one wish of speech recognition, like, one thing that it could do better to make dialogue systems a whole lot better, Like, what is the thing that is, like, most annoying in practice?
[00:24:09] Oliver Shoulson: I mean, I think that one thing that we encounter a ton in practice is the difference between in performance between models that perform really well on short turns and models that perform well on longer turns. And because that's so unpredictable in the conversation, like, you don't know necessary like, you can you can sort of predict if you're asking a yes or no question. You can imagine that the utterances the response is probably gonna be short, but you don't know. And it's frustrating to have to configure different models on every single turn based on whether you're expecting a shorter or longer response, whether you're expecting a yes, no versus, like, kind of numeric or more value driven response. So that's super annoying in practice. Yeah. And then I think also just like like, the reason that these audio models perform so well and we can treat them as the gold standard probably has a lot to do with the way that they are being prompted in addition to just returning sort of blind transcriptions. Um, like, they get to see the conversation context, they get text context, and they evaluate and they understand the audio tokens in the context of all that text. And so I think, like, that's really the direction that this has to go, probably.
[00:25:12] Nikola Mrkšić: 100%. I mean, maybe, like, to to take a quick digression that is very, like, topical. Have you seen the new, like, GPT Live?
[00:25:20] Oliver Shoulson: No. Oh oh, I saw I saw their, like, little promo video for it.
[00:25:22] Nikola Mrkšić: Yeah. Yeah. Yeah. I find it really interesting because, you know, I think that talking about short terms, the quick yeses and noes, I think, you know, it's not as much our world and that we're really focused on, like, real conversations. They drive, like, business sense and and customer needs, but they have this one where, like, let's say that you're talking and I'm going, yes. Yes. Yes. No. Do the it's like, it's almost like artificial. It's close that someone has just, like, fixated on the problem and really wanted to show that, like, full duplex. Like, I'm speaking while listening and kinda, like, inputs are coming in and out, which we as humans do. And when we do, we just tend to be annoying. Right? I think I was curious about how you felt about that piece because I think, like, you know, like, Rave and Omni and our, you know, internal voracious appetite for it, amazement at, like, everything that it can do both in terms of understanding, using all that context to just, you know, fix problems that we never could fix before is one thing. But how do you feel about the full duplex piece? Do you think it's even good as a feature for a conversational system?
[00:26:20] Oliver Shoulson: Yeah. I do. Like, I think I can't count the number of times I've seen, like, a delay in a barge in taking effect such that we only hear the tail end of what someone said and then totally misinterpret it. Like, I think I think just in terms of at least, like, responsiveness to barge in and interruption, it's necessary to have like, maybe that's not full duplex. Like, we just need to know whether they're speaking or not, not necessarily what they're saying. But yeah. Yeah. I I don't know. I would wanna think about it more, but I feel I have to feel like it's necessary because it's such a key part, as you said, of, like, the way humans navigate conversation. Yeah. I mean, I feel like,
[00:26:54] Nikola Mrkšić: you know, they've done one UX thing, you know, a grave, grave, grave crime there, which is because it's built on ChargeGPT, which is used to, you know, sociopathically, uh, inducing more more more interaction from you. Right? It is built and designed for you to interrupt it rather than to have a natural flow of back and forth conversation. Right? Mhmm. Because of that, I kinda feel like they've taken for granted that that should be the interaction mode between two parties. Right? That I just speak and that you will interrupt me, and that's why we have the whole fascination with barging and stuff. Although barging matters. Right? If I'm asking a question, you go, yep, and I don't hear you, like, that's obviously bad. Right? But I feel like now it's almost like turning into a separate, you know, like, academic discipline of just like and look. Like, this model is, like, taking inputs, producing outputs. Like, we as humans, if we're talking and if you try to, like, slot something in, like, I don't really, for the most part, want want you to do it. Right? Like, sure. If it's a critical piece of information, you should interrupt me. But I think, like, for the most part, we see people react very weirdly when they get interrupted by AI because they just feel like they're not being heard. Right?
[00:28:06] Oliver Shoulson: Yeah. Yeah. No. And I I I see what you're saying about how it feels like it feels like they're excusing themselves for the way that the model is just going to monologue at you and, like, not actually design information delivery in a way that, like, makes it makes you not want to interrupt the model as much. And just saying basically, like, we're not concerned with that problem. Like, we'll just let you interrupt. Yep. And as a conversation designer, that does hurt.
[00:28:32] Nikola Mrkšić: Mhmm. Fair. Well, on that note, I think we I think we're out of time for today. But, Oliver, always a pleasure. And Always a pleasure. Please, you know, like, share, subscribe, and we'll see you on the next one.