Vendors quote accuracy rates in the high nineties. In regulated industries at scale, the cost of the remaining percentage is where the economics of Voice AI are actually decided.
Three numbers that change the procurement conversation
Vendors quote accuracy in the high nineties. In regulated industries at scale, the cost of the remaining percentage is not small: it is where the economics of Voice AI are actually decided.
At 100k calls/month and 98% accuracy, that's ~66 confidently wrong answers every day. '95% accurate' at a million calls is 50,000 bad experiences a year.
At plausible mid-range cost inputs (R=£25, L=£180, B=£40, F=£90), a 100k-call operator carries ~£670k of monthly hallucination cost that never appears in the AI budget.
Moving from vector-only RAG to hybrid retrieval with a validation agent typically cuts modelled hallucination cost by this much, far more than equivalent benchmark gains.
The number nobody wants to do
Take a contact centre handling 100,000 calls per month. Assume a best-in-class Voice AI with a genuine 98% accuracy. Now do the maths.
At that rate, the system will produce2,000 factually wrong responses per month, around 66 per day, close to three per hour. That's not a theoretical number. Those are customers who received an answer the systemconfidently believedwas correct, and wasn't.
In a retail context, that's a refund conversation. In a utility context, a complaint. In healthcare or financial services, sometimes a regulatory event.
“"95% accurate" at a million calls a year is fifty thousand wrong answers. The industry's favourite number is actively misleading.”
One bad answer is four layered costs
Hallucination cost isn't a single line. It is a layered sum of four very different types of damage, each with its own half-life.
- Immediate remediation: £8–£60.Agent handoff, customer re-contact, refund or goodwill. £200+ per incident in regulated B2B contexts.
- Customer lifetime value: +3–9%.A single wrong answer on a live line materially changes churn probability for the customer involved. Averages lift churn risk by 3–9% annually.
- Reputational propagation: ~1 in 20.Roughly 1 in 20 negative voice experiences ends up publicly visible. Those compound across the cohort, not the individual.
- Regulatory exposure.A single auditable misstatement on a recorded line can trigger remediation measured in the five- or six-figure range, before any fine is levied.
A useful equation: C = P × (R + L + B + F)
WhereP is probability of incident, R is direct remediation, L is lifetime value at risk, B is brand propagation, Fis regulatory exposure.
Plug in plausible mid-range numbers for a regulated B2C operator (R=£25, L=£180, B=£40, F=£90), and expected cost per hallucination is around£335. At 2,000 hallucinations a month, that's roughly £670,000 of shadow cost that doesn't appear on any AI vendor's P&L projection.
Governance isn't a cost centre: it's the highest-leverage lever
Architectures that combine hybrid retrieval, a validation agent and confidence gating don't eliminate hallucinations. They do three things that matter more.
- Reduce the rate.Hybrid retrieval and validation typically bring confident-wrong rates from 2–4% to well under 1% for grounded domain answers.
- Convert failures to deferrals.Unknown failures become known deferrals. The system hands off with confidence, rather than asserting with false confidence.
- Produce an audit trail.The remediation cost of the ones that slip through is lower, because the incident path is instrumented from day one.
Don't let the vendor pick the scoreboard
If you're piloting Voice AI in a regulated industry, track these three numbers monthly. They tell you more than any benchmark accuracy score ever will.
- Confident-wrong rate.Percentage of responses where the system's confidence was high and the answer was factually wrong. The dangerous number, since the customer will not realise to question it.
- Deferral rate.Percentage of calls where the system gracefully escalated or asked a clarifying question. Rising deferral rate is often a healthy signal, not a weakness.
- Source-attribution rate.Percentage of responses traceable to a specific document, version and policy line. This is the number regulators will ask you for first.
Price the remaining 2%, deliberately
The enterprises who win with Voice AI in the next five years will not be the ones with the highest benchmark scores. They will be the ones who priced hallucinations honestly, and architected for the remaining 2% as deliberately as for the confident 98%.
“Cloudax's production stack is built around the confident-wrong rate, not the headline accuracy number. The full architecture is in our Beyond RAG whitepaper.”
Price the whole picture.
Cloudax's stack is built around the confident-wrong rate, not the benchmark number. Read the full architecture framework in Beyond RAG or speak to our team about your deployment's hallucination posture.




