Skip to main content
    Cloudax
    Partners
    Sign inLet's Talk
    News > Industry Insight

    The three-second rule: why latency is the new accuracy in voice AI

    Cloudax

    10 April 2026

    8 min read

    Share

    All news

    The industry obsesses over accuracy benchmarks. In production voice, the number that decides whether a caller trusts your AI is measured in milliseconds.

    The three thresholds that decide trust on the line

    Accuracy is scored on a benchmark. Presence is scored in milliseconds. Cross the wrong thresholds and it doesn't matter how right the answer is. The caller has already formed a judgement.

    200 ms

    Average gap between human speaker turns across languages

    < 3 s

    Hard ceiling for end-to-end latency before callers stop trusting the AI

    ~ 700 ms

    The silence threshold where listeners read absence, not thought

    How humans actually turn-take

    Callers are not benchmarking your word error rate. They're benchmarking your presence, and they do it in the first second.

    Conversation analysis across languages puts the average gap between human speaker turns at around200 milliseconds. The 90th percentile sits at roughly 500 ms. Anything longer than about 700 ms and the listener interprets the silence as one of three things: confusion, disagreement, or absence.

    A 900 ms wait for the first response frames every subsequent sentence as uncertain, no matter how correct it turns out to be. A sub-500 ms turn signals competence long before the content arrives.

    “Accuracy you can't hear doesn't exist. If the caller has already hung up, your 99% F1 score is a rounding error.”

    Latency is a system problem, not a model problem

    End-to-end voice latency is a stack of bills that nobody wants to pay in full. Add them naïvely and you cross the three-second barrier before breakfast.

    • Capture & codec: 40–150 ms.Audio off the PSTN, through the codec, into the stack.
    • Speech recognition: 150–350 ms.Streaming partials, plus finalisation overhead.
    • Retrieval: 50–400 ms.Vector store, BM25 or knowledge-graph hops.
    • LLM reasoning: 400–1,200 ms.First useful token, depending on model and prompt.
    • Text-to-speech: 200–600 ms.First audio on the wire, including prosody planning.
    • Network egress: 30–80 ms.Back through the media path to the caller.
    The compounding cost

    Cross three seconds often enough and the brand is disposable.

    Every caller after that learns (correctly) that your AI is slower than a human. The response to every subsequent turn is pre-discounted, no matter how good the content is.

    Accuracy and latency are the same dial

    Production voice platforms that hit sub-1-second end-to-end don't split ownership between two teams. They treat speed and correctness as one metric. The patterns that actually work look similar across every high-performing deployment we've inspected:

    • Streaming everywhere.Streaming ASR partials into streaming retrieval into streaming generation into streaming TTS. Any blocking step eats every downstream budget.
    • Speculative retrieval.Fire the first retrieval on the ASR partial, not the final transcript. In our traces, this reclaims 120–180 ms with no accuracy cost when paired with a confidence gate.
    • Small-first LLM routing.Route 70–80% of turns to a fast, tuned smaller model. Escalate to the flagship only on low-confidence paths. The caller hears the small model almost always.
    • Filler tokens."Let me check that" is not a stall; it's a latency-hiding primitive. Used well, it buys 400–700 ms of invisible compute while maintaining presence.

    The single rule that gates every other feature

    Our working rule for any Voice AI roadmap is simple. If P90 end-to-end latency (first-utterance-end to first-audio-out) is not under three seconds, nothing else you ship will matter.

    • Sub-3 s: product.Turns a demo into something you can put in front of real callers without apologies.
    • Sub-1.5 s: conversation.Turn-taking starts to feel natural. Containment rate and CSAT compound.
    • Sub-1 s: human.Indistinguishable from a well-trained agent on presence. Operational gains unlock.

    If you can't produce these four numbers this week…

    …your Voice AI platform is not yet measured on the dimension your callers are actually scoring it against. Don't start with accuracy. Start here, per turn, per call, over a random sample of 1,000 calls.

    1. Barge-in latency.How fast you detect that the caller has started speaking. Sets the ceiling on perceived responsiveness.
    2. Time to first token.ASR finalisation to the LLM's first useful token. The single most controllable budget in the stack.
    3. Time to first audio.LLM first token to audible speech on the wire. Dominated by TTS warm-up and prosody planning.
    4. End-to-end P90.Utterance end to first audio out. The only number that matters to the caller. Own this and the rest follows.

    Build for presence

    Voice AI is won and lost in the first second. Hold P90 under three, engineer for sub-one, and the rest of the roadmap compounds. Skip this and the rest of the roadmap is reshipping a demo.

    “Cloudax's voice stack is engineered latency-first. We routinely hold P90 end-to-end under 0.9 seconds on production inbound traffic, across accents, codecs and carriers.”

    Engineer for the first second.

    Cloudax's voice stack is engineered latency-first. Talk to our team about how to instrument P90 end-to-end on your current deployment and what it takes to close the gap to sub-one.

    Talk to our teamSee it run live

    Keep reading

    Policy · 22 JUL 2026

    The EU AI Act hits voice: what August 2026 means for conversational AI

    The Act's high-risk obligations apply from 2 August 2026. Most voice AI in production is limited risk, but emergency triage, recruitment, credit, insurance, education and public services are not. What actually applies to a voice agent, where the line falls by industry, and the role trap that turns a buyer into a provider.

    Standards · 28 MAY 2026

    Cloudax publishes OHP: the Open Handoff Protocol for voice AI

    OHP: the Open Handoff Protocol. A vendor-neutral specification for handing off live voice conversations between AI agents, IVRs, contact centres and humans, with verified identity, intent, consent and media continuity across competing platforms.

    Partnership · 18 MAY 2026

    Cloudax partners with Gamma to bring voice AI to enterprise UC

    Cloudax joined Gamma's main stage at GX Summit 2026 in Westminster as part of Gamma's AI strategy vision, and ran a live voice AI deployment on the day. Cloudax integrates with existing Gamma systems with no rip-and-replace and no lengthy migrations.

    See how Cloudax's more human AI can transform your business

    No commitment, free consultation included

    Let's Talk
    CloudaxCyber Essentials certifiedCyber Essentials Plus certifiedISO 27001 certified

    Find us

    167-169 Great Portland St.
    London W1W 5PF

    1 Hardman Square
    Manchester M3 3EB

    +44 333 011 1190
    [email protected]

    Solutions

    • Connect
    • Inbound
    • Outbound
    • Chat

    Use cases

    • Property
    • Legal
    • Finance
    • BPOs
    • Utilities & Energy
    • Automotive
    • Travel & Hospitality
    • E-commerce & Retail

    Company

    • About Us
    • Careers
    • Case Studies
    • News
    • Press
    • Contact
    • Security
    • Brand

    Legal

    • Privacy Policy
    • Terms & Conditions
    • Cookies Policy
    • Accessibility Statement
    • Modern Slavery Statement
    • Anti-Bribery Statement
    • Code of Conduct
    • Environmental & Sustainability Policy
    © 2026 Cloudax Ltd. All rights reserved. Cloudax® is a registered trade mark.