New: Voice AI Orchestration Benchmarks — Retell, Vapi, Pipecat, LiveKit & more

Call Deflection: What It Is & How to Increase It (2026)

Atul Jain
Written bySEP 18, 202619 MIN READ
Atul JaininExpert verified
Founding Engineer, CekuraIIT Kanpur

Has stress-tested 5M+ voice agent minutes at Cekura.

Why Trust Cekura on Voice AI Evals

  • Built by engineers from Google, Apple, Microsoft. Backed by Y Combinator.
  • 60K+ voice AI calls evaluated daily.
  • Native integration for every major voice AI stack: LiveKit, Pipecat, Vapi, Retell, ElevenLabs, Telnyx.

Call deflection counts customer contacts that never reached an agent, which means it’s a figure built on estimates of calls that didn't happen. That makes the rate easy to inflate on paper and hard to defend under audit, and HMRC's audited call data and our own repeat-run benchmarks show why.

Here is the formula, where your inbound volume originates, and the factors that move the rate without pushing callers away.

TL;DR: Call deflection rate

  • Call deflection rate is the share of customer contacts resolved without a live agent across your whole support portfolio. Every version of the formula estimates a call that did not happen, so the figure is only as sound as the intent data behind it.
  • Avoidable contact you generate yourself is the largest lever. Process delays, progress chasing, and customer errors drive most inbound volume in audited public-sector data, and a call you never trigger needs no deflecting at all.
  • Deflection shifts your call mix before it cuts your cost. Simple requests leave the queue first, average handle time on what remains climbs, and the savings land smaller than the percentage suggests.

What is call deflection?

Call deflection is the practice of resolving a customer request without a live phone agent. It covers three moves. You remove the reason for the call, you finish the request in another channel, or you complete it on the line with automation.

The unit is a contact, and the scope is your entire support portfolio. A contact counts as deflected once the request reaches an outcome and no person handled it.

What counts as a deflected call

Two properties have to hold at once. The request reached an outcome, and no agent was involved. If you drop either one, the contact belongs somewhere else on your dashboard.

Two variants also get counted together while behaving differently. A hard deflect tells the caller they cannot reach an adviser. A soft deflect, meanwhile, names the digital route and lets the caller keep holding.

What happenedVerdictWhy
Caller checks order status in your app and never dialsDeflectedRequest completed with no agent involved
AI voice agent books the appointment and confirms it backDeflectedCompleted on the line, no person needed
Caller hears a message naming your website, then hangs upUnknownNo evidence the request reached an outcome
Caller waits 12 minutes and abandons the queueNot deflectedDemand went unmet and tends to return
Caller reads a help article, then calls anywayNot deflectedSame contact, arriving later
Automated line answers, then transfers to an agentNot deflectedThis one belongs to your containment number

Call deflection compared with containment and self-service success

Deflection counts contacts across your portfolio, while containment counts sessions inside a single automated channel that ended with no transfer.

Self-service success, on the other hand, counts whether the customer found an answer, which is the weakest of the three claims. In HMRC's 2022 customer survey, 74% of people who used its webpages rated them positively. Among people who used both the web and the phone, 84% had called because they couldn't resolve the issue online.

Useful information is still a long way from a resolved issue. ServiceXRG, which developed the multi-factor approach to measuring deflection, draws that line deliberately, because a customer can find an answer, report success, and call you anyway.

Our guide to containment rate covers the single-channel version of this measurement, including how hang-ups inflate it.

How to calculate call deflection rate

You calculate call deflection rate by dividing contacts resolved without an agent by total contacts that could have reached one, then multiplying by 100. Two versions of that arithmetic circulate, and they give different answers on the same month of data.

The single-division call deflection rate formula

Call deflection rate = (contacts resolved without an agent ÷ total contacts) × 100

Say 40,000 people contacted you in March and 12,000 finished with no agent. That gives a call deflection rate of 30%.

The denominator is a log count you can trust, and the numerator is an inference. Your systems record the end of a session, and whether the request was satisfied sits outside that record.

The multi-factor formula from ServiceXRG

The rigorous version multiplies four inputs together, and typically returns a lower rate than the single-division formula on the same data.

Deflection = self-help events by entitled customers × success rate × intent rate × no-further-action rate

Each input does a specific job:

  • Entitled customers narrows the population to people who had a live-support option and passed it up. A visitor with no support entitlement was never going to call you.
  • Success rate captures whether they found the answer they came for.
  • Intent rate captures whether they would have called you had self-service come up empty.
  • No-further-action rate confirms nothing else was needed to close the request.

Three of those four inputs come from asking customers. Deflection is a counterfactual, so it gets measured by survey or experiment, which is the honest cost of the metric.

Why the denominator decides the answer

HMRC runs one of the largest telephone estates in the UK and publishes its call outcomes in full. Of the 36.5 million calls it handled in 2023-24, on data pro-rated from April 2023 to February 2024, 12.3 million ended after a deflection message.

Another 4.9 million callers abandoned the queue, and 16.3 million reached an adviser. A further 3.0 million heard a busy message. Four different outcomes sit inside one inbound volume figure, and only one of them involves a person answering.

HMRC cannot follow customers across channels, so it does not know whether deflected callers resolved anything. Satisfaction among them ran at 28%, against 63% for callers who spoke to an adviser.

A deflection number built on attempts, answered calls, or ended calls will differ by millions on one estate. Write the inclusion rule down, version it, and start a new series whenever it changes.

Holdout measurement for a deflection claim

Treat the rate as an estimate and test it like one. Withhold the deflection offer from a random slice of matching contacts for two weeks, then compare contact volume across the two groups.

Normalize against a volume driver you already track, such as orders, active accounts, or enrolled members. Contacts per 1,000 accounts moves for real reasons. A raw call count moves every time sales has a good quarter.

Where inbound call volume comes from

Deflection strategy usually starts at the queue, which is the last place the demand was created. Three sources account for most of what arrives there.

Digital investment on its own has a poor record of moving these numbers, and the figures below explain why. The demand was never sitting in the channel you are trying to shift it out of.

Avoidable contact you generate yourself

HMRC attributes 72% of the calls it received in 2023-24 to failure demand, meaning its own process delays, customers chasing progress, and customer errors. The National Audit Office recorded the rise, with that proportion climbing from 65% in 2018-19.

That volume needs the upstream process to stop producing the calls in the first place.

Callers who already tried the digital route

In a December 2023 Gartner survey of 5,728 customers, 73% used self-service at some point, and only 14% of issues were fully resolved there. Gartner's February 2024 cost benchmarks put the median assisted contact (phone, chat, or email) at $13.50, against $1.84 for self-service.

Your queue is largely the residue of self-service that came up short. A deflection program aimed at the phone line arrives after the customer already tried the cheaper option once.

Outbound volume outside deflection scope

Deflection only addresses calls the customer places. Across 58.2 million calls in 2024 and 2025, the telephony vendor Natterbox measured 60% outbound, leaving inbound at 40%.

Those are the vendor's own production records, so read the split as directional for your estate. Cut the same ratio on your own switch before you size any deflection target.

The same dataset shows what happens when routing improves, and demand holds. Hunting time (the time callers spend inside routing menus) dropped 54% year over year while total call volume grew 16.1%.

Faster routing moves callers to a person sooner. That is a service win with no deflection in it.

Three routes to a deflected call

Every tactic below reduces to one of three mechanisms. Pick the mechanism first, because each has a different owner and a different payback period.

Removing the reason for the call

Product and process own this route. A repayment that lands on the promised date removes contacts permanently. So does a status page that answers where an order is.

This is the only route that lowers cost per customer in every channel at once. It is also the slowest, since the fix sits in someone else's roadmap.

Resolving the request before the call starts

Self-service and proactive messaging own this route. An outbound SMS confirming a delivery window removes the inbound call that would have chased it.

Proactive contact competes with the call. It deflects without degrading anyone's access to a person.

Resolving the request on the line without a human

This route is what changed in 2026. An AI voice agent answers the phone, authenticates the caller, executes the transaction, and confirms the outcome inside the channel the caller chose.

Deflection here is more about task completion. The caller gets an answer on the phone, and no agent minutes are consumed.

Eight call deflection strategies for 2026

These strategies are ordered by movement per engineering hour, starting with the cheapest wins.

1. Set an annual reduction target for avoidable contact

Report avoidable contact as its own line, split between causes you own and causes the customer owns. Then set a percentage reduction target for your half and review it monthly.

Volume with no target attached grows by default. A deflection percentage can climb for a full year while the absolute number of avoidable calls climbs with it.

2. Publish a status surface for your top progress intents

Progress chasing is the highest-volume avoidable intent in most support portfolios and the easiest one to remove. Give customers a tracked state for every request that takes more than a day, with a timestamp and an expected date.

A tracking page deflects the same contact repeatedly for years after you ship it, which is rare among deflection tactics.

3. Correct misrouting before adding another channel

A misrouted call becomes a transfer, a repeat explanation, and often a second call. Fixing identity resolution and intent capture at the front of the call removes that duplicate volume.

Adding a self-service channel on top of a misrouting problem adds a hop to the same journey.

4. Attach the answer to the deflection offer

A recorded message naming your website is a redirect, and an SMS carrying a deep link to the specific transaction is a handoff.

Send the link, hold the caller's queue position for five minutes, and let them return with one keypress. Deflection that keeps the exit open converts better and costs you nothing in access.

5. Offer a callback instead of a deflection message on your longest queues

A callback removes hold time without removing access to a person, so it survives the objection that kills most deflection programmes. It does not reduce adviser minutes, which means it belongs in your service budget rather than your deflection target.

Reserve it for intents your automated line cannot complete, and measure it on abandonment rather than on deflection rate.

6. Give the automated line write access to the systems callers ask about

Every intent that requires a state change downstream escalates until the agent can perform that change itself.

Scope tool access by intent volume, starting with the three highest-volume transactional requests in your call reasons report.

7. Instrument the journey across channels

You cannot confirm a deflection you cannot follow. Stitch a customer identifier across web, app, chat, and voice, then check whether a deflected contact returned on the same intent inside 72 hours.

Re-contact inside a short window is the strongest disconfirming signal for any deflection claim. Our voice observability guide covers the tracing layer those cuts depend on.

8. Score the automated path on repeat runs before publishing a rate

A single passing test says nothing about whether the same caller with the same request gets the same outcome tomorrow.

Cekura Bench, which we run and publish, gives the shape of that gap on live platforms.

One brief runs across seven voice AI agent configurations and 82 caller scenarios, with three retained runs each (246 calls per provider). Scoring uses pass^3, and a scenario counts as passing only when all three runs pass.

Calls that never connect stay in the denominator for pass^3 and for infrastructure reliability, while task completion is scored only across calls that produced outcome evidence. Each provider chose its own configuration.

Retell completed the caller's task on 93.88% of scored calls, and passed all three runs on 75.61% of scenarios, or 62 of 82. That 18-point spread is the distance between a pilot number and a production number.

Vapi recorded the highest task completion at 97.56% across its 205 scored calls, while 41 of its 246 retained calls never connected. A call that never connects deflects nothing, so it belongs in your abandonment column.

Call deflection rate benchmarks and reference points

Deflection earns its place when it removes contacts your customers never wanted to make.

Done well, it cuts cost per contact, shortens queues for the callers who still need a person, and frees adviser time for the work that needs judgement. The rest of this section is about how much of that shows up in the numbers.

No audited industry benchmark exists for call deflection rate. The ranges in circulation come from self-reported vendor programs, measured under different entitlement rules, denominators, and re-contact windows.

HMRC's "digital containment" measure counts a user as contained if they don't call any of its five main helplines within seven days. On that rule, 97% of app users and 94% of Personal Tax Account users stayed off the phone in 2023.

Three further reference points hold up well enough to reason from:

  • Avoidable contact dominates inbound volume in audited public-sector reporting, which sets the realistic size of your opportunity.
  • Most customers use self-service somewhere in their journey, so your addressable deflection pool is smaller than your call volume implies.
  • Inbound is the minority of contact center traffic in production telephony data, which caps what any deflection program can touch.

Baseline against your own contact rate per 1,000 accounts, cut by intent. An aggregate percentage borrowed from another company's dashboard says nothing about your estate.

Limits of a higher call deflection rate

Past a point, deflection gains come from making the phone harder to use. Two documented effects are worth planning around before you set a target.

Handle time on the calls that remain

Deflection removes your simplest calls first, which raises the average difficulty of everything left in the queue.

HMRC's advisers answered 22% fewer calls in 2022-23 than in 2019-20, and each one took 21% more time to handle. Average handling time went from 11:24 to 13:48 minutes.

Total adviser time on calls fell just 6% across those three years. Model your saving on adviser hours, since a deflection rate landing on two-minute calls returns a fraction of what it implies.

Deflection that degrades access to a person

The UK Public Accounts Committee reviewed the estate and found 36.7 million calls in 2023-24, with only 66.4% of attempts answered.

In the first 11 months of that year it cut off 43,690 callers who had waited 70 minutes, up from 6,875 the year before. No warning was given and no callback was offered.

The Committee found HMRC too willing to let its phone service deteriorate in the hope of pushing people online. You need to combine every deflection target with the guardrail that would expose it.

Deflection targets and the evidence each one needs

Each target below is only as credible as the evidence beside it. Pair them before you publish a number.

TargetEvidence that keeps it honest
Overall call deflection rateContacts per 1,000 accounts, measured against a holdout group
Per-intent deflectionRe-contact rate on the same intent inside 72 hours
Automated line deflectionRepeat-run pass rate across the same scenario set
Deflection message acceptanceCompletion rate in the destination channel
Quarter-over-quarter gainsSatisfaction among callers who were deflected

Which call deflection lever to start with

Start with avoidable contact if: your call reasons report is dominated by status requests, corrections, and chasing. The demand is self-inflicted and removable, and no channel investment competes with deleting it.

Start with proactive messaging if: your volume spikes predictably around fulfillment, billing, or appointment dates. Those spikes are forecastable, which makes them cheap to intercept.

Start with the automated line if: your callers want transactions completed and queue time is the complaint. This route needs the most testing and pays back fastest on high-volume simple intents.

Start with measurement if: you already report a deflection rate and cannot say what the denominator excludes. Every lever above is unmeasurable until that rule is written down.

How Cekura tests a deflected call path

Cekura is an automated QA and observability platform for voice and chat AI agents. Cekura scores each deflected path against the same success condition you report on, so the rate on your dashboard is the rate the tests measured.

The same success condition runs in pre-production simulation and against live traffic. Your deflection definition never drifts between the test suite and the dashboard.

Simulation (pre-production)

  • Automated scenario generation across every branch of a transactional workflow, weighted toward your highest-volume call reasons.
  • Repeat runs with pass^k scoring, where a scenario counts as passing only when every run passes. The simulation suite drives these through your own stack.

Infrastructure

  • Interruption, background noise, and latency scenarios run against your live configuration, plus keypad and menu traversal coverage for IVR paths.

Evaluation (live traffic)

  • Dropoff scoring on production calls, separating requests that reached an outcome from callers who left mid-flow.
  • Production calls that miss their success condition are converted into regression scenarios and gated in CI.

Native integrations work out of the box for Retell, VAPI, ElevenLabs, LiveKit, Pipecat, Bland, and more. You add a testing and monitoring layer on top of the stack you already run.

Cekura holds SOC 2, HIPAA, and GDPR compliance, and the platform covers transcript redaction, role-based access, and audit logging.

Twin Health, a Cekura customer, runs a voice agent as the clinical front door for member onboarding.

Path count is the testing problem in any deflection workflow. VP Engineering Manoj Ananthapadmanabhan puts the scale of it at "thousands of potential conversational paths where a single logic error could result in a failed clinical enrollment."

Where to start with call deflection

Pull your top ten call reasons and mark each one as avoidable, interceptable, or automatable. The avoidable column is your deflection budget, and it is usually the largest of the three.

Then compute your current call deflection rate twice. Once as your dashboard reports it, and once with 72-hour re-contacts and abandoned sessions pulled out.

The difference between those two figures is your real starting position. From there, the work is ordinary engineering, and our guide to performance monitoring covers the wider metric set these numbers sit inside.

Want to know whether your automated line deflects a call or delays it? Book a demo, and we will run your top call reasons through simulation and score each against your own success conditions. You get the repeat-run pass rate behind the headline number.

Frequently asked questions

What is call deflection?

Call deflection is resolving a customer request without a live phone agent. You remove the reason for the call, complete the request in another channel, or complete it on the phone with automation. A contact counts as deflected only once the request reached an outcome.

How do you calculate call deflection rate?

You calculate call deflection rate by dividing contacts resolved without an agent by total contacts, then multiplying by 100. The stricter multi-factor method multiplies self-help events, success rate, intent rate, and no-further-action rate, which returns a lower figure.

What is a good call deflection rate?

No audited benchmark exists for call deflection rate, so a good rate is one that holds up against your own baseline. Measure contacts per 1,000 accounts by intent, then read re-contact rate alongside it. Any figure that climbs while re-contacts also climb was bought.

What is the difference between call deflection and call containment?

Deflection counts contacts across your whole portfolio that never reached a person. Containment counts sessions inside one automated channel that ended without a transfer. Our containment rate guide covers the single-channel metric in full.

Does call deflection reduce contact center costs?

Yes, though by less than the percentage suggests. Deflection removes your shortest calls first, so average handling time on the remaining volume rises and total adviser hours drop more slowly than call counts do. Model the saving in adviser hours.

Can an AI voice agent count as call deflection?

Yes, an AI voice agent counts as call deflection when it completes the caller's request with no human involved. A call the agent answers and then transfers is containment, and a call that never connects is abandonment.

Ready to ship voice
agents fast? 

Book a demo