A containment rate can climb steadily while customer satisfaction slides, and your dashboard will show only the good half.
The metric counts sessions that ended without a human, which is a narrower claim than it sounds. Here is the formula, the benchmarks that come from federal inspector general and GAO reporting as opposed to vendor blogs, and the seven levers that move the number.
TL;DR: Containment rate
- Containment rate is the number of contained sessions divided by the total number of sessions entering the automated channel. It measures the absence of a transfer, which means a session can be contained and unresolved at the same time.
- The most common measurement defect is counting hang-ups as contained. A federal watchdog review of one 90-million-call IVR found exactly that, then concluded the published number was overstated.
- Durable gains come from transactional tools, repairing your weakest intents, and running multi-turn scenarios repeatedly before launch. One clean pre-production simulation predicts very little about production.
What is containment rate?
Containment rate is the percentage of sessions that enter an automated channel and end without a transfer to a human agent. It applies equally to an IVR, a chatbot, or a voice AI agent.
The metric answers the question, “Did the conversation stay inside the automated channel?” Whether the customer got what they came for is a separate measurement, and most reporting folds the two together without saying so.
How to calculate containment rate
The formula is a single division:
Containment rate = (contained sessions ÷ total sessions entering the channel) × 100
Say your chat agent handles 12,000 sessions in a month and 8,400 end with no handoff request. Your containment rate is 70%.
The denominator decides that answer, and every exclusion in it is a documented choice. See the measurement rules below.
Containment rate vs. the metrics it gets confused with
Four metrics tend to get used interchangeably in automation reviews. They measure different things, and swapping one for another changes the entire business case.
| Metric | What it counts | Scope | What it misses |
|---|---|---|---|
| Containment rate | Sessions ending with no human transfer | One automated channel | Whether the issue was resolved |
| Deflection rate | Contacts that never reached a human anywhere | The whole support portfolio | Which channel did the work |
| Resolution rate | Sessions where the stated issue was closed | One channel or the portfolio | Issues closed on a later contact |
| Escalation rate | Sessions handed to a human | One automated channel | Whether the handoff was correct |
| Abandonment rate | Callers who hung up mid-flow | One automated channel | Why they left |
Contact centers have tracked containment, transfer, and abandonment together for decades. The U.S. Postal Service Office of Inspector General names all three as the standard measures of IVR performance.
Containment and escalation only sum to 100% when abandonment is reported separately. If you blend abandonment into containment, the number stops meaning anything at all.
Why containment rate is the easiest agent metric to overstate
Containment is cheap to game because the default definition rewards silence, which means a caller who gives up and hangs up produces the same log signal as a caller who got a perfect answer.
Hang-ups counted as contained sessions
The Postal Service runs one of the largest IVR deployments in the country. Its 1-800-ASK-USPS line handled more than 90 million calls in FY2020, and its own inspector general reviewed the containment methodology in a September 2021 white paper.
Every call not transferred to an agent was counted as contained, and that included calls where the customer hung up before reaching any resolution. Customer-ended calls made up the majority of contained calls.
The OIG's conclusion was blunt. It found the agency was overstating its containment rate, and recommended reporting complete contained calls separately from incomplete abandoned ones.
The survey data alongside it shows what the overstatement conceals. In FY2020, 26% of IVR survey respondents were very dissatisfied, and nearly four in ten reported that their issue went unresolved inside the system. Containment that year was climbing.
Confident wrong answers counted as containment
A caller who receives a wrong answer and accepts it leaves without escalating. Your metric records a win.
Gartner surveyed 5,728 customers in December 2023 and found that only 14% of customer service issues fully resolve in self-service. Even for issues customers themselves called very simple, the figure was 36%.
Set that against the 70% to 90% containment ranges circulating in vendor marketing. The gap between those two numbers is the size of the problem.
The same survey named the mechanism. Forty-five percent of customers who started in self-service said the company did not understand what they were trying to do.
In 43% of cases, they could not find content relevant to their issue. Both scenarios end in a session with no transfer request.
The denominator moves the number
Two engineers can compute containment from the same log table and land ten or more points apart, depending on what each one excludes. Every exclusion is an editorial decision.
Write the inclusion rule down and version it. Any change to the rule starts a new time series, and a containment trend built on a shifting denominator is really a chart of your reporting choices.
Containment rate benchmarks for 2026
There is no audited industry benchmark for containment rate. The USPS OIG says so directly, noting that because IVR systems differ so widely across industries, no universal benchmarks exist for containment, transfer, or abandonment.
The tidy 70% to 90% range quoted everywhere traces back to vendor blogs citing each other. Three figures are documented well enough to reason from.
- Postal Service IVR: Containment moved from 62% of more than 72 million calls in FY2018 to 78% between October 2020 and May 2021, against a stated target of 80%.
- IRS phone lines: Automation answered 41% of filing-season calls in 2026, up from 34% in 2025, on a total volume of about 28 million calls, per the GAO. Average wait time over the same period went from 3 minutes to 8.
- Self-service resolution: 14% of issues fully resolve without a human, on Gartner's customer-reported measure.
Treat all three as calibration points. None of them is a target you should adopt, and the IRS figures in particular show containment and customer experience moving in opposite directions inside one fiscal year.
Containment variation across intents inside one system
The most useful table in the Postal Service report is the per-intent split, covering October 2020 through May 2021.
| Call type | Containment rate |
|---|---|
| General inquiry | 93% |
| Passports | 91% |
| Post office lookup | 90% |
| Package tracking | 72% |
| International tracking | 63% |
| Redelivery | 62% |
| Hold mail request | 60% |
| Change of address | 56% |
That is a 37-point spread across eight common intents inside a single system, on a single platform, in a single reporting window, and specialized request types in the same report sit below 5%.
The same pattern shows up wherever intents are split out. Intents that only relay information contain well. Intents that require the caller to supply a tracking number, an address, or a date contain far worse, because every extra turn is another chance to lose the thread.
Aggregate containment hides all of it. A support organization reporting 78% overall could be running 93% on store lookups and 56% on address changes, and the address change customers are the ones churning.
Setting containment targets per intent
Set the target where the work happens. A single company-wide goal punishes whichever product owner inherited the hardest intents and rewards the one who inherited store hours.
Rank your intents by volume, then by containment, then by revenue exposure. High volume and low containment is the only combination that earns a roadmap slot this quarter.
IVR containment rate, chatbot containment rate, and voice AI agent containment rate
The formula is identical across all three channels. What changes is the session boundary and the reason sessions leak.
IVR containment rate
An IVR containment rate measures menu traversal. The caller navigates a scripted tree using speech or the keypad, and containment records whether that tree resolved the request.
Define the session boundary before you count anything. A session starts when the caller reaches the automated greeting and ends at hang-up or transfer, and repeat calls inside your re-contact window belong to the same session, not a new one
Because the paths are finite, the diagnosis is usually mechanical. A menu node with a high exit rate is a design problem, and you can find it by walking every branch.
Keypad entry deserves its own coverage. DTMF behaves differently from speech input, and it is often the weakest link in an otherwise healthy flow.
Keypad behavior deserves its own coverage, and our guide to conversational IVR sets out where scripted trees end and language models begin.
Chatbot containment rate
Chatbot containment rate is the percentage of chat sessions closed with no handoff to a live agent. The hard part is defining when a session ends.
Scoring whether a chat actually resolved is its own problem, and our guide to chatbot evaluation covers the metric set.
Chat has no dial tone. A customer who closes the tab produces the same log entry as one who got their answer. Meanwhile, a customer who returns forty minutes later may or may not count as a new session, depending entirely on your timeout.
Set the timeout explicitly, then pair containment with a re-contact check before anyone reports it.
Voice AI agent containment rate
For a voice AI agent, there is no menu to traverse, so containment rides on conversational mechanics. Turn-taking, barge-in handling, response latency, and tool execution all determine whether the caller stays.
Callers now expect the agent to do things. Gartner's February to March 2026 survey of 3,566 customers found 58% of GenAI users have used it to complete a task on their behalf, rising to 74% in B2B.
An agent that only answers questions has a hard limit on what it can contain. The metrics that predict voice agent quality sit upstream of that limit.
How to measure containment rate accurately
Getting the number right takes five decisions, made once and enforced in code.
Define the qualifying session
Decide what counts as a session before you count anything. The workable standard is one meaningful customer message or one completed menu selection.
Exclude test traffic, zero-message opens, and automated pings. Log the rule next to the metric so anyone reading the dashboard can see what was counted.
Score the outcome directly
Stop inferring resolution from the absence of a transfer. Score it directly: a success condition written in plain English per scenario, returning Pass, Review Required, or Failed.
Write that condition the way your business would state it. For a scheduling agent, it reads as an appointment booked, confirmed back to the caller, and written to the calendar system.
A session that ends with none of those and no transfer is a contained session that resolved nothing. It should be visible on your dashboard as exactly that.
Separate the four session endings
One containment number collapses four different outcomes. Split them before you report anything.
- Resolved: The success condition passed.
- Escalated: The agent handed off, correctly or otherwise.
- Abandoned: The customer left mid-flow. Cekura's Dropoff Node reports which stage they left from, and the Appropriate Termination check separately scores whether the agent closed the call correctly when it did end the session.
- Re-contacted: The customer came back about the same issue inside your window.
Only the first bucket is containment worth celebrating. The abandoned bucket is your highest-value debugging queue, because it holds customers who wanted help and left without it.
Set a re-contact window
Pick a window, usually 24 hours, and reclassify any contained session followed by a new contact on the same issue. Requiring an intent match keeps that reclassification honest.
Expect the corrected number to land below the raw one. The drop is the measurement working.
Report both figures side by side for a full quarter. It is the fastest way to get leadership to trust the corrected one.
Segment by intent, caller condition, and language
Aggregate containment is a management number and a poor engineering one. Cut it by intent using a topic classifier, then cut it again by the conditions that vary across your caller base.
Background noise, non-native accents, and code-switching all move containment independently of your prompt. For example, a model change that lifts containment overall can drop it for Spanish-language callers.
Only segmented reporting surfaces that. Our voice observability guide covers the tracing layer that makes those cuts possible.
How to improve containment rate for AI agents
Seven levers, ordered by how much movement they produce per engineering hour.
1. Give the agent transactional tools
Read-only agents plateau early. Every intent that requires a state change downstream will escalate until the agent can make that change itself.
The Postal Service data makes the point cleanly. Information relay intents contained at or above 90%, while task completion intents like change of address sat at 56%.
Wiring in the write path is usually a larger containment gain than any prompt work. It is also the capability customers now arrive expecting.
2. Repair your lowest-containment intents first
Sort intents by volume multiplied by the containment gap, then work the top of that list. A 30-point gain on an intent carrying 2% of volume moves your overall rate by 0.6 points.
Ignore the aggregate while you do this. It is a lagging indicator of work you shipped three sprints ago.
3. Test the multi-turn path
Agents that handle a question perfectly in isolation come apart when the same information arrives across six turns. Researchers at Salesforce and Microsoft ran more than 200,000 simulated conversations across 15 leading models.
Splitting a fully specified task across turns cost an average of 39% of performance. Their diagnosis matters more than the headline figure.
The degradation came mostly from unreliability, with models committing to an assumption early and never recovering. A wrong turn on turn two costs you the session on turn seven, and single-turn evaluation cannot see it coming.
4. Measure repeatability before publishing a rate
One passing simulation run tells you the path is possible, but not whether it’s reliable.
The tau-bench authors built the pass^k metric for this. Their state-of-the-art function-calling agent succeeded on under 50% of tasks on a single attempt, and its pass^8 in the retail domain fell below 25%.
Cekura's voice agent benchmark, which we run and publish, shows the same shape on live platforms.
Across 82 caller scenarios run three times each, Vapi recorded 97.56% task completion on the calls that connected, but only 59.76% of scenarios passed all three runs. That 37-point gap is the distance between your pilot number and your production number.
5. Harden the infrastructure layer
Containment losses often have nothing to do with language. Calls that never connect, agents that talk over the caller, and multi-second response gaps all end sessions.
The benchmark scores this separately, measuring the share of calls completing without a provider-side or connection problem. Results ranged from 100% down to 72.36% across the same cohort, depending on configuration.
Two identical prompts on different infrastructure can differ by more than 25 points on calls that simply completed. Cekura's infrastructure suite runs latency, noise, and interruption scenarios as a standing pre-release gate.
6. Build the escalation path as a product feature
An agent that hides the exit will contain more sessions and lose more customers. Gartner's August 2026 release reports that 87% of customers consider human access essential when a company uses GenAI in service.
In the same survey, only 50% said GenAI makes their interactions easier. Design the handoff so it carries context, fires on low confidence, and takes one request from the caller.
A clean escalation protects the containment you earned elsewhere. Customers who know the exit exists are far more willing to try the agent again.
7. Red-team the containment number
Adversarial sessions produce false containment. A caller who talks the agent into an off-policy answer and leaves satisfied registers as contained and resolved.
Run jailbreak, prompt-injection, and off-script personas against the same scoring you use for happy paths. Cekura's multi-turn red teaming covers the attack patterns that survive across a conversation, which single-turn safety checks miss.
Limits of a higher containment rate target
Past a certain point, containment gains come from making the exit harder to find. The federal data carries that trade openly. IRS automation rose seven points in 2026 while average hold time for callers who still needed a person went from 3 minutes to 8.
Gartner's projection that agentic AI will autonomously resolve 80% of common service issues by 2029 often gets read as a containment target. However, it’s actually a forecast about resolution, which is the harder number by a wide margin.
Pair every containment goal with a guardrail, and review the pair together.
| Containment target | Guardrail metric |
|---|---|
| Overall containment rate | Re-contact rate within 24 hours |
| Per-intent containment | Intent-level CSAT |
| Voice agent containment | Time to human once escalation is requested |
| Chatbot containment | Mid-session abandonment rate |
| Quarter-over-quarter gains | Share of contained sessions scoring negative sentiment |
If containment rises and any guardrail moves the wrong way, you bought the number. You did not earn it. Our guide to agent performance monitoring covers the wider metric set these guardrails belong inside.
How Cekura keeps your containment number honest
Cekura is a testing and observability platform for voice and chat AI agents. The same success condition runs in pre-production simulation and on live traffic, so the definition never drifts between your test suite and your dashboard.
Pre-production
- Automated scenario generation across every branch of a workflow, weighted toward your weakest intents.
- LLM-judge metrics scoring resolution against a plain-English success condition across the full session.
- Repeat runs with pass^k scoring, where a scenario counts as passing only when every run passes.
Infrastructure
- Interruption, background noise, and latency scenarios run against your actual stack.
- Keypad and menu traversal coverage for IVR paths, plus voicemail detection (currently in beta) on outbound calls.
Observability
- Dropoff Node and Appropriate Termination scoring on live sessions, separating resolved endings from abandoned ones.
- Observability alerts that fire to Slack when a metric crosses a threshold.
- Production sessions that miss their success condition are converted into regression scenarios and gated in CI through GitHub Actions.
Confido Health, a Cekura customer, runs the suite in CI. Co-founder and CPO Vichar Shroff described simulating 30+ service workflows in one click, scanning every node and edge.
Native integrations cover Retell, VAPI, ElevenLabs, LiveKit, Pipecat, Bland, and more, so the measurement layer sits on top of the stack you already run.
Cekura supports SOC 2, HIPAA, and GDPR compliance, covering transcript redaction, role-based access, and audit trails.
Where to start with containment rate
Pick your three highest-volume intents and compute containment separately for each of them this week. Then compute it again with abandoned sessions and 24-hour re-contacts pulled out.
The gap between those two numbers is your starting position, and it’s almost always wider than the aggregate on your current dashboard suggests.
From there, the work is ordinary engineering. Score the outcome, run every multi-turn scenario three times, and gate each prompt change on the result.
Want to see what your containment rate looks like once abandoned sessions stop counting as wins? Book a demo, and we will run your workflows through simulation, score resolution against your own success conditions, and show the per-intent split alongside the guardrails that keep it honest.
Frequently asked questions
What is a good containment rate?
There is no audited industry benchmark, and the USPS Office of Inspector General states plainly that no universal benchmarks exist. Set a target per intent against your own baseline. Treat any figure above 90% as a prompt to check your abandonment and re-contact rates.
What is the difference between containment rate and deflection rate?
Scope. Containment measures one automated channel, counting sessions that entered it and ended with no human transfer, while deflection covers your whole support portfolio, counting contacts that never reached a human at all.
How do you calculate IVR containment rate?
Divide the calls that ended without transfer to an agent by total calls entering the IVR, then multiply by 100. Report abandonment rate alongside it. Callers who hang up mid-flow otherwise count as contained and inflate the result.
Can a containment rate be too high?
Yes, it can, when the gain comes from making escalation harder to reach. Gartner found 87% of customers consider access to a human essential when companies use GenAI in service. Containment bought by hiding the exit tends to cost repeat usage.
How do you measure containment rate in pre-production?
You measure containment in pre-production by running simulated sessions across your real intents and scoring each against a defined success condition. Run every scenario at least three times and report the share passing on all runs. Single-run pass rates overstate what production will deliver.
Does containment rate apply to chat and voice equally?
Yes, it applies to both chat and voice, though the session boundary differs. Voice sessions end on hang-up or transfer. Chat sessions need an explicit inactivity timeout, and choosing that timeout materially changes the chatbot containment rate you report.
