New: Cekura Voice AI BenchmarksView results

First Call Resolution (FCR): Definition, Formula & Benchmarks (2026)

Tarush Agarwal
Written byOCT 6, 202618 MIN READ
Tarush AgarwalinExpert verified
Co-founder & CEO, Cekura

Has stress-tested 5M+ voice agent minutes at Cekura.

First call resolution measures how many customer contacts are fully resolved during the first interaction, with no callback or follow-up about the same issue. It is also a contact center metric an AI agent cannot be trusted to report on itself, because the agent's claim that the job is done is the thing being scored.

A voice agent that says a refill is submitted, then sends nothing to the pharmacy, produces a call that reads as resolved on any dashboard that trusts the transcript.

TL;DR: first call resolution

  • FCR is eligible calls resolved on first contact divided by total eligible calls. How you define eligibility and resolution moves the number: SQM Group finds internally measured FCR tends to run 10% to 20% higher than customer-reported FCR.

  • SQM Group's 2026 call center benchmark average is 71%, against a world-class mark of 80% that roughly 5% of centers reach. Industry averages run from 56% in telco to 77% in retail.

  • AI agents push the reported number up by asserting actions the backend never executed. Across 9,876 customer-service trajectories in a 2026 ICML workshop paper, that pattern accounted for 45% to 47% of unsuccessful runs in two of three domains.

What is first call resolution?

First call resolution is the percentage of eligible customer contacts fully resolved during the first interaction, with no callback, transfer, or follow-up required about the same issue. The customer decides the outcome, and they express that decision by not calling back.

SQM Group links 9 in 10 cases of customer dissatisfaction to a first call that did not resolve the issue, and counts FCR among the most reliable predictors of satisfaction.

First contact resolution is the same measurement widened past the phone. First call resolution usually names the voice or IVR touchpoint, and first contact resolution names the version that spans chat, email, and web.

One contact resolution counts the issue as resolved only if the customer used a single touchpoint to get there, and SQM calls it the tougher measure, because it counts every touchpoint the customer used on the same issue.

First call resolution compared with the metrics beside it

Four metrics get swapped for one another in automation reviews. What separates them is who gets to declare the issue closed.

MetricWho decides the outcomeObservation windowWhat it cannot see
First call resolutionThe customer, by not calling backFirst contact plus a re-contact windowIssues carried to another channel outside the window
Containment rateThe routing logOne sessionWhether anything got resolved
Resolution rateThe agent or a scoring modelOne session or a ticket lifetimeRepeat contacts about the same issue
Repeat contact rateThe contact logA fixed window after first contactWhich contact should have closed it
Transfer rateThe telephony layerOne sessionWhether the transfer was correct

Containment and FCR are easy to conflate, and they answer different questions. Our guide to containment rate covers the session-level metric; FCR is what happens after the caller hangs up.

The first call resolution formula

The first call resolution formula is a single division:

FCR = (eligible calls resolved on first contact ÷ total eligible calls) × 100

The numerator depends on how you define resolution, and the denominator depends on which calls you deem eligible to be resolved at all.

Call eligibility rules that set the denominator

An eligible call is one where the agent had a fair opportunity to resolve something. Billing questions, account changes, order inquiries, and support requests qualify.

Wrong numbers, misdials, internal test calls, and calls dropped before the intent was captured do not. Every exclusion you write is an editorial choice, and it needs to live in version control next to the metric.

Loose criteria are one of the main reasons a reported rate flatters. SQM's benchmarking work finds that internally measured FCR tends to run 10% to 20% above customer-measured FCR, because many call centers use criteria or calculation methods that inflate it.

A worked FCR calculation for a voice AI agent

Take an inbound refill agent handling 4,800 calls in a month. This example is illustrative, and the arithmetic is the part to copy.

Pull out 210 wrong numbers, 145 calls that ended before any intent was captured, and 60 internal test calls. That leaves 4,385 eligible calls.

Of those, 3,112 met the written success condition and produced no second contact about the same issue within 24 hours. FCR is 3,112 ÷ 4,385, or 71%.

Now run the lazy version. Count every call that ended without a transfer, skip the exclusions, and skip the callback check, and 4,090 of 4,800 calls qualify. That reports 85%.

This is the same month and the same agent, with 14 points of difference. The gap is entirely definitional. A dashboard that counts every call without a transfer as resolved is reporting the second number.

First call resolution benchmarks for 2026

SQM Group's call center industry benchmark average is 71%, based on post-call surveys and AI-assisted QA evaluation across the 500-plus centers it benchmarks each year. A 71% rate means 29% of customers call back.

The industry standard for a healthy rate is between 70% and 79%. World-class is 80% or higher, and about 5% of centers hit it. Observed performance spans roughly 40% to 91%.

FCR benchmarks by industry and call complexity

Call complexity separates the industry averages: SQM finds lower-complexity industries such as retail post higher FCR than higher-complexity ones such as telco and tech support.

IndustryAverage FCRWhat the number reflects
Retail77%Highest average; lower call complexity
Insurance75%Most consistent, with a lower bound near 69%
Not-for-profit73%Lower call complexity
Tech support64%Greater call complexity
Telco56%Lowest average of any benchmarked industry

Within-industry spread matters more than the averages. Health insurance ranges from 51% to 91%, and financial services from 53% to 90%, which puts two organizations in the same regulated vertical 40 points apart.

Adopt none of these as a target. Instead, compute your own baseline per intent, then measure movement against it.

Why the gap between FCR and CSAT has widened since 2013

FCR and CSAT track each other closely, and the distance between them has grown. SQM measured that at roughly 4 percentage points in 2013 and about 8 points in 2026.

SQM attributes the widening to self-service. As simple contacts move to websites, chat, apps, and IVR, the calls reaching a human concentrate into harder problems, which lengthens handle time, adds transfers, and makes FCR harder to achieve. SQM dates the acceleration to 2020.

Deploying an AI front door accelerates exactly this shift. Your automated channel takes the easy intents and posts a strong number, while your human channel inherits a harder mix and posts a weaker one. Neither figure is comparable to last year's.

Why AI voice agents overstate first call resolution

Human FCR measurement assumes the person who handled the call is not also the person grading it. Automated stacks violate that assumption three ways.

Agents that assert an action the backend never executed

An agent says the refund is processed, but the database has no record of it. The caller believes the issue is closed and calls back a week later when the money never arrives.

Researchers characterized this pattern across 9,876 customer-service trajectories drawn from 8 frontier model families on tau2-bench. False completion claims accounted for 45% of unsuccessful runs in the airline domain and 47% in retail.

Per-model rates spanned 13% to 79%. Reasoning models offered no protection, and the highest rate in the corpus belonged to an explicitly reasoning-trained model whose traces rationalized completion in place of verifying it.

In the telecom domain, where an independent process could verify state, the same pattern dropped to 3%. The authors treat independent verification as a hypothesis worth testing, not a proven cause.

LLM judges that grade the closing sentence

The standard solution is an LLM judge reading the transcript and scoring resolution from what it reads.

In the same study, no judge configuration exceeded 0.65 AUROC across five judge models, five prompt strategies, and a baseline supplied with the full ground-truth task specification. Judges anchor on confident closing language, which is precisely the signal a false completion produces.

A lightweight structural detector recovered 72% of false completions at a 10% review budget, against 13% for the strongest judge. The judges were scoring a surface signal of completion, not a verified change in system state.

One passing run reported as a resolution rate

A single successful simulation still doesn’t say much about how often the path holds.

Cekura's voice agent benchmark runs 82 caller scenarios three times each against eight platform configurations, with each provider choosing its own models, speech components, and settings. Retell ranked first on repeatable reliability, with 62 of 82 scenarios passing all three runs, or 75.61%.

That leaves 20 scenarios, roughly one in four, that failed at least one of their three runs on the top-ranked configuration. Publish a pass rate from a single run and you are publishing a first call resolution target the agent has not shown it can hold.

The benchmark's provider notes show what a miss looks like. In one configuration, the transcript captured a phone number correctly, but a different number was sent to the tool. In another, the agent narrated a tool call and continued with an invented result.

How to measure first call resolution accurately

There are five decisions, made once and enforced in code.

Choose a measurement method and state its limits

Post-call surveys put the resolution judgment with the customer, which is the definition working as intended. Coverage is the constraint, since only a fraction of callers respond.

AI-assisted QA evaluation covers everything, trading human judgment for inference. SQM reports that its own AI model predicted 72% FCR against 71% from post-call surveys, and matched the survey distribution at 93%.

Run both for a quarter and publish the delta. Calibrating the automated score against real customer responses is what earns the automated score its credibility.

Set the re-contact window and require an intent match

Pick a window per call reason, and reclassify any call followed by a new contact about the same issue. SQM's internal method looks 1 to 30 days out, from about 5 days for a status inquiry to 30 days for a claim that needs investigation. Requiring an intent match keeps the reclassification honest, since a caller with a new problem is not a resolution miss.

Expect the corrected rate to land below the raw one. That drop is the measurement doing its job.

Score resolution against a written success condition

Stop inferring resolution from the absence of a transfer. Write the condition in plain English, one per scenario, returning pass, review, or miss.

For a scheduling agent, it reads as an appointment booked, confirmed back to the caller, and written to the calendar system. A call ending with none of those and no transfer is contained and unresolved at once.

Cekura scores resolution against a written success condition and the tool calls behind it, so a confident closing line carries no weight on its own.

Verify the backend state behind every completion claim

Assert on the tool call. Check that the arguments the agent sent match the values the caller gave, and that the downstream write actually landed.

Twin Health's onboarding agent normalizes spoken dates and unit measurements before passing them to clinical tools, converting "March 20, first" into a real date and feet plus inches into total inches.

A silent normalization error produces a confident call and a wrong record. Twin Health runs its entire simulation library before every deployment, so a prompt change in one agent does not break a medical check elsewhere.

Segment FCR by intent, channel, and caller condition

Aggregate FCR is a management number and a poor engineering one. Cut it by intent first, then by the conditions that vary across your callers.

Segment by background noise, accent, and language. In one noisy-audio run in Cekura's benchmark, a long pause was followed by lost digits and a skipped tool action. A model change that lifts FCR overall can drop it for Spanish-language callers.

Only segmented reporting surfaces that. Our guide to voice observability covers the tracing layer those cuts depend on.

How to improve first call resolution

There are seven levers here. The first five each target a cause of repeat calls that SQM Group documents, and the last two keep a fixed intent from breaking again.

1. Fix the policy causing the callback before touching the prompt

SQM's root-cause work splits resolution loss three ways. Organizational policies, processes, and procedures account for 49%, agent errors for 38%, and customer miscommunication for 13%.

When a callback comes from an approval threshold, a handoff rule, or a refund limit your agent is correctly following, prompt engineering cannot move it.

2. Confirm the outcome to the caller and write it to the system of record

Customers calling to verify the status of an unresolved issue is a top repeat-call cause. State the confirmation number or the completed action out loud, then persist it where the customer can check it without dialing again.

Send the confirmation by SMS. A caller who can check the status on their phone has one less reason to dial.

3. Connect the knowledge source the agent keeps missing

Agents lacking the knowledge or system access needed to resolve an issue is a second documented cause. Cluster your unresolved calls by intent and read what the agent reached for and did not find.

Wire the missing source in as retrieval before you rewrite instructions around the gap.

4. Assert the tool result on every task call

A requested action that is never completed is another of SQM's documented causes of repeat calls. On an AI agent it can hide behind a confident confirmation, so the caller hangs up believing the job is done.

Every scenario that changes state needs a tool-level assertion, checking that the call fired, the arguments were correct, and the response came back clean. Cekura's tool call testing validates the argument values themselves.

5. Remove the third-party redirect from the flow

Sending a caller to a separate organization to finish the job counts against you. The customer experienced one problem and two phone calls.

Where a third party genuinely owns the resolution, do the transfer warm and carry the context across. A cold redirect with a phone number guarantees a repeat contact somewhere.

6. Run every resolution scenario three times before publishing a rate

One clean run predicts very little. Score a scenario as passing only when all repeats pass, and report that figure to leadership as your FCR expectation.

Cekura's GitHub Action applies the same rule in CI: set each scenario to run three times, and the job passes only when every run passes.

7. Convert every repeat call into a regression scenario

A caller who came back has handed you a labeled example. Replay it against your current agent, confirm the miss reproduces, then keep it in the suite permanently.

Repeat calls are test cases your callers have already labeled for you. Gate merges on the suite so a fixed intent stays fixed.

First call resolution tradeoffs against handle time and scope

Past a point, FCR gains come from narrowing what the agent will attempt. An agent restricted to three simple intents posts a beautiful rate while routing everything else away, and the routed calls become somebody else's repeat contact problem.

Treat a reported rate above 90% as a prompt to audit, since the highest rate in SQM's benchmark is 91%. Check your eligibility exclusions, your re-contact window, and whether abandoned calls are sitting in the numerator.

FCR also pulls against average handle time. If resolving on the first call takes six minutes and a transfer takes two, an AHT target set independently will push agents toward the transfer.

Pair the FCR goal with a guardrail you review alongside it. Re-contact rate within 72 hours, intent-level CSAT, and time-to-human once escalation is requested each catch a different way of buying the number.

Our guide to agent performance monitoring sets out the wider scorecard those guardrails belong inside.

How Cekura keeps a first call resolution number honest

Cekura is a testing and observability platform for voice and chat AI agents, backed by Y Combinator.

One resolution metric, written as an LLM Judge metric, scores both pre-production simulations and production calls, so the definition of resolution stays fixed between your test suite and your dashboard.

Pre-production

  • Scenario generation from a plain-language description or a real call, so your lowest-resolution intents get test coverage of their own.

  • LLM judge metrics that score resolution against a plain-English success condition and can assert on the tool calls in the transcript, so the verdict rests on what the agent did rather than how it signed off.

  • Repeat runs, with each scenario set to run as many times as you choose, so a resolution rate reflects every attempt rather than one.

Infrastructure

  • Interruption, background noise, and latency scenarios run against your own stack.

  • Keypad and menu traversal coverage for IVR paths, so DTMF entry gets tested alongside speech.

Observability

  • Production calls scored by the same LLM Judge metrics, with Deep Research auditing a window of traffic for error patterns nobody wrote a metric for.

  • Unresolved production calls converted into regression scenarios and gated in CI through GitHub Actions.

Native integrations cover Retell, Vapi, ElevenLabs, LiveKit, Pipecat, Bland, and more, so the measurement layer sits on top of the stack you already run.

Cekura is SOC 2, HIPAA, and GDPR compliant, with PII redaction for transcripts and audio recordings, a signed BAA and DPA from the Startup plan, and audit logs on Enterprise.

Where to start with first call resolution

Pick your three highest-volume intents and compute first call resolution separately for each one this week. Then compute it again with ineligible calls removed and 72-hour repeat contacts pulled out of the numerator.

The distance between those two numbers is your real starting position. It shows how much of your current figure comes from the definition rather than from resolved calls.

From there it is ordinary engineering. Write the success condition, assert the tool result, run every scenario three times, and gate each prompt change on the outcome.

Want to see what your first call resolution rate looks like once unverified completions stop counting as wins?

Book a demo, and we will run your workflows through simulation, score resolution against your own success conditions, assert the tool calls behind each one, and show the per-intent split alongside the guardrails that keep it honest.

Frequently asked questions

What is a good first call resolution rate?

A good first call resolution rate is 70% to 79%, and world-class is 80% or higher. SQM Group's 2026 call center benchmark average is 71%, with roughly 5% of centers reaching the world-class mark. Set your target per intent against your own baseline, since telco averages 56% and retail averages 77%.

How do you calculate the first call resolution formula?

You calculate FCR by dividing eligible calls resolved on first contact by total eligible calls, then multiplying by 100. Exclude wrong numbers, internal test calls, and calls that ended before an intent was captured. Document the exclusion rule, because any change to it starts a new time series.

What is the difference between first call resolution and first contact resolution?

The main difference between first call resolution and first contact resolution is channel scope. First call resolution usually measures the phone or IVR touchpoint, while first contact resolution usually names the same outcome on chat, email, and web, and many teams use the two terms interchangeably.

Can an AI voice agent improve first call resolution?

Yes, an AI voice agent can improve first call resolution when it can execute the actions callers request and those actions are verified. An agent that can only answer questions has to escalate any intent that needs a backend change, and an agent that claims an action it never executed sets up a repeat call the transcript does not show.

How do you measure FCR before an agent goes live?

You measure pre-production FCR by running simulated calls across your real intents and scoring each against a written success condition. Run every scenario at least three times and report the share passing on all runs. In Cekura's benchmark, even the top-ranked configuration failed at least one of three runs on 20 of its 82 scenarios.

Does a high FCR always mean better customer satisfaction?

No, a high FCR does not always mean better satisfaction, though the two correlate strongly. The gap between them widened from about 4 points in 2013 to about 8 points in 2026, which SQM attributes to self-service absorbing simpler contacts and leaving harder problems for live agents.

Ready to ship voice
agents fast? 

Book a demo