New: Cekura Voice AI BenchmarksView results

Average Handle Time (AHT): How to Measure & Reduce It (2026)

Atul Jain
Written bySEP 25, 202616 MIN READ
Atul JaininExpert verified
Founding Engineer, CekuraIIT Kanpur

Has stress-tested 5M+ voice agent minutes at Cekura.

Why Trust Cekura on Voice AI Evals

  • Built by engineers from Google, Apple, Microsoft. Backed by Y Combinator.
  • 60K+ voice AI calls evaluated daily.
  • Native integration for every major voice AI stack: LiveKit, Pipecat, Vapi, Retell, ElevenLabs, Telnyx.

Average handle time (AHT) is the average time an agent spends on one customer contact: talk time plus hold time plus after-call work, divided by handled contacts.

The six-minute figure quoted as the industry standard averages staffing-calculator inputs, so set targets from your own AHT by call reason, read next to first contact resolution.

TL;DR

  • AHT equals talk time plus hold time plus after-call work, divided by handled contacts. Queue time sits outside it, so AHT can fall while callers wait longer.
  • Published benchmarks are planning assumptions. Set targets per call reason from your own median and 90th percentile.
  • Cut the minutes that resolve nothing (wrap-up typing, dead air, repeated confirmations, slow AI turns) and protect the minutes that close the issue on the first contact.

What is average handle time (AHT)?

Average handle time (AHT) is the mean time an agent spends on one customer contact, from answer through the wrap-up work that closes it. Contact center planners use AHT to size staffing, price each contact, and find call types that run long for a fixable reason.

The clock starts when the agent picks up. Time in the queue is excluded, so a queue can grow while AHT drops.

Talk time, hold time, and after-call work

Three components make up every AHT figure:

  • Talk time: The live conversation, including identity checks.
  • Hold time: Minutes the customer waits on the line while the agent checks a system or asks a colleague.
  • After-call work (ACW): Notes, dispositions, tickets, and follow-up tasks completed after the customer hangs up.

ACW is the part customers never hear. An agent who talks for 4 minutes and types for 2 carries a 6-minute AHT, and the customer experienced only two-thirds of it.

Handle time on AI voice agent calls

An AI voice agent can write its own summary and disposition inside the call session. ACW falls toward zero, and handle time becomes the call's connected duration.

Duration is also the billing unit. Retell's pay-as-you-go rate runs $0.07 to $0.31 per minute, tracked to the second, and silence and hold time are billed because speech-to-text keeps listening.

AHT gets confused with four neighboring metrics, and each one answers a different operational question.

๐Ÿ“ Metricโฑ๏ธ What it measures๐Ÿงฎ Formula๐ŸŽฏ Decision it informs
AHTAgent effort per contact(talk + hold + ACW) รท handled contactsStaffing and cost per contact
Average talk timeLive conversation onlyTotal talk time รท handled contactsScript length and verification steps
Average speed of answer (ASA)Wait before an agent answersTotal queue wait รท answered contactsQueue health and service level
First contact resolution (FCR)Share of issues closed without a repeat contactResolved on first contact รท total contactsWhether the handle time produced an outcome
Call duration (AI agents)Connected time on the lineHang-up timestamp minus answer timestampPer-minute platform bill

Handle time measures duration. It says nothing about whether the agent followed instructions or completed the task, so pair it with the task-level QA metrics that score instruction following and tool-call accuracy.

Average handle time formula with worked examples

AHT = (total talk time + total hold time + total after-call work) รท total handled contacts

Use one unit throughout, seconds or minutes, and count only contacts an agent handled. Abandoned calls never reached an agent and stay out of the denominator.

Take a week with 1,200 handled calls, 5,400 minutes of talk time, 900 minutes of hold, and 1,500 minutes of ACW. The total is 7,800 minutes, and 7,800 รท 1,200 gives an AHT of 6.5 minutes, or 6 minutes 30 seconds.

AHT formulas for voice, chat, and email

๐Ÿ“ž Channel๐Ÿงฎ Formulaโš ๏ธ What to watch
Voice(talk + hold + ACW) รท handled callsHolds and transfers inflate it fastest
Live chat(active chat time + after-chat work) รท handled chatsAgents running parallel chats look slower than they are
Email(time worked on the ticket + follow-up work) รท resolved emailsElapsed time and worked time differ widely

Chat AHT overstates effort when agents run parallel conversations. Economists measuring AI-assisted chat support tracked chats per hour alongside AHT because it captures that multitasking. Track both for any chat queue.

Blended AHT vs AHT by call reason

A blended AHT can rise while every call type runs at its old speed, as this illustrative example shows.

๐Ÿ“‹ Call reason๐Ÿ“ž Calls beforeโฑ๏ธ AHT๐Ÿ“ž Human calls after an AI agent takes 300 order-status calls
Order status4003 min100
Billing dispute6008 min600
Blended human AHT1,0006 min 0 s7 min 17 s

Human AHT rose about 21% after the short calls moved to automation, and nobody got slower. Report AHT by call reason, or a deflection win reads as a productivity loss.

Median and 90th percentile handle time

Handle time is right-skewed. A handful of 40-minute escalations drags the mean up while the typical call barely moves.

Report the median for the typical contact and the 90th percentile for the long tail. Then read the tail by call reason, since that is where process defects cluster.

AHT benchmarks and where they come from

Treat every published AHT benchmark as a planning input. None of them reflects your call mix, and the most quoted ones are older or softer than they look.

๐Ÿ“Š Source๐Ÿ”ข Figure๐Ÿ” What it measures๐Ÿ“… Date
Call Centre Helper6 min 3 s blendedAverage of 190,702 entries typed into its Erlang staffing calculatorPublished April 2018, modified January 2026
Cornell U.S. Call Center Industry Report6.1 min average. Business and IT services 8.8 min, large business 8.7 min, telecom 5.4 min, financial services 4.7 min, retail 4.7 minManager survey of 472 US call center worksites2004

The 6-minute figure describes what planners typed into a calculator, so it is an assumption about calls as opposed to a measurement of them.

The sector split is older still: it comes from Cornell's 2004 survey of 472 US call centers, and some republished versions attach the figures to the wrong sectors.

Build your baseline from 8 to 12 weeks of your own data, split by call reason and channel. Your trend line is the benchmark that predicts next quarter's staffing.

How AHT drives staffing and cost

AHT is one of the two inputs to every Erlang staffing model, alongside call volume. Hourly workload in Erlangs equals calls per hour multiplied by AHT in seconds, divided by 3,600.

The illustrative table below uses the standard Erlang C model for 300 calls per hour at an 80% answered-within-20-seconds target.

โฑ๏ธ AHT๐Ÿ“ฆ Workload (Erlangs)๐Ÿ‘ฅ Agents needed
6 min 0 s30.036
5 min 30 s27.533
5 min 0 s25.030

Every 30 seconds of AHT moves this queue by 3 seats per interval, before shrinkage for rest periods, training, and absence.

The median hourly wage for US customer service representatives was $21.53 in May 2025. That works out to about 36 cents per paid minute in wages alone, before benefits, software, and management overhead.

On AI calls, minutes on the line are minutes on the invoice. Trimming 30 seconds across 100,000 monthly calls removes 50,000 billable minutes, worth $3,500 to $15,500 a month at the Retell rate range above.

Run your own volume and handle time through an ROI calculator before you commit to a savings target, since automation rates from demos rarely hold on live traffic.

Signs a low average handle time is hiding a problem

A falling AHT is good news only while repeat contacts stay flat. SQM Group's research finds that each 1% gain in first contact resolution cuts operating cost by 1% and lifts customer satisfaction by 1%, so a rushed call that triggers a second one erodes its own saving.

Watch for these four patterns after any AHT push:

  • Repeat contacts rise within a week. The first call ended before the issue did.
  • Transfers climb. A handoff shortens one agent's handle time and lengthens the customer's journey.
  • Short calls end at the same step. A cluster of calls ending at identity checks or payment points to a defect in that step.
  • Quality scores slide on your fastest agents. Speed that comes from skipped steps appears in QA reviews first.

AI agents produce the same illusion. Nurix, a Cekura customer, traced a false-silence bug that disconnected callers who gave one-word answers. Those calls looked short and efficient, and every one of them was a lost conversation.

What adds handle time to AI voice agent calls

Slow turns, repeated information, padded replies, dead air, and retries after tool errors add avoidable minutes to AI voice agent calls. Each one leaves a measurable trace in the recording and transcript.

๐Ÿข Time driver๐Ÿ”Ž How it sounds on the call๐Ÿ“ Metric that catches it
Slow agent turnsA pause after every caller sentenceLatency, reported at P50 through P99
RepetitionThe agent re-confirms details that have not changedUnnecessary Repetition Count
Padded or overloaded repliesRestated questions, stacked disclaimers, options nobody asked forVerbosity, scored 1 to 5
Dead airBoth sides silent past a threshold, 10 seconds by defaultDetect Silence in Conversation
Unresponsive agentNo reply within the timeout after the caller finishesInfrastructure Issues
Tool errorsThe agent stalls, apologizes, or re-asks after a lookup returns an errorTool Call Success

Cekura scores each driver with a pre-defined metric, measuring latency from the end of the caller's speech to the agent's first audible reply on the stereo recording. That audio-level clock includes endpointing, model time, and speech synthesis in one number.

Turn speed compounds over a call. On Cekura Bench, providers chose their own models, speech components, and settings, and mean agent response time ranged from 1.27 seconds (ElevenLabs) to 3.08 seconds (Vapi) across 8 configurations and 82 scenarios.

Over an illustrative 20-turn call, a 1.81-second gap per turn adds about 36 seconds of waiting, and across 10,000 calls a month, that is 6,000 minutes billed for silence.

How to reduce average handle time without hurting resolution

Cut the minutes that close nothing and protect the minutes that close the issue, on human queues and AI agents alike.

1. Segment AHT by call reason before setting targets

A single AHT target punishes the agents who draw the hardest calls. Tag every contact with a reason and set a target band per reason from its own median and 90th percentile.

On AI calls, topic classification can tag every conversation against the reasons you define, which gives you per-intent duration without manual coding.

2. Automate after-call work

Generate the summary, disposition, and ticket from the transcript while the call is still live, and have the agent confirm the draft before closing. Keep a human review step on anything that feeds billing or compliance records.

3. Give human agents in-call AI assistance

In a field study of 5,172 customer support agents, AI assistance raised issues resolved per hour by 15% on average. Less experienced agents improved in both speed and quality, while the most experienced saw small speed gains and small quality declines.

Roll assistance out to newer hires first. Then audit quality scores for your veterans during the first month after launch.

4. Move routine intents to a tested AI voice agent

Order status, appointment changes, and balance checks follow short, fixed paths. An AI voice agent can take them end to end and free human minutes for calls that need judgment, and our guide to voice AI automation compares the platforms that build them.

Simulate those intents with off-script callers before launch. A deflected call that loops back to a human creates a second contact for the same problem, and expect human AHT to rise afterward, as the call-reason example above shows.

5. Trim verbosity and repeated confirmations

Audit prompts and scripts for confirmation loops and stacked disclaimers. Confirm each detail once, read back only what the caller changed, and keep each reply to what the caller asked.

6. Tighten response latency and endpointing

Endpointing decides when the caller has finished speaking, and a cautious setting adds dead air to every turn. Shorten it in small steps and watch for callers being cut off, especially mid-date or mid-account-number.

Latency work belongs in the same pass. Speech, model, tool, and telephony layers each add delay, so measure per turn and fix the slowest layer first.

7. Replay production calls after every prompt or model change

One added confirmation step adds seconds to every call on that path.

Turn a fixed set of production calls into test scenarios, run them against each new version, and compare pass rates and per-metric deltas side by side before rollout. Cekura's A/B testing runs two agent versions against the same evaluators and shows which metrics moved.

8. Coach agents on their longest call reasons

Pull the 90th-percentile calls for each reason and review them with the agent who handled them. Long calls usually trace to a missing knowledge-base article, a slow lookup, or an unclear policy, and fixing one shortens every future call of that type.

Route each reason to the agents trained on it so transfers stop adding a second handle time. Coaching costs supervisor hours, so start with the reasons that carry the most total minutes.

How Cekura helps measure and reduce AHT on AI agent calls

Cekura is a QA platform that tests, monitors, and improves voice and chat AI agents. It ties extra call minutes to a cause you can fix before launch or after it.

Pre-production

  • Scenario simulations with expected outcomes: Every simulated call is graded on whether the agent reached the goal, so a shorter call counts as a win only when the task completed.
  • Personas that go off-script: Simulated callers interrupt or speak with strong accents, which exposes the loops and timeouts that stretch real calls.

Infrastructure

  • Latency and stability testing: The infrastructure suite runs pre-built scenarios that validate latency, stability, and failure handling across voice infrastructure stacks.

Observability

  • Per-call duration drivers: Latency percentiles, repetition, verbosity, silence, and tool-call errors are scored on every production call.
  • Topic and drop-off tagging: Calls are grouped by reason and by the stage where they ended, which turns raw minutes into conversational analytics you can act on.

Native integrations work out of the box for Retell, Vapi, ElevenLabs, LiveKit, Pipecat, Bland, and more. You add a testing and monitoring layer on top of the stack you already run.

It's SOC 2-, HIPAA-, and GDPR-compliant for transcript redaction, role-based access, and audit trails, with documentation in the trust center.

Tracking average handle time in 2026

AHT earns its place when you track it by call reason, read it next to first contact resolution, and treat every cut as a hypothesis until repeat contacts confirm it.

On AI calls, trace each extra minute to its source in latency, repetition, verbosity, silence, or tool errors, then fix the largest one first.

Want to see which of those minutes your AI agent could give back?

Book a demo and see Cekura score latency, repetition, silence, and task completion across simulated and live calls.

Frequently asked questions

What is a good average handle time?

A good average handle time is one that holds steady or falls while first contact resolution holds or rises. Blended figures around 6 minutes circulate widely, but they come from staffing-calculator inputs and a 2004 survey, so compare each call reason against its own history.

Does average handle time include queue time?

No, average handle time excludes queue time. It starts when an agent answers and ends when after-call work is complete. Wait time before the answer is tracked separately as average speed of answer (ASA).

How do you calculate AHT for live chat?

Chat AHT equals active handling time plus after-chat work, divided by handled chats. Pair it with resolutions per hour, because agents running several chats at once look slower on AHT than they are.

Can average handle time be too low?

Yes, average handle time can be too low when calls end before the problem is solved. Rising repeat contacts, more transfers, or clusters of calls ending at the same step all point to rushed or defective handling.

What is the difference between AHT and average talk time?

The main difference between AHT and average talk time is scope. Talk time covers only the live conversation, while AHT adds hold time and after-call work.

How do AI voice agents change average handle time?

AI voice agents change average handle time in two ways. They take short, routine calls off the human queue, which raises human AHT, and their own call duration becomes a per-minute cost driven by latency, repetition, and silence.

Ready to ship voice
agents fast?ย 

Book a demo