Average speed of answer (ASA) is the mean time a caller waits in queue before a person answers, counted only for calls that were answered. You calculate it by dividing total queue wait for answered calls by the number of answered calls. It is a staffing signal, not a picture of the typical caller's wait.
Most explanations stop at the formula, but ASA is easy to misread. It leaves out the callers who gave up, it averages a skewed distribution, and its clock starts at a point each platform defines for itself. This guide covers the formula, the choices hidden inside it, the targets that are actually written into regulation, and what happens to the metric once an AI agent answers the phone.
What is average speed of answer?
ASA is a queue metric. The clock starts when a call enters the queue for a human agent and stops when an agent picks up. Time in the IVR menu before the queue is normally excluded, and so is any call that ends before an agent answers.
As an ASA call center metric, it sits beside service level and abandonment rate on almost every wallboard. Your phone platform produces it automatically, which is part of why it gets so much attention. The catch is that "automatically" hides several decisions about what counts.
The Genesys Cloud definition is a good example of how specific those decisions get. In the Genesys Cloud glossary, ASA "includes only interactions that agents answered", "does not include the time spent before entering the queue", and "is recorded in the interval in which the agent answered the interaction". That last rule matters for interval reports. A call that waited from 9:58 to 10:03 counts in the 10:00 interval, not the interval where the wait began.
The average speed of answer formula
The average speed of answer formula is:
ASA = total queue wait time for answered calls ÷ number of answered calls
Both halves of the fraction use the same population: calls an agent answered. Take five answered calls that waited 5, 10, 15, 40 and 180 seconds. Total wait is 250 seconds. Divide by 5 and the ASA is 50 seconds.
Notice what that number says about the five callers. Four of them waited 40 seconds or less, and one waited three minutes. The 50-second average describes nobody's actual experience. Calls answered immediately count too, with a wait of zero, so a quiet interval pulls ASA down.
Two arithmetic errors show up often enough to name:
-
Mixing denominators. Dividing answered-call wait time by calls offered, instead of calls answered, drags ASA down whenever callers abandon. The worse your abandonment, the better your ASA looks.
-
Multiplying by 100. ASA is a duration in seconds, not a percentage. At least one widely read definition page presents the formula with "x100" at the end, which turns a 30-second wait into 3,000.
What to include before you calculate
The formula is simple. The inputs are where teams disagree, and two contact centers with identical callers can report different ASA because of these choices.
| Decision | Common default | Why it changes the number |
|---|---|---|
| IVR time | Excluded | A long menu adds real waiting the caller feels but ASA never shows |
| Ring time at the agent's phone | Varies by platform | Some platforms stop the clock at routing, others at pickup |
| Abandoned calls | Excluded | The longest waits are often the ones that ended in a hang-up |
| Short abandons (a few seconds) | Often filtered out | Removes misdials, but the threshold is a choice someone made |
| Transferred or overflowed calls | Varies | A call can be counted once per queue it entered, or once in total |
| Interval attribution | Interval of answer | Long waits land in the next interval, flattering the busy one |
Write your choices down next to the metric. If you change platforms, compare the new platform's definition line by line before you compare the numbers.
Why ASA hides the callers who waited longest
ASA counts only answered calls. A caller who waits four minutes and hangs up adds nothing to it. Under heavy load, that exclusion stops being a footnote and starts deciding what the metric says.
Court records from Missouri's SNAP call centers show how far apart the numbers can drift. A 2026 preprint by Daw, Pache and Zhou fit a queueing model to daily reports disclosed in Holmes v. Knodell (Due Process on Hold). In the four centers it summarizes, average daily abandonment fractions were 2.5%, 58%, 57% and 57%.
At one center, observed ASA was 117.8 minutes, while the average wait across all callers, including those who abandoned, was 38.8 minutes. The two figures describe the same queue and differ by a factor of three.
The same paper names a second blind spot: redials. A caller who abandons and calls back is a new arrival, so one person can generate several abandoned calls and one answered call. The authors argue that staffing guidance which assumes abandoned callers leave for good "fundamentally understaffs" a system where they cannot.
So read ASA beside abandonment rate for the same interval, every time. A falling ASA with rising abandonment usually means callers stopped waiting, not that service got faster.
How staffing drives average speed of answer
ASA responds to staffing in a sharply non-linear way. Near capacity, one agent can move it by tens of seconds. The standard tool for seeing this is Erlang C, the M/M/s model that Ger Koole calls the basic model for a call center.
Under Erlang C, a call waits with probability C. Ger Koole's overview of contact center performance models gives the rest: a delayed call's wait is exponentially distributed with rate sμ − λ, where s is agents, μ is the service rate and λ is the arrival rate. ASA is therefore C ÷ (sμ − λ).
We ran that calculation for one half-hour interval with 250 calls and a 210-second (3.5-minute) average handle time. That workload is about 29.2 Erlangs, so 30 agents is barely more than the traffic itself.
| Agents | Share of calls that wait | ASA | Answered in 20 s | Waited over 60 s | 95th percentile wait |
|---|---|---|---|---|---|
| 31 | 65.2% | 74.7 s | 45.2% | 38.6% | 294 s |
| 32 | 50.7% | 37.6 s | 61.3% | 22.6% | 172 s |
| 33 | 38.8% | 21.3 s | 73.0% | 13.0% | 112 s |
| 34 | 29.3% | 12.7 s | 81.5% | 7.4% | 77 s |
| 35 | 21.8% | 7.8 s | 87.5% | 4.1% | 53 s |
| 36 | 15.9% | 4.9 s | 91.7% | 2.3% | 36 s |
Table inputs: Erlang C, 250 calls per 30 minutes, 210-second AHT, no abandonment. Our calculation.
Two things stand out. First, the step from 31 to 32 agents halves ASA. Second, at 34 agents the ASA is 12.7 seconds, yet one caller in twenty still waits more than 77 seconds. The average sits far below the tail, which is the same pattern p99 latency analysis finds in voice agent response times.
Erlang C is also optimistic about real queues. Koole, Li and Ding validated nine staffing models against 237 days of multi-skill call center logs (Call center data analysis and model validation). All nine models were, in the authors' words, slightly optimistic: on most days they predicted lower ASA than the actuals, so real waits ran longer than planned.
The best three models missed ASA by about 7 seconds on average (6.1 to 7.6 seconds). Ignoring agents' paid breaks raised that error from 6.71 to 17.34 seconds, the largest single effect they measured.
The practical lesson: if your forecast says 15 seconds, plan for worse, and check whether your staffing model knows about breaks.
What is a good average speed of answer?
There is no universal good number. A 28-second global average is repeated on ranking pages without a cited study, so treat it as folklore rather than a target. What you can anchor to is your own history, your service level target, and any rule that applies to your line.
Some lines do have a rule. Medicare Part D sponsors must run a toll-free customer call center that meets fixed telephone standards under 42 CFR 423.128. For coverage beginning on or after January 1, 2022, the regulation requires that the call center:
-
limits average hold time to 2 minutes, where hold time is "the time spent on hold by callers following the interactive voice response (IVR) system, touch-tone response system, or recorded greeting, before reaching a live person"
-
answers 80 percent of incoming calls within 30 seconds after the IVR, touch-tone system or recorded greeting
-
limits the disconnect rate of all incoming calls to 5 percent
Notice that the rule sets an average and a percentile-style target together. That pairing is the right design for any line, regulated or not. An average alone can pass while a large share of callers wait far longer.
Outside regulated lines, a reasonable process is:
-
Set a service level target first (for example, a share of calls answered within a set number of seconds).
-
Run Erlang C for your own volume and AHT to see what ASA that target implies.
-
Track ASA by interval, not by day, so a bad lunch hour cannot hide inside a good daily average.
-
Treat a sudden ASA improvement as a question, and check abandonment before you celebrate.
ASA vs service level, abandonment rate and AHT
ASA rarely means much alone. Each neighboring metric covers a gap ASA leaves.
| Metric | What it measures | What it catches that ASA misses |
|---|---|---|
| ASA | Mean queue wait for answered calls | Nothing on its own; it is the baseline |
| Service level | Share of calls answered within a threshold | The share of callers who waited too long |
| Abandonment rate | Share of callers who hung up in queue | The callers ASA leaves out |
| Longest wait or 95th percentile | The tail of the wait distribution | The callers an average hides |
| Average handle time | Agent effort per contact, after answer | Why queues build: longer calls, fewer free agents |
AHT and ASA are linked through staffing. Under Erlang C, every extra second of handle time adds load, and near capacity that load shows up as queue time. A process change that adds 20 seconds of verification to every call can raise ASA without a single extra caller.
What happens to average speed of answer when an AI agent answers
When a voice AI agent answers every call on the first ring, ASA for the front door falls to roughly the ring time. That looks like a solved metric. It is really a moved one, and three measurements take its place.
The queue moves behind the handoff. Callers who need a person still wait for one, but now after the AI agent transfers them. If your platform counts the AI pickup as the answer, that second wait vanishes from ASA. Report ASA separately for transferred calls, starting the clock at the transfer.
On regulated lines, check the definition first: the Medicare hold-time rule counts the wait "before reaching a live person", so ask your compliance team how an AI answer is treated before you report against it.
Speed of answer becomes a per-turn number. A caller talking to an AI agent waits after every sentence, not just once at the start. Cekura's benchmarks measured agent response time across a frozen study of 8 voice agent configurations, 82 scenarios and 3 retained repeats. Mean main-agent response time ranged from 1.27 to 3.08 seconds across the eight configurations (Cekura's voice agent workflow benchmark).
Two caveats travel with those figures. Providers chose their own models, speech components and settings, and the timing is Cekura's main-agent response measure, not provider-native component latency.
Calls that never connect behave like abandons. A call that fails to connect never reaches an answer, so it cannot appear in ASA. In Cekura's workflow benchmark, one configuration had 41 of its 246 calls not connect. The benchmark keeps those calls visible in its infrastructure reliability figures rather than removing them, which is the same discipline ASA reporting needs for abandoned calls.
The transfer itself is the step most likely to break quietly. A transfer that fires to the wrong queue, or drops the context the caller already gave, creates a second wait and a repeat explanation. Call transfer and IVR handoff testing covers how to check that path before callers find the failures.
How to improve average speed of answer
Every lever below has a cost. Pick the one whose cost you can carry.
-
Fix the schedule before you add headcount. Match agent start times and breaks to the interval forecast. The Koole, Li and Ding validation shows breaks alone more than doubled ASA prediction error, so a plan that ignores them is understaffed by design. Cost: planning time and less flexible shifts.
-
Cut handle time where it adds no value. Shorter calls free agents sooner and lower the load behind ASA. Cost: cut the wrong step, such as verification, and repeat calls rise.
-
Shorten or reroute the IVR. This does not lower reported ASA, since IVR time is usually excluded. It lowers the wait callers actually feel. Cost: fewer self-service completions if menus are cut carelessly.
-
Offer a callback. A callback removes the caller from the live queue. Cost: you must count callbacks in reporting, or ASA improves on paper while callers still wait.
-
Automate the first answer with an AI agent. Front-door ASA drops close to zero. Cost: you now own per-turn response time, transfer reliability and the wait after handoff, and each needs its own measurement.
-
Watch the tail, not just the mean. Add a 95th percentile wait or longest wait to the same report. Cost: a harder conversation with stakeholders used to one number.
How Cekura measures speed on AI-answered calls
Once an AI agent answers, the useful questions shift from "how long did callers queue" to "how long did the agent take to respond, every turn, under load". Cekura tests, monitors and improves voice and chat agents, and measures that shift directly.
Cekura's latency metric measures the gap from the end of the caller's utterance to the start of the agent's first audible response. It is taken from the call audio, so the same method works for cascaded and speech-to-speech pipelines. Cekura tracks per-turn latency, including the 95th percentile, across test runs and production calls. When an agent speaks first, Cekura's latency data gives the start time of the agent's first message, which is the AI version of speed of answer.
For staffing-style questions, Cekura's load testing runs each evaluator several times at once, so 10 evaluators at a frequency of 5 place 50 concurrent calls. Latency and Infrastructure Issues are applied to every load test run by default. A response time that holds at one call and stretches at fifty is the AI equivalent of a queue building at peak.
When a test finds a failure, Cekura's Improve Agent, currently in beta for agents on Vapi, Retell, ElevenLabs and Bland, runs an optimization loop on a private copy of the agent. It changes the agent's instructions and closely related settings and reports before-and-after results per metric. Transfer destinations and phone numbers are never changed by that loop. If you want to see these measurements on your own agent, you can book a Cekura demo.
Frequently asked questions
What is the formula for average speed of answer?
ASA equals total queue wait time for answered calls divided by the number of answered calls. Use answered calls in both parts of the fraction. Dividing by calls offered makes ASA look better whenever callers abandon. The result is in seconds, so never multiply it by 100.
Does average speed of answer include abandoned calls?
No. Standard ASA counts only calls an agent answered, so a caller who hangs up in queue adds nothing. That is why ASA must be read beside abandonment rate. In a 2026 preprint on Missouri SNAP call centers, one center reported an ASA of 117.8 minutes while its average wait including abandoned calls was 38.8 minutes.
Does ASA include IVR time?
Usually not. Most platforms start the ASA clock when the call enters an agent queue, after the IVR. Genesys Cloud, for example, excludes "the time spent before entering the queue". Medicare Part D rules likewise measure hold time after the IVR or recorded greeting. Long menus therefore add waiting that ASA never reports.
What is the difference between ASA and service level?
ASA is the mean wait for answered calls. Service level is the share of calls answered within a set threshold, such as 80 percent within 30 seconds. Service level shows how many callers waited too long. ASA cannot, because a few very long waits and many short ones can produce the same average.
What is the difference between ASA and average wait time?
It depends on the platform, because the two terms are not defined the same way everywhere. In the Missouri SNAP queueing preprint, average wait includes callers who abandoned and ASA does not. Call Centre Helper instead describes average wait time as the period before being connected to an advisor. Check your platform's definition before comparing the two.
Is there a legal limit on average speed of answer?
Some regulated lines have one. Medicare Part D sponsors must limit average hold time to 2 minutes and answer 80 percent of calls within 30 seconds, both measured after the IVR, under 42 CFR 423.128. Medicare Part D is the only federal answer-time rule this guide cites; outside it, targets come from contracts and internal service levels.







