Queue management is the practice of controlling how people wait for service: how they join a line, how long they wait, who serves them next, and what happens when they give up. It applies to a bank lobby, a clinic front desk and a contact center phone line, and the same math governs all three.
Last updated: October 2026 · By Shashij Gupta
What is queue management?
Every service operation has more demand at some moments than it has people to meet it. A queue is what forms in the gap. Queue management is the set of decisions you make about that gap: how customers enter the line, what order they are served in, what they are told while they wait, and how many staff you schedule so the line stays short.
A queue management system is the software and hardware that carries out those decisions. In a branch or clinic it is usually a ticket kiosk or a mobile check-in link, a display or SMS that calls the next number, and a dashboard of wait times. In a contact center the equivalent is the automatic call distributor (ACD), which holds callers in a call queue, plays announcements and routes each call to an available agent.
The two settings look different, but they share one structure. Customers arrive at some rate, each one takes some time to serve, and a limited number of servers do the serving. Change any of those three numbers and the wait changes, often by far more than you expect.
Types of queues and queue disciplines
Most guides sort queues by what the customer experiences. That is a useful first cut.
| Queue type | How customers wait | Typical setting | Main weakness |
|---|---|---|---|
| Structured physical line | In a fixed line with barriers | Airports, stadiums, retail tills | Takes floor space, and visible length deters arrivals |
| Unstructured physical | Wherever they choose, served by memory or a ticket | Small shops, pharmacies | Fairness disputes when order is unclear |
| Virtual or remote | Join by app, kiosk or SMS and wait elsewhere | Clinics, government offices, restaurants | No-shows when the call-up text arrives late |
| Appointment | Booked slot, short queue on arrival | Healthcare, banking advice | Idle capacity when slots go unfilled |
| Call queue | On hold, unable to see the line | Contact centers, help desks | Callers cannot judge their wait, and abandonment is invisible |
The second way to sort queues is by discipline, meaning the rule that decides who is served next. First in, first out (FIFO) is the default and the one people perceive as fair. Priority queues serve some customers ahead of others, such as a VIP tier or an urgent clinical case. Skills-based routing sends each customer to a server qualified for their issue, which shortens service time but splits one large queue into several small ones.
Fairness matters more than most operators assume. Richard Larson's 1987 paper in Operations Research, "Perspectives on Queues: Social Justice and the Psychology of Queueing", argues that customers may become infuriated by "social injustice, defined as violation of first in, first out." A shorter wait that someone else visibly skipped can feel worse than a longer, orderly one.
Larson makes two further points. The environment customers wait in, and the information they get about the likely delay, shape their attitude to the queue. And the outcome of a wait can vary nonlinearly with the delay, which makes the average wait a weaker measure than it looks. A queue management system that posts an honest estimate changes the experience without changing the wait.
How a queue management system works
Whether the queue is in a lobby or on a phone line, a queue management system runs the same six steps.
-
Intake. The customer takes a ticket, checks in by phone, or dials in. The system records the arrival time and, where it can, the reason for the visit.
-
Classification. The system assigns a queue based on that reason, the customer's tier or the language they chose. In a contact center this is often an IVR menu.
-
Estimation. The system predicts a wait from the current queue length and recent service times, and may announce it.
-
Holding. The customer waits, in a line, elsewhere with SMS updates, or on hold with music and position announcements. Good systems offer a callback instead of holding.
-
Routing. When a server frees up, the system picks the next customer according to the discipline and hands them over.
-
Measurement. The system logs wait time, service time and whether the customer was served or left, which feeds staffing for the next shift.
Many failures happen at the boundaries between steps: an estimate that is wrong, a callback that never fires, or a routing rule that sends a caller to an empty queue.
What a queue management system changes, and what it costs
Each benefit a queue management system brings has a matching cost.
-
Shorter perceived waits. Virtual check-in lets customers wait somewhere other than the line. The cost is no-shows when the call-up message arrives late.
-
Measured service. Every arrival, wait and service time is logged, which is the input the staffing math needs. The cost is the kiosk, the SMS budget and someone who reads the dashboard.
-
Fair order. A ticket number or position announcement removes disputes about who was first. The cost is that priority tiers must be explained, or they feel unjust.
-
Staffing by evidence. Logged arrival rates feed a staffing model instead of guesswork. The cost is that the forecast is only as good as the data behind it.
The math behind queue management
Why does a queue that was fine at 2 p.m. explode at 2:15? Three results from queueing theory explain it, and they apply to a lobby and a call queue alike.
Little's Law
Little's Law says the average number of customers in a system equals their average arrival rate multiplied by the average time each spends there: L = λW. John Little proved it in 1961 and revisited it in his 2011 Operations Research paper, "Little's Law as Viewed on Its 50th Anniversary".
Little's 2011 paper says the law "holds under remarkably general conditions" and is independent of the queue discipline, which is why it is so useful. If your phone line receives 120 calls an hour (2 a minute) and callers wait 3 minutes on average before an answer, then on average 6 callers are on hold at any moment. If you know any two of the three numbers, you know the third.
Utilization makes waits nonlinear
Utilization is the share of time your servers are busy. Waits do not grow in a straight line with it. In the simplest single-server model (M/M/1, with random arrivals and service times), the average number of customers in the system is ρ / (1 − ρ), where ρ is utilization. At 80% busy that is 4 customers. At 90% it is 9, and at 95% it is 19.
So the last few points of utilization cost far more waiting than the first eighty. This is why a small rise in arrivals, or one agent going on break, can double a hold time. Larger teams soften the curve, because pooling servers absorbs random bursts, but they never remove it.
Erlang C, and what it leaves out
Contact centers have staffed with the Erlang C formula for decades. Given an arrival rate, an average handle time and a number of agents, it predicts the probability that a caller waits and the average wait. Its weakness is its assumptions. Gans, Koole and Mandelbaum's review of call center research in Manufacturing & Service Operations Management states that "the Erlang C model ignores busy signals, customer impatience, and services that span multiple visits."
Impatience is the big one. Real callers hang up, and every caller who abandons shortens the wait for those behind them. A model that assumes everyone waits forever overstates how bad the queue is. The review notes a common complaint from call center managers that workforce management systems "consistently recommend overstaffing", and says a better approach is to model abandonment in the first place. It recommends Erlang A ("A" for abandonment) "as the standard to replace the prevalent Erlang C model."
The review also traces the square-root staffing rule to Erlang himself, who described the relationship as early as 1924. The rule says the safety capacity you need above your base load grows with the square root of that load, not in proportion to it. Doubling call volume does not require doubling your buffer of spare agents.
The stakes are financial. The same review puts human resource costs at 60 to 70 percent of a call center's operating expenses, so a staffing model that is off by a few agents per shift is the largest controllable cost in the building.
Call queue management: what changes on the phone
Call queue management uses the same theory, but three things make it harder than managing a lobby.
Callers cannot see the line
In a branch, a customer who sees ten people ahead can decide to leave or stay with full information. On hold, callers decide with no information except what you announce. Position and estimated-wait announcements are the only signal they get, so an inaccurate estimate is worse than none.
Abandonment is partly invisible
A caller who hangs up shows up in the ACD's abandoned-call count. In chat and messaging queues, many customers simply walk away without closing the session. A 2025 study by Castellanos, Yom-Tov, Goldberg and Park found that 71.3% of abandoning customers left silently in one company's chat contact center, where "the system is unaware that the customers left, and the agent's time is wasted." The authors estimate that silent abandonment cost that system 15.3% of its capacity.
Configuration interacts in ways no one tests
Call queue platforms expose a dozen settings, and their defaults are rarely right for your volume. In Microsoft Teams, for example, the call queue documentation sets the maximum number of calls that can wait at a default of 50, configurable from 0 to 200. The timeout can be set anywhere from 0 seconds to 45 minutes. Each exception (overflow, timeout, no agents) can disconnect the caller or redirect them to another queue, an auto attendant or voicemail.
The platform also decides which agent takes the next call. Teams offers attendant routing (ring everyone at once), serial routing (ring agents in a fixed order), round robin and longest idle, and each one changes how evenly work spreads across the team.
The same page documents a trap worth knowing. If callers become eligible for a callback after 60 seconds, but the queue times out at 120 seconds and the default hold music runs two minutes, "the call queue timeout occurs first, and callback isn't offered." Every setting is valid on its own. Together they silently remove the callback feature.
The metrics that matter for a call queue
| Metric | What it measures | Why it matters |
|---|---|---|
| Average speed of answer (ASA) | Mean time from entering the queue to reaching an agent | The wait callers actually feel |
| Service level | Share of calls answered within a target time | The contractual or regulatory promise |
| Abandonment rate | Share of callers who hang up before an answer | Lost demand, and a sign your staffing model is off |
| Longest wait | The worst wait in an interval | Averages hide the caller who waited 40 minutes |
| Occupancy | Share of logged-in agent time spent handling contacts | High occupancy means waits are about to spike |
| Callback completion | Share of offered callbacks that reach the customer | Callbacks that never connect are abandonment in disguise |
Queue time sits outside average handle time, so a queue can grow while AHT falls. Track both, or one will flatter the other.
Service level is also becoming law. Spain's Ley 10/2025 on customer service, in Article 10.3, requires companies within its scope to answer 95% of phone calls in under three minutes on average. Article 8.4 bars them from cutting a call because the wait is long. Article 8.2 requires that automated answering systems and conversational bots let the customer ask for a person at any point in the interaction, from the start. Companies have twelve months from the law's entry into force on 28 December 2025 to comply.
Common queue management failures
These failure modes show up in call queues more than in lobbies, because no one can see them happening.
-
The callback that never fires. Timeout, hold music length and callback eligibility interact so the offer is never made, as Microsoft's own Teams documentation shows. Callers then hang up, and the logs show them as abandonments.
-
The redial loop. A caller who abandons often calls straight back, so one frustrated customer appears as several arrivals. Daw, Pache and Zhou's 2026 working paper on a public benefits phone line argues that standard Erlang A staffing understaffs when abandoned callers redial.
-
Priority starvation. Strict priority rules serve high-priority work first regardless of how long lower-priority customers have waited. Without an age-based escalation, a low-priority caller can wait indefinitely during a busy hour.
-
Overflow into a dead end. Overflow and no-agent rules that redirect to a queue that is also closed, or to an unmonitored voicemail box, look correct in the admin console and fail only after hours.
-
Announcements that lie. An estimated wait computed from the morning's handle time can be wrong by minutes in the afternoon, and callers act on what they hear.
Where voice AI agents fit in queue management
A common change to call queues now is putting a voice AI agent in front of them. The agent answers every call immediately, resolves routine requests and hands the rest to a human queue. Done well, that cuts arrivals into the human queue, which is where the nonlinear math works in your favor: shaving 10% off arrivals at 90% utilization removes far more than 10% of the wait.
It also creates a new queue to manage. An AI agent has no hold music, but it has a concurrency limit, a response time and a failure rate under load. Little's Law still applies: the number of simultaneous calls the agent must carry equals the arrival rate multiplied by average call duration. Montanari, Scarsini and Perchet's 2026 paper on hybrid chatbot and human queues states the trade-off plainly: "relying more on automation reduces human congestion but increases chatbot costs, while insufficient automation may overload the human agent."
Reliability is not uniform across platforms. Per Cekura's voice agent workflow benchmark, a frozen study of 8 configurations, 82 scenarios and 3 repeats, the share of calls that connected and completed cleanly ranged from 72.36% to 100%. That figure counts all 246 retained calls per configuration, with failed connections kept visible rather than removed. Mean main-agent response time ranged from 1.27 to 3.08 seconds. Two caveats travel with those numbers: providers chose their own models, speech components and settings, and response time is Cekura's main-agent measure, not provider-native component latency.
A voice agent that drops calls under load does not reduce your queue. It sends the same callers back to redial, adding arrivals to the very queue it was meant to shrink.
How to test a call queue before callers do
Most queue failures come from settings that interact, so the only reliable check is to place real calls through the whole path. Cekura runs simulated callers against voice agents and phone flows, and each check in this list maps to a Cekura test. Each one costs test minutes, so focus them on the paths your callers use most.
-
Load test at and above peak. Use Little's Law to find your peak concurrency, then test above it. Cekura's load testing runs each evaluator multiple times in one cycle, so 10 evaluators at a frequency of 5 put 50 concurrent calls on your agent. Cekura schedules those calls at 5 calls per second and scores every run for Infrastructure Issues, Latency and Talk Ratio.
-
Walk every IVR branch. Cekura's simulated callers send DTMF keypresses when the test runs from a Cekura-provided or Plivo number, so you can confirm each menu option lands in the right queue, including after-hours branches.
-
Test the handoff, not just the bot. The moment a voice agent transfers to a human queue is where context gets lost. Cekura tests call transfer and IVR handoff end to end, including whether the agent verifies identity before it transfers.
-
Make the caller wait. Long holds, silent lookups and hold music break agents that expect a reply. A Cekura test can insert a fixed hold or raise the caller's idle timeout, and its silence detection metric flags stretches where neither side speaks.
-
Rehearse the exceptions. Fill the queue to its maximum, let it time out and log every agent out. Then confirm that overflow, timeout and no-agent rules each send the caller somewhere a person will answer.
One limit to be clear about: Cekura measures the agent and the call path, not your ACD's statistics. Queue wait time, ASA, abandonment and service level still come from your contact center platform's reporting. Use both. The platform tells you how long callers waited, and Cekura tells you whether the agent and the routing behaved correctly once they got through.
FAQ
What is the difference between queue management and a queue management system?
Queue management is the practice of deciding how customers wait, who is served next and how many staff you schedule. A queue management system is the software that carries those decisions out, such as a ticket kiosk with SMS alerts in a branch, or an automatic call distributor in a contact center.
What is call queue management?
Call queue management is the setup and monitoring of the queue that holds phone callers until an agent is free. It covers routing rules, maximum queue size, timeouts, overflow destinations, callback offers and announcements, plus the staffing that keeps wait times within your service level.
How do you calculate how many people are waiting in a queue?
Use Little's Law: the average number waiting equals the arrival rate multiplied by the average wait. A line that receives 2 calls a minute with a 3-minute average wait has 6 callers on hold on average. The law holds under very general conditions, whatever order customers are served in.
Why is Erlang C often wrong for call centers?
Erlang C assumes every caller waits until answered and that the queue has no size limit, so it ignores abandonment, busy signals and repeat calls. Because real callers hang up, Erlang C usually overstates waits and leads to overstaffing. The Erlang A model adds abandonment and is the recommended replacement in the research literature.
Can a voice AI agent replace a call queue?
A voice AI agent can answer every call at once and resolve routine requests, which shrinks the human queue. It does not remove the need for one, since complex calls still go to people. The agent also has its own concurrency limit and failure rate, so test it under peak load before routing your queue through it.
Summary
Queue management comes down to three numbers: how fast customers arrive, how long each takes to serve and how many servers you have. Little's Law ties them together, utilization makes waits rise far faster than load, and Erlang A captures the abandonment that Erlang C ignores. On the phone, most damage comes from settings that interact and are never tested end to end.
If a voice AI agent now sits in front of your call queue, Cekura can load test it, walk its IVR branches and verify its handoffs before your callers find the gaps. You can book a demo to see a load test run against your own agent.







