Customer service quality assurance is the practice of scoring a sample of customer conversations against a defined rubric, then feeding the results into coaching, process changes and compliance evidence. A QA program has three parts: a scorecard that defines good, a sampling rule that decides what gets reviewed, and a loop that turns scores into changes.
Most programs get the first part roughly right and the other two badly wrong. Teams argue for weeks about scorecard wording, then review two percent of contacts, and never check whether two reviewers scoring the same conversation would agree.
This guide covers the parts that decide whether a QA score means anything: what belongs on a scorecard, how large a sample must be to detect the problem you are hunting, how to measure reviewer agreement with a number instead of a meeting, and what changes once some of your agents are AI.
What is customer service quality assurance?
Customer service quality assurance, often shortened to customer service QA, evaluates individual customer interactions against written standards, so that service quality becomes something you measure rather than something you sense. It covers email, chat, phone, messaging and any other channel where a customer talks to your company.
It answers a narrow question well: on this conversation, did the agent do what the company decided a good agent does? It is not the same thing as customer satisfaction. A survey tells you how a customer felt. A QA score tells you what the agent did. The two often disagree, and the disagreement is informative: a conversation that scores 95 and produces a detractor usually means the standard is wrong, not the agent.
Three activities get grouped under the same heading and should be budgeted separately.
Quality assurance scores conversations against a rubric and feeds coaching.
Quality management covers the surrounding workflow: assigning reviews, tracking disputes, scheduling calibration, reporting trends to the business.
Compliance monitoring checks that legally required things happened, such as a disclosure, a consent capture, or a recording notice. Compliance is not a section of the scorecard. It is a separate check with a different failure cost, because a missed disclosure is a regulatory exposure rather than a coaching moment.
Teams that buy scoring, call it quality management and assume compliance came along with it tend to discover the gap during an audit.
What belongs on a QA scorecard
A scorecard is a list of criteria, a scale for each, and a weighting. Most working scorecards run to eight or twelve criteria across three groups.
Resolution. Was the customer's actual problem solved? Was the answer correct? Was the correct process followed? These carry the heaviest weight, because everything else is decoration on a wrong answer.
Communication. Tone, clarity, acknowledgement of the customer's situation, absence of jargon. Real but softer, and the group where reviewer disagreement concentrates.
Process and compliance. Notes recorded, ticket categorised, disclosure read, consent captured. These are binary and should be scored binary.
Four rules keep a scorecard usable.
Write each criterion as something observable in the transcript. "Showed empathy" is not observable. "Acknowledged the customer's stated problem before offering a solution" is.
Score binary criteria as yes or no. A five-point scale on a yes-or-no question invents variance that reviewers then have to argue about.
Keep one auto-fail, not five. Auto-fails are for conduct that makes the rest of the score irrelevant, such as a data-protection breach. Every additional auto-fail turns the score into a coin flip.
Cap the scorecard at what a reviewer can hold in their head. Beyond roughly twelve criteria, reviewers stop reading the rubric and start scoring on impression.
How many conversations do you need to review?
This is the question QA programs answer worst. The common answer, review two to five percent of contacts, or five per agent per week, is a staffing decision dressed up as a statistical one. It tells you nothing about what that sample can detect.
Work it the other way round. Decide the smallest problem you want to catch, then compute the sample that catches it.
Suppose a defect shows up on three percent of contacts: a specific refund policy stated wrongly, say. Review twenty random conversations and the chance that not one of them contains the defect is 0.97 raised to the power of 20, which is 0.54. More than half the time, a three percent defect is completely invisible to a twenty-conversation sample. Reviewing fifty gets you to 0.97 to the power of 50, or 0.22, so a defect on one call in thirty still escapes about a fifth of the time.
To be ninety-five percent confident of seeing at least one instance of a three percent defect, you need 99 conversations. For a one percent defect, 299. The formula is n greater than or equal to ln(0.05) divided by ln(1 minus p), and it is worth putting in a spreadsheet before the next headcount conversation.
Two consequences follow.
Random sampling is the wrong default for small samples. If you can only review thirty conversations a week, spend them on a targeted slice: one queue, one intent, one new hire, one policy change. A random thirty across everything detects nothing reliably. A targeted thirty on refund conversations detects refund problems.
Stratify rather than spread. Sample separately by channel, by queue and by tenure, because the defect rates differ and a pooled average hides the queue that is failing.
How do you know your reviewers agree?
Calibration is the practice of getting reviewers to score the same way. Almost every guide recommends running calibration sessions. Almost none says how to tell whether calibration worked, and the usual answer, percentage agreement, is close to worthless on a QA scorecard.
The reason is chance. Artstein and Poesio's survey of agreement measurement in Computational Linguistics sets out the arithmetic: raw observed agreement cannot be compared across studies because part of it is produced by chance alone, and how much depends on how skewed the categories are. Their worked example is directly relevant to QA. If 95 percent of items belong to one category and 5 percent to another, two reviewers assigning labels at random in those proportions would agree on 0.95 squared plus 0.05 squared, which is 90.5 percent of items. Against that baseline, the paper notes, an observed agreement of 90 percent is actually worse than chance.
QA scorecards are exactly that skewed. Most conversations pass most criteria. A calibration session that ends with "we agreed on 92 percent of the items" has demonstrated nothing.
The fix is to report a chance-corrected coefficient instead. Cohen's kappa for two reviewers, Krippendorff's alpha where reviewers vary between items or the scale is ordinal. Both subtract the agreement expected by chance before dividing. Artstein and Poesio record the thresholds the field settled on, taken from Krippendorff by way of Carletta: above 0.8 counts as good reliability, and between 0.67 and 0.8 supports only tentative conclusions, with Krippendorff later describing even 0.8 as a low standard.
Practical version for a support team:
- Pick ten conversations spanning the score range, not ten easy ones.
- Have every reviewer score them independently, with no discussion first.
- Compute kappa or alpha per criterion, not for the scorecard total. The total hides which criterion is broken.
- Rewrite any criterion scoring below 0.67. The reviewers are not the problem. The wording is.
- Repeat quarterly, because agreement decays as the business changes and nobody updates the rubric.
Reviewers disagreeing is not a discipline issue. It is the scorecard telling you a criterion is ambiguous, and rewriting it is cheaper than arguing about it every quarter.
Which metrics should a customer service quality assurance program track?
A QA program produces one number of its own and borrows the rest.
| Metric | Who produces it | How to read it |
|---|---|---|
| Internal quality score | The QA program | The only metric QA fully controls, and the only one that moves when the rubric changes rather than when service changes |
| Customer satisfaction and net promoter score | Customers | Track beside the quality score, never averaged into it |
| First contact resolution | Operations | The outcome measure closest to what a scorecard rewards, and the first to degrade when handle-time targets tighten |
| Compliance pass rate | The compliance check | Reported on its own line, on the full population |
| Reviewer agreement | Calibration | Kappa or alpha; the number that says whether the other four can be trusted |
Two of those carry a reason rather than a rule. Compliance pass rate stays separate because mixing it into the quality score lets a strong tone score conceal a missed disclosure. Reviewer agreement gets reported alongside the rest because without it, nobody reading the quality score knows whether it would survive a second reviewer.
What changes when your agents are AI?
Once part of the queue is handled by an AI agent, most of the QA program above stops applying, and the parts that break are not the obvious ones.
Coaching disappears. You cannot coach a model. The equivalent action is editing a prompt, a tool definition or a retrieval source, and that edit changes behaviour on every conversation at once rather than on one agent's next shift. So the loop shortens and the blast radius grows.
Sampling stops being the constraint. Automated scoring can read every conversation, so the hard problem moves from choosing which conversations to review to trusting the scores. Anyone selling automated scoring should be asked for a measured agreement rate against your own human reviewers on your own conversation types, computed the chance-corrected way described above. Cekura publishes the measurement rule behind each score on its public benchmark, which is the level of disclosure worth demanding from any scoring vendor. Our guide to LLM-as-a-judge scoring covers where model-graded evaluation is reliable and where it is not.
Repeatability becomes the thing to measure, and this is where scoring a conversation once misleads badly. Cekura's public voice agent benchmark scores each configuration on a fixed set of 82 caller scenarios and runs every scenario three times, then reports pass-cubed, the share of scenarios that passed on all three retained runs. In the frozen v1 release on Cekura's benchmarks site, the top-ranked configuration recorded 93.88 percent single-run task completion but only 75.61 percent pass-cubed. Read those two numbers together and the point is unavoidable: an agent that looks 94 percent reliable when each scenario is checked once is repeatably reliable on three quarters of them. Two caveats travel with those figures wherever they are quoted. Providers chose their own models and settings for the configurations tested, and calls that failed to connect stay in the denominator rather than being dropped from the release.
Cekura runs that same structure against a customer's own agent before it reaches production, which is the part conventional QA tooling has no equivalent for. Cekura simulates the conversations, scores each run against the customer's evaluators, and repeats the scenario set on every prompt change, so a regression shows up as a failed scenario rather than as next month's scorecard dip. Cekura then monitors the live conversations after deployment against the same evaluators, which keeps the pre-deployment standard and the production standard identical.
If you are evaluating tooling for a contact centre that runs both human and AI agents, our buyer's guide to call center quality assurance software covers the vendor questions, pricing shapes and the checklist in detail.
Voice QA inherits every transcription error
Quality assurance on phone conversations has a dependency that chat QA does not: the transcript. Reviewers read transcripts, automated scorers read transcripts, and both inherit whatever the speech recognition system got wrong.
Word error rate is the standard measure of that gap, counting every substitution, insertion and deletion against a reference transcript. The Open ASR Leaderboard, a community benchmarking effort that compares 86 open-source and proprietary systems across 12 datasets, standardises word error rate alongside inverse real-time factor so that accuracy comparisons hold across different architectures and toolkits. The point worth carrying into a QA program: transcription accuracy varies between systems and conditions, so it is a measurable property of your stack rather than a constant.
The practical consequences are small and specific. A misheard account number reads as an agent error. A dropped negation inverts the meaning of a policy statement. Overlapping speech during a busy call produces transcripts that no reviewer scores consistently. Before trusting voice QA scores, sample a set of conversations, correct the transcripts by hand, and rescore them. The difference between the two scores is your transcription tax. Our walkthrough of call transcript QA for a voice bot sets out how to run that review end to end.
What does compliance QA actually require?
"Ensures compliance" appears on every QA vendor page and almost never comes with a rule attached. For contact centres making outbound calls in the United States, at least one rule is specific, recent and directly about AI.
In a declaratory ruling adopted on 2 February 2024 and released on 8 February 2024, the Federal Communications Commission confirmed that the Telephone Consumer Protection Act's restrictions on the use of an "artificial or prerecorded voice" cover current AI technologies that generate human voices, citing 47 U.S.C. section 227(b) and 47 CFR section 64.1200(a)(1) and (3). The ruling states that calls using such technologies therefore require the prior express consent of the called party, absent an emergency purpose or exemption, and that calls carrying an advertisement or telemarketing require prior express written consent.
For a QA program the implication is concrete. If an AI voice agent places outbound calls, consent state is a compliance criterion on every one of those calls, not a sampled one. Sampling is a technique for estimating quality. It is not a defence, because the exposure attaches to the individual call rather than to the average. Compliance checks belong on the full population, scored automatically, with the evidence retained.
Turning scores into change
A QA program that produces scores and stops is an expensive reporting exercise. Three habits separate the ones that change anything.
Route findings by cause, not by agent. A defect appearing across many agents is a process, policy or documentation problem, and coaching individuals for it is both ineffective and resented. Sort defects by how many agents produced them before deciding who gets feedback.
Give agents the rubric and the right to dispute. Reviewers score against a standard the agent can read, and disagreements go through a defined process. QA earns a punitive reputation quickly when scores arrive without the rubric behind them and with no route of appeal, and that reputation is expensive to reverse.
Close the loop in public. When a scorecard criterion changes because reviewers could not agree on it, or a macro is rewritten because thirty conversations tripped over the same wording, say so. Programs that visibly change things get honest engagement. Programs that only produce scores get gamed.
If part of your queue is already handled by an AI agent, the measurement above is worth running before the next prompt change rather than after it. Cekura runs a fixed scenario set against your own agents, repeats it, and reports where the results stop holding, on your evaluators rather than a vendor's. You can talk to the team about running it on your agents.
Frequently asked questions
What is the difference between quality assurance and quality control in customer service?
Quality control inspects output after the fact, which is what conversation scoring does. Quality assurance is the wider system meant to stop defects being produced: the standards, the training, the process design and the review loop. Most teams use "QA" for both. The distinction matters when a defect appears across many agents, because that is a design failure rather than an inspection failure.
How many conversations should we review per agent?
Set the number from the smallest defect rate you need to detect, not from a fixed percentage. Detecting a defect that occurs on three percent of contacts with 95 percent confidence takes about 99 conversations. If your capacity is well below that, narrow the scope to one queue or one intent rather than sampling randomly across everything.
How often should QA scores be calibrated?
Quarterly for a stable program, monthly while a scorecard is new or the business is changing quickly. Measure the result with a chance-corrected coefficient such as Cohen's kappa rather than percentage agreement, because on a skewed scorecard high raw agreement can be no better than chance.
Can automated QA replace human reviewers?
Automated scoring can read every conversation, which no human team can, so it changes what coverage costs. It does not remove the need to establish that its scores match a human standard on your own conversation types. Ask any vendor for a measured agreement rate on your data before signing. Cekura scores conversations against evaluators the customer defines, so that standard stays yours rather than the vendor's.
Does customer service quality assurance work for AI agents?
The scoring part transfers and the coaching part does not. Testing an AI agent means running fixed scenarios repeatedly and checking that they pass every time. Cekura's benchmark shows why repetition matters: the top-ranked configuration recorded 93.88 percent single-run task completion but 75.61 percent pass-cubed across three runs of 82 scenarios. Providers chose their own configurations for that benchmark, and failed connections stay in the denominator.
What metrics should a customer service quality assurance program report?
Internal quality score, customer satisfaction, first contact resolution, compliance pass rate, and reviewer agreement. Report compliance separately from the quality score so a strong communication score cannot mask a missed disclosure, and report reviewer agreement alongside everything else, because it is what makes the other numbers interpretable.







