New: Cekura Voice AI BenchmarksView results

Word Error Rate: Formula, 2026 Benchmarks and How to Reduce It

Atul Jain
Written byOCT 6, 202620 MIN READ
Atul JaininExpert verified
Founding Engineer, CekuraIIT Kanpur

Has stress-tested 5M+ voice agent minutes at Cekura.

Word error rate (WER) counts a speech recognizer's mistakes per word actually spoken: substitutions, deletions, and insertions divided by the length of a verified reference transcript. It is the standard speech-to-text accuracy metric, and it describes a deployment, not a model: on Artificial Analysis, Whisper large-v3 scores 4.1% to 10.1% across four hosted endpoints.

TL;DR: Word error rate

  • WER counts three error types against a reference transcript. Substitutions, deletions, and insertions, divided by the number of words the caller actually said.

  • A published WER describes the vendor's audio, the vendor's reference transcript, and the vendor's text normalization. Change any of the three and the number moves.

  • A lower WER can produce worse understanding. Microsoft researchers cut slot understanding error by as much as 17% while running a recognizer with 46% higher WER.

What is word error rate?

Word error rate is the percentage of words a speech recognition system gets wrong when its output is compared against a verified transcript of the same audio. It is the standard accuracy measure for automatic speech recognition.

The calculation runs on edit distance at the word level. Character error rate runs the same calculation on characters, and WER runs it on whole words.

An algorithm lines up the machine transcript against the reference, then counts the smallest number of edits that would turn one into the other.

A substitution swaps one word for another, a deletion drops a word the caller said, and an insertion adds a word nobody said.

The word error rate formula

The word error rate formula is a single division.

WER = (S + D + I) ÷ N

  • S is the count of substitutions.

  • D is the count of deletions.

  • I is the count of insertions.

  • N is the total word count of the reference transcript.

N is the reference length. The denominator counts what the caller said, and the numerator can count errors from a transcript of any length at all.

A word error rate calculation, worked end to end

Take a rescheduling request where the caller says twelve words.

Reference: I need to reschedule my appointment to the fifteenth at two thirty

The agent's speech layer returns this.

Hypothesis: I need reschedule my appointment to the fiftieth at two thirty please

Three edits separate them.

Error typeWhat happenedCount
DeletionThe first "to" is missing1
Substitution"fifteenth" became "fiftieth"1
Insertion"please" was added1

Run the numbers. (1 + 1 + 1) ÷ 12 = 0.25, so the word error rate is 25%. That is the figure a dashboard would show you.

Now read the transcript as an engineer. The dropped 'to' and the added 'please' are cosmetic. The 'fifteenth'/'fiftieth' swap turned a real date into one no calendar has, so the booking cannot go through as heard, and WER scored all three the same, which is the central limitation of the metric and the reason the rest of this guide exists.

from jiwer import wer, mer, cer

reference = "I need to reschedule my appointment to the fifteenth at two thirty"
hypothesis = "I need reschedule my appointment to the fiftieth at two thirty please"

wer(reference, hypothesis)  # 0.25
mer(reference, hypothesis)  # 0.23 (bounded: never exceeds 1.0)
cer(reference, hypothesis)  # 0.18

Word accuracy and error rates above 100%

Word accuracy is the complement of WER, so a 25% error rate reads as 75% accuracy. That relationship holds while the numbers stay reasonable, and it comes apart at the edges.

Insertions have no upper limit, and N is fixed by the reference. A transcript longer than the audio can push WER above 100%, which drives word accuracy below zero.

The arithmetic is easy to hit. Take a one-word reference where the recognizer emits nine words of invented text. S=1 and I=8 against N=1 gives a WER of 900% on that segment.

Word error rate compared with CER, MER, and WIL

WER has three close relatives, and each one answers a question WER cannot. All four run on the same Levenshtein alignment, and jiwer computes every one of them from the same reference and hypothesis pair.

MetricWhat it divides byUse it when
WERReference word countEnglish and other space-delimited languages
CERReference character countChinese, Japanese, and morphologically rich languages where word boundaries move
MERTotal aligned positions, including insertionsYou want a bounded score between 0 and 1
WILReference and hypothesis lengths togetherYou care about information lost, weighted both directions

Match error rate is the one worth adding first. Because its denominator grows with insertions, MER stays bounded at 1.0 no matter how much text a model invents.

The 2004 Interspeech paper that introduced MER sets the range at 0 with no errors and 1 with no hits. A WER of 900% and an MER of 1.0 describe the same segment, and only one of them plots on a dashboard.

Character error rate matters for a different reason. CER counts a one-letter slip as one error, where WER charges the full price of a wrong word. That suits any language in which a single character carries a whole morpheme.

Embedding error rate (EmbER), introduced at Interspeech 2022, softens the flat cost of a substitution. In the original paper, a substituted word costs 0.1 instead of 1 when its embedding sits close to the reference word's, at a cosine similarity above 0.4. Close in meaning is not the same as harmless, so check how any semantic metric scores your dates and numbers before you trust it.

What counts as a good word error rate in 2026

There is no universal target. The honest answer is a range, and the range is wide.

The Artificial Analysis leaderboard, checked on 6 October 2026, spans 1.7% at the top and 31.2% at the bottom across the models it evaluates. Its AA-WER v2 index is an audio-duration-weighted average over roughly eight hours drawn from three datasets.

Here is where the widely deployed models sit on that index.

ModelAA-WERSpeed factorPer 1,000 min
MAI-Transcribe-2 (Microsoft AI)2.0%304.4x$1.67
Scribe v2 (ElevenLabs)2.2%79.7x$3.67
Gemini 3.5 Transcribe (Google)2.6%89.0x$5.00
Universal-3 Pro (AssemblyAI)3.1%96.0x$3.50
GPT Transcribe (OpenAI)3.3%40.0x$4.50
Nova-3 (Deepgram)5.2%621.8x$4.30

Figures verified against the Artificial Analysis leaderboard on 6 October 2026. Speed factor is a rolling seven-day median, so expect that column to move.

Accuracy and price hold steadier than speed. Nova-3 posts the highest WER in that group and the highest speed factor on the whole leaderboard, 621.8x. Speed factor measures batch transcription of a 10-minute file, so for a voice agent that has to answer inside a turn, check the streaming leaderboard Artificial Analysis publishes separately, which measures latency instead.

Two caveats travel with every figure above. All of them come from one benchmark harness, and none of them were measured on 8 kHz telephone audio with a caller in a moving car.

The same model on two hosts, six points apart

Whisper large-v3 appears four times on that leaderboard across two providers. fal.ai lists it at 4.1% and its Wizper endpoint at 4.7%. Replicate lists Incredibly Fast Whisper at 5.7% and its standard entry at 10.1%.

Artificial Analysis lists the same large-v3 version on all four rows. Chunking strategy, audio resampling, decoding parameters, and post-processing all sit outside the checkpoint, and all of them move the score.

This is the single most useful thing to understand about published WER. The number describes a deployment, and a vendor benchmark therefore tells you very little about the WER your own stack will produce.

Five reasons a published word error rate does not transfer to your calls

1. The reference transcript carries its own errors

WER assumes the reference is correct. Human transcribers miss words, disagree on formatting, and mishear the same audio your recognizer mishears.

Artificial Analysis hit this hard enough to act on it. Two of the three datasets in its index are cleaned versions of VoxPopuli and Earnings22, built by removing transcription errors from the reference text. The site publishes the cleaned and original scores side by side.

When a model is more accurate than the human who wrote the reference, WER penalizes the model. Ground truth is itself a measurement, and every reference set carries its own error rate.

2. Normalization choices move the number

Before any comparison happens, both transcripts get normalized. Casing, punctuation, contractions, numerals, and filler words all have to be resolved to a single convention.

Every one of those is a decision, and the decisions are not standardized. "Healthcare" against "health care" and "$40" against "forty dollars" each count as two errors under a strict rule. A looser rule makes both of them vanish.

Artificial Analysis shows how much rides on these choices. It runs OpenAI's Whisper normalizer, adds its own rules for phone numbers and IDs, and extended those rules again in its April to May 2026 update. Write the normalization rule down, version it in the repo, and treat any change to it as the start of a new time series.

3. WER is a distribution across speakers

A single WER is an average, and averages hide the callers you are most likely to lose. Stanford and Georgetown researchers ran 19.8 hours of matched sociolinguistic interviews through five commercial recognizers from Amazon, Apple, Google, IBM, and Microsoft.

The published result was an average WER of 0.35 for Black speakers against 0.19 for white speakers, on audio matched for speaker age and gender. All five systems showed a substantial racial disparity.

Report WER segmented by accent, language, and channel, or you are reporting a number that describes nobody. This guide to multilingual testing covers the accent and code-switching coverage those cuts require.

4. WER weights a filler word and an account number the same

The formula assigns every error a cost of 1. A dropped "um" and a corrupted member ID are arithmetically identical.

In the twelve-word rescheduling example, a 25% WER contained exactly one error worth engineering attention, and a 99% word accuracy score means nothing if the missing word was "cancel".

Names, dates, dollar amounts, and identifiers carry the intent. Everything else is padding, and WER has no way to tell them apart.

5. Insertions have no upper limit

Recognizers built on language-model decoders, Whisper among them, can generate fluent text when the audio contains no speech at all.

Cornell, University of Washington, NYU, and University of Virginia researchers audited Whisper transcriptions. They found that roughly 1% contained entire invented phrases absent from the underlying audio.

Of those hallucinations, 38% carried explicit harms, including fabricated violence and invented authority claims.

Long silences correlated with the effect, which is exactly what a hold, a lookup, or a caller checking a card number produces. Since jiwer 4.0.0, released 19 June 2025, an empty reference has defined behavior: every hypothesis word counts as an insertion, so you can score silent audio for invented text.

How to measure word error rate on your own voice agent

1. Build the reference set from your own traffic

Pull 200 to 500 real call segments weighted toward your highest-risk intents. Verification steps, payment capture, and scheduling belong in the set before anything else does.

Transcribe them twice with different human reviewers and adjudicate every disagreement. The disagreement rate is your reference error floor, and you should know it before you quote a WER.

2. Write the normalization rule and version it

Decide casing, punctuation, numeral format, contraction expansion, and filler handling once. Commit the rule as code next to the scoring script.

Beyond the base rules, add a domain lexicon of accepted variants. Your drug names, plan names, and SKUs need an explicit spelling authority.

3. Compute WER, MER, and CER together

Run the three metrics in one pass. WER gives you the headline, MER stays bounded when the model invents text, and CER catches near-misses inside otherwise correct words.

Log per-utterance scores, never a single corpus average. The distribution is where the diagnosis lives, and a corpus mean flattens it. Cekura's Transcription Accuracy metric reports a 0 to 100 score and the raw WER percentage for every test run and production call, and its explanation lists each mistranscription that drove the score.

4. Segment every score before you report it

Cut the results by accent, language, background noise profile, codec, and caller device. Then cut them again by intent.

A model swap that improves your aggregate WER by a point can degrade Spanish-language calls by five. Only segmented reporting will surface that, and the aggregate will look like a win the whole time.

5. Score intent-bearing words separately

Extract the entity spans from each reference. Compute a second error rate over names, numbers, dates, and identifiers alone.

That entity-level figure is the one to put on a dashboard. A 4% WER sitting next to a 12% entity error rate is a damaged workflow wearing a good score.

6. Gate the number in CI

Set a threshold per segment and per entity class, then run the suite on every prompt, model, and provider change. Cekura triggers scenario runs from GitHub Actions on pull requests and schedules, so a regression is caught at review time.

Seven ways to reduce word error rate

1. Repair the reference set first

Fixing reference errors lowers your measured WER with no model change, because it removes errors the model never made. The cost is a second human pass and the time to adjudicate it. Adjudicate every disagreement between your two human passes before you tune anything downstream.

2. Add your vocabulary through keyterm prompting

Keyterm prompting tells the recognizer which domain terms to expect, such as drug names, plan names, and SKUs. Deepgram's Nova-3 accepts up to 100 key terms per request, capped at 500 tokens, and Deepgram recommends focusing on the 20 to 50 that matter most. Keyterm prompting is billed as an add-on, at $0.0013 per minute on Pay As You Go.

Keep the list to the terms the model actually misses. Deepgram's own guidance is to leave out common words that are rarely misrecognized.

3. Fix the audio path before swapping models

Packet loss removes syllables, and a syllable lost before the recognizer is audio no model ever hears. Check jitter, codec negotiation, and regional routing before you run a provider bake-off.

4. Tune endpointing so words survive the turn

Aggressive turn detection clips the end of an utterance, and clipped audio produces deletions the recognizer never had a chance to avoid. Digits at the end of a spoken account number are exposed, since a pause between digit groups can read as the end of the turn. The fix costs response time: every millisecond added to the end-of-turn wait is a millisecond the caller waits for an answer.

Cekura's endpointing guide covers the three signals modern agents use to decide a turn has ended.

5. Test each acoustic condition on its own

Run the same scenario suite across clean audio, background noise, packet loss, and non-native accents as separate conditions. Attribute the accuracy loss to the condition, which turns one unusable average into a ranked list of fixes.

Cekura sets background noise and network simulation, covering packet loss, jitter, and latency, per personality, so the same scenario runs under each condition and returns its own score.

6. Choose the model against your own audio mix

Treat leaderboards as a shortlist. Run your top three candidates over your own reference set, in your own languages, at your own sample rate, then decide.

Weigh latency alongside accuracy, and read it from the streaming leaderboard Artificial Analysis publishes separately: the batch speed factor measures a 10-minute file, not a single turn.

7. Confirm low-confidence spans in the conversation

Deepgram and AssemblyAI both return a confidence score for each word. Route anything below your threshold into a read-back before the agent acts on it.

A confirmed digit costs one extra turn, but a wrong digit costs a callback, a refund, or a missed appointment.

Where word error rate stops being the right target

Optimizing WER and optimizing understanding are different objectives, and they can point in opposite directions. Ye-Yi Wang, Alex Acero, and Ciprian Chelba demonstrated this at ASRU 2003 with a result that has aged well.

They swapped the recognizer's language model for one trained on the understanding objective. Word error rate went 46% higher, and slot understanding error dropped by as much as 17%, which is documented in the Microsoft Research publication.

The lesson transfers directly to voice agents. Transcript fidelity and task completion are different objectives, and a speech layer tuned for one is not automatically tuned for the other.

Cekura scores the two separately, with Transcription Accuracy for the speech layer and Expected Outcome for whether the caller's task got done, so a drop in one is visible even while the other holds.

Cekura's evaluation metrics guide sets out the full library, and Mean Opinion Score covers the perceived-quality side WER never touches.

A completed call is also not a reliable agent. On Cekura Bench, eight complete voice agent configurations ran the same 82 caller scenarios three times each.

Telnyx completed the caller's task on 97.56% of its 246 calls, yet only 56 of its 82 scenarios, 68.29%, passed on all three runs. Providers chose their own models, speech components, and settings.

How Cekura scores transcription accuracy on test and production calls

Cekura is a testing, monitoring, and self-improvement platform for voice and chat AI agents. Transcription accuracy is scored as its own layer. Because Transcription Accuracy is its own metric, a speech-side error shows up in its own score instead of disappearing inside a call-level verdict.

On a test run, Cekura generated every word the simulated caller said, so it scores your provider's transcript of the caller against an exact ground truth. On production calls, Cekura builds the ground truth with two separate transcription models and compares both against your provider's transcript. Either way, the output carries a 0 to 100 score alongside the word error rate percentage.

The weighting is where it diverges from textbook WER. Errors in names, common nouns, and numbers count in full, and so does a flipped negation, such as "can" heard as "can't". Verb errors count as half, and articles, pronouns, and fillers barely move the score.

Pre-production

  • Persona-driven simulation in the caller's own language, with accent personalities and code-switching callers such as Spanglish and Hinglish.

  • Transcription Accuracy scored on every run, with each mistranscribed name, number, or noun listed in the explanation.

  • Repeat runs through the frequency setting, which runs each scenario N times so one lucky pass cannot hide a flaky agent.

Audio conditions

  • Background noise and network simulation, covering packet loss, jitter, and latency, set per personality so each condition gets its own run.

  • The AI Interrupting User metric uses voice activity detection to count each time the agent starts talking before the caller finishes, the signature of endpointing that cuts callers off.

Observability

  • The same Transcription Accuracy metric scores production calls after each call ends.

  • Alerts route to Slack when a metric crosses its threshold, and Create Evaluator from Call turns a failed production call into a regression test.

Native integrations cover Retell, Vapi, ElevenLabs, LiveKit, Pipecat, Bland, and more, so the measurement layer sits on the stack you already run.

Cekura supports SOC 2, HIPAA, and GDPR compliance. HIPAA and GDPR come on every plan, PII redaction strips sensitive details from transcripts and recordings, and Enterprise adds SSO, SCIM, and audit logs.

Where to start with word error rate

Pull a first fifty calls from your two highest-risk intents this week, transcribe them twice by hand, and compute word error rate over the set. Then compute a second figure over the entity spans alone.

The gap between those two numbers is your starting position. A wide gap points straight at the workflow that needs attention first.

From there, the work is ordinary engineering. Version the normalization rule, segment every report, and gate the number on every change that touches the audio path.

Want to see what your agent's transcription accuracy looks like under noise, accents, and a caller who interrupts?

Book a demo, and we will run your scenarios through simulation, score transcription accuracy on every run with each mistranscribed name and number listed, and show the per-condition split next to task completion.

Frequently asked questions

What is a good word error rate?

A good word error rate depends on the audio and the workload. Across the Artificial Analysis leaderboard, scores range from 1.7% to 31.2%. None of that benchmark audio is documented as 8 kHz phone audio, so set your target against your own reference set.

How do you calculate word error rate?

You calculate word error rate by aligning the machine transcript against a verified reference, then counting substitutions, deletions, and insertions. Divide that total by the reference word count and multiply by 100. A twelve-word reference with three errors gives 25%.

What is the difference between WER and CER?

The main difference between WER and CER is the unit being counted. Word error rate divides errors by reference words, and character error rate divides them by reference characters. CER suits Chinese, Japanese, and morphologically rich languages where word boundaries shift.

Can word error rate be higher than 100%?

Yes, word error rate can exceed 100%, because insertions are unbounded while the denominator stays fixed at the reference length. A recognizer that generates text over silence produces more errors than there are reference words. Match error rate stays bounded at 1.0.

Does a lower word error rate mean better understanding?

No, a lower word error rate does not guarantee better understanding. Microsoft researchers reduced slot understanding error by up to 17% while running a recognizer with 46% higher WER. Score transcription accuracy and task completion separately, then read them together.

How often should you measure word error rate?

Measure word error rate on every change that touches the audio path, the speech provider, or the model version, plus continuously on production traffic. Gating the metric in CI catches a regression at review time.

Ready to ship voice
agents fast? 

Book a demo