- Word error rate is substitutions plus insertions plus deletions, divided by the number of words in the reference transcript.
- Because insertions are unbounded, the figure can pass 100 percent. Speechmatics shows a transcript scoring 125 percent that a human reads as nearly correct.
- Normalisation decides several points on its own. AssemblyAI describes a case where New York against new york produced 60 percent error on a transcript that was otherwise perfect.
- A quoted accuracy figure without its test set and its normalisation rules is not comparable to any other vendor figure.
- One model does not have one accuracy. Whisper publishes its error rates broken down per language across two separate datasets, and the spread is wide.
- The metric weights every word the same, so a wrong number and a dropped filler cost you identically, which is not how your users experience it.
- Fifty files of your own audio, scored with your own rules, beats any vendor benchmark for deciding what to buy.
Three suppliers quote you 95 percent, 97 percent and 92 percent accuracy. All three numbers can be honest, and the ranking between them can reverse entirely on your own recordings. Understanding why takes about ten minutes and saves you from buying the wrong one.
The metric behind those claims is word error rate, and the trouble is not that it is a bad metric. It is a reasonable one. The trouble is that it is a measurement of a system against a specific reference transcript under specific comparison rules, and vendors quote the output while omitting both inputs.
How is word error rate calculated?
Count three kinds of mistake against a reference transcript, add them up, divide by the length of the reference.
A substitution is a word transcribed as a different word. A deletion is a spoken word the system missed. An insertion is a word the system added that nobody said. AssemblyAI sets out the same three terms and the formula, giving the simplest possible worked example: the spoken phrase "Hello there" transcribed as "Hello bear" is one substitution across a two word reference, so 50 percent.
Work a slightly longer one by hand and the mechanics stick. Reference: "send the invoice to Maria on Friday", seven words. Suppose the system produces "send the invoice to Mariah on a Friday". Maria became Mariah, one substitution. The word "a" appears from nowhere, one insertion. Nothing was dropped. Two errors against seven reference words gives 28.6 percent, and a reader looking at that output would call it very nearly right.
The denominator is the reference length, not the output length. That detail is the source of most confusion about the metric, and it leads directly to the thing nobody tells buyers.
Can word error rate go above 100 percent?
Yes, and the fact that it can is proof that it is not a percentage of correctness.
Substitutions and deletions are each capped by the number of reference words. Insertions are not capped by anything. A system that hallucinates a paragraph over a three word utterance can rack up more errors than there are words to be wrong about, and the arithmetic runs past 100 without complaint.
This is not a theoretical curiosity. Speechmatics gives a concrete pair of examples that ought to be printed on every vendor datasheet: one transcript scoring 125 percent that a human reads as essentially correct, and another scoring 20 percent where a single substitution creates real misinformation about the status of a message. Same metric, opposite verdicts on usefulness.
Why do two vendors measuring the same thing disagree?
Because they are not measuring the same thing. Five factors move the number, and only some of them belong to the vendor.
| Factor | Typical effect on the number | Who controls it | Ask the vendor |
|---|---|---|---|
| Audio quality and microphone distance | Large. The dominant factor on real recordings | You | What audio conditions was this measured on |
| Domain vocabulary and proper nouns | Large on names, products and jargon | Shared, if custom vocabulary is supported | Can I supply a word list, and does it affect your quoted figure |
| Accent and dialect | Large, and uneven across your customer base | Vendor, through training data | Per accent breakdown, not an average |
| Overlapping speakers | Severe on meetings and calls | Shared | Was the test single speaker read audio |
| The normalisation pipeline | Several points, silently | Whoever ran the test | The exact normalisation rules used |
The last row is the one buyers never raise and the one that moves quoted figures most cheaply. Before scoring, both transcripts get put into a comparable form: lowercased, punctuation stripped, numerals converted to words or the reverse. Skip a step and identical content scores as wrong. AssemblyAI describes exactly this failure, where comparing New York against new york produced a 60 percent error rate on a transcript that was correct in substance.
Speechmatics makes the same point from the other side. Normalisation is required for the arithmetic to work, and yet punctuation and capitals carry real meaning for a reader, so the metric discards information that determines whether a transcript is usable. Numbers are worse: 1,000 against 1000, or eleven against 11, are the same to a listener and different to the scorer.
What does a vendor benchmark actually tell you?
That the system performed a certain way on audio that is not yours, scored by rules you have not seen.
The honest version of this is visible in open documentation. The Whisper repository does not publish a single accuracy figure. It publishes a per language breakdown of word error rates, with character error rates for languages where words are the wrong unit, evaluated across two separate datasets, Common Voice 15 and Fleurs, for two model versions. Its own text says performance varies widely depending on the language, and it points readers to the paper appendices for the numbers on other models and datasets.
That is what a real accuracy claim looks like: a matrix, not a number. A vendor quoting one figure has collapsed that matrix, and the collapse is where the marketing lives.
There is a further layer for anyone transcribing meetings or calls. Even the definition of the metric is contested once more than one person is speaking. A 2022 paper on word error rate definitions and their efficient computation shows that the commonly used definitions for multi speaker scenarios are specialisations of a more general formulation, each tuned to a particular application, with different answers about how to handle speaker assignment. Two vendors scoring the same meeting audio can use different definitions and both be correct.
What should you measure instead?
Your own audio, your own rules, enough files to mean something, weighted by what actually costs you money.
Fifty files is a workable floor for a small business. Pull them from real recordings rather than clean samples: the noisy ones, the ones with an accent your staff struggle with, the ones where two people talk over each other. Transcribe them by hand once. That reference set is an asset you keep, and you can rescore every future vendor against it in an afternoon.
Fix your normalisation rules and write them down. Whether you lowercase, whether you strip punctuation, how you render numbers and currency. Apply the identical rules to every vendor. The rules matter less than their consistency.
Then weight the errors. A flat word error rate says a missing "um" and a wrong invoice number cost the same. For nearly every commercial use they do not. Score entity errors and number errors separately, or apply a multiplier to them, and you get a figure that predicts user complaints instead of one that predicts nothing. Speechmatics argues the same case: an ideal metric would distinguish a trivial mistake from one that changes the sentiment or creates misinformation.
Building a reference set that is worth having
The reference transcripts are the expensive part, and the temptation is to shortcut them. Do not, because every future comparison you run inherits their quality.
Sample deliberately rather than conveniently. Take files across the conditions you actually encounter: the quiet office recording and the one from a phone in a warehouse, the customer who speaks quickly and the one whose first language is not yours, the call where someone interrupts. A set drawn only from your best recordings will rank vendors by how well they handle easy audio, which is a question you did not need answered.
Transcribe by hand, then have somebody else check a sample. Speechmatics is right that human transcribers introduce their own errors and apply different specifications, and the second reader catches the systematic ones: how you handled a half finished word, whether you wrote out numbers, what you did with crosstalk. Write those decisions down as you go. They are your normalisation rules, discovered rather than invented.
Fifty files is enough to separate vendors that differ meaningfully and not enough to detect small differences. If two candidates land within a point or two of each other on your set, treat them as tied and decide on price, latency, custom vocabulary support or data handling instead. Chasing a difference smaller than your measurement noise is how buyers talk themselves into the wrong supplier.
Why the metric and your users disagree
Word error rate treats a transcript as a bag of tokens. Your users read it as meaning, and meaning is concentrated in a small fraction of the words.
Take a customer service call of six hundred words. A system that drops twenty filler words scores about 3.3 percent and produces a transcript nobody complains about. A system that gets every filler right but mistranscribes one account number and one date scores about 0.3 percent and produces a transcript that causes a refund to go to the wrong place. The better number belongs to the worse system.
The fix is not exotic. Tag the tokens that carry consequence in your domain, which is usually names, numbers, dates, currencies and product identifiers, and report their error rate as a second figure alongside the flat one. Two numbers per vendor is still a comparison a person can hold in their head, and the second one is the one that predicts trouble.
This is the same failure the industry keeps rediscovering with aggregate scores in other places. An average hides a distribution, and the distribution is where the product breaks. It is worth being sceptical of any single figure offered as a summary of quality, whether it describes transcription or anything else, which is the habit we tried to build into a decision grid for comparing language models rather than a single leaderboard position.
What is a good word error rate?
The question has no answer without the conditions, which is unsatisfying and also the most useful thing in this article.
On clean read speech in a well resourced language, low single digit figures are ordinary and unremarkable. On a noisy phone call with two speakers and product names, the same system may land several times higher. Neither number is the system's accuracy. Both are measurements of a pairing between a system and a condition.
There is also a floor you cannot get under. Speechmatics points out that human transcribers make errors too, so the reference itself is imperfect, creating an intrinsic floor that makes zero percent unreachable. Any vendor implying otherwise is describing a test, not a capability.
The practical framing is to ask what error rate your workflow tolerates rather than what is good in the abstract. A searchable archive tolerates a lot. A transcript that feeds an automated action tolerates very little, which is the same reasoning we applied to what disappears from automatically generated meeting notes, where the errors that hurt are the ones a summary then repeats with confidence.
Where this shows up in a small business
Four places, and the tolerance is different in each.
Captions on product video are forgiving of small errors and unforgiving of wrong product names, which is a vocabulary problem rather than an accuracy problem. Ask about custom word lists before you ask about the headline figure.
Customer call transcripts feed disputes and refunds, so the cost of a wrong number is high and the cost of a dropped filler is zero. This is the clearest case for a weighted score.
Voice ordering and phone agents sit under a latency constraint as well as an accuracy one, and the two trade against each other, a tension we went through in detail on how a phone agent spends its latency budget. A more accurate model that answers a second later can be the worse product.
Multilingual content adds a second measurement problem on top of the first, because transcription error compounds into whatever translation follows it. The methods for judging that second stage are their own subject, covered in how to measure machine translation quality, and the same warning applies: a single quality number hides the distribution.
How to run the comparison without wasting a week
Keep it mechanical. Send every candidate the same fifty files. Score every output with the same normalisation script. Produce two numbers per vendor, a flat word error rate and a weighted one that penalises entities and numbers. Then look at the ten worst files per vendor by hand, because the aggregate hides the failure mode and the failure mode is what your customers will meet.
Ask each vendor, in writing, for the test set behind their published figure and the normalisation rules they applied. The answer itself is informative. A supplier who can produce both quickly has done the work. A supplier who cannot has a marketing number. We take the same view of our own published credit pricing: a figure you cannot reproduce from stated rules is not a figure a buyer should rely on.
Word error rate is not the only speech metric. Character error rate is used where word boundaries are unreliable, which is why the Whisper documentation reports it for some languages instead. Diarisation quality, which decides who said what, is measured separately again, and a system can be excellent at one and poor at the other.
One last habit worth adopting: rescore your incumbent supplier on the same set once a year. Models get replaced underneath an API without a version bump, and the vendor you chose on merit two years ago may no longer be the one you would pick today. The reference set makes that check cheap, which is the whole reason to build it.
The short version
Word error rate counts three kinds of mistake and divides by the length of the reference. It can exceed 100 percent, it treats a lost filler and a wrong figure as equal, and it changes by several points depending on comparison rules nobody publishes.
None of that makes it useless. It makes a quoted figure uninterpretable on its own. Demand the test set, demand the normalisation rules, and score fifty of your own files before you sign anything. The vendor with the second best headline number is often the one that wins on your audio.