How to calculate word error rate, with a worked example
·5 min read
Word error rate is the standard way to score a transcript against a correct one. The formula fits on one line, and most explanations stop there. The trouble is that two people can apply the same formula to the same pair of transcripts and get different numbers, because the formula leaves the interesting decisions to you.
Here is the whole calculation, by hand, and then the decisions.
The formula
You need two texts: a reference (what was actually said, checked by a person) and a hypothesis (the transcript you're scoring). Line them up word by word and count three kinds of mistake:
- Substitutions (S) — a word in the reference replaced by a different word.
- Deletions (D) — a word in the reference that's missing from the hypothesis.
- Insertions (I) — a word in the hypothesis that isn't in the reference.
Then divide by the number of words in the reference (N):
WER = (S + D + I) / N
The denominator is the reference only. That detail matters later.
A worked example
Reference:
We did not approve the budget for fifteen new hires.
Hypothesis:
We did approve the budget for fifty new hires today.
Line them up so that as many words as possible match, and mark what's left:
REF: we did NOT approve the budget for FIFTEEN new hires ***
HYP: we did *** approve the budget for FIFTY new hires TODAY
D S I
- "not" is missing: one deletion.
- "fifteen" became "fifty": one substitution.
- "today" was added: one insertion.
The reference has 10 words. So:
WER = (1 + 1 + 1) / 10 = 30%
The alignment step is the part people get wrong by hand. You aren't comparing word 3 with word 3; you are finding the cheapest set of edits that turns one text into the other — the same edit distance used for spelling, counted in words instead of letters. A deletion early on shifts everything after it, and a naive position-by-position comparison would score every one of those shifted words as an error. Tools do this with dynamic programming; by hand, line up the long runs of matching words first and count what's left between them.
Why WER can be more than 100%
Because insertions are counted but the denominator is only the reference, a transcript can make more mistakes than there were words. Reference: "Yes." Hypothesis: "Yes yes yes yes." One word in the reference, three inserted:
WER = (0 + 0 + 3) / 1 = 300%
That isn't a bug in the formula. It's what happens with a recogniser that gets stuck repeating itself, and it's why "accuracy = 100% − WER" stops making sense at the bad end.
The decisions the formula leaves to you
Before counting anything, both texts get normalised — reduced to a form where "the same word" has a clear meaning. Every choice here moves the number. Some common cases, with what our word error rate calculator does in each:
Capital letters. "The" and "the" are the same word for almost every purpose. Lowercase both sides. (Our calculator does.)
Punctuation. A missing comma isn't a word error. Strip it. (Ours does, including the hyphen in "e-mail", so "e-mail" and "email" match.)
Contractions. Is "don't" the same as "do not"? Word by word, "I don't know" against "I do not know" is one substitution and one insertion: 67% WER on a three-word sentence that means exactly the same thing. Ours drops apostrophes, so "don't" matches "dont", but it does not expand contractions. If your reference and hypothesis follow different styles, expand them yourself on both sides before comparing.
Numbers. "15" against "fifteen" is a substitution unless you normalise numbers, and ours doesn't: a reference written with digits and a transcript that spells numbers out will score worse than it deserves. Pick one convention and convert both texts to it first.
Fillers and false starts. If the reference includes every "um" and the transcript leaves them out, each one is a deletion. Decide whether you're scoring a verbatim transcript or a clean one, and make the reference match that decision.
Languages without spaces. Split Chinese, Japanese or Thai on spaces and a whole sentence is one "word", so a single wrong character scores 100%. Research in those languages usually reports character error rate instead. Our calculator splits those scripts into dictionary words using the browser's own word segmenter, so it still reports a word rate — but two spellings of the same word can be split differently, which moves the result a little.
None of these choices is wrong. What's wrong is comparing two WER figures that made different ones. A vendor's benchmark number and the number you measure on your own recording are only comparable if the normalisation is the same, and it usually isn't stated.
What the number doesn't tell you
Look at the worked example again. Of its three errors, one is harmless — an extra "today" — and two reverse the meaning: the budget was not approved, and it was for fifteen people, not fifty. WER scores all three the same.
That's why a single rate is a poor way to choose a transcription tool, which what "99% accurate" actually means goes into at length. Our calculator reports the headline number and then breaks the errors down by category — numbers, likely proper nouns, negations and everything else. For the example above it shows the one negation and the one number both wrong and nothing else, which tells you far more than "30%". It's an approximation: a proper noun is guessed from a capital letter in the middle of a sentence, and the negation and number-word lists are English (digits count in any language).
Doing it yourself
- Get a real reference. Transcribe a few minutes of your own audio by hand, carefully. A few minutes that represent your recordings beat an hour that doesn't.
- Normalise both sides the same way — case, punctuation, numbers, contractions, fillers — and write down what you chose.
- Align, count, divide. By hand for a sentence; with a tool for anything longer. Ours runs in the browser and scores the first 3,000 words of the reference; past that it cuts both texts and says so.
- Read the errors, not just the rate. The number tells you how many mistakes there are. Only the list tells you whether any of them matter.