Skip to content
Blog

ASRMetrics

What Are WER and WRR in Speech Recognition Systems

Key ASR metrics: Word Error Rate (WER) and Word Recognition Rate (WRR). Formulas, worked examples and how to interpret the results.

Maksim3 min read

When working with ASR systems, one of the first questions that arises is: "How well does this model perform?" To answer objectively, we need quantitative metrics. The industry standard for evaluating ASR accuracy is Word Error Rate (WER). Let's explore what it is, how to calculate it, and how to interpret the results correctly.

What Is Word Error Rate (WER)?

Word Error Rate (WER) is a metric that measures the discrepancy between the text generated by an ASR system (the hypothesis) and the reference transcription verified by a human (the reference). The lower the WER, the more accurate the model.

WER is based on the Levenshtein algorithm, adapted to work with words instead of characters.

To calculate WER, three types of errors must be identified:

  • Substitutions (S): Words the system recognized incorrectly. For example, "will" instead of "was".
  • Deletions (D): Words present in the reference transcription but missed by the system.
  • Insertions (I): Extra words the model "invented" that were not in the original audio.

Error Types in WER Calculation

Loading diagram…

Formula and Worked Example

The WER formula:

Where — substitutions, — deletions, — insertions, — total number of words in the reference transcription.

Информация

Because of insertions (I), WER can theoretically exceed 100%.

Worked Example

Let's calculate WER for a specific case:

  • Reference: the weather will be nice today (N = 6 words)
  • Hypothesis: the weather was be nice today yes

Aligning the words:

ReferenceHypothesisResultSDI
thethematch000
weatherweathermatch000
willwassubstitution100
bebematch000
nicenicematch000
todaytodaymatch000
—yesinsertion001

Total: , , ,

What Is Word Recognition Rate (WRR)?

Word Recognition Rate (WRR), sometimes called Word Accuracy, is the "inverse" metric to WER. It indicates the proportion of correctly recognized words.

For our example:

An alternative formula uses hits () directly:

For our example:

Внимание

Different implementations may yield slightly different results. Always verify which formula is used when comparing models.

How to Interpret Results?

WER evaluation depends heavily on context: audio quality, domain, presence of accents. However, these general benchmarks apply:

WER Interpretation Scale

Loading diagram…
  • 0–5% WER — excellent, comparable to human transcription quality.
  • 5–10% WER — great quality, text requires almost no corrections. Production-ready.
  • 10–20% WER — acceptable quality, some post-editing may be needed.
  • 20–30% WER — fair quality, noticeable errors. Model needs tuning.
  • 30% and above — poor quality, transcription is hard to use. Significant improvement required.

Limitations of WER

Despite its popularity, WER is not a perfect metric:

  • All words are equal. WER treats replacing "in" with "on" the same as replacing "not" with "now", even though the second error completely changes the meaning.
  • No punctuation awareness. Standard WER ignores punctuation, capitalization and formatting.
  • Does not measure readability. Two texts with identical WER can have vastly different readability.

Conclusion

WER and WRR are fundamental tools for evaluating ASR system performance. They provide a fast, standardized accuracy assessment, enabling model comparison and training progress tracking.

However, a deep analysis requires looking beyond the final number. It's essential to analyze the errors themselves — are substitutions, insertions, or deletions dominant? Which specific words does the model confuse? Answering these questions is key to further improving speech recognition quality.1

Footnotes

  1. For industrial ASR evaluation, it is recommended to use WER alongside other metrics such as Character Error Rate (CER), Sentence Error Rate (SER) and Match Error Rate (MER), and to test across various acoustic conditions and domains. ↩

Tags:ASRMetrics

Author

Maksim

Found it useful? Share it with your colleagues