AI GlossaryㄷTechnical words in the news
Word Error Rate
An accuracy metric showing what percentage of words a speech-to-text system gets wrong.
In plain words
Word Error Rate (WER) is a percentage that shows how accurately a speech-to-text program transcribes spoken words. It works much like scoring a school test by dividing the number of wrong answers by the total number of questions. You take the sentence a person actually said as the correct answer, then count how many words the program dropped, wrongly inserted, or swapped for a different word, and divide that count by the total number of words in the original sentence.
For example, if a ten-word sentence is transcribed with one word missing and another word wrongly substituted, that's two errors, giving a Word Error Rate of 20 percent. The lower the number, the fewer mistakes were made and the higher the accuracy; the higher the number, the worse the transcription.
This metric is commonly used by speech recognition companies to compare whether a new model performs better than an older one. The rate tends to rise easily in situations involving accents or dialects, multiple languages mixed together, or heavy background noise.
How it shows up in the news
An article reports that a system "improved multilingual performance and Word Error Rate (the rate of mistranscription)." The easy mistake here is thinking that an "improvement" means the number went up—it actually means the number went down. Since Word Error Rate measures the error rate, a lower number is a better score.
Try it yourself
You can get a feel for this yourself.
- Say a short sentence out loud and use your smartphone or an app's voice-to-text feature to transcribe it.
- Place the original sentence and the transcribed result side by side, and count each missing word, wrongly inserted word, and substituted word.
- Divide the number of errors by the total number of words in the original sentence and multiply by 100—that's the Word Error Rate for this transcription.
- Try comparing what happens when you speak slowly and clearly versus quickly and mumbling, and see how the error rate changes.
See also
Stories using this term
- Apple scales up a diffusion-style language model to 1.7 billion parametersAI · 2026.08.11
- ElevenLabs launches 'v3 Conversational' voice model for real-time dialogueAI · 2026.08.21
- Netflix Pilots In-House Language Model GenRec in Recommendation EngineAI · 2026.08.22
- Gemini 3.5 Transcribe teases custom vocabulary registration featureAI · 2026.08.27
- Hume AI Measures Benchmark Memorization in Speech Recognition ModelsAI · 2026.08.22
- AWS Unveils Automated Improvement Feature for Bedrock Automated Reasoning PoliciesAI · 2026.08.09
