Ask a vendor how accurate their transcription is and you will hear a number like "99 percent." It sounds settled. It is not. Accuracy is not one number that a tool either has or does not; it is a range that shifts with your audio, your accents, your jargon and your language. The same tool can be near perfect on a clean English podcast and frustrating on a noisy four person call in German.
This guide explains the one metric worth knowing, tells you which values are actually good, and shows you how to measure a tool on your own meetings in about ten minutes. By the end you will trust your own test more than any headline figure.
The metric that matters: word error rate
Word error rate, or WER, is the fairest way to compare transcription quality. It counts three kinds of mistake against a correct reference transcript: substitutions (a word heard wrong), deletions (a word missed) and insertions (a word added that nobody said). Add them up, divide by the number of words actually spoken, and you get a percentage. Lower is better.
A worked example makes it concrete. Say ten words were spoken and the transcript has one wrong word, one missing word and one invented word. That is three errors in ten words, a WER of 30 percent, which is poor. The mirror image of WER is accuracy: a 5 percent WER is a 95 percent accurate transcript. WER is the number professionals quote because "95 percent accurate" hides whether the missing 5 percent is the occasional filler word or every second name.
Which WER values are actually good?
Here is the part vendors rarely spell out. A useful transcript and a publishable one are different bars, and both are reachable. Use these bands as your yardstick:
| WER | Quality | What it feels like in practice |
|---|---|---|
| Below 5% | Excellent | Near human. Clean English audio reaches this. Ready to publish with a light read through. |
| 5 to 10% | Good | Reliable for notes and summaries. A few names or terms need a quick fix. |
| 10 to 15% | Usable | The gist survives, but check every name, number and decision before you trust it. |
| Above 15% | Poor | You fix more than you save. Change the tool, the language model or the audio setup. |
For everyday meeting notes, aim for 10 percent or better. If you publish transcripts, quote them in contracts, or work in medicine or law where a wrong word carries weight, hold out for 5 percent and plan a human review on top. Clean, single speaker English routinely lands below 5 percent with a good tool. A crowded multilingual call can sit above 20 percent even on the same tool, which is exactly why one headline number is misleading.
What drags accuracy down
Before you blame the tool, know what actually moves the number. Most of it is about the sound reaching the microphone, not the AI behind it:
- Audio quality. A cheap laptop mic in a hard walled room is the single biggest killer. Echo and low volume confuse every model.
- Overlapping speech. When two people talk at once, both lines suffer. This is why panel discussions transcribe worse than interviews.
- Accents and dialects. Models are trained mostly on standard accents. Strong regional speech raises WER, sometimes a lot.
- Jargon and names. Product names, acronyms and people's names are not in the model's everyday vocabulary, so they get guessed.
- Language. English is the best supported language almost everywhere. Coverage and quality for other languages vary widely between tools.
What you can actually control
The good news: you can often improve accuracy more by fixing your setup than by switching tools. In rough order of impact:
- Use a real microphone. A headset or external mic beats a laptop's built in mic by a wide margin. This is the cheapest, biggest win.
- One speaker at a time. A light meeting norm ("please don't talk over each other") does more for the transcript than most settings.
- Cut the noise. A quiet room, closed windows and muted keyboards remove the background the model has to fight through.
- Add custom vocabulary. Most serious tools let you feed in names, products and acronyms up front, which sharply cuts errors on exactly the words you care about most.
Measure it yourself in ten minutes
Never take an accuracy claim on faith when you can test it. The method is simple and works for any language:
- Take one real recording of a typical meeting, ideally two or three minutes with your normal audio and accents.
- Run it through your two or three shortlisted tools.
- Pick the same 200 word passage in each transcript and read it against what was said.
- Count the substitutions, deletions and insertions, then divide by 200. That is your WER for each tool on your audio.
- Note where the errors cluster. Errors on filler words matter little. Errors on names, numbers and decisions matter a lot.
Ten minutes of this tells you more than a week of reading marketing pages, because it measures the tool on the exact conditions you will use it in.
When testing yourself is essential
For clean English meetings, most well known tools are close enough that other factors decide. Do your own test when any of these apply: your meetings are not in English, they are full of domain jargon, you have strong accents or several languages in one call, or the transcript feeds something high stakes like a contract or a medical note. In those cases the gap between tools is real and worth an afternoon to find.
Key takeaways
- Accuracy is a range, not a fixed number. Judge it on your own audio.
- Word error rate is the fair metric. Aim below 10 percent for notes, below 5 percent for publishing or regulated work.
- A good microphone and one speaker at a time often beat switching tools.
- Custom vocabulary fixes the errors you care about most: names and terms.
Ready to compare? Our comparison table shows which tools cover the most languages, and the tool finder narrows the field to the ones that fit your setup. If you are still choosing, start with our guide on how to choose an AI meeting tool.