7 August 2026 · Updated 24 August 2026 · By Heni Hazbay
How Accurate Is AI Transcription? An Honest Answer
What the “99% accurate” claims actually measure, how word error rate works, and what to realistically expect from a recording of a real conversation.
The short answer: Accuracy is set mostly by your recording, not your software. On clean audio of one clear voice, modern transcription is very good. On a four-person meeting recorded from across a table, expect to correct names, jargon and anything said over the top of someone else.
“99% accurate” appears on almost every transcription product’s homepage, ours included in spirit. It is worth understanding what that number measures, because the gap between the claim and your experience is usually not the software being dishonest — it is your audio being different from the audio it was measured on.
How accuracy is actually measured
The standard metric is word error rate (WER). Take a recording, produce a careful human transcript as the reference, run the software over the same audio, and count what went wrong:
- Substitutions — a word transcribed as a different word.
- Deletions — a word spoken but missing.
- Insertions — a word appearing that nobody said.
Add those together and divide by the number of words actually spoken. A WER of 5% means one word in twenty is wrong somehow, which is another way of saying 95% accurate.
Two things about WER are worth knowing before you compare any two products.
It treats every word equally. Getting “the” wrong costs the same as getting a surname or a dosage wrong. But those errors are not equally costly to you — which is why a transcript with a good WER can still need careful checking in exactly the places that matter.
It says nothing about who was speaking. Attribution is scored separately, by diarization error rate. A transcript can have excellent words and useless speaker labels at the same time; they are genuinely different problems.
Why headline numbers do not survive contact with your recording
Benchmark figures are measured on curated datasets: usually one speaker, close to a decent microphone, in a quiet room, speaking a well-represented accent, using ordinary vocabulary. Under those conditions the best systems are genuinely excellent, and the marketing claims are broadly fair.
Your recording is probably not that. In rough order of how much damage each one does:
- Distance from the microphone. The single biggest factor. Sound falls off fast, and a phone at the far end of a table is capturing a much weaker signal than one in the middle.
- Background noise. Air conditioning, a fan, traffic, a café. Steady noise you have stopped noticing is still competing with every word.
- Overlapping speech. When two people talk at once, both the words and the speaker labels degrade together. This is the hardest problem in the field.
- Accents and dialects. Model performance varies with how well represented an accent is in training data — a real and well-documented effect.
- Specialist vocabulary. Names, medical and legal terms, product names, acronyms. These are what transcripts get wrong most, and unfortunately often what you needed.
Notice that four of those five are decided before you press record.
What to realistically expect
Some honest bands, for ordinary use rather than benchmark conditions:
One voice, phone on the desk, quiet room. Very good. Light editing, mostly proper nouns. This is the case Apple’s free built-in transcript already handles well.
Two people, phone between them, normal room. Good. Some corrections, mostly at the moments people interrupted each other.
Four or more people, one phone on a meeting table. Usable and a large time saving, but plan to read it through. Expect a handful of misattributed short interjections and some wrong names.
Anything recorded from a bag, a pocket, or across a large room. Poor, and no software fixes it. This is the case where people conclude transcription “doesn’t work” — the recording, not the model, is the problem.
How to get more accuracy for free
Because the recording dominates, technique buys you more than switching products does:
- Move the microphone closer and centre it. Nothing else on this list comes close.
- Face up on a hard surface, never under paper or in fabric.
- Turn off the fan, shut the window, move away from the coffee machine.
- Say unusual names clearly, once, early — it gives you a reliable place to correct from.
- Let people finish their sentences, at least roughly.
Placement matters most of all when several people are talking — we cover that case in how to transcribe audio with multiple speakers, and the nine changes that move the number most in how to improve speech-to-text accuracy.
The part no tool removes
Every transcript needs a read-through before you rely on it. That is true of ours, and of every alternative, and any product implying otherwise is overselling. What good transcription actually buys you is not perfection — it is going from re-listening to an hour of audio to skimming and correcting a page of text, with speaker labels telling you who said what.
That is a large, genuine saving. It is just a different promise from “you will never check it again”. If you want to see where the line falls on your own audio, the free tier includes three transcriptions and needs no card — testing it on a real recording of yours will tell you more than any accuracy figure will.
Accuracy is also what you are paying for when you pay more: what transcription costs explains why a human service charges by the audio minute, and when that is worth it.
The companion to accuracy is time. How long it takes to transcribe an hour of audio sets out the real numbers for both manual and automatic transcription — and makes the case that the bottleneck has moved from typing to checking, which is exactly the read-through described above.