Record and Transcribe

7 August 2026 · Updated 24 August 2026 · By Heni Hazbay

How Accurate Is AI Transcription? An Honest Answer

What the “99% accurate” claims actually measure, how word error rate works, and what to realistically expect from a recording of a real conversation.

The short answer: Accuracy is set mostly by your recording, not your software. On clean audio of one clear voice, modern transcription is very good. On a four-person meeting recorded from across a table, expect to correct names, jargon and anything said over the top of someone else.

“99% accurate” appears on almost every transcription product’s homepage, ours included in spirit. It is worth understanding what that number measures, because the gap between the claim and your experience is usually not the software being dishonest — it is your audio being different from the audio it was measured on.

How accuracy is actually measured

The standard metric is word error rate (WER). Take a recording, produce a careful human transcript as the reference, run the software over the same audio, and count what went wrong:

  • Substitutions — a word transcribed as a different word.
  • Deletions — a word spoken but missing.
  • Insertions — a word appearing that nobody said.

Add those together and divide by the number of words actually spoken. A WER of 5% means one word in twenty is wrong somehow, which is another way of saying 95% accurate.

Two things about WER are worth knowing before you compare any two products.

It treats every word equally. Getting “the” wrong costs the same as getting a surname or a dosage wrong. But those errors are not equally costly to you — which is why a transcript with a good WER can still need careful checking in exactly the places that matter.

It says nothing about who was speaking. Attribution is scored separately, by diarization error rate. A transcript can have excellent words and useless speaker labels at the same time; they are genuinely different problems.

Why headline numbers do not survive contact with your recording

Benchmark figures are measured on curated datasets: usually one speaker, close to a decent microphone, in a quiet room, speaking a well-represented accent, using ordinary vocabulary. Under those conditions the best systems are genuinely excellent, and the marketing claims are broadly fair.

Your recording is probably not that. In rough order of how much damage each one does:

  1. Distance from the microphone. The single biggest factor. Sound falls off fast, and a phone at the far end of a table is capturing a much weaker signal than one in the middle.
  2. Background noise. Air conditioning, a fan, traffic, a café. Steady noise you have stopped noticing is still competing with every word.
  3. Overlapping speech. When two people talk at once, both the words and the speaker labels degrade together. This is the hardest problem in the field.
  4. Accents and dialects. Model performance varies with how well represented an accent is in training data — a real and well-documented effect.
  5. Specialist vocabulary. Names, medical and legal terms, product names, acronyms. These are what transcripts get wrong most, and unfortunately often what you needed.

Notice that four of those five are decided before you press record.

What to realistically expect

Some honest bands, for ordinary use rather than benchmark conditions:

One voice, phone on the desk, quiet room. Very good. Light editing, mostly proper nouns. This is the case Apple’s free built-in transcript already handles well.

Two people, phone between them, normal room. Good. Some corrections, mostly at the moments people interrupted each other.

Four or more people, one phone on a meeting table. Usable and a large time saving, but plan to read it through. Expect a handful of misattributed short interjections and some wrong names.

Anything recorded from a bag, a pocket, or across a large room. Poor, and no software fixes it. This is the case where people conclude transcription “doesn’t work” — the recording, not the model, is the problem.

How to get more accuracy for free

Because the recording dominates, technique buys you more than switching products does:

  • Move the microphone closer and centre it. Nothing else on this list comes close.
  • Face up on a hard surface, never under paper or in fabric.
  • Turn off the fan, shut the window, move away from the coffee machine.
  • Say unusual names clearly, once, early — it gives you a reliable place to correct from.
  • Let people finish their sentences, at least roughly.

Placement matters most of all when several people are talking — we cover that case in how to transcribe audio with multiple speakers, and the nine changes that move the number most in how to improve speech-to-text accuracy.

The part no tool removes

Every transcript needs a read-through before you rely on it. That is true of ours, and of every alternative, and any product implying otherwise is overselling. What good transcription actually buys you is not perfection — it is going from re-listening to an hour of audio to skimming and correcting a page of text, with speaker labels telling you who said what.

That is a large, genuine saving. It is just a different promise from “you will never check it again”. If you want to see where the line falls on your own audio, the free tier includes three transcriptions and needs no card — testing it on a real recording of yours will tell you more than any accuracy figure will.

Accuracy is also what you are paying for when you pay more: what transcription costs explains why a human service charges by the audio minute, and when that is worth it.

The companion to accuracy is time. How long it takes to transcribe an hour of audio sets out the real numbers for both manual and automatic transcription — and makes the case that the bottleneck has moved from typing to checking, which is exactly the read-through described above.

Written by Heni Hazbay, the independent developer of Record and Transcribe. These guides come from building the recording and transcription pipeline they describe.

Frequently asked questions

What does word error rate mean?

Word error rate, or WER, is the standard measure of transcription accuracy. It counts three kinds of mistake against a human reference transcript — words substituted, words deleted, and words inserted that were never said — then divides the total by the number of words actually spoken. A WER of 5% means one word in twenty is wrong in some way. Lower is better.

Is AI transcription really 99% accurate?

On the audio those claims are measured against, often yes. That audio is typically a single speaker, close to a good microphone, in a quiet room, speaking a common accent without jargon. Real recordings rarely look like that, and accuracy on a four-person meeting in a room with background noise is meaningfully lower. The number is not dishonest, but it describes conditions rather than a guarantee.

What makes transcription accuracy worse?

In rough order of impact: distance from the microphone, background noise, people talking over each other, strong or unfamiliar accents, and specialist vocabulary such as names, drug names and product names. Almost all of these are decided when you record, not when you transcribe.

Is AI transcription more accurate than a human?

On clean audio the gap has largely closed, and software is far faster and cheaper. Humans remain better on genuinely difficult audio — heavy crosstalk, strong accents, poor recordings — because they use context and world knowledge that speech models handle less reliably. For work where an error carries real consequences, the usual approach is automatic transcription followed by a human check.

Does transcription accuracy differ between languages?

Yes, substantially. Widely spoken languages with large amounts of training data are transcribed considerably more accurately than less represented ones, and regional accents within a language vary too. If you work in a less common language, test on your own audio rather than trusting a headline figure, which will almost certainly have been measured in English.

Ready when you are.

Free to start — 3 transcriptions included, no card needed. Apple Watch app comes with it.

Get the beta on TestFlight Free while in beta. Needs Apple’s TestFlight app — it installs it for you.