1 September 2026 · By Heni Hazbay
How to Clean Up a Transcript Without Changing What Was Said
An automatic transcript is a draft. The fastest order to fix it in, the errors worth hunting, and the edits that quietly turn a record into a misquotation.
The short answer: Fix it in passes, in this order — speaker attribution, then names and jargon, then punctuation and paragraphs, then a read against the audio. Attribution errors first, because everything downstream depends on knowing who said what. And decide your verbatim level before you start, since removing filler words cannot be undone without the recording.
An automatic transcript is a first draft that arrived quickly. Treating it as finished is how misquotations get published; treating it as worthless is how people waste an afternoon retyping something that was 95% right. The useful posture is in between, and it is mostly a matter of knowing where the errors actually live.
Pass 1: speaker attribution
Do this first. Every other correction is easier when you know who is talking, and attribution errors are the ones that change meaning rather than polish it.
Speaker errors do not scatter randomly — they cluster at turn boundaries, the moments where one person stops and another starts. So do not read for wrong names. Read for wrong splits:
- A single labelled block that contains both a question and its answer is two speakers merged into one. This is the most common failure by a wide margin.
- A speaker who changes mid-paragraph without a new label.
- A short interjection (“mm”, “right”, “exactly”) absorbed into the other person’s turn.
Fix the boundary, then relabel the whole block in one go rather than line by line.
Two situations make this worse and are worth knowing in advance. Similar voices — two speakers of the same gender and similar pitch — are genuinely hard for any system, human or automatic. And platform transcripts label by account, not by voice: a Zoom or Teams transcript names whoever’s account the audio came from, so two people sharing a laptop appear as one person throughout, and a dial-in shows as a phone number. Check the attendee list against the labels before you trust any of it.
If you are wondering why the transcript uses “Speaker A” rather than names at all, how speaker diarization works explains the distinction — the system separated the voices without being told whose they are. Renaming them is your job, and it is a two-minute one once the boundaries are right.
Pass 2: proper nouns and jargon
This is where automatic transcription reliably fails, and it fails on exactly the words you cared about: people’s names, company names, product names, technical terms, acronyms, place names, numbers.
The efficient method is find-and-replace, not reading. Before you start, write down the names you expect to appear — attendees, organisations, products discussed. Then search for each one, including its likely mishearings. A name transcribed wrongly is usually transcribed wrongly consistently, so one replacement fixes twenty instances.
Numbers deserve their own pass if any of them matter. Dates, figures, dosages, clause references and prices are high-consequence and easy to mishear, and they do not look wrong on the page the way a mangled name does.
Pass 3: punctuation and paragraphs
Automatic transcripts tend towards long, under-punctuated runs, because the model is segmenting on pauses rather than on syntax. Sentence boundaries land in odd places and questions frequently arrive without a question mark.
Break the wall of text into paragraphs at topic changes. This is the change that does most for readability and the one least likely to alter meaning — but it is genuinely an editorial act, so keep it structural. Moving a sentence break can change emphasis; moving a paragraph break rarely does.
If you need timestamps for citation, add them at paragraph level rather than per line. Per-line timestamps make a transcript hard to read for a precision almost nobody uses.
Pass 4: read it against the audio
Not all of it. Read the passages you intend to quote, the numbers, and anything that reads oddly — because “reads oddly” is a reliable signal that the transcript diverged from the recording.
Two things you will only catch here:
- Confident errors. Recognition does not flag uncertainty in the output. A misheard phrase arrives looking exactly as authoritative as a correct one, and the only tell is that it does not quite make sense.
- Overlapping speech. Where two people talked at once, the transcript will usually have attributed the passage to one of them or lost it. Mark it, transcribe it by hand, and note the overlap rather than presenting a tidy version of something that was not tidy.
What not to change
The edits below feel like improvements and are not.
Do not silently fix grammar in anything that will be quoted. If a speaker said “there was three of them”, that is what they said. Correcting it makes them sound different from how they sounded, and you have now put words in someone’s mouth on their behalf. If readability genuinely matters more than fidelity, produce an intelligent verbatim transcript and say so.
Do not delete a passage because it is unclear. Mark it [inaudible] with a timestamp. A gap you have flagged is honest; a gap you have closed is a fabrication.
Do not tidy contradictions. People change their minds mid-sentence and say things they do not mean. That is what the record says.
Do not guess at a name. [inaudible name] is better than a plausible wrong one, which will propagate into everything downstream.
Do not remove filler words if you might need full verbatim. This one is irreversible without going back to the audio, which is why the verbatim level is a decision to make before transcription rather than during cleanup. For research this belongs in the protocol rather than in the transcription session — see what to decide before a qualitative study’s first interview.
Reduce the work at the source
Most of a cleanup is caused by the recording rather than the software. Distance from the microphone, a table between the phone and the speaker, an air conditioner, and people talking over each other account for a large share of the errors you will spend the afternoon fixing.
Improving transcription accuracy covers the recording-side habits that pay for themselves within one session. And an app that separates speakers as it transcribes removes pass 1 almost entirely — that is what our own recorder does on every memo, which is less a sales point than an explanation of why we think attribution belongs at the start of the pipeline rather than the end of it.
The order, one more time
| Pass | Look for | Why first or last |
|---|---|---|
| 1. Attribution | Merged turns at speaker boundaries | Everything else depends on it |
| 2. Names and numbers | Consistent mishearings, figures | Find-and-replace is fast |
| 3. Structure | Paragraphs, question marks | Big readability gain, low risk |
| 4. Against audio | Quotes, oddities, overlaps | Only catches what the eye cannot |