Record and Transcribe

1 September 2026 · By Heni Hazbay

How to Clean Up a Transcript Without Changing What Was Said

An automatic transcript is a draft. The fastest order to fix it in, the errors worth hunting, and the edits that quietly turn a record into a misquotation.

The short answer: Fix it in passes, in this order — speaker attribution, then names and jargon, then punctuation and paragraphs, then a read against the audio. Attribution errors first, because everything downstream depends on knowing who said what. And decide your verbatim level before you start, since removing filler words cannot be undone without the recording.

An automatic transcript is a first draft that arrived quickly. Treating it as finished is how misquotations get published; treating it as worthless is how people waste an afternoon retyping something that was 95% right. The useful posture is in between, and it is mostly a matter of knowing where the errors actually live.

Pass 1: speaker attribution

Do this first. Every other correction is easier when you know who is talking, and attribution errors are the ones that change meaning rather than polish it.

Speaker errors do not scatter randomly — they cluster at turn boundaries, the moments where one person stops and another starts. So do not read for wrong names. Read for wrong splits:

  • A single labelled block that contains both a question and its answer is two speakers merged into one. This is the most common failure by a wide margin.
  • A speaker who changes mid-paragraph without a new label.
  • A short interjection (“mm”, “right”, “exactly”) absorbed into the other person’s turn.

Fix the boundary, then relabel the whole block in one go rather than line by line.

Two situations make this worse and are worth knowing in advance. Similar voices — two speakers of the same gender and similar pitch — are genuinely hard for any system, human or automatic. And platform transcripts label by account, not by voice: a Zoom or Teams transcript names whoever’s account the audio came from, so two people sharing a laptop appear as one person throughout, and a dial-in shows as a phone number. Check the attendee list against the labels before you trust any of it.

If you are wondering why the transcript uses “Speaker A” rather than names at all, how speaker diarization works explains the distinction — the system separated the voices without being told whose they are. Renaming them is your job, and it is a two-minute one once the boundaries are right.

Pass 2: proper nouns and jargon

This is where automatic transcription reliably fails, and it fails on exactly the words you cared about: people’s names, company names, product names, technical terms, acronyms, place names, numbers.

The efficient method is find-and-replace, not reading. Before you start, write down the names you expect to appear — attendees, organisations, products discussed. Then search for each one, including its likely mishearings. A name transcribed wrongly is usually transcribed wrongly consistently, so one replacement fixes twenty instances.

Numbers deserve their own pass if any of them matter. Dates, figures, dosages, clause references and prices are high-consequence and easy to mishear, and they do not look wrong on the page the way a mangled name does.

Pass 3: punctuation and paragraphs

Automatic transcripts tend towards long, under-punctuated runs, because the model is segmenting on pauses rather than on syntax. Sentence boundaries land in odd places and questions frequently arrive without a question mark.

Break the wall of text into paragraphs at topic changes. This is the change that does most for readability and the one least likely to alter meaning — but it is genuinely an editorial act, so keep it structural. Moving a sentence break can change emphasis; moving a paragraph break rarely does.

If you need timestamps for citation, add them at paragraph level rather than per line. Per-line timestamps make a transcript hard to read for a precision almost nobody uses.

Pass 4: read it against the audio

Not all of it. Read the passages you intend to quote, the numbers, and anything that reads oddly — because “reads oddly” is a reliable signal that the transcript diverged from the recording.

Two things you will only catch here:

  • Confident errors. Recognition does not flag uncertainty in the output. A misheard phrase arrives looking exactly as authoritative as a correct one, and the only tell is that it does not quite make sense.
  • Overlapping speech. Where two people talked at once, the transcript will usually have attributed the passage to one of them or lost it. Mark it, transcribe it by hand, and note the overlap rather than presenting a tidy version of something that was not tidy.

What not to change

The edits below feel like improvements and are not.

Do not silently fix grammar in anything that will be quoted. If a speaker said “there was three of them”, that is what they said. Correcting it makes them sound different from how they sounded, and you have now put words in someone’s mouth on their behalf. If readability genuinely matters more than fidelity, produce an intelligent verbatim transcript and say so.

Do not delete a passage because it is unclear. Mark it [inaudible] with a timestamp. A gap you have flagged is honest; a gap you have closed is a fabrication.

Do not tidy contradictions. People change their minds mid-sentence and say things they do not mean. That is what the record says.

Do not guess at a name. [inaudible name] is better than a plausible wrong one, which will propagate into everything downstream.

Do not remove filler words if you might need full verbatim. This one is irreversible without going back to the audio, which is why the verbatim level is a decision to make before transcription rather than during cleanup. For research this belongs in the protocol rather than in the transcription session — see what to decide before a qualitative study’s first interview.

Reduce the work at the source

Most of a cleanup is caused by the recording rather than the software. Distance from the microphone, a table between the phone and the speaker, an air conditioner, and people talking over each other account for a large share of the errors you will spend the afternoon fixing.

Improving transcription accuracy covers the recording-side habits that pay for themselves within one session. And an app that separates speakers as it transcribes removes pass 1 almost entirely — that is what our own recorder does on every memo, which is less a sales point than an explanation of why we think attribution belongs at the start of the pipeline rather than the end of it.

The order, one more time

PassLook forWhy first or last
1. AttributionMerged turns at speaker boundariesEverything else depends on it
2. Names and numbersConsistent mishearings, figuresFind-and-replace is fast
3. StructureParagraphs, question marksBig readability gain, low risk
4. Against audioQuotes, oddities, overlapsOnly catches what the eye cannot
Written by Heni Hazbay, the independent developer of Record and Transcribe. These guides come from building the recording and transcription pipeline they describe.

Frequently asked questions

How do I clean up an automatic transcript?

Work in passes rather than line by line. Fix speaker attribution first, then proper nouns and jargon, then punctuation and paragraphing, and read it against the audio last. Editing in one pass means re-reading the whole transcript every time you find a new class of error.

How do I fix wrong speaker labels in a transcript?

Find the boundaries rather than the labels. Speaker errors cluster at turn changes, so scan for places where a single labelled block contains a question and its answer — that is nearly always two speakers merged into one. Fix the split, then relabel the whole block at once.

How do I clean up a Zoom or Teams transcript?

The same order applies, with one addition: platform transcripts label speakers by account name, so anyone sharing a device or dialling in appears under the wrong name throughout. Check who was actually in the room before trusting the attribution.

Should I remove filler words from a transcript?

It depends on the verbatim level you need. Clean verbatim removes them and is the normal default; full verbatim keeps them because hesitation is data. Decide which you are producing before you start, because removing fillers is not reversible without the audio.

Is it acceptable to fix someone's grammar in a transcript?

Only if you are producing an intelligent verbatim transcript and are not presenting the result as a direct quotation. Silently correcting grammar in a document that will be quoted misrepresents the speaker, even when the intention is generous.

How do I transcribe overlapping speech?

Automatic transcription generally cannot separate simultaneous speakers and will attribute the overlap to one of them or drop it. Mark the passage, return to the audio, and transcribe it by hand — noting the overlap explicitly rather than pretending it resolved cleanly.

Ready when you are.

Free to start — 3 transcriptions included, no card needed. Apple Watch app comes with it.

Get the beta on TestFlight Free while in beta. Needs Apple’s TestFlight app — it installs it for you.