Record and Transcribe

28 August 2026 · By Heni Hazbay

Verbatim Transcription: Clean, Full and Intelligent Explained

Verbatim is not one style but several, and picking the wrong one wastes hours or destroys your data. What each level captures, and who actually needs which.

The short answer: Verbatim transcription means word-for-word rather than summarised — but it comes in levels. Full verbatim keeps every stammer, filler and pause. Clean verbatim keeps every word but removes the disfluency. Intelligent verbatim additionally tidies grammar for readability. They are not interchangeable, and the right one is decided by what you will do with the transcript.

If you have ever ordered transcription and been asked to choose a “verbatim level”, or read a methods section that specified one, this is what was being asked. Getting it wrong in either direction is costly: too much detail and you have paid for hours of unusable noise, too little and you have quietly deleted your own evidence.

What verbatim transcription means

Verbatim transcription is the practice of recording exactly what was said, word for word, rather than paraphrasing or summarising it. The transcript is a record of the speech itself rather than of its content.

That is the whole of the definition, and it is why the term needs qualifying. “Exactly what was said” turns out to be ambiguous, because speech contains a great deal that is not words.

Full verbatim (also called strict or true verbatim)

Full verbatim captures everything audible, including the parts speakers are not aware of producing. That means filler words, repetitions, false starts, stammers, and non-speech sounds noted in brackets.

Speaker A: So I — I think, um, the, the second option is — [laughs] — sorry, is probably better. [pause]

It is the slowest level to produce and the hardest to read. It is also the only level that preserves how something was said, which for some purposes is the entire point.

Who needs it: conversation analysts, discourse analysts, some clinical and forensic work, and legal transcription where the manner of an answer may matter as much as its content.

Clean verbatim (also called clean read)

Clean verbatim keeps every substantive word that was spoken but removes the disfluency around it. Filler words go, stammers and repeated words are resolved to one, false starts are dropped. Nothing that carries meaning is changed, and grammatical errors the speaker actually made are left in place.

Speaker A: So I think the second option is probably better.

This is the default for most professional transcription and, unless somebody specified otherwise, the level people usually mean when they say “verbatim”.

Who needs it: most interview research, journalism, business meetings, podcasts, and anything intended to be read by a human rather than analysed as speech.

Intelligent verbatim

Intelligent verbatim applies light editing on top of clean verbatim: grammar is tidied, redundant phrases are cut, and sentences that ran together are separated. The speaker’s meaning is preserved; their exact wording is not.

Speaker A: I think the second option is better.

Who needs it: published interviews, marketing and content work, executive summaries, and any case where readability matters more than fidelity.

The caution: because the wording changes, an intelligent verbatim transcript should not be quoted as though it were the speaker’s exact words. If your output is a direct quotation, do not use this level.

Non-verbatim, for completeness

Non-verbatim transcription summarises rather than records — it captures what was discussed and decided without preserving the wording. Meeting minutes are the everyday example: minutes are deliberately not a transcript, because their job is to record decisions rather than conversation.

Choosing, in one table

LevelKeeps disfluencyKeeps exact wordingBest for
Full verbatimYesYesConversation analysis, legal, clinical
Clean verbatimNoYesResearch interviews, journalism, meetings
Intelligent verbatimNoNoPublication, readability
Non-verbatimNoNoMinutes, summaries

Decide before you transcribe, not after

This is the practical point the definitions exist to serve. You can always go from a fuller level to a lighter one — dropping fillers from a full verbatim transcript is straightforward, and most of a transcript cleanup is exactly that work.

You cannot go the other way. The disfluencies are not recoverable from a clean transcript; you would have to return to the audio and start again. So if there is any chance your analysis will care about hesitation, keep the audio, and think about the level before the transcription happens rather than after.

In qualitative research this is a methods decision that belongs in your protocol, not a preference — and one that is much cheaper to make before fieldwork than after it. It is one of three decisions to settle before the first interview, alongside anonymisation and where the audio is processed.

What automatic transcription gives you

Automatic speech recognition produces something close to clean verbatim by default. Most engines drop the majority of filler words on their own, and they do not mark pauses, laughter or overlapping speech — the information simply is not in the output.

That is convenient if clean verbatim is what you wanted, and it is a genuine limitation if you needed full verbatim. No automatic transcript is a substitute for full verbatim work, and any tool claiming otherwise is worth checking against your own audio before you rely on it. Where automatic transcription does hold up well is the other axis entirely — knowing who was speaking, which human transcription charges extra for and which our own app includes on every recording.

If you are weighing whether an automatic transcript is faithful enough for your purpose at all, how accurate AI transcription really is is the honest version of that answer.

Written by Heni Hazbay, the independent developer of Record and Transcribe. These guides come from building the recording and transcription pipeline they describe.

Frequently asked questions

What does verbatim transcription mean?

Verbatim transcription means writing down exactly what was said, word for word, rather than summarising or tidying it. In practice the term covers several levels that differ in how much of the non-word content — stammers, filler words, pauses, laughter — is also recorded.

What is full verbatim transcription?

Full verbatim captures everything audible: every "um", repeated word, false start, stutter, laugh, cough and audible pause, usually with non-speech sounds noted in brackets. It is the most complete and the most time-consuming level to produce.

What is clean verbatim transcription?

Clean verbatim records every word that was spoken but removes the noise around them — filler words, stammers, false starts and repeated words are dropped, and obvious slips are left as spoken. The meaning and the wording are preserved; the disfluency is not.

What is intelligent verbatim transcription?

Intelligent verbatim goes a step further than clean verbatim and lightly edits for readability — tidying grammar, removing redundant phrases, and smoothing sentences that ran together. It is the most readable level and the least faithful to the exact wording.

What is the difference between clean verbatim and full verbatim?

Full verbatim keeps the disfluencies; clean verbatim removes them. Both keep every substantive word. A sentence transcribed as "I — I think, um, we should wait" in full verbatim becomes "I think we should wait" in clean verbatim.

Which verbatim level should I use for research?

It depends on your analysis. Conversation analysis and discourse analysis need full verbatim, because pauses and hesitations are data. Thematic analysis of interview content is normally done on clean verbatim. Decide before transcription starts, not after.

Ready when you are.

Free to start — 3 transcriptions included, no card needed. Apple Watch app comes with it.

Get the beta on TestFlight Free while in beta. Needs Apple’s TestFlight app — it installs it for you.