4 September 2026 · By Heni Hazbay
Transcription for Qualitative Research: What to Decide First
Verbatim level, anonymisation and file handling are methods decisions, not clerical ones. What to settle before the first interview, and why order matters here.
The short answer: Three decisions belong in your protocol before the first interview: which verbatim level your analysis requires, how transcripts will be anonymised, and where the audio is allowed to be processed. All three are hard to change afterwards, and the first one cannot be reversed without returning to the recordings.
Transcription gets treated as the clerical bit between fieldwork and analysis. It is not — it is where several irreversible methods decisions get made, usually by default and usually by whoever is doing the typing at 11pm.
Scope note: this describes common practice in qualitative research. It is not ethics or data-protection advice, and where this page and your ethics committee or institutional data policy disagree, they win. Ask them before the fieldwork, not after.
Decision 1: the verbatim level
This is the one people get wrong most often, because it looks like a formatting preference and is actually a decision about what counts as data.
If your analysis treats how something was said as meaningful — conversation analysis, discourse analysis, some narrative approaches — you need full verbatim. Pauses, overlaps, hesitations and repairs are the phenomena you are studying, and a transcript that tidies them has deleted your dataset.
If your analysis is about what was said — thematic analysis, framework analysis, content analysis — clean verbatim is the normal choice, and full verbatim adds hours of work and a much harder document to read for no analytic gain.
The asymmetry is what matters: you can always reduce a full verbatim transcript to a clean one, and you can never do the reverse without the audio. If you are genuinely unsure, or your analytic approach may change, transcribe fuller than you think you need and keep the recordings.
Decision 2: anonymisation, and when it happens
Anonymising at the transcription stage rather than afterwards is standard practice for a practical reason — it means no identifying version of the transcript ever exists to be mishandled.
The usual approach is a consistent pseudonym or participant code applied as the transcript is produced, with the mapping between code and real identity held separately and more securely than the transcripts themselves. Names are the obvious target; the ones that get missed are employers, job titles that are unique in a small field, place names, distinctive life events, and the names of third parties the participant mentions who never consented to anything.
A caution worth stating: anonymisation is not deletion. In a small population, a combination of unremarkable details can identify someone. Where your sample is a specific organisation or a small professional community, generalise rather than merely rename.
Decision 3: where the audio is processed
If you use automatic transcription — and many researchers now do — the recording of your participant leaves your device and is processed by a third party. That is a data-handling decision, not a tooling one.
The questions your committee will ask, and which you should ask the provider before you commit:
- Where is the audio processed, and in which jurisdiction is it stored?
- Is it retained after processing, and for how long?
- Is it used to train models?
- What does your consent form actually tell participants about this?
None of these has a universally right answer. What is not acceptable is discovering them after the interviews, because at that point the consent has already been taken on a description that may not match what happened.
Should you transcribe it yourself?
There is a genuine methodological argument for doing so, and it is not just thrift. Transcribing forces a slow, complete re-encounter with the data, and researchers commonly report that this is where they first notice a pattern — the thing the participant hesitated before saying, the question that landed badly.
The honest counter-argument is time. At roughly four hours per hour of clear audio, a study with twenty interviews is a substantial fraction of the project, and the returns fall off after the first few: by interview fifteen you are typing, not noticing.
A defensible middle path, and a common one: transcribe the first two or three yourself, then use automatic transcription corrected against the audio for the rest, and keep listening to the recordings during analysis rather than treating the transcript as the data. Cleaning up an automatic transcript is a much faster job than typing from scratch, and the correction pass preserves some of the immersion benefit.
Format: boring and consistent
Qualitative analysis software — NVivo, MAXQDA, Atlas.ti and the rest — imports plain, consistently structured text reliably and heavily formatted documents badly. Some can auto-detect speakers if your labels follow a strict pattern, which is worth checking before you transcribe twenty files in a pattern it does not recognise.
Practical rules that survive contact with any of them:
- One consistent speaker label format throughout the study, not per transcript. Speaker label conventions covers the standard forms.
- Timestamps at a stated granularity if you need them at all — paragraph level is enough for citation, and per-line timestamps make the document unreadable.
- A short header on each transcript: participant code, date, duration, interviewer, verbatim level.
- Mark unclear passages
[inaudible 00:14:32]rather than guessing. A flagged gap is data; a plausible invention is not.
Focus groups are harder, and it is worth planning for
Everything above gets more difficult when several people are in the room. Attribution becomes the dominant cost, overlapping speech is constant rather than occasional, and automatic transcription struggles precisely where the interaction is liveliest.
Two things help disproportionately: a recorder placed centrally on the table rather than near the moderator, and the moderator naming participants as they come in (“Yes, go ahead, P4”) — which gives both a human and an automatic transcript an anchor. Transcribing audio with multiple speakers covers the recording side in more detail.
Where our app fits
We make an iPhone and Apple Watch recorder that returns speaker-labelled transcripts, which removes the most tedious part of multi-speaker transcription. For interview and focus-group work that is a real saving.
It is a first-pass tool for research, not a replacement for a corrected transcript, and we would rather say that than imply otherwise. It also processes audio off the device, so it belongs in the Decision 3 conversation above — check it against your data-management plan before using it on participant recordings, the same as any other service.