28 July 2026 · By Heni Hazbay
How to Label Speakers in a Transcript (Formats That Work)
Speaker labels are the difference between a transcript and a record. The standard formats, when to use names instead of letters, and how to fix wrong labels.
The short answer: Put the speaker identifier at the start of the line, followed by a colon, and use the same label for the same person every single time. Names beat letters when your readers know the people; consistency beats everything.
Speaker labels look like formatting. They are actually the thing that turns a recording into a record — the difference between text you can quote from and text you have to re-listen to. Here is how to label a transcript properly, whether the labels came from software or you are adding them by hand.
The standard format
Across journalism, research, and legal transcription, the same basic shape has settled in:
INTERVIEWER: When did you first notice the problem? PRIYA: About a week before we launched. Maybe ten days. INTERVIEWER: And you raised it then?
The elements that matter:
- Identifier first, then a colon. This is what lets a reader’s eye run down the left edge and follow one person’s thread.
- Visually distinct labels. Capitals or bold. The point is that labels read as structure, not as speech.
- One speaker per paragraph. A new speaker always starts a new line, however short their contribution.
- The same label every time. “PRIYA”, then “Priya”, then “P.” reads as three people to anyone skimming, and breaks any search you run later.
That last point is the one people get wrong, and it costs the most.
Names or Speaker A?
Software hands you generic labels because it genuinely cannot know names. Diarization clusters voices by how they sound; nothing in that process reveals identity. So the default output is Speaker A, Speaker B, Speaker C, consistently applied.
Whether you replace those with names is a judgement call:
Use real names when the transcript is for people who were in the room or know the participants — meeting minutes, a team’s own notes, an interview you are quoting in a piece with attribution. Names are dramatically easier to read, and the whole point of a meeting record is knowing who committed to what.
Keep generic labels when the transcript will circulate beyond the people involved, when a participant was promised anonymity, or when you are doing research where neutral identifiers are the convention. In that case, be systematic: PARTICIPANT 1, PARTICIPANT 2, and a separate key held somewhere else if you need to map them back.
Use role labels when the role matters more than the person — INTERVIEWER and RESPONDENT, CHAIR and MEMBER. This is common in qualitative research and in formal minutes, and it survives being read years later by someone who never knew the names.
Fixing labels that came back wrong
Automatic labelling gets the structure right and the edges wrong. Rather than reading suspiciously line by line, look for the three failure patterns — they cluster, so finding one usually finds several.
Two people merged into one label. The most common failure, and it almost always has a physical cause: two similar voices, or one person sitting much further from the microphone. Search for the merged label and check the passages where the conversation moves fastest.
Stray single-word lines. “Yeah”, “Right”, “Mm” carry almost no acoustic information and frequently land on the wrong speaker. They are also usually harmless — decide whether they are worth fixing at all for your purpose.
A label that switches partway. Occasionally one person picks up a new identifier mid-recording, typically after a long silence or a move around the room. Fix the first occurrence and the pattern is usually obvious from there.
The efficient method is one pass from the top with the audio ready to scrub, correcting the first instance of each pattern rather than hunting every instance individually.
Conventions worth deciding up front
Before a long transcription job — especially a set of interviews you will compare later — settle these and write them down:
- Verbatim or cleaned? Do you keep “um”, false starts and repetitions, or tidy them away? Both are legitimate; mixing them across a set of transcripts is not.
- How to mark overlap. Square brackets around overlapping words, or a [crosstalk] note. Pick one.
- How to mark inaudible passages. [inaudible] with a timestamp is standard, and the timestamp is what makes it recoverable later.
- How to handle non-speech. Laughter, long pauses, someone leaving the room — note them only if they change the meaning.
Consistency is not pedantry here. It is what makes two transcripts comparable, and what stops a reader misreading a convention as content.
Getting labels without doing it by hand
Adding labels manually to an hour of conversation is slow work, and it is the part software genuinely does well. Record and Transcribe (our app) applies speaker labels automatically to every recording as part of transcription, with no voice enrolment and nothing to configure — you get a labelled transcript and rename the speakers once.
The honest boundary: if you are recording only yourself, none of this applies and Apple’s free built-in transcript is all you need. Labels earn their keep the moment there is a second voice — which is also the moment a transcript stops being a convenience and starts being a record. For the practical side of capturing several voices well, see how to transcribe audio with multiple speakers.