Record and Transcribe

28 July 2026 · By Heni Hazbay

How to Label Speakers in a Transcript (Formats That Work)

Speaker labels are the difference between a transcript and a record. The standard formats, when to use names instead of letters, and how to fix wrong labels.

The short answer: Put the speaker identifier at the start of the line, followed by a colon, and use the same label for the same person every single time. Names beat letters when your readers know the people; consistency beats everything.

Speaker labels look like formatting. They are actually the thing that turns a recording into a record — the difference between text you can quote from and text you have to re-listen to. Here is how to label a transcript properly, whether the labels came from software or you are adding them by hand.

The standard format

Across journalism, research, and legal transcription, the same basic shape has settled in:

INTERVIEWER: When did you first notice the problem? PRIYA: About a week before we launched. Maybe ten days. INTERVIEWER: And you raised it then?

The elements that matter:

  • Identifier first, then a colon. This is what lets a reader’s eye run down the left edge and follow one person’s thread.
  • Visually distinct labels. Capitals or bold. The point is that labels read as structure, not as speech.
  • One speaker per paragraph. A new speaker always starts a new line, however short their contribution.
  • The same label every time. “PRIYA”, then “Priya”, then “P.” reads as three people to anyone skimming, and breaks any search you run later.

That last point is the one people get wrong, and it costs the most.

Names or Speaker A?

Software hands you generic labels because it genuinely cannot know names. Diarization clusters voices by how they sound; nothing in that process reveals identity. So the default output is Speaker A, Speaker B, Speaker C, consistently applied.

Whether you replace those with names is a judgement call:

Use real names when the transcript is for people who were in the room or know the participants — meeting minutes, a team’s own notes, an interview you are quoting in a piece with attribution. Names are dramatically easier to read, and the whole point of a meeting record is knowing who committed to what.

Keep generic labels when the transcript will circulate beyond the people involved, when a participant was promised anonymity, or when you are doing research where neutral identifiers are the convention. In that case, be systematic: PARTICIPANT 1, PARTICIPANT 2, and a separate key held somewhere else if you need to map them back.

Use role labels when the role matters more than the person — INTERVIEWER and RESPONDENT, CHAIR and MEMBER. This is common in qualitative research and in formal minutes, and it survives being read years later by someone who never knew the names.

Fixing labels that came back wrong

Automatic labelling gets the structure right and the edges wrong. Rather than reading suspiciously line by line, look for the three failure patterns — they cluster, so finding one usually finds several.

Two people merged into one label. The most common failure, and it almost always has a physical cause: two similar voices, or one person sitting much further from the microphone. Search for the merged label and check the passages where the conversation moves fastest.

Stray single-word lines. “Yeah”, “Right”, “Mm” carry almost no acoustic information and frequently land on the wrong speaker. They are also usually harmless — decide whether they are worth fixing at all for your purpose.

A label that switches partway. Occasionally one person picks up a new identifier mid-recording, typically after a long silence or a move around the room. Fix the first occurrence and the pattern is usually obvious from there.

The efficient method is one pass from the top with the audio ready to scrub, correcting the first instance of each pattern rather than hunting every instance individually.

Conventions worth deciding up front

Before a long transcription job — especially a set of interviews you will compare later — settle these and write them down:

  • Verbatim or cleaned? Do you keep “um”, false starts and repetitions, or tidy them away? Both are legitimate; mixing them across a set of transcripts is not.
  • How to mark overlap. Square brackets around overlapping words, or a [crosstalk] note. Pick one.
  • How to mark inaudible passages. [inaudible] with a timestamp is standard, and the timestamp is what makes it recoverable later.
  • How to handle non-speech. Laughter, long pauses, someone leaving the room — note them only if they change the meaning.

Consistency is not pedantry here. It is what makes two transcripts comparable, and what stops a reader misreading a convention as content.

Getting labels without doing it by hand

Adding labels manually to an hour of conversation is slow work, and it is the part software genuinely does well. Record and Transcribe (our app) applies speaker labels automatically to every recording as part of transcription, with no voice enrolment and nothing to configure — you get a labelled transcript and rename the speakers once.

The honest boundary: if you are recording only yourself, none of this applies and Apple’s free built-in transcript is all you need. Labels earn their keep the moment there is a second voice — which is also the moment a transcript stops being a convenience and starts being a record. For the practical side of capturing several voices well, see how to transcribe audio with multiple speakers.

Written by Heni Hazbay, the independent developer of Record and Transcribe. These guides come from building the recording and transcription pipeline they describe.

Frequently asked questions

What is the proper format for a speaker label?

The widely used convention is the speaker’s name or identifier, followed by a colon, at the start of the line, with the speech on the same line or the line below — for example "INTERVIEWER: How did you first hear about it?". Labels are conventionally in capitals or bold so the eye can skip down them, and the same label must be used for the same person throughout.

Should I use names or Speaker A and Speaker B?

Use names whenever you know them and the transcript is for people who know them too — it is far easier to read. Keep generic labels when the transcript will be shared outside the room, when a participant was promised anonymity, or in research where identifiers are deliberately neutral. Software gives you Speaker A and Speaker B by default because it has no way to learn names on its own.

How do I fix speaker labels that are wrong?

Work through the transcript once from the top, because errors cluster rather than scatter. Watch for two people merged into one label — usually caused by similar voices or unequal distance from the microphone — and single-word lines like "Yeah" attached to the wrong person. Fixing the first occurrence of each pattern usually resolves most of the rest.

How do you label an unknown speaker?

Give them a stable, descriptive placeholder rather than leaving them blank: UNKNOWN 1, or something functional like SECOND JOURNALIST. The rule is consistency — one label per voice for the whole transcript — so a reader can follow a thread even without a name attached to it.

How should overlapping speech be marked?

Most conventions give each speaker their own line and mark the overlap explicitly, commonly with square brackets around the overlapping words or a bracketed note such as [crosstalk]. What matters is picking one convention and applying it identically throughout, so a reader can tell an overlap from an interruption.

Ready when you are.

Free to start — 3 transcriptions included, no card needed. Apple Watch app comes with it.

Get the beta on TestFlight Free while in beta. Needs Apple’s TestFlight app — it installs it for you.