Record and Transcribe

28 July 2026 · By Heni Hazbay

How to Transcribe Audio With Multiple Speakers

Transcribing a conversation is a different job from transcribing one voice. How to get a transcript that separates speakers, and what decides whether it works.

The short answer: Transcribing several people is a different task from transcribing one, and it needs a tool that does speaker diarization. Get the microphone equidistant from everyone, avoid crosstalk, and the transcript comes back attributed line by line.

Most transcription advice quietly assumes one person talking into a microphone. The moment a second voice appears, the job changes — and the tools that were perfectly good at dictation start producing something you cannot use. Here is what actually changes, and how to get a transcript of a conversation that is worth having.

Why a conversation is a harder problem

A single voice needs only speech recognition: turn sound into words. A conversation needs that plus attribution — working out which stretch of audio belongs to which person. That second job is called speaker diarization, and it is what separates a usable transcript from a wall of text.

The difference shows up immediately. Without attribution you get this:

right so we can’t ship both features by the 14th we could if we cut the onboarding flow no let’s not do that again fine then the 21st

With it, you get a record:

Speaker A: Right, so we can’t ship both features by the 14th. Speaker B: We could, if we cut the onboarding flow. Speaker A: No, let’s not do that again. Speaker B: Fine. Then the 21st.

Identical words. Only the second one tells you who agreed to what.

The recording decides most of it

This is the part people skip, and it matters more than which app you choose. A diarization model separates voices by their acoustic fingerprint, so anything that blurs those fingerprints costs you accuracy.

Put the microphone in the middle. The single biggest predictor of a good multi-speaker transcript is that everyone is roughly the same distance from the microphone. One person twice as far away as everyone else is the classic cause of a speaker who keeps getting merged into someone else.

Get it off soft surfaces. On the table, face up. Not on a notepad, not in a bag, not in a shirt pocket. Fabric muffles consonants, and consonants carry most of the information.

Kill the steady noise you stopped noticing. Air conditioning, a projector fan, a fridge. You have tuned it out; the model has not.

Ask people to leave gaps. Crosstalk is the single hardest case in the field. When two voices overlap, both the words and the attribution degrade at once. You do not need a formal turn-taking protocol — just not everyone talking at once.

Say names early. “Thanks for joining, Priya” near the start gives you an anchor for relabelling later, and helps you spot immediately if two people have been merged.

What to use

Apple’s built-in transcript is free and right there on iOS 18 and later, and for a recording of yourself it is genuinely enough. For a conversation it returns one unbroken block of text with no attribution, which is where it stops being useful.

Meeting bots — Otter, Fireflies and similar — handle attribution well, because a video call hands them each participant on a separate audio stream. That is a real structural advantage, and if your conversations happen on Zoom or Teams they are the right tool. They cannot help you in a room. Our comparison with Otter covers where each one wins.

A recorder that diarizes on device audio is what you need for in-person conversation. Record and Transcribe (our app) transcribes every recording automatically when you stop, with each line labelled by speaker and no upload step. That covers the cases a bot cannot reach — meetings in a room, interviews, a conversation over coffee.

What to expect, honestly

Good audio of a four-person conversation with normal turn-taking produces a transcript that needs light tidying: fix a few names, merge a couple of stray label switches, and it is ready. That is the realistic best case, and it is very good.

Three things will still need your eye:

  • Very short interjections. “Yeah.” and “Right.” carry almost no acoustic information, and often get attached to whoever was speaking around them.
  • The moments people talk over each other. Expect both words and labels to be shakiest exactly where the conversation was liveliest.
  • Names, jargon and acronyms. These are the most common word-level errors in any transcript, regardless of how many people are speaking.

Nothing on the market removes the final read-through. Any tool promising a perfect transcript of a real meeting is selling you something, and the honest version of the promise is much better than it sounds: you go from re-listening to an hour of audio to skimming and correcting a page of text.

Turning it into something useful

A transcript is raw material, not the deliverable. For a meeting, what people actually want is the decisions and the actions with an owner — we cover that method, with a template, in how to write meeting minutes from a recording. For an interview, it is accurate attributed quotes.

Either way, speaker labels are what make the step possible at all. Without them you are back to replaying audio to work out who committed to what, which is the exact job you recorded the conversation to avoid.

If you want to judge it on your own audio rather than take our word for it, the free tier includes three transcriptions and needs no card.

Written by Heni Hazbay, the independent developer of Record and Transcribe. These guides come from building the recording and transcription pipeline they describe.

Frequently asked questions

Can AI transcription tell who is speaking?

It can tell speakers apart, but not name them. The process is called speaker diarization: the software groups the audio by voice and labels each one consistently as Speaker A, Speaker B and so on. Attaching real names is something you do once, afterwards, unless you have enrolled each person’s voice in advance.

What is the most accurate way to transcribe multiple speakers?

Record everyone from one central microphone at conversational distance, ask people not to talk over each other, and use a transcription tool that does diarization rather than plain speech-to-text. Accuracy is set far more by the recording than by the software — the same tool will produce an excellent transcript of a well-placed phone and a poor one of a phone in a bag.

Why does my transcript merge two people into one speaker?

Usually because their voices are acoustically similar, one of them speaks only in short bursts, or one sits much further from the microphone than the other. All three make the voice fingerprints harder to separate. Moving the microphone to an equal distance from everyone fixes more of this than any setting.

Does Apple’s built-in transcript separate speakers?

No. The transcript in Voice Memos on iOS 18 and later returns one continuous block of text with no attribution, which is fine for your own notes and hard to use for a conversation. Speaker separation is what a dedicated transcription app adds.

How many speakers can be transcribed at once?

Modern systems handle a normal meeting of a handful of people comfortably. Accuracy tends to fall as the number grows, because more voices mean more chances that two of them sound alike and more crosstalk overall. Four people around a table is routine; twelve people in a large room is genuinely hard for any tool.

Ready when you are.

Free to start — 3 transcriptions included, no card needed. Apple Watch app comes with it.

Get the beta on TestFlight Free while in beta. Needs Apple’s TestFlight app — it installs it for you.