Record and Transcribe

24 July 2026 · Updated 28 July 2026 · By Heni Hazbay

How to Transcribe an Interview on iPhone — Every Method

Type it, use Apple’s built-in transcript, pay a human, or let one app do it all: real costs of every interview transcription method, plus keeping quotes accurate.

The short answer: Four ways: type it yourself (an evening per hour of tape), Apple’s free built-in transcript (no speaker labels), a human service (highest accuracy, priced per minute), or an app that records and transcribes in one step. Whichever you choose, proof-read quotes before they’re published.

An hour-long interview produces roughly 7,000–9,000 spoken words. How those words become text determines your evening: typing them yourself takes four to six hours; a human service takes your money and a day; an app on the phone that recorded the interview takes minutes. Here’s the full picture, including the part most guides skip — keeping the quotes accurate.

First: record it properly, or nothing else matters

Transcription quality is capped by recording quality.

  1. Get consent on the record. Ask before recording — it’s law in many places and good practice everywhere. Start the recorder, ask again, and the “yes” is part of the file.
  2. Put the phone between you, on the table, not in a pocket. Distance to the quietest speaker is what hurts accuracy most. For walking interviews, a wrist-worn recorder beats a pocketed phone.
  3. Kill the failure modes. Do Not Disturb on, storage checked, and use a recorder that saves audio continuously to disk — an interview is the recording you cannot re-schedule.

Choosing the room

The room does more damage than most people expect, and you usually get a say in it. Avoid cafés — background music is the single hardest thing to transcribe around, because it occupies the same frequency range as speech and never holds still. Avoid large hard rooms, where reflections arrive a fraction of a second after the voice and smear the consonants.

What you want is small, soft and boring: a meeting room with carpet and a closed door beats an atmospheric location every time. If you have no choice, get the microphone dramatically closer to compensate.

The four ways to get text

1. Type it yourself. Free, and brutal: professional transcribers run about 4× real time; most people run 5–6×. An hour of tape is an evening gone. Worth it only when you need to internalize every word — some writers deliberately transcribe key interviews by hand for exactly that reason.

2. Use Apple’s built-in Voice Memos transcript. Also free, and already on your phone: on recent iPhones, open the recording in Voice Memos and tap the transcript view. The catch for interviews specifically is that there are no speaker labels — your questions and their answers arrive as one unbroken block of text, which you’ll untangle by hand for anything longer than a few minutes. Full walkthrough and limits: how to transcribe voice memos on iPhone.

3. Pay a human service. Typically $1–2.50 per audio-minute, so $60–150 for an hour, with hours-to-days turnaround. Highest accuracy on difficult audio (heavy accents, crosstalk, bad rooms). The trade-offs: cost at volume, and your source’s words leaving your custody — check the service’s confidentiality terms if the material is sensitive.

4. Use an app that records and transcribes in one step. This is the workflow Record and Transcribe (our app) is built around: the memo you recorded transcribes itself — no export, no upload to a second tool — and comes back with every line labelled by speaker, so your questions and their answers are already separated. Minutes instead of hours, from $0 for the first transcriptions. The honest catch applies to every AI option on the market: machine transcripts need proof-reading before publication.

MethodCost per hourYour timeSpeaker labels
Type it yourselfFree4–6 hoursYes, you add them
Apple Voice Memos transcriptFreeMinutes, plus untanglingNo
Human service~$60–150Minutes of adminUsually
Record-and-transcribe appFree tier, then paidMinutesYes, automatic

(The interview workflow page has the full comparison table.)

Remote and phone interviews

One limitation worth knowing before you plan around it: iOS does not let any app record phone call audio. That is a platform restriction Apple applies to every app in the App Store, not something a particular recorder has failed to implement, and there is no workaround.

For remote interviews, the practical routes are to conduct them on a video platform that records and often transcribes on its own, or — where your interviewee agrees — to have them recorded at their end and sent to you. Recording a phone call by putting it on speaker and pointing another device at it technically works and produces the worst audio you will ever try to transcribe. Treat it as a last resort, and expect to correct heavily.

Before you quote: the accuracy pass

Whatever method you used, do this before a quote reaches print:

  • Re-listen to every quoted passage against the text. Errors cluster exactly where quotes come from — emotional, fast, overlapping speech.
  • Mark uncertainty honestly. Unclear audio gets [inaudible], not a guess. A guessed word in a published quote is how corrections columns get written.
  • Decide verbatim vs. cleaned-up, once. Journalism usually lightly cleans false starts (“um”, repeated words) without changing meaning; qualitative research often needs true verbatim for coding. Pick the convention before editing, not per quote.
  • Names and numbers get a second check. Speech models are weakest on proper nouns — the CEO’s name, the drug dosage, the village. Verify each against your notes.
  • Check attribution at every handover. Where two people spoke over each other, the speaker labels are least reliable — and a quote attributed to the wrong person is a worse error than a wrong word.

If the recording came out badly

It happens, and it is recoverable more often than people assume. Work in this order:

  1. Transcribe it anyway. A poor transcript with timestamps is still a map of where things were said, which beats scrubbing an hour of audio by ear.
  2. Fix only what you need. You almost certainly need a handful of passages accurately, not all 9,000 words. Identify the quotes and work on those.
  3. Slow the playback rather than raising the volume. Reduced speed recovers far more from difficult audio than turning it up does.
  4. Go back to your source for the critical line. Asking someone to confirm a sentence you could not hear clearly is normal, professional practice — and much better than publishing a guess.

For researchers specifically

Speaker-labelled transcripts make coding faster — participant speech is already separated from your prompts. Keep your ethics approval’s storage requirements in mind: with this app, audio stays on your device and is deleted after transcription unless you keep it, which maps cleanly onto most data-handling protocols.

Two conventions worth settling before the first interview rather than during the fifth: how you anonymise (replace names at transcription time, and keep any key that maps back to participants stored separately from the transcripts), and how you mark pauses and overlaps. Note your cleaning convention in the methods section, and cite times from the recording so a passage can be located again.

The interview you conducted well deserves a transcript you can trust. Record it safely, pick the method that fits your deadline and budget, and spend the saved hours on the part machines can’t do — deciding what the words mean.

Written by Heni Hazbay, the independent developer of Record and Transcribe. These guides come from building the recording and transcription pipeline they describe.

Frequently asked questions

How do you transcribe an interview verbatim?

A verbatim transcript keeps every word as spoken, including false starts, repetitions and filler words like "um". Start from an automatic transcript, then play the audio back at reduced speed and restore what the model tidied away — speech recognition systems are trained to produce clean readable text, so they routinely drop exactly the disfluencies a verbatim transcript needs.

How do you transcribe an interview for qualitative research?

Research transcription needs two things ordinary transcription does not: consistent speaker attribution throughout, and a decision made up front about how much to clean up. Agree your conventions before you start — verbatim or intelligent verbatim, how to mark pauses, how to anonymise names — and apply them identically across every interview, because inconsistency between transcripts undermines any comparison you draw later.

How do you cite an interview in APA format?

Under APA style, an interview you conducted yourself is treated as personal communication: it is cited in the text with the person’s initial, surname and the date, and it does not appear in the reference list, because readers cannot retrieve it. An interview published in a journal, newspaper or podcast is different and is referenced normally, following the format for that source type. Check your institution’s own guidance, which sometimes adds requirements.

How long does it take to transcribe an interview?

Typing one out by hand realistically runs to several hours of work per hour of audio, depending on audio quality and your typing speed. Automatic transcription reduces that to the time it takes to read through and correct — usually a fraction of the recording length, concentrated on names, jargon and any passage where people talked over each other.

Ready when you are.

Free to start — 3 transcriptions included, no card needed. Apple Watch app comes with it.

Get the beta on TestFlight Free while in beta. Needs Apple’s TestFlight app — it installs it for you.