Record and Transcribe

19 August 2026 · Updated 24 August 2026 · By Heni Hazbay

How Long Does It Take to Transcribe an Hour of Audio?

AI transcribes an hour of audio in minutes; a person needs four to six hours. Realistic timings by recording length, and what makes each of them slower.

The short answer: A person typing takes about four to six hours per hour of audio. Automatic transcription returns the same hour in a few minutes, and then needs fifteen to twenty minutes of review before you can rely on it. The ratio that matters is not the machine time — it is the checking time.

If you are planning a project around transcription, or deciding whether to do it yourself, the numbers below are the ones to plan with.

Manual transcription: the 4:1 ratio

The standard planning figure in transcription work is four hours of typing per hour of audio, for a competent typist working on clear, single-speaker recordings. That is where “4:1” comes from, and it is realistic rather than optimistic — it assumes someone who does this often, with a foot pedal and transcription software.

For everyone else it is worse. Someone typing in a text editor, rewinding by hand, working on a real meeting rather than dictation, will land closer to six to eight hours per hour.

Audio lengthExperienced (4:1)Typical non-specialist (6:1)Difficult audio (10:1)
10 minutes40 min1 hr1 hr 40
30 minutes2 hrs3 hrs5 hrs
1 hour4 hrs6 hrs10 hrs
2 hours8 hrs12 hrs20 hrs

“Difficult audio” is not an edge case. It means several speakers, background noise, accents unfamiliar to the transcriber, overlapping speech, or a poorly placed microphone — which describes most real meetings.

Automatic transcription: minutes, not hours

Automatic transcription of an hour of audio typically completes within a few minutes. The exact figure depends on the service, the load on it, and whether speaker labelling is running, but the order of magnitude is minutes rather than hours.

What takes the time:

  • Reading and segmenting the audio. Fast, and roughly proportional to length.
  • Running the speech model. The main cost, and on modern hardware this runs considerably faster than real time.
  • Speaker labelling. Diarization is a second pass that works out how many distinct voices are present and which segments belong to which. It adds meaningfully to the total, and it cannot start until the whole recording is available, because deciding there are three speakers rather than four requires having heard all of it.

That last point explains something people find odd: a transcript can appear quickly and the speaker labels arrive after. The two jobs are not the same job.

The number nobody plans for: review time

This is where estimates go wrong. Automatic transcription is fast, and then you still have to read it.

Budget a quarter to a third of the recording length — roughly fifteen to twenty minutes per hour of audio — if the transcript is going to be relied on for anything.

AudioMachine timeReview timeRealistic total
10 minutesunder a minute3 min~4 min
30 minutes1–2 min8–10 min~12 min
1 hour2–5 min15–20 min~25 min
2 hours5–10 min30–40 min~50 min

Even at the pessimistic end, an hour of audio is about twenty-five minutes of your time rather than four hours of it. That is the actual change automatic transcription made — not perfect text, but a tenfold reduction in the work.

You do not need to read every word. Errors concentrate predictably in names, numbers, technical terms, and the first few seconds after a speaker changes. Checking those four categories catches most of what matters at a fraction of the effort of a full proofread. How accurate is AI transcription covers what the error rates actually look like, and why the “99% accurate” figure in marketing does not survive contact with a real meeting.

What makes any of it slower

The same factors slow down humans and machines, which is a useful thing to know — if a recording is hard for you to follow, the transcript will be poor.

  • Multiple speakers, especially talking over each other. Overlap is the single hardest thing in transcription, and neither humans nor machines handle it well.
  • Distance from the microphone. Every metre costs you.
  • Background noise — air conditioning, traffic, a café.
  • Accents and dialects the system or transcriber is less exposed to.
  • Specialist vocabulary. Medical, legal and technical terms are transcribed phonetically and wrongly unless the system knows them.
  • Poor recording quality at source — a compressed phone line, a bad video call connection, a microphone under a jacket.

Most of these are decided in the first thirty seconds of the recording, before anyone has said anything important. Improving speech-to-text accuracy is largely a list of things to do at that point.

When you need it faster than any of this

The fastest transcript is the one that started before you asked for it.

Record and Transcribe (our app, for iPhone and Apple Watch) runs transcription automatically as part of recording rather than as a separate job you queue afterwards — you stop the recording, and the text is on its way without an upload step or an export. Every line arrives labelled by speaker, which removes most of the work of turning a two-person conversation into something readable.

Three transcriptions free, no card needed.

The summary

Time and money trade against each other here, and the trade is steep in both directions — how much transcription costs covers the three pricing models and which one your deadline can actually afford.

  • By hand: 4:1 if you are good at it, 6:1 realistically, 10:1 for hard audio.
  • Automatically: minutes for the machine, then 15–20 minutes of review per hour.
  • The bottleneck moved. It used to be typing. It is now checking — and checking is the part worth doing properly.
Written by Heni Hazbay, the independent developer of Record and Transcribe. These guides come from building the recording and transcription pipeline they describe.

Frequently asked questions

How long does it take to transcribe 1 hour of audio?

A person typing it takes roughly four to six hours for clear single-speaker audio, and longer for difficult recordings. Automatic transcription typically returns an hour of audio within a few minutes, then needs review time on top before you can rely on it.

How long does it take to transcribe 30 minutes of audio?

Around two to three hours by hand, or a couple of minutes automatically plus about fifteen to twenty minutes of checking. Manual transcription scales almost linearly with length, so halving the audio roughly halves the time.

What is the standard transcription ratio?

Four to one is the widely used planning figure for a competent typist working on clear audio: four hours of work per hour of recording. Difficult audio with several speakers, accents or background noise runs to eight or ten to one.

Why does AI transcription still take a few minutes?

The audio has to be read, segmented, run through a speech model and, where speakers are labelled, analysed a second time to work out who spoke when. Speaker labelling is the step that adds the most, and it cannot begin until the audio is complete.

How much review time should I budget after automatic transcription?

Plan for roughly a quarter to a third of the recording length if the transcript matters — about fifteen to twenty minutes per hour of audio. Names, numbers and technical terms are where errors concentrate, so those are what you check.

Ready when you are.

Free to start — 3 transcriptions included, no card needed. Apple Watch app comes with it.

Get the beta on TestFlight Free while in beta. Needs Apple’s TestFlight app — it installs it for you.