24 August 2026 · By Heni Hazbay
How to Transcribe a YouTube Video to Text, Free
YouTube already writes a transcript for most videos and it is three clicks away. How to open and copy it, and what to do for the videos that have none.
The short answer: Do not reach for a tool first. YouTube auto-captions most videos and shows the result as a copyable transcript panel — three-dot menu under the video, then “Show transcript”. Turn timestamps off before copying. Only videos with captions disabled, no speech, or an unsupported language need anything else.
There is an entire industry of “YouTube transcript generator” sites, and for most videos they are re-selling you a transcript YouTube already made and already shows you for free. Start where the text already is.
1. Open YouTube’s own transcript
On desktop, under the video, click the three-dot menu (…, next to Share) and choose Show transcript. A panel opens to the right of the video with the whole spoken content as timestamped lines. Click any line to jump the playhead there.
In the mobile apps this is less consistent — expand the description and look for a transcript button. It is present for many videos and missing for others, so desktop is the reliable route.
Before you copy anything, turn off the timestamps. The transcript panel has its own three-dot menu with a Toggle timestamps option. With it on, every line arrives prefixed with a time and you will spend longer cleaning that up than you spent getting the transcript. With it off, you can select the panel contents and copy clean paragraphs in one go.
2. Know what you have actually got
YouTube’s automatic captions are good and they are not a finished transcript. Specifically:
- No speaker labels. A two-person interview arrives as one undifferentiated stream of text. Nothing marks where one person stops and the other starts, which is the single biggest difference between captions and a usable interview transcript.
- Little punctuation logic. Sentences run together, and question marks are inconsistent.
- Names and jargon go wrong. Proper nouns, product names and technical terms are where automatic captioning reliably fails, and they are often exactly the words you wanted.
- It is a caption track, not a document. Line breaks fall where the caption timing fell, not where the sentences end.
None of that matters if you are searching a video for the bit where somebody said a particular thing. All of it matters if you are quoting for an article, a paper or a report. Budget a cleanup pass, and fix the speaker boundaries before anything else — how to clean up a transcript sets out the order to work in, and transcripts with speaker labels covers the conventions worth following once you do.
3. If the creator uploaded real captions, you are in luck
Not all captions are automatic. Where a creator has uploaded a human-written caption file — common for well-funded channels, educational content and anything with accessibility obligations — the transcript panel shows that instead, with proper punctuation and often speaker names. You cannot tell from the panel which you are looking at, but the quality gives it away immediately.
4. When there is no transcript at all
Three situations produce a video with no transcript option:
Captions are disabled by the uploader. Their choice, and there is no way around it from the viewer side.
There is no speech. Music, ambient footage and silent screen recordings have nothing to caption.
The language is not well covered. Automatic captioning quality varies enormously by language, and for some it is not offered at all.
In all three cases you are producing the transcript yourself, and the practical route is the same one that works for podcasts: play the video and record the audio on a second device — a phone on the desk next to the speakers is genuinely enough — then transcribe that recording. It sounds crude and it works, because speech transcription is far more tolerant of room audio than people expect. What it will not survive is a noisy room, so close the door.
5. The copyright line, stated plainly
Transcribing a video so you can search it, quote it or study it is ordinary personal use and nobody minds.
Publishing the full transcript is a different act. The video is somebody’s copyrighted work, and a complete transcript reproduces the whole of its spoken content. Republishing it — even with a link, even with credit — is a reproduction, and “I transcribed it myself” is not a defence. Quote the passage you need with attribution, link the video at the timestamp, and if you want the whole thing published, ask the creator. Educational channels in particular often say yes.
Where our app fits, and where it does not
We make an iPhone and Apple Watch recorder that transcribes with speaker labels. For YouTube specifically it is the fifth option, not the first, and only for step 4 above: a video with no transcript, played aloud and captured on the phone. If YouTube’s own transcript exists, use that — it is free, instant, and we would rather say so.
Where the difference is real is the thing YouTube’s captions cannot do at all: telling you who was speaking. For a two-hour panel discussion or a multi-person interview, a stream of unattributed text is close to unusable, and transcripts with speaker labels are a different document entirely. That is worth knowing when the video is an interview rather than a monologue.
For episodes rather than videos, transcribing a podcast to text follows the same principle — check Apple Podcasts and Spotify before making anything. And if you are weighing whether an automatic transcript is accurate enough to quote from, how accurate is AI transcription sets a realistic expectation.
The order worth trying
| Step | Cost | Works when |
|---|---|---|
| YouTube transcript panel | Free | Captions are enabled — most videos |
| Creator’s uploaded captions | Free | Channel provides them; best quality |
| Auto-translate the transcript | Free | You accept compounding errors |
| Record it aloud and transcribe | App or service | No captions exist at all |
Four out of five people reading this need the first row.