Skip to content

How do I get the transcript of a YouTube video?

Paste a YouTube link. Where the caption track can be reached, this returns it as paragraphs with timestamps rather than the two-word fragments captions are actually made of, and you can copy it or download it as text, Markdown, SubRip or WebVTT. Automatic captions repeat the end of each line at the start of the next so the text scrolls smoothly on screen, which means a plain dump says everything roughly twice; that duplication is removed here, and sentences are joined on punctuation and pauses rather than chopped every few lines.

No signup. The result is yours before we ask for anything.

Run it on your own domain

Paste captions instead

Any caption file works: SubRip, WebVTT, a player payload, or plain lines copied from the platform. The cleaning, paragraphing and exports all run on whatever you paste.

How this works

  1. 1The link is read first, so a Share link, a Shorts link, a mobile link or a link with a playlist and a start time all resolve to the same video.
  2. 2The watch page is read for its caption track list, and a caption track written by a human is preferred over an automatic one where both exist.
  3. 3The rolling duplication in automatic captions is removed only where the overlap can be proved by matching words, so a phrase genuinely said twice is left alone.
  4. 4Fragments are joined into paragraphs on the evidence of sentence punctuation and pauses between cues, rather than on a fixed number of lines.
  5. 5Nothing is stored. The transcript is returned to whoever asked for it and the request ends.

What it cannot tell you

  • YouTube blocks caption downloads from server addresses. Measured on 12 September 2026: the page still lists the caption tracks, and fetching one returns an empty body with a success code. That is a block on where the request came from, not a statement about the video, and it is why a hosted transcript tool needs either a provider or a residential route.
  • A video with no captions of any kind has nothing to return. Speech recognition is a separate job with a bill attached.
  • It reads what the caption track says. Automatic captions mishear names, jargon and anything spoken over music.

Questions people ask

How do I get a transcript from a YouTube video?

Paste the video URL into a transcript tool, or open the video on YouTube, use the three-dot menu under it and choose Show transcript. YouTube publishes a caption track for most videos, automatic where the creator uploaded none, and a transcript tool reads that track rather than listening to the audio.

Why do transcript tools fail on some videos?

Usually for one of two reasons. The video genuinely publishes no caption track, or YouTube is refusing the request because of where it came from: caption fetches from server addresses return a success code with an empty body, which is easy to misreport as "no captions available".

Can I get a transcript of a YouTube Short?

Yes, where the Short has a caption track. A Shorts URL points at the same video as a normal watch URL, so any tool that reads the link properly handles both.

Why is a raw YouTube transcript so hard to read?

Because captions are a display format rather than text. They arrive as two to five word cues timed to speech, automatic ones repeat the previous line so the on-screen text scrolls, and many carry no punctuation or capitals at all. Reading well requires removing the duplication and rebuilding sentences.

Does this store the transcripts?

No. The transcript is returned to the person who pasted the link and nothing is kept.

The short version, if you would rather not run it

A transcript tool is two jobs and only one of them is interesting. The first is fetching a caption track, which is a commodity and increasingly a blocked one. The second is turning that track into something a person would read, which almost nothing does well.

The reason is that captions are not text. They are a display format for a video player. Automatic captions arrive as cues of two to five words, timed to speech, and many formats repeat the tail of the previous cue at the head of the next so the text scrolls smoothly rather than jumping. Concatenate those naively and the result says everything close to twice. Split them every few lines, which is what most tools do, and sentences get cut in half at random.

The approach here is deliberately conservative. Duplication is removed only where the overlap between two cues can be proved by matching words at the join, so somebody repeating a word for emphasis keeps it. Paragraphs break on the evidence available: a sentence-ending mark followed by a pause, then a long pause on its own, then a length cap so a monologue does not become one block. No punctuation is invented and no word is rewritten, because the point of a transcript is that it is what was said.

On the fetching half, the honest position is that this is harder than it was. Measured from a server address in September 2026, YouTube lists a video’s caption tracks and then returns an empty body when the track itself is requested, with a success code rather than an error. A tool that does not check for that reports "no captions available" and blames the video for its own address. Where the track cannot be reached, the processing here still applies to a caption file you already have, which every platform will give its own creator.

The other tools

How do I get the transcript of a YouTube video? · InstinctGTM