A transcript is not a summary
A lot of tools hand you the same words in a different box and call it summarization. Here's what actually has to happen in between.

Search for a YouTube summarizer and most of what you'll find does one thing: fetches the caption track and prints it. Sometimes with paragraph breaks. Occasionally with an AI-generated opening sentence stapled on top.
This is genuinely useful for some things — searching for a phrase, checking a quote. It is not a summary, and the gap between the two is bigger than it looks.
What a raw transcript is actually like
Speech is not writing. A transcript of a competent, engaging speaker reads like this:
so the thing is — and I'll come back to this — um, when you look at the second case, the second case, right, it's it's basically the same as the first except, and this is the important bit, except the constraint is reversed
That's roughly 45 words carrying about 12 words of content. Multiply by ninety minutes and you have 12,000 words that take nearly an hour to read and leave you assembling the argument yourself. You've converted a video you can't skim into text you won't finish.
The arithmetic is worth doing once, because it kills the intuition that text is inherently faster than video. Twelve thousand words is a long magazine feature, and it's dense, unstructured, unedited magazine feature. Skimming it doesn't work either: skimming relies on topic sentences and paragraph shape to tell you what to skip, and a transcript has neither. You end up reading most of it or none of it, and in practice it's none.
Worse, the structure is invisible. A speaker signals a new section with a pause, a slide change, a shift in tone. None of that survives into text. The transcript is a flat wall where the original had shape.

And that's the good case
The example above assumes a clean caption track. Most of what you'll actually get is YouTube's automatic transcription, which adds its own layer of damage: no punctuation worth the name, no paragraph breaks, no speaker labels. A two-person interview arrives as a single undifferentiated stream in which the question and the answer are the same voice.
Then there are the substitutions. Automatic captioning is at its worst on exactly the words carrying the most meaning — technical terms, proper nouns, anything the speaker says once. These don't come back flagged as uncertain. They come back as confident, plausible, wrong words sitting in otherwise reasonable sentences, which is a good deal harder to notice than a gap would have been.
Any tool printing that back to you is passing along every one of those errors with a layer of formatting on top, and the formatting makes them look more settled than they are.
The four things that have to happen
1. Disfluency has to go, meaning has to stay
Stripping the "um"s is easy. The hard part is that speakers repeat themselves for emphasis, and repetition-for-emphasis looks identical to repetition-because-they-lost-their-thread. Cut the wrong one and you've removed the thing the speaker most wanted you to notice.

2. Structure has to be recovered
The sections were real — the speaker had an outline. They're just not marked in the text. Recovering them means noticing topic shifts from the content itself, which is a genuine inference rather than a formatting pass. This is what a chapter map is: the outline the speaker was working from, reconstructed.

3. Claims have to be separated from scaffolding
Most of a talk isn't claims. It's setup, analogies, jokes that buy thinking time, recaps of the previous section. The analogies matter for understanding and don't belong in a takeaway list. Telling "here's a claim" from "here's a vivid way of explaining the previous claim" is most of the work.
4. Timings have to survive all of it
After three transformations, each output line has to still know where it came from. This is the step most tools skip, because it makes every other step harder and its absence isn't visible in the output — you only miss it the first time you want to check something.

Why so many tools stop at step zero
Not because it's hard. Because fetching a caption track costs nothing, returns instantly, and produces something that looks unmistakably like output. It is long, it is clearly derived from your video, and it arrives fast enough to feel impressive. Every incentive points at shipping it.
The four steps above cost real money per video and take a few seconds, and the benefit is visible only to someone who already knows what they were supposed to get. Nobody opens a reformatted transcript and thinks "this tool skipped topic segmentation". They think the video was less interesting than they hoped.
How to tell them apart
Two quick tests on any tool claiming to summarize video:
- The length test. Feed it a ninety-minute video. If what comes back takes more than a few minutes to read, it's reformatting, not summarizing. Real compression is dramatic — 12,000 words to maybe 400.
- The order test. Check whether the most important point appears first. A transcript is ordered by when things were said; a summary is ordered by what matters. If the output still follows the video's chronology exactly, nothing has been prioritised — and prioritising is the job.
- The stranger test. Hand the output to someone who hasn't seen the video and ask what it argued. If they can tell you in a sentence, the structure survived. If they have to read the whole thing and then hedge, you're looking at raw material rather than a result — the work of working out what matters has simply been passed back to the reader, which is where it started.

When you actually do want the transcript
None of this makes transcripts bad — they're the right tool for a narrow set of jobs, and worth knowing which. If you need the exact wording of a quote, only the transcript will do; a summary has by definition paraphrased it. If you're searching for whether a term came up at all, ctrl-F over the raw text beats any amount of compression. If you're captioning, translating, or feeding the text to something else, you want the source, not an interpretation of it.
The failure isn't offering a transcript. It's offering one in answer to "what's in this video?" — a question it structurally cannot address, since answering it means deciding what matters, and a transcript is the artefact produced by declining to decide.
Why this is worth caring about
Because the failure is invisible. A reformatted transcript looks like output. It's long, it's relevant, it's clearly derived from your video. You only discover it didn't help when you're staring at 12,000 words with no more idea of what the video argued than before.
The cost isn't just the wasted time. It's that the failure teaches you the wrong lesson. Having read four thousand words that went nowhere, the natural conclusion is that summarizing video doesn't really work, or that this particular video had less in it than you thought. The tool's limitation gets filed as a fact about the material — which is the most expensive kind of bug, because nobody reports it.
The question worth asking isn't "did this tool produce text about my video". It's "can I now decide whether to watch it, and find the part I need?" If not, you got a transcript with better typography.
Apply the length test
Take the longest video in your watch-later and see how much actually comes back.
Summarize a video — free