Why every summary should carry timestamps
A summary you can't check is a rumour with good formatting. The link back to the source isn't a nice extra — it's the thing that makes the summary worth anything.

Here's a summary of a talk: "The researchers found that the intervention improved outcomes by 40%."
Sounds useful. Now here's what the speaker actually said: "In the first cohort — and I want to be careful here, this didn't replicate in the second — we saw something like a 40% improvement."
The summary isn't lying. Every word in it corresponds to something in the talk. It has simply dropped the two clauses that determine whether the number means anything. And there is no way to tell that from reading it, because a confident summary of a hedged claim looks exactly like a confident summary of a confident claim.
Notice that no reviewer catches this. Someone checking the summary against the talk would have to already suspect that clause existed. Someone reading the summary alone has nothing to go on — it is internally consistent, it is specific, it names a real finding. Every signal a careful reader normally uses to detect a shaky claim has been stripped out by the same process that produced the claim.

Three kinds of claim that reliably break
Not everything compresses equally badly. In practice the damage concentrates in three places, and they're worth knowing because they're the ones to check.
- Numbers. A figure is short, quotable and feels like the substance of a finding, so it survives compression almost every time — while the conditions attached to it, which are long and boring, do not. The number arrives intact and unaccompanied.
- Attributions. "One team argues X, though the field mostly disagrees" reduces very naturally to "X". A speaker describing a position is hard to tell from a speaker holding it, and summaries resolve that ambiguity in whichever direction reads more cleanly.
- Negations and conditionals. "This doesn't work unless you already have the infrastructure" carries its entire meaning in two small words that a shorter sentence has every incentive to drop.
Compression always loses something
This isn't a solvable problem. Taking ninety minutes down to six bullets means discarding roughly 99% of the words, and some of what gets discarded will turn out to have mattered. Any summarization system has this property — a person taking notes, or a YouTube summarizer working from the transcript.
What differs is what happens next. If the summary is a sealed unit, you either trust it or you go and watch ninety minutes of video. Since nobody does the second one, in practice you trust it — and the tool has become a machine for manufacturing unearned confidence.
If every claim carries the second it came from, the calculation changes completely. Checking costs one click. And crucially, you only have to check the claims you're about to act on, which is usually one or two out of the six.
This is the part worth being precise about: timestamps don't make the summary more accurate. The same clause still went missing. What they change is that the loss is now recoverable in one second instead of ninety minutes, and that difference is the whole argument. A tool that cannot be wrong usefully is worse than one that can.

The fluency problem makes it worse
There's a nasty interaction between summarization and how people judge reliability. We treat fluent, well-organised prose as more trustworthy than hesitant prose. Summarizers produce extremely fluent prose by construction — they've stripped out exactly the hesitations, false starts and self-corrections that a careful speaker uses to signal uncertainty.
So the output of a summarizer reads as more authoritative than its source, while containing strictly less information. That's an unusually bad failure mode, and it's why "our summaries are very accurate" is not a sufficient answer. The question isn't whether the summary is usually right. It's what you can do when it isn't.
You can watch this happen in a single sentence. A speaker says "I think — and I could be wrong about this — that the effect is mostly driven by selection." The hedges are doing real work: they mark a live hypothesis, offered for argument. Compress it to "the effect is driven by selection" and you have not shortened the claim, you have promoted it. The speaker's own uncertainty was information, and it was the first thing discarded because it was the least efficient thing to keep.

What this looks like in practice
Three things follow from taking this seriously:
- Timings survive the whole pipeline. It's tempting to fetch a transcript, throw away the timings, and summarize the text — the text is all the model needs. But then you can never reattach a claim to a moment. The timings have to be carried through every step, even though they make everything slightly harder.
- Hedges are content, not noise. "We think this might generalise" and "this generalises" are different claims. A summarizer that treats the first as a verbose way of saying the second isn't compressing, it's editorialising.
- Failure has to be visible. When there's no transcript, the honest output is "there's no transcript", not a plausible summary assembled from the title. The failure cases are where a tool's honesty gets tested, and they're the easiest place to quietly cheat.

What it costs to keep them
It is worth being honest that carrying timings through is not free, which is why so many tools quietly drop them. The transcript arrives as thousands of short fragments with times attached. Every useful transformation — merging fragments into sentences, sentences into topics, topics into a claim — is a step where a piece of output stops corresponding to exactly one piece of input, and the mapping has to be maintained by hand at each one.
The cheap version is to summarize the text and then guess: search the transcript for words resembling the output and cite whatever matches. It works often enough to look fine in a demo, and it fails precisely on the claims that were most heavily rewritten — which are, inevitably, the ones most worth checking. A timestamp that is right when the summary is right and wrong when the summary is wrong is not a safeguard. It's decoration.
There's a reasonable objection here: if you have to check, what did the summary buy you? The answer is that it moved the expensive part. Without one you must watch ninety minutes to find the four claims worth having. With one you read six bullets, identify the one you're about to rely on, and spend ten seconds confirming it. You haven't eliminated the verification, you've narrowed it from the whole video to the single sentence that matters — which is the only form of the problem anyone actually solves.
The test to apply
Whatever summarization tool you use, try this: find a video you already know well, one with a genuinely contested or carefully qualified claim in it. Summarize it. Then see whether the qualification survived.
If it did, good. If it didn't, ask what you'd have had to do to notice — and if the answer is "rewatch the video", then you don't have a summary you can rely on. You have one you can only believe.
The same test applies to the timestamps themselves, and it's worth running once. Click three of them. If they land within a few seconds of the claim they're attached to, the mapping is real. If they land vaguely in the right region — the right topic, a minute early — they were reconstructed after the fact, and they'll be furthest off exactly where you need them most.

Run the test yourself
Pick a video you know well and check whether the caveats survive the compression.
Summarize a video — free