Guides

Interview Video Summarizer

An interview has two people in it. Automatic captions do not record which one is talking, and that single gap shapes everything about summarizing this format.

Two chairs facing each other with microphones between them.
Photo by Haberdoedas II on Unsplash

Interviews are the format where summarizing is most useful and most dangerous at once. Most useful because a ninety-minute interview contains maybe twelve minutes of substance and no way to find it. Most dangerous because the substance is attributable to a named person who can object to being misquoted.

Both of those follow from the same property: an interview is a conversation, and conversation is what long-form audio summarizing handles least cleanly.

The attribution problem, stated precisely

Automatic captions are a stream of text with no speaker labels. The W3C's captions guidance treats identifying the speaker as part of what proper captioning provides, and automatic captioning does not provide it.

So a summary of an interview infers who said what from context — from turn-taking patterns, from question-shaped sentences, from names mentioned. That inference is usually right and sometimes wrong, and it is never flagged as inference.

The specific failure that matters: an interviewer proposes a position to test it ("so you'd say the whole sector is overvalued?"), the guest declines it, and the summary records the sector being called overvalued. Attributed to the guest. Who said the opposite.

Two blank name plates on a table.
Photo by Natasha Fernandez on Pexels

What survives well

Despite that, a great deal does come through reliably, and it is worth being specific about which parts.

Interview content and how well it summarizes
Element Reliability Notes
Topics coveredHighThe main use — an index to the conversation
The guest's argumentGoodGuests talk most; their position dominates
Specific claims and figuresGoodVerify before repeating
Who said whatInferredAlways check at the timestamp
DisagreementWeakOften flattened into apparent agreement
Tone and hedgingPoor"Possibly" becomes a flat statement

The top half of that table is why interviews are worth summarizing. The bottom half is why the summary is a finding aid rather than a source.

A list of questions written on a sheet of paper.
Photo by Annie Spratt on Unsplash

Hedging is where the meaning lives

A specific loss worth understanding, because it changes claims rather than losing them.

People being interviewed hedge constantly, and the hedges are load-bearing. "I think, though I'd want to check this, that we're looking at something like thirty percent" is a different statement from "thirty percent", and summaries compress toward the second.

The hedge was the speaker marking their own uncertainty. Removing it converts a tentative estimate into an assertion they never made, and it is the most common way a summary misrepresents someone who was being careful.

This matters most for exactly the claims you would want to repeat — the specific, quotable ones. Those are the ones to open the timestamp for.

Quotation marks printed large on a page.
Photo by https://kaboompics.com/ on Pexels

Using it properly: triage, then verify

The workflow that makes interviews worth summarizing has two phases and people skip the second.

  1. Summarize to find the twelve minutes. A ninety-minute interview has a handful of substantive passages. The summary locates them.
  2. Read the summary for topics, not for claims. At this stage you are deciding what is worth your attention.
  3. Open the timestamps for anything you care about. Listen to the actual exchange.
  4. Take quotes from the audio, never from the summary text.
  5. Confirm who was speaking before attaching a name.
  6. Keep the hedges when you write it up.

Steps three through six take about ten minutes for a typical interview. That is still eighty minutes saved, with none of the risk of publishing something the speaker did not say.

Interview formats differ more than you would expect

"Interview" covers several things that behave very differently once summarized.

The promotional interview. A guest with something to sell, asked friendly questions. These summarize cleanly and contain almost nothing, because the guest is delivering prepared material. The summary correctly reflects that there was no substance.

The long-form conversation. Two or three hours, wandering, with real content in unpredictable places. The best case for summarizing and the format where the timestamps matter most.

The adversarial interview. The worst case. The meaning lives in what the subject declines to answer, in hesitation, in a question asked four times. A summary records what was said and cannot record an evasion, so it systematically reads as though the questions were answered.

The panel interview. Multiple guests, which multiplies the attribution problem by the number of people in the room.

Knowing which you have before you read the summary tells you how much of it to trust without checking.

Why long interviews are the best case

The longer the interview, the better the trade. A three-hour conversation is genuinely unsearchable — there is no way to know whether the thing you need is at minute eighteen or minute two hundred.

A timestamped summary turns it into a document with a contents page. That is the single largest gain available, and it applies to exactly the format that is otherwise impossible to use as a source: the very long recording nobody will scrub through.

An audio waveform displayed on a screen.
Photo by Techivation on Unsplash

Reading a summary for what is absent

A skill worth developing, because interviews carry meaning in omissions and summaries record only presences.

If you know a subject was interviewed about a controversy and the summary contains no mention of it, that is information — either the question was never asked, or it was asked and deflected so thoroughly that nothing summarizable came back. Both are worth knowing and neither appears in the text.

The same applies to proportion. A summary flattens ninety minutes into equal-weight sections, so a topic the guest spent forty minutes avoiding and a topic they covered in two look identical on the page.

So when an interview matters, read the summary against what you expected it to contain, not only for what it does contain. The gaps tell you where to open the timestamps, and those passages are usually the ones worth the listening time.

What it will not do

Label the speakers. Not reliably, and not at all where the conversation overlaps.

Capture visual reaction. A raised eyebrow, a long pause, a laugh that reframes what was just said.

Handle crosstalk. Where two people speak at once, automatic captions produce interleaved fragments and the summary inherits the mess.

Serve as a citation. A summary is a paraphrase of a paraphrase. Cite the video and the timestamp.

A microphone with its cable unplugged.
Photo by Jason Morrison on Pexels

Frequently asked questions

Does it identify who said what in an interview?

Only by inference, since automatic captions carry no speaker labels. It is usually right and never flagged when wrong, so verify at the timestamp before attributing anything.

Can I quote from the summary?

No. Summaries paraphrase and tend to drop hedging, which can turn a tentative estimate into a flat assertion. Take quotes from the audio at the timestamp.

How does it handle two people talking over each other?

Badly. Automatic captions produce interleaved fragments during crosstalk, and the summary reflects that. Those passages need listening to directly.

Is a three-hour interview worth summarizing?

That is the best case for it. Very long conversations are otherwise unsearchable, and a timestamped summary turns one into a document with a contents page.

Find the twelve minutes that matter

Summarize to locate the substance, then listen to it before you quote it.

Summarize an interview — free

Keep reading