Scriber
← Blog

How I Summarized a 75 Minute Podcast With Timestamps

July 21, 2026 · 5 min read

The transcript took two minutes. Getting from there to a summary I would put my name on took three hours. Here is where that time went, and why the timestamps mattered more than the speed.

Works with youtube.com, youtu.be, and Shorts links.

A 75 minute podcast episode holds maybe six ideas I actually want. The problem is never finding them the first time. It is finding them again three days later, when I am writing something and I remember roughly what was said but not where.

Last week I wrote up an episode of Lenny's Podcast with Judd Antin, who built the research practice at Facebook and led research at Airbnb. I wanted a summary where every claim pointed back at the moment it came from. Here is what it actually cost.

The transcript took two minutes

I pasted the YouTube link into Scriber, turned on Speaker labels and turned on Timestamps.

Speaker labels matter on interviews for a reason that does not show up elsewhere. Without them a two person conversation arrives as one continuous block and you cannot tell a question from an answer. With them it comes back split into turns.

Two minutes later I had 75 minutes of conversation as text, with a timestamp on every turn.

That is the entire contribution of the tool. Everything after this point was me.

Reading took two hours

Longer than the episode. Someone will point out that listening at double speed would have taken 37 minutes, and they are right, so it is worth being clear about what those two hours bought.

Listening produces comprehension. It does not produce anything you can quote, search, or hand to somebody. At the end of 37 minutes of listening I would have understood the episode and owned nothing.

I read the speaker turns rather than scanning for keywords. In an interview the useful material sits where one person makes a claim and the other pushes on it. Searching for the word research in this episode would have returned almost every paragraph.

What I was looking for: claims specific enough to disagree with.

Taking the quote and the timestamp together

When something was worth keeping I copied the line with its timestamp attached. This costs nothing at the moment you do it, and going back later to find where a quote came from costs far more.

Three that survived:

[24:10] "one of my big kind of mantras was, we don't validate, we falsify, right? We are looking to be wrong."

[31:41] "my metric for success is when they won't have that meeting without you."

[33:03] "Good research doesn't slow us down, it speeds us up."

The markdown export writes each timestamp as a link to that second of the video. Someone reading the finished summary can jump to the moment and hear it in context, which is the difference between a summary they have to trust and one they can check.

Writing took one hour

I wrote last, with the quotes already in the document. Each section became a claim followed by the evidence for it. The finished piece ran about 400 words with six citations.

The structure that worked:

  • The argument in one line: the version of user research built over the last fifteen years is ending and the discipline has to change shape.
  • The framework at [7:52], splitting research into macro, middle range and micro, with the claim that most teams are stuck in the middle.
  • The concrete example at [34:46], the story he calls the multimillion dollar button, where changing seven characters of button text moved conversion by roughly one percent.
  • The uncomfortable part at [24:10], where asking for a study at the end of a project is described as checking a box rather than trying to learn something.

Anyone who disagreed could click a timestamp and argue with the source instead of with me.

Where it went wrong

Speaker labels are not speaker names. You get Speaker A and Speaker B and you map them to people yourself, which takes ten seconds.

The mapping is also not perfect. At [26:40] the transcript has this:

Speaker A: "I know there's this quote in your post I'm going to read."

That is Lenny, not Judd. It is Judd's post being read from, so the line belongs to Speaker B. The label slipped for exactly one turn. The same thing happens at [3:14], where a sponsor read gets split across both speakers.

Names take damage too. The transcript renders Judd Antin as "Judd Anton" throughout, and as "Jed" once at [1:08:51].

None of that is fatal and all of it is invisible unless you look. Which is the point. If you are going to put a quote next to a named person in public, read the surrounding turns first. The timestamp makes that a ten second check rather than a hunt.

What I would tell someone starting

Turn timestamps on before you start reading, not after you find a line worth keeping. The whole workflow depends on the citation being attached at the moment you notice something.

And do not expect the transcript to save you three hours. It saves you the two hours you would otherwise spend scrubbing back through audio hunting for a half remembered sentence. The reading is the work, and it should be.

If you want to try it, Scriber does the transcript part free for the first three videos and does not ask you to create an account. For lectures rather than interviews, the four pass method in this post fits better.

Frequently asked questions

Can I get timestamps on any YouTube video?

Yes. Turn on the Timestamps toggle before you export. Each timestamp becomes a link to that second of the video in the Markdown export, so a quote in your summary can be checked in one click.

Does it tell me the speakers' names?

No. Speaker labels come back as Speaker A and Speaker B, because the labels are worked out from the audio and nobody tells the tool who is in the room. You map them to real people yourself, which takes a few seconds at the start.

How long does a long episode take to transcribe?

A 75 minute episode with speaker labels took two minutes. Videos that already have captions and do not need speaker labels come back in seconds, because the captions are read rather than the audio being transcribed.

Is the transcript accurate enough to quote from?

For ordinary conversation, yes. Proper nouns are where it slips, and speaker labels can swap for a turn. Read the surrounding turns before you attribute a quote to a named person in public.