Scriber
Built for multi-speaker audio

Transcribe a YouTube Podcast or Interview

Podcasts and interviews are the hardest videos to get a usable transcript from. They are long, they often have no captions, and more than one person is talking. Scriber transcribes the audio directly and can label each speaker, so the text reads as a conversation instead of one long block. Paste the link below.

Works with youtube.com, youtu.be, and Shorts links.

Why podcasts and interviews are harder than other videos

Two things go wrong. Most transcript tools only copy captions the creator already uploaded, and conversational uploads frequently have none, so those tools return nothing at all. Then, even when you do get text, it arrives as an undifferentiated wall with no indication of who said what. For a solo talk that is fine. For a two-hour conversation it makes the transcript hard to quote from and slow to read.

Speaker labels tell you who is talking

Turn on speaker labels and Scriber transcribes the audio with speaker separation, then hands back the text split into turns: Speaker A, Speaker B, and so on. It does not know their real names, because it works from the audio alone and nobody tells it who is in the room. What it gives you is the structure of the conversation, which is the part that makes a transcript quotable. Labels come from the audio, so they work on videos with no captions.

What people do with an interview transcript

Pull an accurate quote without scrubbing back through the audio to find it. Write show notes or a summary from the text. Turn a long conversation into a newsletter or an article. Search a back catalogue for every time a topic came up. Researchers use transcripts to code interviews, and journalists use them to check what was actually said before publishing it.

Frequently asked questions

Does it work if the podcast has no captions?

Yes. When there are no captions Scriber transcribes the audio directly, which is the usual case for podcasts and recorded interviews.

Can it tell the speakers apart?

Yes. Turn on speaker labels and the transcript comes back split by speaker, as Speaker A, Speaker B and so on. It labels who changes, not who they are, so it will not put real names to the voices.

How long a video can I transcribe?

Up to three hours of audio, which covers almost every podcast episode and interview.

How accurate is it on two people talking?

Good on clear conversation. Speakers who talk over each other are the hard case for any transcription tool, so check those passages against the audio before you quote them.

Is it free?

You can transcribe a few videos free to try it. After that you can buy credits or go unlimited monthly.