PodcastsToText
Back to articles
Guides

Podcast Transcript Format: Examples and a Simple Style Guide

PodcastsToTextPodcastsToTextSeptember 28, 20265 min read

A podcast transcript has one job: let someone read the episode as easily as they could listen to it. The standard format that does this is simple. Name the speaker, put a colon, write what they said, and start a new paragraph whenever the speaker changes. Everything else in this guide (timestamps, subtitle files, how to mark laughter or a word you cannot hear) builds on that.

Below are real examples taken from one short episode, the same words in four formats, followed by the rules that keep a transcript readable.

The basic podcast transcript format

This is the layout almost every published podcast transcript uses:

  • Speaker name in bold, then a colon. Use real names once you know them, or roles such as Host and Guest.
  • A new paragraph for every change of speaker. A long answer can be split into several paragraphs, but a speaker change always starts a new one.
  • The name appears only when the voice changes. Repeating it on every paragraph of the same answer adds noise.
  • Optional timestamps at the start of each paragraph, so a reader can jump to that point in the audio.

Here is that layout with two speakers. This block only illustrates the layout; the real examples further down come from a one-speaker episode.

**Host:** Welcome back. Today we are talking about why transcripts matter.

**Guest:** Thanks for having me. The short answer is that people search in text.

**Host:** Say more about that.

Example 1: a clean read-along transcript

This is the real transcript of the opening of our own launch episode, exported as plain text with the speaker renamed to Host. It reads like prose: no timestamps, one speaker, paragraphs where the thought changes.

Host: Have you ever listened to a podcast and wish you could read what was being said? Maybe to quote it, study it, or just to find that one line again without re-listening to the whole podcast.

podcasttotext.com.

Why did I build this up? You may ask. I built it because I listened to a lot of podcasts. And at the end of each episode, I was forgetting what was being said.

Two things to notice. The words are left as spoken ("wish" rather than "wished", "build this up"); whether to tidy grammar like that is a style decision, covered below. And the one-line paragraph is the product's own name misheard: the address is podcaststotext.com, with an "s" the transcript dropped. Names are where automatic transcripts go wrong most often, which is why checking them is one of the rules further down.

Example 2: a timestamped transcript

The same passage with a timestamp on every line. This is the format to use when readers need to find the moment in the audio: researchers quoting a guest, editors marking cuts, or show notes that link to chapters.

[00:00] Have you ever listened to a podcast and wish you could read what was being said?
[00:05] Maybe to quote it, study it, or just to find that one line again without re-listening to the whole podcast.
[00:17] podcasttotext.com.
[00:29] Why did I build this up? You may ask.
[00:32] I built it because I listened to a lot of podcasts.

Minutes and seconds ([mm:ss]) are enough for most episodes. Switch to [hh:mm:ss] for anything over an hour so the timestamps sort correctly.

Example 3: SRT subtitles

SRT (SubRip) is the subtitle format video editors expect. Each cue has a number, a start and end time with milliseconds, and the text. Use it when the episode also goes out as video, on YouTube or as clips.

1
00:00:00,980 --> 00:00:05,060
Have you ever listened to a podcast and wish you could read what was being said?

2
00:00:05,500 --> 00:00:12,060
Maybe to quote it, study it, or just to find that one line again without re-listening to the whole podcast.

3
00:00:17,720 --> 00:00:19,280
podcasttotext.com.

4
00:00:29,260 --> 00:00:32,260
Why did I build this up? You may ask.

Example 4: WebVTT captions

VTT is the web's version of SRT: a WEBVTT header, a dot instead of a comma before the milliseconds, and no cue numbers required. HTML5 audio and video players read it directly.

WEBVTT

00:00:00.980 --> 00:00:05.060
Have you ever listened to a podcast and wish you could read what was being said?

00:00:05.500 --> 00:00:12.060
Maybe to quote it, study it, or just to find that one line again without re-listening to the whole podcast.

00:00:17.720 --> 00:00:19.280
podcasttotext.com.

00:00:29.260 --> 00:00:32.260
Why did I build this up? You may ask.

Which transcript format should you use?

If you want toUseWhy
Publish it on your episode pagePlain text or MarkdownReads like an article and gives search engines the full episode to index.
Edit it or hand it to someoneDOCXSpeaker names and timestamps stay as real formatting in Word or Google Docs.
Share or archive a fixed copyPDFPaginated and cannot be changed by accident.
Caption a video versionSRTWhat Premiere, Final Cut, DaVinci Resolve and YouTube take.
Caption audio or video on a websiteVTTNative to HTML5 players.
Feed it to your own code or an AI toolJSONEvery segment with its start, end and speaker as data.

If your podcast host supports the Podcasting 2.0 <podcast:transcript> tag, you can also attach the transcript to the episode in your RSS feed. The tag accepts SRT, VTT, JSON, HTML and plain text, and apps that support it show the transcript alongside the audio.

Style rules for a readable transcript

Clean verbatim or full verbatim

Full verbatim keeps every "um", false start and repeated word. It is the right choice for legal, research or linguistic work, where how something was said matters. Clean verbatim removes those and fixes small slips so the text reads smoothly, without changing meaning. Most published podcast transcripts are clean verbatim. Pick one and apply it to the whole transcript.

Speakers

  • Use the same name for a speaker throughout: "Sarah", not "Sarah" in one place and "Dr. Lee" in another.
  • If you do not know a speaker's name, use a role (Host, Guest, Caller) rather than Speaker 1 in a published transcript.
  • For crosstalk, give each person's words their own line and add [crosstalk] where they overlap.

Non-speech sounds and gaps

  • Put sounds that matter to meaning in square brackets: [laughs], [music], [phone rings].
  • Mark words you cannot make out as [inaudible 12:04], with the timestamp, so someone can check the audio later.
  • Leave out background noise that does not change what was said.

Numbers, names and ads

  • Write numbers the way your publication already does, and keep it consistent.
  • Check every proper noun (guests, companies, places) before publishing. Names are where automatic transcription is most often wrong.
  • Decide whether sponsor reads stay in. Many shows drop them from the published transcript or mark them [ad break].

How to get a transcript in this format

Typing a transcript by hand takes several times the length of the episode. The examples above were produced automatically: paste an episode link on the PodcastsToText homepage (or use the Spotify and Apple Podcasts tools, or upload an audio file) and the transcript comes back with speakers separated and every line timed. Rename the speakers once in the editor, then export it as plain text, Markdown, SRT, VTT, JSON, DOCX or PDF, with speaker names and timestamps switched on or off.

Ready to transform your content?

Turn your podcasts into accurate text in minutes. Start transcribing for free and see how PodcastsToText can save you hours of work.

Get started now

Transcription tools