Captions and transcripts

Captions follow audio in time on a video; a transcript is the whole recording as readable text you can scan without pressing play.

do I need subtitles and a written versiontext below the podcastwords timed to the videoread the video instead of watching itfull script under the playerthe auto captions are full of mistakessubtitles for deaf viewerstransript and captions

See it

Live demo coming soon

What it is

Captions are timed text synchronized with video. They identify speech, speakers, and meaningful sounds at the moment each happens. A transcript is one complete, readable text version of the media that can be skimmed, searched, copied, or read without playing anything.

Use WebVTT captions for video and provide an HTML transcript for podcasts, interviews, lectures, and other content people may want to read at their own pace. A good transcript preserves speaker changes and sound cues, then adds headings or links where they make a long recording easier to navigate.

One does not automatically replace the other. A transcript below a video does not give a Deaf viewer synchronized text while watching, and a caption track is awkward as a stand-alone document. Captions are also not translation subtitles: streaming services label the caption track SDH, subtitles for the deaf and hard of hearing, precisely because it carries sound cues and speaker names that a translated subtitle track drops. Automatic speech-to-text is a draft: YouTube's auto-captions are the everyday proof, fluent on plain speech and wrong on names, jargon, punctuation, speaker changes, and meaningful sounds, all of which still need human correction.

Ask AI for it

Using the supplied media and verified dialogue, create both closed captions and a full transcript. Author a WebVTT file with accurate cue timings, speaker changes, and meaningful non-speech sounds, then attach it with an HTML track element using kind='captions', src, srclang, and a clear label. Add a semantic HTML transcript below the player with a heading, speaker labels, paragraphs, sound cues, and links from any chapter list. Correct names and jargon against the source; do not invent words or timestamps. Keep the transcript available without JavaScript and make the caption toggle keyboard operable.

You might have meant

captionsscreen readerwcagsemantic htmlassistive technology

Go deeper