Transcript-based editing
Editing video by editing its transcript: delete a sentence in the text and the matching footage vanishes from the timeline.
See it
What it is
The editor runs speech-to-text with word-level timestamps, then binds every word in the transcript to its slice of the timeline. Highlight a sentence, hit delete, and the footage goes with it. Descript built its whole product on this; Premiere Pro ships it as Text-Based Editing, and most AI video tools now open on a transcript rather than a timeline.
Reach for it on anything talk-driven: interviews, podcast video, webinars, talking-head explainers. The rough cut that used to take an afternoon of scrubbing becomes twenty minutes of reading, and the one-click 'remove filler words' pass strips the ums, uhs, and 'you knows' across a two-hour recording in one go.
Gotcha: it edits the words, not the pictures. Every deletion is a jump cut, so you still need B-roll, cutaways, reaction shots, or a morph-cut style blend to hide the seams. Cuts also land on the word boundary rather than the breath, which chops off tails and consonants unless you nudge the audio handles by a few frames. And transcript accuracy collapses on crosstalk, heavy accents, and domain jargon, so read before you delete.
Ask AI for it
Edit the linked footage through its word-timed transcript. Remove every filler word (um, uh, like, you know, I mean), every false start, and every repeated phrase, then cut the answers down to the strongest 90 seconds while keeping full sentences intact. Pad each remaining segment with 6 frames of audio handle at the head and tail so no consonant gets clipped, and cover every visible seam with the marked B-roll. Return the edited sequence itself, plus an EDL of in and out timecodes and the revised transcript.