Auto-transcription and text-based editing
Turning speech into editable text, so deleting a word or sentence from the transcript also cuts the matching audio.
See it
What it is
Automatic transcription turns speech into time-coded words, then a text-based editor keeps every word linked to its exact stretch of audio. Delete a sentence in the transcript and the matching waveform section disappears too. Descript built its editing workflow around this idea, and Adobe Premiere Pro now calls its version Text-Based Editing.
Reach for it on interviews, podcasts, and voiceovers where the first job is shaping what was said. Searching text is much faster than scrubbing a long waveform, and repeated words, false starts, and whole answers can be removed like lines in a document. The timeline still matters for music, overlapping speakers, breaths, and the exact rhythm of each cut.
Gotcha: a clean transcript is not automatically a clean edit. Speech recognition can attach the wrong word or speaker, and deleting text can clip a breath, expose a room-tone jump, or make the cadence sound inhuman. Correct the transcript before editing, then listen across every cut and adjust the audio boundaries by ear.
Ask AI for it
Edit this podcast in Descript using automatic transcription and text-based editing: identify the speakers, correct names and obvious transcription errors first, then remove false starts, repeated phrases, and off-topic answers through the transcript. Keep intentional pauses and natural breaths, review every text edit in the waveform, add short crossfades where a cut clicks, and fill exposed gaps with matched room tone. Export the revised transcript and a 48 kHz WAV of the final edit.