Ducking
The music automatically dips when the voice comes in, then rides back up in the gaps. Radio has done it for eighty years.
See it
What it is
A compressor sits on the music, but instead of listening to the music it listens to the voice track. Voice arrives, music drops. Voice stops, music climbs back. That keying arrangement is sidechain compression, and it is the same trick radio DJs have used since the 1940s. You can also draw it by hand with volume automation, or hit the one-click version: Auto Ducking in Adobe Audition, Essential Sound in Premiere, Duck in Descript, Smart Ducking in Final Cut.
The numbers that matter: dip the bed roughly 8 to 15 dB under the dialogue, attack fast (10 to 50 ms), release slow (500 ms to 1 second) so it swells back like a decision rather than a glitch. Between sentences, hold the dip so the music does not bounce during breaths. Attack alone cannot get ahead of the voice: the detector only reacts once the first syllable has arrived. If you want the music already out of the way before the word lands, give the sidechain 20 to 50 ms of lookahead or pre-roll, or nudge a duplicate of the voice track earlier and key off that. Without it, accept that the opening transient beats the dip.
Gotcha: ducking is a mix trick, not a rescue. If the bed is a busy track with its own vocals and a fat 3 kHz midrange, ducking just makes it quiet and still fighting; pick a sparser bed or carve a midrange dip with EQ instead. The same effect used deliberately and hard, with a kick drum as the key, is the pumping you hear all over dance music.
Ask AI for it
Duck the music bed under the voiceover: sidechain-compress the music bus keyed off the dialogue track so the music drops 10 to 12 dB whenever speech is present. Use 20 to 50 ms of lookahead or pre-roll on the sidechain detector plus a fast attack (about 20 ms) so the dip lands before the first syllable instead of chasing it, a hold so the music does not jump back up between sentences, and a slow release (about 800 ms) so it swells back smoothly at phrase ends. The dip should feel invisible: no pumping, no audible stair-stepping, and the voice should stay clearly on top throughout.