DUCKER

How to duck music under dialogue

Turning the music down under a voice is the single change that makes amateur audio sound professional. Here is how much to turn it down, when, and why level alone is usually not enough.

What ducking actually is

Ducking is automatic volume control: when one track has sound in it, another track gets quieter. In practice that means the music drops when someone speaks and comes back up when they stop, without anyone drawing volume curves by hand. Engineers reach it through a sidechain compressor — a compressor on the music that listens to the dialogue instead of to itself — but the idea is older and simpler than the tool.

Radio has done this since the 1930s. The reason it still matters is that speech and music occupy the same space, and a listener can only follow one of them. If the music never moves, every word is a small effort to hear. People stop watching, and they do not know why.

How far down: numbers that work

The useful range is narrower than people expect. Duck too little and the words are still a struggle; duck too much and the music vanishes and comes back like a light switch, which is worse than not ducking at all.

Kind of thingMusic drops byWhy
Talking-head video, tutorial−12 to −18 dB Speech is the content. The music is only there so silence does not feel dead.
Podcast with a music bed−15 to −20 dB Listeners are often in a car or on earbuds in the street, with the worst possible noise floor.
Documentary, narration over score−8 to −12 dB The score is doing emotional work. It should retreat, not disappear.
Trailer, promo−4 to −8 dB The music is the point. The voice has to cut through it instead.
Ambience and room tone−2 to −6 dB Barely at all. If the room disappears when someone speaks, the edit sounds fake.

Timing is what makes it invisible

Two settings decide whether anyone notices the ducking happening:

A good release is longer than the gap between words and shorter than the gap between sentences. That is the whole trick.

The part most guides leave out: carve, do not just duck

A human voice does most of its intelligibility work between roughly 1 kHz and 4 kHz — not in the warm low end people think of as "voice", but in the consonants: the t, k, s and f sounds that separate cat from cap. Full-range music covers that band too, and when both are present the ear has to choose.

So instead of pulling the whole music track down by 15 dB, pull it down by 8 and take another 3–5 dB out of that one band while the voice is speaking. The result is a music bed that still sounds full — the bass and the air are untouched — but that no longer competes for the frequencies the words live in. This is sometimes called dynamic EQ ducking or spectral ducking, and it is why a professional mix can keep the music noticeably louder than an amateur one and still be easier to follow.

A quick test. Play your mix at a level where you can just barely make out the words. If the music is what you hear most, you have a carve problem, not a level problem. Turning everything down further will not fix it.

Doing it without buying anything

Every mainstream tool can do this: Premiere Pro through Essential Sound's auto-ducking, DaVinci Resolve's Fairlight page, Audition and Logic through a sidechain compressor, Audacity through envelope points by hand. All of them assume you already own them and already know where the sidechain input is.

Ducker's Master page does the same job in a browser. You load your dialogue, music, effects and ambience as separate files; it reads where the dialogue is, ducks and carves the other three around it, and hands back one mixed file at the loudness target you pick. The three presets — Dialogue Forward, Balanced Mix and Trailer — are the three settings above, already dialled in. Nothing is uploaded: the files are decoded in your browser and never leave it.

Open Ducker and try it → Free, no account, nothing uploaded. Load a voice track and a music track and hear the difference in about a minute.

Common questions

What is the difference between ducking and sidechain compression?

Sidechain compression is the mechanism; ducking is the result. A compressor on the music, triggered by the dialogue instead of by the music itself, ducks the music. Some tools give you a dedicated ducking control that hides the compressor underneath.

How much should music be ducked under a voiceover?

Between 12 and 18 dB for most spoken-word video. Less for a documentary score, more for a podcast that people listen to in noisy places. If it sounds obvious, the release is too fast rather than the amount being wrong.

Can I duck audio for free without installing anything?

Yes. Ducker runs in the browser, needs no account, and processes your files locally — they are never uploaded. Audacity is the usual free desktop option, but it has no automatic ducking, so the volume changes have to be drawn in by hand.

Why does my music still bury the voice even after I turn it down?

Because level and clarity are different problems. The music is probably still full in the 1–4 kHz range where consonants live. Take a few dB out of that band while the voice is speaking, and the same music will feel quieter without actually being quieter.

Does ducking work on a music bed that has vocals in it?

Poorly. Two voices in the same frequency range cannot both be followed, however carefully you duck. Instrumental beds exist for this reason.