Videos · Glossary
What is audio ducking?
Audio ducking is lowering one sound, usually background music, automatically whenever another sound, usually speech, is playing, then bringing it back up when the speech stops. The music ducks out of the way of the voice. It is defined by three things: how far the music drops, how quickly it fades, and what triggers it.
Music that sits at one level for a whole video is either too quiet in the gaps or too loud under the words. Ducking is the reason a voice stays clear without the music disappearing from the rest of the film.
Nuwan Madhusanka · Co-founder
5 min read · Published
| Approach | Depth | Ramp | Trigger |
|---|---|---|---|
| WCAG 1.4.7 guideline (AAA) | Background at least 20 dB below foreground speech, about four times quieter | Not specified | Any prerecorded speech; occasional sounds of one or two seconds exempt |
| Premiere Pro auto ducking | Duck amount chosen in the Essential Sound panel | Fade duration and fade position chosen by the editor | Clips tagged as Dialogue, then keyframes generated on the music |
| Video builder here, narration | Music drops to 22 percent of its set level, about 13 dB | 12 frames each side, 0.4 seconds at 30 fps | Narration only |
| Video builder here, uploaded clip audio | No ducking | None | Keep and duck is offered but not wired |
Making room for speech
Every mix has a foreground and a background, and this technique is how the background makes room without being switched off. Between lines, music carries mood and pace; under a line, it has to step back far enough that each word is understood. Doing that by hand means drawing volume changes around every sentence, which is slow and breaks the moment the script changes. Ducking automates the same shape. The editor decides the depth and the ramp once, marks which track is speech, and the software lowers the music wherever that track has sound.
Depth, ramp and trigger
Depth is how much quieter the music becomes, usually stated in decibels or as a percentage of its original level. A drop to 50 percent is about 6 dB, to 25 percent about 12 dB, to 10 percent 20 dB. Ramp is how long the change takes. Too fast and the music audibly lurches down at the first syllable; too slow and the opening words land on full volume music. Many tools start the fade slightly before the voice so the dip is complete when speech begins. The trigger is whatever the software listens to: a tagged dialogue track, a sidechain input, or a list of narration times. Premiere Pro, for example, asks you to tag clips as Dialogue and Music, then generates volume keyframes on the music track.
When it matters
Ducking matters most where speech carries information and music runs underneath: explainers, promos with voiceover, tutorials and interviews. It matters for accessibility as well. The W3C guideline on background audio sets a stricter bar than most marketing mixes meet, asking that background sound sit at least 20 dB below foreground speech, roughly four times quieter, so that a listener who is hard of hearing can separate the words from the bed. Even where that level is not required, getting close to it is a reliable way to make narration readable on a phone speaker in a noisy room.
Ducking and the loudness figure
Ducking changes the balance inside a mix, not its delivery loudness. EBU R 128 measures a programme in its entirety, speech, music and effects together, so a video with deep ducking and one with none can report a similar integrated figure while being very different to follow. That is why the two jobs are done in order. First the balance: speech clearly on top, music stepping down under every line and returning between them. Then the level: the finished file is measured and, where a platform or broadcaster sets a target, adjusted as a whole, so the balance built in the first step survives the second.
Common mistakes
The first is a shallow duck. A drop of 3 or 4 dB sounds like a mistake rather than a choice, and busy music still masks consonants. The second is pumping: short gaps between sentences make the music rush back up and down, which is distracting, so a longer release or ignoring very short gaps sounds smoother. The third is ducking against the wrong thing, such as sound effects, so the music dips for a door slam. The fourth is judging the result on studio monitors. Loudness is measured over the whole programme, but ducking is heard line by line, so check it on the speaker your audience will actually use.
Where it shows up in the product
In the video builder here, ducking is automatic and follows the narration. Wherever a scene's voiceover plays, the background music drops to 22 percent of its set level, ramping down over 12 frames before the line and back up over 12 frames after it, which is 0.4 seconds at the 30 fps every template uses. There is no depth or ramp control. Music also starts from a default volume of 25 percent, so it sits low before any ducking. One limit is worth planning around: an uploaded video clip never ducks the music. Its keep and duck audio setting is offered but not wired, so if a clip carries speech, lower the music track or mute the clip. The music page covers the library and the fades.
Questions people ask
What is sidechain ducking?
Sidechain ducking uses a compressor on the music track that listens to a different signal, the voice, instead of the music itself. When the voice gets louder than a threshold, the compressor turns the music down by a set ratio, and it recovers when the voice drops. It reacts to the actual level of speech rather than to a fixed timeline, which is why radio and live streams favour it.
Is ducking the same as a fade?
No. A fade takes a sound from silence to full level or back to silence, usually at the start or end of a piece. Ducking lowers music to a reduced level while it keeps playing, then returns it to where it was. A single video can use both: a fade in on the opening scene, ducking under every line of narration, and a fade out at the close.
Should the music stop completely under a voice?
Rarely. Cutting music to silence under every sentence makes the bed start and stop constantly, which draws more attention than a steady lower level. A deep duck keeps the rhythm and mood running while leaving the words clear. Silence works as a deliberate moment, such as the one line that carries the offer, rather than as the default treatment for all speech.
Does ducking work for music with vocals?
It works less well. A sung vocal sits in the same frequency range as speech, so even a deep duck leaves competing words that make narration harder to follow. Instrumental tracks are the safer choice under voiceover. If a track with vocals is essential, keep narration to the instrumental sections or lower the music far more than you would for an instrumental bed.
How do I check whether my ducking is deep enough?
Play the video on a phone speaker at a moderate volume, in a room with some background noise, and listen for any word you have to replay. Then read along with the captions switched off. If a product name, price or date is hard to catch, the music is still too loud under that line, whatever the settings say.
Make one with videos
The button opens the generator with this use case already described. Change the wording to match your own.
Create a video with OneCraftRelated questions
- What is LUFS?What is LUFS? The loudness unit broadcasters and podcast apps use, how it is measured, and the targets set by EBU R 128, ATSC A/85 and Apple Podcasts.
- Voiceover vs narration: is there a difference?Voiceover vs narration: how the two words overlap, where film, advertising, e-learning and accessibility use them differently, and what to ask for in a brief.
- What is B-roll?What is B-roll? The supporting footage cut over the main recording, how it differs from A-roll, and the shots that make good cover for interviews and ads.
Step by step in the builder: Add voiceover, captions and music to a video.
Written and checked by the OneCraft team. Last checked .