How to · Audio

Picking music that doesn't fight your voiceover

The right music makes the voiceover sound smarter - the wrong music makes the voiceover sound like it's shouting.

8 min read · Audio

Most explainer mixes lose the fight at the music selection stage, not at the mix stage. You can sidechain, EQ, and compress all day, but if the Bed Music occupies the same frequency space as the human voice, the listener is going to feel like the two are wrestling.

This is the working approach for picking music that supports the voiceover instead of competing with it.

The frequency rule

The human voice lives between roughly 100 Hz and 4 kHz, with the most intelligibility-critical range concentrated between 1 kHz and 3 kHz. If your music has dominant elements in that range - lead vocals, lead piano, lead guitar - they're going to fight the voiceover no matter what you do at the mix stage.

The cleanest bed music sits around the voice, not in it:

  • Subs and lows (below 80 Hz) - bass, kick, sub-pads. The voiceover doesn't live here.
  • Lower mids (80–500 Hz) - fine in moderation; this is where music feels warm.
  • Voice range (500 Hz–4 kHz) - keep music elements sparse or absent here.
  • Highs (4 kHz–12 kHz) - cymbals, hi-hats, sparkly synth pads, all fine.
  • Air (above 12 kHz) - ambient air noise, very high pads. Fine, and adds polish.
A music bed that's heavy in lows and highs and sparse in mids is engineered to coexist with narration. A music bed with a lead vocal or a busy piano is engineered to be foreground.

Tempo math

A second rule: music tempo should match the speech tempo, not contend with it.

  • Slow narration (130 WPM) - music at 70–90 BPM.
  • Standard explainer (150 WPM) - music at 90–110 BPM.
  • Fast / energetic (175 WPM) - music at 110–130 BPM.
Faster than that and the music is racing the voice; slower and it drags. Test by reading three sentences over the bed - if the voice naturally lands on beats, the tempo is right. If the voice and the beat keep stepping on each other, the tempo is wrong.

The four-archetype shortlist

Most explainer music collapses into four archetypes. Pick the archetype before you scroll the library - otherwise you'll spend forty minutes auditioning and end up with whatever you played last.

Cinematic Underscore. Slow, pad-heavy, no drums or very light percussion. For somber explainers, deep dives, founder stories. Risks feeling pretentious if the content is light.

Lo-Fi / Beats. Mid-tempo, soft drums, mellow chords. For casual explainers, walk-throughs, behind-the-scenes. Risks feeling generic if the content is high-stakes.

Indie / Acoustic. Real instruments, mid-tempo, builds gently. For lifestyle, craft, product features. Risks fighting the voice if the acoustic instruments live in the same range as the speech.

Synth / Modern Electronic. Heavy lows, sparkly highs, drum machine, often instrumental. For tech, futurism, fast cuts. Risks feeling cold for warm content.

If you're hesitating between two archetypes, pick the one a level quieter in personality. The music's job is to support, not to perform.

Sting and Musical Sting placement

A Musical Sting is a short musical accent - a hit, a chord, a swell - used to mark a transition or punctuate a moment. Stings are the single biggest difference between a video that feels edited and one that feels uploaded.

Three rules for stings:

  • One per chapter, max two. A video with seven stings is anxious. A video with two is paced.
  • Land them on cuts, not on speech. A sting that lands while someone's talking competes with the talking; a sting that lands on a hard cut feels like a frame.
  • Match the sting to the bed. If your bed is acoustic indie, your sting should be an acoustic chord, not an orchestral hit. Mixed-genre stings call attention to themselves.

Ducking - automatic or manual

Ducking is the trick of lowering the music whenever the voice is present. Two ways to do it:

  • Sidechain compression - automatic, ties the music level to the voice signal. Set the threshold so the music drops 4–6 dB when the voice is above −24 dB, and recovers over 200–400ms when the voice drops out.
  • Manual automation - draw a volume curve under the music track. Slower, more precise, and lets you taper the music up gently between sentences rather than yanking it.
For explainers with fast cuts, sidechain is fine. For long-form essay videos, manual usually sounds better - the gentler the music's return, the more cinematic the cut.

Length math

A 60-second video needs a 60-second music piece, ideally with a 5-second intro and a 5-second outro that fade naturally. Stock libraries usually offer 30s, 60s, 90s, and 2:30 versions of the same track. Pick the version that's slightly longer than your video and fade the end - never try to extend a track by looping it. Loops are audible.

For longer videos, look for tracks with built-in arc - they start sparse, build, then resolve. A flat music bed across 5 minutes is exhausting to listen to even at low volume; an arc'd bed gives the voiceover something to ride against.

Licensing reality

A reminder that gets people in trouble: a track being on YouTube doesn't mean you can use it. A track being in a stock library doesn't mean you can use it without crediting it. A track being "free" usually means free for personal use - commercial use is a separate license.

Three reliable options:

  • Epidemic Sound / Artlist / Musicbed - subscription, broad commercial license, predictable.
  • YouTube Audio Library - free, but no commercial license for anything beyond YouTube itself.
  • Direct licensing from artists - for hero pieces. Sites like Bandcamp let you contact musicians directly; small artists are often happy to license at reasonable rates.

A test that takes 30 seconds

Drop the music in your timeline, mute the voiceover, listen to 10 seconds of music alone. Then unmute the voiceover and listen to the same 10 seconds. If the music feels different when the voice is on it - quieter, hollower, somehow recessed - the music's frequency profile is conflicting with the voice's. Pick a different track.

If the music feels the same with or without the voice, you've found a bed that occupies different territory than the voice does. That's the bed you want.

Bring it to VideoCue

The platform's music library is filtered by archetype, tempo, and frequency profile - the four-archetype shortlist above maps to the categories in the picker. For the underlying audio theory, the Bed Music entry covers the role-of-music-in-narration tradition, Musical Sting catalogs the punctuation patterns above, and Ducking walks through the sidechain math in more depth.

Related Filmwiki terms

Your first cut
is on us.

+ New videoFree plan · no card · 1080p export