How Stem Separation Quality Is Measured

Flat vector illustration of a music track splitting into vocal, drum, and bass stems, with a signal-quality meter showing separation quality

Stem separation quality is measured two ways at once: objectively, as a single number called SDR (signal-to-distortion ratio, in decibels), and subjectively, by how much of the other instruments leak into a stem and how many strange, watery artifacts you can hear. When you run a track through an AI stem splitter, the model is trying to hit both targets — a clean number and a clean listen — and the two don't always agree. This explainer covers what "good" actually means, what the number measures, and how to judge the results yourself.

What "good separation" actually means

The goal of stem separation is easy to state and hard to reach: take a finished mix and split it back into its sources — vocals, drums, bass, and everything else — so each source lives cleanly in its own stem. Perfect separation would mean the vocal stem contains only the vocal and the drum stem only the drums, with nothing leaking across and nothing smeared or distorted along the way.

That ideal never fully happens. A mix is a single stereo file where every instrument overlaps in time and frequency, and no algorithm can perfectly undo that sum. So "good separation" is really a question of how close you get, measured by how much unwanted sound leaks in and how much the wanted sound gets damaged on the way out. Those two failure modes, bleed and artifacts, are exactly what the objective score decomposes into.

Separation quality as an objective score: SDR

The standard objective metric is SDR — signal-to-distortion ratio, measured in decibels (dB). It compares a separated stem against the true, original stem (the "ground truth") and expresses how much of the output is the correct signal versus everything that shouldn't be there. Higher is better: more SDR means a cleaner stem with less unwanted content. Because it needs the original stems to compare against, SDR is measured on reference datasets where those stems are known — not on your own song, where the true parts are the very thing you're trying to recover.

SDR is the metric public source-separation benchmarks are built on. Community evaluations like the Music Demixing challenges and the earlier SiSEC campaigns rank models by their average SDR across a hidden test set, which is why it's the number researchers quote when they line one model up against another.

SDR is the headline figure, but it belongs to a family of related measures (often called BSS Eval) that break the same idea into parts:

  • SIR (signal-to-interference ratio) captures bleed — how much of other sources leaks into this stem.
  • SAR (signal-to-artifacts ratio) captures artifacts — the distortion the algorithm itself introduces.

That decomposition is where it gets useful: bleed and artifacts aren't separate from the score, they are the two things the score is made of. As a rough, directional rule, the catch-all "other" stem tends to score lowest, while vocals, drums, and bass score higher — a grab-bag of leftover instruments is far harder for a model to reconstruct than a single, identifiable source. Exact figures vary enormously by model and by song, so treat any single number with skepticism — the relative comparison tells you more than the absolute value.

Bleed: when other parts leak in

Bleed (also called interference or leakage) is when sound from one source shows up in the wrong stem. The classic example: you solo the vocal stem and hear a faint snare ghosting underneath every backbeat, or a smear of cymbal in the quiet gaps between phrases. The vocal is mostly there, but the drums didn't fully leave.

Bleed happens because sources overlap in frequency. A snare and a vocal both put energy in the low-mids, so when the model carves out "the vocal," it can't cleanly avoid pulling a little snare with it. This is also why older subtractive methods struggle: techniques that try to cancel the center-panned vocal rely on phase cancellation, and anything sharing that center position — kick, snare, bass — gets partially cancelled or partially left behind. AI models like the one behind our splitter learn the shape of each instrument instead of just its position, which is why they handle bleed better than subtractive tricks, though never perfectly.

Artifacts: the watery, smeared sound

Artifacts are distortions the separation process creates that weren't in the original at all. Where bleed is real audio in the wrong place, artifacts are new, unnatural sound. They show up as a watery or underwater quality, warbling, a faint robotic or metallic ring, or a "pre-echo" smear just before a drum hit. Engineers often call the swirly, granular version "musical noise."

Artifacts appear because of masking and overlapping frequencies. When two sounds occupy the same frequency at the same moment, the model has to guess how to divide that shared energy — and when it guesses imperfectly, the reconstruction warbles. Pushing harder makes it worse: straining to remove every last trace of the other instruments tends to gouge holes in the frequency content, and those holes are what your ear hears as artifacts. It's the same underlying reason vocal removal leaves artifacts — there's an unavoidable tradeoff between removing more bleed and introducing more distortion.

Why the number and your ears can disagree

A high SDR doesn't guarantee a stem that sounds good, because SDR is an average computed over the entire file. A track can score well on average while hiding one ugly, glaring artifact in an exposed section — a warble on a solo vocal line, or a burst of musical noise in a quiet breakdown.

To a listener, that one moment can wreck the whole take even though it barely moves the average. The reverse happens too: a stem can score modestly yet sound perfectly usable for what you need. SDR measures mathematical closeness to the original; your ears measure perceived quality, and they weight exposed, quiet, and solo moments far more heavily than a whole-file average does. That gap is the main reason objective and perceptual assessments of stem separation quality can point in different directions, and why you should always listen rather than trust a benchmark on its own.

What determines the quality you get

The same model produces very different results on different songs, because separation quality depends heavily on the input:

  • Arrangement density. A sparse arrangement — voice, one guitar, light drums — separates far more cleanly than a wall of layered synths, doubled vocals, and dense percussion all fighting for the same frequencies.
  • Mix and source quality. A clean, well-recorded, high-bitrate file gives the model more to work with. Heavy compression, distortion, or a lossy source that has already thrown away detail all hurt the result.
  • Genre and production. Music with clearly separated, distinct instruments tends to split better than material built on dense, overlapping textures.

The full picture of why some songs separate cleaner than others goes deeper, but the short version is that the cleaner and simpler the source, the cleaner the stems.

How to judge separation quality yourself

You don't need a benchmark dataset to evaluate your own results — you need your ears and a couple of habits:

  1. Solo each stem and listen for bleed. Play the vocal alone and listen in the gaps between phrases; play the drum stem and listen for vocal remnants. The quiet moments expose leakage best.
  2. Listen for artifacts on exposed parts. Focus on solo vocal lines and quiet sections, where watery or robotic textures are easiest to catch.
  3. Look at a spectrogram. A spectrogram shows leftover energy as visible smears or ghost traces where a stem should be silent — a fast visual check that backs up what you hear.
  4. A/B against the mix. Play a stem next to the original to confirm it captured the whole part and didn't drop notes.

Because our AI-based audio stem splitter runs entirely in your browser — HTDemucs separation via onnxruntime-web, producing four stems (vocals, drums, bass, and other) that you can solo, mute, mix, and download in full stereo — the most honest test is to run it on your own file and listen. The first run downloads the AI model to your browser and the separation takes real time on your own CPU, but your audio never leaves your device. For the full walkthrough, see our guide on how to split a song into stems, and for the underlying mechanics, how AI vocal removal works.

Frequently asked questions

Is SDR the same as sound quality?

No. SDR measures how mathematically close a separated stem is to the original, averaged across the whole file. Perceived sound quality is what your ears judge, and they weight exposed or quiet moments far more than an average does. A stem can score well yet contain one obvious artifact that ruins it, so always listen rather than trusting the number alone.

What is a good SDR score?

Higher is better — more decibels means a cleaner stem with less bleed and fewer artifacts. There's no single "good" threshold, because scores vary widely by stem and by song: the catch-all "other" stem typically scores lowest while vocals, drums, and bass score higher, and a simple arrangement scores far above a dense one. Compare results relative to each other rather than chasing an absolute figure.

Why does my separated vocal still have drums in it?

That's bleed. The vocal and drums share frequency ranges, so when the model isolates the vocal it can't perfectly avoid pulling some percussion along. It's most audible on center-panned elements like snare and kick, and in the quiet gaps between vocal phrases. A cleaner, sparser source usually reduces how much leaks through.

Can separation ever be perfect?

No. A mix is a single file where instruments overlap in time and frequency, and no algorithm can perfectly undo that sum. Some degree of bleed and some artifacts always remain — the goal is to get close enough for your purpose, not to reach a flawless split. AI models get meaningfully closer than older subtractive methods, but perfection isn't achievable.