What Is Audio Source Separation? How AI Unmixes Sound

AudioShake
October 6, 2026

Almost every recording you've ever heard is a mix. A song is vocals, drums, bass and guitar blended into one file. A movie soundtrack is dialogue, music and sound effects mixed together. A podcast recorded in a café is two voices plus clinking cups, an espresso machine and someone else's conversation.

‍

Source separation is the process of taking that single mixed recording and pulling it back apart into its individual sources. In music the separated parts are usually called stems, so you'll also hear it called stem separation, demixing or unmixing.

‍

Humans do a version of this all the time. At a noisy party you can focus on one person's voice and tune out the rest. Researchers even call it "the cocktail party problem." For machines, it has been one of the hardest problems in audio.

‍

A short history: from clever tricks to deep learning

‍

People have tried to separate audio for decades. Early methods relied on signal-processing tricks. For example, karaoke machines often removed vocals by cancelling out whatever sat in the center of a stereo mix. That worked sometimes, but it also removed the bass and kick drum, and it failed completely on mono recordings. Later statistical methods did better in controlled conditions, but they struggled with real-world, messy audio.

‍

Doing it by hand was possible but brutally slow. A member of the iconic Bomb Squad once told us that it would take him days to isolate just a small sample of audio.

‍

Deep learning changed that. Instead of hand-writing rules about what a voice or a guitar sounds like, you can train a model on enormous amounts of audio and let it learn those characteristics itself. Today a well-trained model can separate a full song in seconds.

‍

Why we can explain audio with a picture

‍

There are various approaches to source separation, and we don't use just one at AudioShake. But one of the clearest methods to explain starts with a visual. It may seem odd to explain something you hear with an image, but one of the clearest ways to understand source separation is to look at a spectrogram, the "image" of audio.

A spectrogram is made by chopping audio into tiny slices of time and measuring how much energy is present at each frequency in each slice. Plot that out and you get a picture:

‍

  • The x-axis is time.
  • The y-axis is frequency, from low (bass, kick drums) at the bottom to high (cymbals, the "s" sounds in speech) at the top.
  • Brightness represents intensity: how much energy is present at each frequency at each moment.
  • ‍

What you don't get? What is in that sound, or its recording characteristics. This spectrogram could be of a song, a film, people talking – but you and the machine don't know. And to make things more difficult, most likely, there is a lot of overlap in sound in both time and frequency. Source separation helps us make sense of that.

‍

‍[IMAGE: real spectrogram]

‍

Look at a real spectrogram and you can't tell what you're looking at. It could be a song, a film scene or two people talking. Every sound in the recording is layered into the same picture, overlapping in both time and frequency.

‍

In the video above, we use a Lite-Brite as a very low-resolution spectrogram. Each peg is a "pixel" of sound: blue at the bottom for bass, pink for vocals, orange for guitar. Separating the guitar means figuring out which pegs are orange and then pulling out every peg that isn't. What's left is just the guitar. Real audio is far messier than a toy (the colors overlap constantly), but the idea is the same.

‍

How source separation works

‍

At a high level, many separation systems do three things:

‍

  1. Learn what sounds look like. A model is trained on huge amounts of audio where the answer is known: recordings where we have both the full mix and the isolated sources. Over time our model learns the patterns that distinguish a voice from a violin, or dialogue from a crowd.
  2. Identify which parts of the mix belong to the target. Given a new recording, the model estimates, for every point in time and frequency, how much of the energy belongs to the sound the model is targeting. This estimate is often called a mask. Importantly, it's rarely all-or-nothing. When a voice and a guitar share the same frequency at the same moment, the model has to split that energy between them.
  3. Keep the target, suppress everything else. The mask is applied to the original recording, keeping what belongs to the target and silencing or reducing the rest. The result is converted back into audio you can hear.

‍

That's the classic spectrogram-based approach, and it's the easiest one to picture. Many modern systems also work directly on the raw waveform, or combine both views. The core idea is the same: identify what's in the recording, isolate what you want, and suppress what you don't.

‍

Why source separation is a hard problem

‍

If separation were just "find the guitar pixels," it would have been solved long ago. A few things make it particularly difficult:

‍

  • Sounds overlap. Instruments and voices constantly occupy the same frequencies at the same time. In a real mix, the pixels don't belong neatly to one source.
  • Sounds resemble each other. A piano and a guitar can share a lot of characteristics. So can two people with similar voices, which makes separating one speaker from another especially tricky.
  • The real world is messy. Reverb, background noise, cheap microphones and compression all blur the picture.
  • Audio has no obvious boundaries. Text has words and images have edges. Raw audio is a continuous signal, and multiple sources can occupy it simultaneously, so a model has to infer which parts of that signal belong to which source.
  • Clean training data is scarce. The world doesn't naturally produce neatly labeled, isolated recordings. You can't just scrape your way to a great separation model. In AudioShake’s case, we have spent years building our training database, which includes sound from production libraries, media partners, and our own created sounds. 

‍

Traditional source separation vs. generative AI

‍

We're often asked what the difference is between AudioShake and generative AI. In fact, many times we're asked whether we've built a wrapper around ChatGPT or Claude! (To be clear, that exclamation point is our way of saying "no, not the same tech, totally different, no.")

‍

Modern source separation is typically performed with AI, but not usually with generative AI.

‍

Source separation asks: what's in this recording, and can we isolate it? It's a subtractive process. It works with the audio that was actually captured and removes what you don't want.

‍

[IMAGE: Lite-Brite — full spectrogram, then an arrow, then the guitar isolation image]

‍

Generative AI asks: can we create audio that sounds like this? It's an additive process. If you ask a generative model to produce the guitar part of a song, it will make its best guess at what that guitar should sound like. The result can sound great. But because the model is generating audio, it can also add things that were never in the original, or put notes or sounds in the wrong places.

‍

Neither approach is "better" in the abstract. For creative work — exploring new sounds, sketching ideas, remixing for fun — a generative approach may be exactly what you want. But when faithfulness to the original matters, separation is usually the right tool. Examples include:

‍

  • News and editorial, where you need to clean up audio without altering what was said
  • Forensics and legal review, where the recording is evidence
  • Archives and restoration, where the goal is to preserve the original performance
  • Structured training data, where you want clean inputs that reflect what was really recorded

‍

Separation isn't magic. A model can leave traces of other sounds behind, distort parts of the target, or remove too much. But unlike a generative model, its job isn't to synthesize a plausible replacement for missing audio; it's to recover the target from the recorded mixture.

‍

You can read more about the differences between generative AI and targeted source separation in our blog discussing Meta's SAM Audio model.

‍

What source separation is used for

‍

The same core technology shows up across a surprising range of industries:

‍

  • Music: creating stems for sync licensing (film, TV and ads), remixes, remasters, immersive and spatial audio, and karaoke from original recordings rather than re-recorded soundalikes. It can also be used for metadata extraction as an input into playlisting and recommendations. 
  • Film and TV: separating dialogue, music and effects for dubbing and localization, dialogue cleanup, and removing copyrighted music from clips before they're distributed.
  • Speech and voice AI: isolating voices to improve transcription and captioning, powering real-time voice applications, and preparing clean, structured training data.

‍

Across all of these, the value is similar. Separation unlocks audio that was previously unusable, whether because the stems were never saved, the original tapes are sitting in a warehouse, or the recording was made somewhere noisy.

‍

How separation quality is measured

‍

Researchers typically score separation with metrics such as signal-to-distortion ratio (SDR). SDR compares a model's output with the true isolated source, and higher is better. AudioShake typically leads the SDR leaderboards in the models it builds. That said, we'll be the first to say that SDR has its drawbacks. Numbers only tell part of the story. Real-world performance on messy, unfamiliar audio, and how the result actually sounds to a trained ear, matter just as much.

‍

FAQ

‍

Is source separation the same as stem separation?

Mostly, yes. "Stem separation" is the term usually used in music. "Source separation" is the broader term, and it also covers dialogue, sound effects, individual speakers and noise. You'll also hear different terms used in different industries as proxies for these terms — such as channel separation, separated tracks, or separated streams.

‍

Is source separation magic?

Yes! Absolutely. Though if you'd prefer a more technical explanation than "magic," we'd suggest reading this blog :)

‍

Does source separation create new audio?

Traditional, non-generative source separation doesn't generate a replacement for the source. Its goal is to isolate the audio already present in the recording. Some newer research does use generative models for separation.

‍

Can source separation separate two people talking at the same time?

Yes, though it's one of the harder cases, especially when the voices are similar or overlap heavily. This is often called speaker separation. Read more about AudioShake's multi-speaker technology.

‍

Can source separation run in real time?

Modern models can be efficient enough for low-latency, real-time use, such as live captioning or voice applications. Read more about AudioShake's real-time models.

‍

Is source separation a type of AI?

Modern source separation is built on AI, specifically deep learning models trained to recognize sounds. AudioShake's approach is different from generative AI: it analyzes and isolates rather than creates. That said, there are generative approaches that can also be used to perform source separation.