How AudioShake's Stem Separation Makes Content More Adaptable for Global Audiences

AudioShake
August 18, 2026

Content today rarely stays in the market where it was made. A documentary produced in English often needs to be dubbed into Spanish or Japanese later on. A corporate training video needs a separate version for every regional office its parent company operates in. In both cases the audio arrives as one mixed track, which forces a localization team to either dub on top of it, duck or remove it, or separate out its individual elements. Copyright compliance adds another layer, since licensed music has to be identified and cleared before content can move across countries and platforms. Here is how AI stem separation solves these problems, and how teams across film, dubbing, and sports media are already putting it to work.

TL;DR

  • Dubbing requires clean dialogue stems and an intact M&E (music and effects) track before any translation work can start.
  • Many productions, especially older ones, no longer have access to their original multi-track audio files.
  • AudioShake’s AI separates a single mixed track into clean dialogue, music, and effects, even when the original session files are gone.
  • Our copyright compliance system helps teams detect and clear licensed music so content can be distributed across different countries and platforms.
  • Teams including CrunchLabs, cielo24, and ESPN already use AudioShake across dubbing, film, sports, and broadcast.

Create M&E tracks without the original stems

Everything in a mix besides the spoken dialogue — score, ambience, sound design, crowd noise — is what the industry calls the M&E track, short for music and effects. A clean M&E track is what lets a dub sound natural instead of hollow, because the original music and effects carry through while only the dialogue changes language.

How AudioShake works

Traditionally, a clean M&E track comes from the original session files, where the production kept the dialogue, music, and effects on separate tracks. For older or archival titles those files are often lost, leaving only the finished broadcast mix. AudioShake’s separation model reconstructs the stems from that mix alone, pulling the fully mixed audio apart into a clean dialogue stem to remove and an intact M&E track to dub over, with no original source files required.

Production studio Pandastorm Pictures ran into exactly this when it set out to dub the original 1960s-era Doctor Who into German. The only surviving source was the fully mixed English broadcast, with dialogue, music, and effects fused into a single track. Layering German dialogue over that mix meant clashing directly with the English dialogue underneath.

Working with AudioShake, Pandastorm separated the audio into individual stems, isolating the original music and effects from the dialogue. With a clean M&E track reconstructed, they pulled the English dialogue and dubbed new German audio in its place.

As Dennis vom Berg, Senior Product Manager at Pandastorm Pictures, put it, the alternative was an expensive from-scratch re-score or living with gaps in the audio. Stem separation made the localization both faithful to the original and economically viable.

Clean dialogue for improved translation and transcription accuracy

In localization and dubbing workflows, an isolated dialogue stem does more than support a single title. It is what makes localizing an entire back catalog realistic. Legacy libraries are often delivered as finished mixes with no separate stems, so every show a studio wants to take international traditionally needs manual audio post-production before dubbing can even begin. Separating the dialogue first with AudioShake’s dialogue isolation removes that bottleneck, giving teams clean dialogue and M&E stems on demand instead of one painstaking title at a time.

Antenna Group, the Greek broadcaster behind ANT1 TV, faced exactly this with its catalog of classic sitcoms, dramas, and entertainment. Much of that legacy content was delivered without separate audio stems, which kept it locked to the domestic market, and creating M&E stems and replacing dialogue by hand could not scale across the whole library. By bringing AudioShake into its localization workflow, Antenna isolates dialogue cleanly from decades-old master recordings while preserving the original music and sound design, and generates the M&E stems dubbing studios need for international versions. Classic shows that once served only Greek audiences now earn revenue through global licensing, without Antenna having to expand its post-production teams.

“AudioShake now allows us to isolate dialogue and music stems quickly, so our teams can focus on creative and editorial decisions rather than technical constraints.”
— Costas Colombus, Group CTO, Antenna Group

Global copyright compliance

Clean audio does not by itself mean the content can be distributed. Music licensing and distribution rules vary by country and platform, and a track cleared for one market can cause legal trouble in another. That is the purpose of AudioShake’s Copyright Compliance System, which flags and removes copyrighted music so teams can clear content before it goes out to a new region.

Mark Rober’s CrunchLabs ran into this exact problem as its science and engineering videos scaled to a global audience of millions. Much of the archival library was fully mixed, which made it difficult to isolate dialogue from music and effects for localization, or to produce any alternate versions at all.

Using AudioShake’s music removal and stem separation, the team scaled its ability to produce localized versions in multiple languages. They ran content through AudioShake to lift out a clean dialogue stem for translation and dubbing, and to remove commercial music where licenses had expired or were never cleared for international release. That let them produce versions suitable for sponsorships, brand partnerships, and platforms with strict copyright policies, while retaining the original energy and reactions from the video. Years of archived videos that had been stuck in a single mixed track became usable again.

ESPN is solving the same problem with AudioShake. As one of the leading sports media companies, ESPN manages a vast library of live, archival, and digital content that has to navigate a wide set of rights guidelines across its platforms. With AudioShake, ESPN can isolate or remove copyrighted music from sports highlights, preserve authentic commentary and crowd energy, and speed up the turnaround of rights-cleared assets for distribution.

In one case, ESPN used AudioShake to pull Phil Simms’ vocal cleanly from a 35-year-old mixed master of his “I’m going to Disney World” moment for a 2026 Super Bowl ad, without needing to license the music in the original recording. Read the full ESPN case study.

“Working with AudioShake to leverage their innovative audio separation lets us unlock more content for fans by accelerating and modernizing workflows and ensuring we can deliver more high-quality sports content to fans wherever they are.”
— Kevin Lopes, VP, Business Development and Innovation, ESPN

Whether you are preparing a single production or an entire content library for international distribution, AudioShake’s AI-powered stem separation makes existing content easier to adapt for new languages, new markets, and new audiences. Have a project you are not sure how to approach? Reach out and tell us where you are stuck. We would love to hear about it.

Get started with AudioShake now

Ready to try it on your own content? Dialogue, music, and effects separation is available on AudioShake Indie, our web-based platform for creators of all sizes. Those looking for enterprise solutions can use AudioShake Live by getting in touch or get started with our API today.