AudioShake Releases Multi-Speaker 2.0 to Isolate Overlapping Dialogue into Discrete Stems

September 22, 2026

AudioShake today released Multi-Speaker 2.0, a dialogue separation tool that takes a single mixed conversation and splits every voice onto its own clean audio stem. AudioShake’s separation technology is already trusted across major post-production pipelines, including ESPN, NFL Films, Warner Bros. Discovery, Paramount, and Deluxe.

‍

Crosstalk is the oldest unsolved headache in dialogue editing. Two guests finish each other’s sentences, someone laughs over a punchline, or a lav mic catches the host sitting next to them. From that point on, the voices are married to a single signal. Dialogue editors are usually left with bad options: carve around it, gate it aggressively, or bury it in the mix.

‍

Multi-Speaker 2.0 isolates the voices without destroying the natural delivery.

‍

Isolated Dialogue and Ambience in One Pass

‍

Using AudioShake’s Studio platform or API, editors can upload a mixed file and export:

‍

  • Discrete speaker tracks matched directly to the source sample rate, from 8 kHz phone audio up to 48 kHz broadcast files.
  • An isolated ambience stem containing the room tone, traffic, and background noise, allowing editors to duck, sweeten, or rebuild the soundstage with total control.
  • Targeted confidence markers: Instead of listening through an hour of dialogue to audit where the separation held, the tool flags low-confidence moments down to the 20-millisecond frame. Editors can jump straight to the few seconds that need an ear check.

‍

Traditional auto-transcription tools only provide speaker labels—marking who spoke when while leaving the underlying audio married on one track. Multi-Speaker 2.0 pulls the actual speech apart, cleanly isolating brief interjections, backchannel agreements, and rapid-fire overlapping lines.

‍

Built for Unscripted, Dubbing, and Archival Audio

‍

The update targets the standard failure points across post-production workflows:

‍

  • Unscripted TV, Documentaries, and Podcasts: Fixes bleed from shared room mics and side-by-side lavalieres, allowing mixers to level, EQ, and clean up one speaker without affecting the others.
  • Localization and Dubbing: Provides pristine, single-speaker dialogue stems required for clean foreign-language replacement tracks and voice-over beds.
  • Subtitling and Captions: Prevents the drops, garbled phrasing, and timing errors that occur when transcription engines hit simultaneous speech.
  • Archival and Field Audio: Salvages mono mixes, legacy radio call-ins, and degraded field recordings where no original session tracks or multitracks exist.

‍

Key Technical Upgrades in 2.0

‍

  • Sharper Isolation in Noisy Environments (32% Less Bleed vs. 1.0): Cuts significantly more ambient contamination alongside vocal crosstalk, delivering isolated stems that don’t drag audible room noise or voice bleed behind separated words.
  • Full-Band Coverage (8 kHz to 48 kHz): A single unified engine handles everything from low-bandwidth mobile calls to hi-res production masters without resampling artifacts.
  • Language-Agnostic Processing: Operates as a purely acoustic model rather than a predictive language model, preserving accented delivery, multilingual conversations, and overlapping vocalizations without hallucinated edits.

‍

“Some of our richest, most human moments come at the point of overlap — an interjection, a laugh, finishing someone else’s sentence,” said Jessica Powell, co-founder and CEO of AudioShake. “In the editing suite, those moments have always forced a compromise between performance and audio quality. Multi-Speaker 2.0 lets editors keep the performance: every voice on its own track, with clear visual flags and metadata showing the exact seconds that need human review.”

‍

Multi-Speaker 2.0 is available today in AudioShake Studio and via the AudioShake API.