AudioShake Releases Multi-Speaker 2.0 to Isolate Overlapping Dialogue into Discrete Stems
AudioShake today released Multi-Speaker 2.0, a dialogue separation tool that takes a single mixed conversation and splits every voice onto its own clean audio stem. AudioShake’s separation technology is already trusted across major post-production pipelines, including ESPN, NFL Films, Warner Bros. Discovery, Paramount, and Deluxe.
Crosstalk is the oldest unsolved headache in dialogue editing. Two guests finish each other’s sentences, someone laughs over a punchline, or a lav mic catches the host sitting next to them. From that point on, the voices are married to a single signal. Dialogue editors are usually left with bad options: carve around it, gate it aggressively, or bury it in the mix.
Multi-Speaker 2.0 isolates the voices without destroying the natural delivery.
Isolated Dialogue and Ambience in One Pass
Using AudioShake’s Studio platform or API, editors can upload a mixed file and export:
- Discrete speaker tracks matched directly to the source sample rate, from 8 kHz phone audio up to 48 kHz broadcast files.
- An isolated ambience stem containing the room tone, traffic, and background noise, allowing editors to duck, sweeten, or rebuild the soundstage with total control.
- Targeted confidence markers: Instead of listening through an hour of dialogue to audit where the separation held, the tool flags low-confidence moments down to the 20-millisecond frame. Editors can jump straight to the few seconds that need an ear check.
Traditional auto-transcription tools only provide speaker labels—marking who spoke when while leaving the underlying audio married on one track. Multi-Speaker 2.0 pulls the actual speech apart, cleanly isolating brief interjections, backchannel agreements, and rapid-fire overlapping lines.
Built for Unscripted, Dubbing, and Archival Audio
The update targets the standard failure points across post-production workflows:
- Unscripted TV, Documentaries, and Podcasts: Fixes bleed from shared room mics and side-by-side lavalieres, allowing mixers to level, EQ, and clean up one speaker without affecting the others.
- Localization and Dubbing: Provides pristine, single-speaker dialogue stems required for clean foreign-language replacement tracks and voice-over beds.
- Subtitling and Captions: Prevents the drops, garbled phrasing, and timing errors that occur when transcription engines hit simultaneous speech.
- Archival and Field Audio: Salvages mono mixes, legacy radio call-ins, and degraded field recordings where no original session tracks or multitracks exist.
Key Technical Upgrades in 2.0
- Sharper Isolation in Noisy Environments (32% Less Bleed vs. 1.0): Cuts significantly more ambient contamination alongside vocal crosstalk, delivering isolated stems that don’t drag audible room noise or voice bleed behind separated words.
- Full-Band Coverage (8 kHz to 48 kHz): A single unified engine handles everything from low-bandwidth mobile calls to hi-res production masters without resampling artifacts.
- Language-Agnostic Processing: Operates as a purely acoustic model rather than a predictive language model, preserving accented delivery, multilingual conversations, and overlapping vocalizations without hallucinated edits.
“Some of our richest, most human moments come at the point of overlap — an interjection, a laugh, finishing someone else’s sentence,” said Jessica Powell, co-founder and CEO of AudioShake. “In the editing suite, those moments have always forced a compromise between performance and audio quality. Multi-Speaker 2.0 lets editors keep the performance: every voice on its own track, with clear visual flags and metadata showing the exact seconds that need human review.”
Multi-Speaker 2.0 is available today in AudioShake Studio and via the AudioShake API.