Separate overlapping speakers into clean voices
What is Multi-Speaker Separation?
From podcast interviews and reality TV, to phone calls, interviews, and field recordings—much of the world’s audio is filled with conversational audio and overlapping speakers.
What makes AudioShake's Multi-Speaker Separation different


Measurably better performance
The greater the overlap, the greater the problem
Speaker separation built for real-world audio
Recordings don't arrive clean. AudioShake's Multi-Speaker Separation holds up across the range of audio you actually work with — different resolutions, channel formats, and messy acoustic conditions.
8kHz - 48kHz
One model for everything from phone calls and meetings to podcast, radio, and film/TV
Far field recordings
Far-field captures with room reverb, uneven distance, and constant crosstalk.
Language agnostic
An acoustic model, not a language model — it works on how voices sound, not what they’re saying.
Dialogue in the wild
Isolate location and archival sound buried under traffic, music, and crowd.
You get the labels, not just the audio
Alongside the separated tracks you get a structured results file — an overlap-aware timeline of who spoke when. Meaning when two people speak at the same moment, both speakers are active in the timeline instead of one of them being dropped.
That's what lets the output go straight into captioning, subtitling, dubbing and data prep, where losing an interjection is the failure mode you can least afford.
What teams do with the output
Built for high-stakes audio workflows
How to use AudioShake's Multi-Speaker Separation
Processing audio at scale?
Separate your audio, then filter by confidence to keep only the cleanest outputs — a programmatic way for labs and enterprises to QA large volumes and build reliable datasets.
Frequently Asked Questions
Yes. AudioShake’s multi-speaker separation is built specifically for this. It detects and isolates each speaker in a recording — even when several people talk over one another — so you get a clean, individual track for every voice, with the overlapping moments kept rather than dropped.
Yes. Confidence scores let you route high-confidence segments straight through your pipeline and hold lower-confidence ones for review, so quality control that used to mean spot-checking a sample becomes a ranking you can apply across thousands of hours of audio. The scores are a relative signal rather than a fixed threshold — the right cutoff depends on your material.
Yes. Teams building conversational AI turn real multi-party recordings into structured data: speaker-isolated tracks with turn boundaries and confidence scores attached. That lets you filter a corpus by confidence instead of by ear, and characterize your training data rather than assume it’s clean — at the scale of millions of hours.
Yes. Roundtable interviews, panel recordings, and multi-guest podcasts often have crosstalk that’s difficult to edit around. Multi-speaker separation extracts each speaker into their own track, so editors can adjust levels, trim, or remove one voice without affecting the others.
Yes. Multi-Speaker 2.0 runs one model across the full range from 8 kHz to 48 kHz — the same model and the same API call or AudioShake Studio upload whether it’s a 1970s phone recording, a 16 kHz meeting capture, or a 48 kHz studio master, with no conversion step and no variant to choose. That brings call-center recordings, field audio, and archive tape into scope, and it’s built to hold up on audio that isn’t clean.
Yes. Diarization is built in and aligned to the separated tracks themselves — who spoke when, and how many people are in the recording, returned on every job. Because the timeline comes with the audio, there’s no second pass with a standalone diarizer and nothing to reconcile afterward.
It keeps it. When two people talk at once, diarization-only tools have to pick one speaker and discard the other — one of the most common failure points in transcription and captioning. AudioShake returns a clean track for each person including the overlapped moment, and labels overlap as overlap (both speakers marked active in the same moment), so interjections, mm-hms and crosstalk are preserved. Accuracy holds up to about 20% sustained overlap — a fifth of total speech time with two or more people talking at once — and begins to degrade past that.
Multi-Speaker 2.0 holds up in complex, real-world audio, where general-purpose tools introduce artifacts and leave voices bleeding into one another. It reduces bleed between voices by up to 32% versus the previous version, produces 4.1× fewer downstream transcription errors than the off-the-shelf open-source separator we benchmarked, and records lower diarization error than every system we tested, open source and proprietary — 30% lower on average than pyannote community-1, the most widely deployed open-source diarizer, and 38% lower on telephone audio.
Voice isolation separates speech as a whole from background music, noise, or effects. Multi-speaker separation goes a step further: once speech is isolated, it splits that speech into an individual stream for each person talking. Noise reduction only cleans up unwanted sound — it doesn’t identify or separate distinct speakers.
Yes. You can integrate multi-speaker separation into your own applications through the AudioShake API. It’s not currently available via the AudioShake SDK, which is reserved for real-time, on-device products.
Confidence scores measure how much to trust each moment of a result. Every job returns two — one for whether speech was assigned to the right speaker, one for whether overlapping voices were pulled apart cleanly — because audio can be separated well and labeled wrong, or labeled right and still carry bleed. Each is a value from 0 to 1, returned per 20 ms frame and as a whole-file score weighted by speech activity, so silence can’t inflate it. Instead of a single pass/fail, you see exactly where a recording was difficult, so you can process the confident material automatically and send a reviewer to the moments that need an ear.
Multi-Speaker is an acoustic model, not a language model — it works on how voices sound, not what they’re saying — and it’s trained across a wide range of languages, so accented, code-switched, and non-English audio holds up rather than merely being tolerated.
Diarization identifies who is speaking and when — but it can’t recover voices that overlap. AudioShake separates overlapping speakers into isolated tracks, improves speaker attribution, and provides confidence scores so you can automate quality control or focus review where it matters.
