Separate overlapping speakers into clean voices

Multi-Speaker 2.0 is the world’s first low- and high-resolution multi-speaker voice technology. It pairs high-accuracy diarization with benchmarked sound separation, turning conversational audio and overlapping voices into clean, isolated tracks.
TRY MULTI-SPEAKER NOW FOR FREE
Loading cliptwo speakers · 53 s excerpt
Original mix
0:00 / 0:53
Switch stems while it plays

Check out more samples on Hugging Face

One mixed recording in, two labelled tracks out
Assignment confidence

Did this moment go to the right speaker? It falls where two voices land on top of each other, and anywhere a stretch of speech could plausibly belong to either person — the places a diarizer is most likely to hand a word to the wrong one. It also dips for a fraction of a second at every handover, which says nothing about the speech on either side, so the row below marks only sustained uncertainty inside a turn.

Separation confidence

How cleanly was this voice pulled out of the mix? It falls where the isolated stem is most likely to carry traces of someone else, so it tells you which passages will still sound crowded on their own.

Output
Confidence

What is Multi-Speaker Separation?

Multi-Speaker Separation takes a fully mixed track and pulls each person onto their own isolated, clean track.

From podcast interviews and reality TV, to phone calls, interviews, and field recordings—much of the world’s audio is filled with conversational audio and overlapping speakers.
Separates overlapping speech
Diarizes different speakers
Returns confidence scores

What makes AudioShake's Multi-Speaker Separation different

01
Isolates individual speaker tracks, even in moments of overlap
Multi-Speaker removes 75% more noise and interference than our previous model, with 32% better separation between voices and 4.1× fewer transcription errors than the best open-source separator.
02
Probably better diarization than the one you're using
30% lower diarization error than pyannote community-1, and lower error than the paid service we replaced — in every domain we test: meetings, telephone, broadcast and web video.
03
Confidence scores allow for automated quality control
The separated audio comes back scored per frame and per file, so you can process high-confidence material automatically. More on confidence scores.
01

Measurably better performance

Our latest Multi-Speaker model combines high-accuracy diarization with AudioShake's sound separation. Here's what that changes, measured against Version 1.0 and the strongest open-source systems available.
75%
more noise and interference removed than Multi-Speaker 1.0
32%
better separation than Multi-Speaker 1.0
4.1×
fewer transcription errors than the best open-source separator
30%
lower diarization error than pyannote community-1

The greater the overlap, the greater the problem

Send a conversation straight to a transcription tool and the errors climb as people interrupt one another. Run Multi-Speaker 2.0 first and that climb flattens — at the heaviest overlap, half the errors.
Transcription errors on LibriCSS, a standard set of recorded group conversations. Same recordings, same transcription tool, in every case.
SEE THE FULL BENCHMARKS
02

Speaker separation built for real-world audio

Recordings don't arrive clean. AudioShake's Multi-Speaker Separation holds up across the range of audio you actually work with — different resolutions, channel formats, and messy acoustic conditions.

8kHz - 48kHz

One model for everything from phone calls and meetings to podcast, radio, and film/TV

Far field recordings

Far-field captures with room reverb, uneven distance, and constant crosstalk.

Language agnostic

An acoustic model, not a language model — it works on how voices sound, not what they’re saying.

Dialogue in the wild

Isolate location and archival sound buried under traffic, music, and crowd.

03

You get the labels, not just the audio

Alongside the separated tracks you get a structured results file — an overlap-aware timeline of who spoke when. Meaning when two people speak at the same moment, both speakers are active in the timeline instead of one of them being dropped.

That's what lets the output go straight into captioning, subtitling, dubbing and data prep, where losing an interjection is the failure mode you can least afford.

Example: One mixed recording in, two labeled tracks out
Speaker 1
Speaker 2
Both talking at once
Input
OUTPUT
Speaker 1
Speaker 2
both active
both active
Confidence
0:00
0:15
0:30
0:45
1:00

What teams do with the output

01
Send a reviewer to the areas that need an ear
The timeline and the per-frame scores point at exact timestamps, so a QC pass is minutes on the moments that matter instead of on the whole file.
02
Process the confident material automatically, route the rest
File-level scores rank a batch, so the clean majority moves through the pipeline untouched and only the uncertain files queue for attention.
03
Filter a training corpus by score instead of by ear
The same scores work as a dataset filter, which is what makes it possible to keep real conversation — interruptions, laughter, backchannels — instead of discarding it as unusable.
04

Built for high-stakes audio workflows

01
Film, TV & post-production
Isolate overlapping dialogue for the edit
Pull apart crosstalk and overlapping dialogue into separate speaker tracks for cleaner editing, ADR, and re-mixing. Recover a single line buried under a crowd, or lift one actor's voice out of a busy scene without touching the rest of the mix.
02
Podcasting
A clean, separate track for every guest
Turn multi-mic or single-track recordings into per-speaker tracks for editing, leveling, and cleanup. Fix one guest's audio, remove crosstalk between hosts, or balance a remote guest against the room — without re-recording.
03
Dubbing & localization
A clean track for every speaker to dub
Get an isolated track per speaker across the whole recording, so localization teams can re-voice and translate speaker by speaker instead of wrestling with a mixed track. Clean separation makes each voice easier to replace, sync, and balance for global distribution.
04
Transcription & ASR
Cleaner input, better machine intelligibility
Feed cleaner per-speaker audio into speech-to-text so a single-speaker ASR system can attribute words to the right person. At up to 20% crosstalk, our separated tracks transcribe as accurately as individually mic'd speakers. Use the built-in confidence scores to route high-confidence segments straight through and flag the rest for review.
05
VOICE AI & TRAINING DATA
Real conversations, turned into structured data
Turn real recordings into speaker-isolated tracks with turn boundaries and confidence scores attached. Filter a corpus by confidence instead of by ear, and keep the overlapped regions that diarization alone forces you to discard.
06
ACCESSIBILITY
Isolate and preserve an individual voice
Extract a single, clean voice from a noisy multi-speaker recording — useful for captioning, assistive workflows, and voice-banking where fidelity to the original speaker matters.
05

How to use AudioShake's Multi-Speaker Separation

01
Upload your recording
Upload or stream your recording as it is. Anything from 8 to 48 kHz — no pre-processing or format conversion required.
02
Run Multi-Speaker
Select Multi-Speaker Separation to access individual recordings per speaker in any piece of audio.
03
Download separated audio
Export the tracks and metadata, or send them into your editing, transcription, or data pipeline.
WEB PLATFORM
Upload and process recordings directly
AudioShake Live is an intuitive, drag-and-drop web platform designed for companies, film studios, and media production teams to create high-quality tracks on demand.
GET STARTED NOW
API
Integrate Multi-Speaker Separation into your workflow
Connect Multi-Speaker Separation to your own product or pipeline, with diarization and confidence scores included.
TRY OUR API

Processing audio at scale?

Separate your audio, then filter by confidence to keep only the cleanest outputs — a programmatic way for labs and enterprises to QA large volumes and build reliable datasets.

Explore Data Services
06

Frequently Asked Questions

Can AI separate multiple people talking at the same time?

Yes. AudioShake’s multi-speaker separation is built specifically for this. It detects and isolates each speaker in a recording — even when several people talk over one another — so you get a clean, individual track for every voice, with the overlapping moments kept rather than dropped.

Can I filter results by confidence score for large-scale projects?

Yes. Confidence scores let you route high-confidence segments straight through your pipeline and hold lower-confidence ones for review, so quality control that used to mean spot-checking a sample becomes a ranking you can apply across thousands of hours of audio. The scores are a relative signal rather than a fixed threshold — the right cutoff depends on your material.

Can this be used to build training data for speech AI models?

Yes. Teams building conversational AI turn real multi-party recordings into structured data: speaker-isolated tracks with turn boundaries and confidence scores attached. That lets you filter a corpus by confidence instead of by ear, and characterize your training data rather than assume it’s clean — at the scale of millions of hours.

Can this be used to clean up overlapping speakers in podcast interviews?

Yes. Roundtable interviews, panel recordings, and multi-guest podcasts often have crosstalk that’s difficult to edit around. Multi-speaker separation extracts each speaker into their own track, so editors can adjust levels, trim, or remove one voice without affecting the others.

Does it work on low-quality, phone, or call-center audio?

Yes. Multi-Speaker 2.0 runs one model across the full range from 8 kHz to 48 kHz — the same model and the same API call or AudioShake Studio upload whether it’s a 1970s phone recording, a 16 kHz meeting capture, or a 48 kHz studio master, with no conversion step and no variant to choose. That brings call-center recordings, field audio, and archive tape into scope, and it’s built to hold up on audio that isn’t clean.

Does multi-speaker separation include diarization (who spoke when)?

Yes. Diarization is built in and aligned to the separated tracks themselves — who spoke when, and how many people are in the recording, returned on every job. Because the timeline comes with the audio, there’s no second pass with a standalone diarizer and nothing to reconcile afterward.

Does multi-speaker separation keep overlapping speech, or does it drop a speaker?

It keeps it. When two people talk at once, diarization-only tools have to pick one speaker and discard the other — one of the most common failure points in transcription and captioning. AudioShake returns a clean track for each person including the overlapped moment, and labels overlap as overlap (both speakers marked active in the same moment), so interjections, mm-hms and crosstalk are preserved. Accuracy holds up to about 20% sustained overlap — a fifth of total speech time with two or more people talking at once — and begins to degrade past that.

How accurate is AudioShake compared to other speaker-separation tools?

Multi-Speaker 2.0 holds up in complex, real-world audio, where general-purpose tools introduce artifacts and leave voices bleeding into one another. It reduces bleed between voices by up to 32% versus the previous version, produces 4.1× fewer downstream transcription errors than the off-the-shelf open-source separator we benchmarked, and records lower diarization error than every system we tested, open source and proprietary — 30% lower on average than pyannote community-1, the most widely deployed open-source diarizer, and 38% lower on telephone audio.

How is multi-speaker separation different from voice isolation or noise reduction?

Voice isolation separates speech as a whole from background music, noise, or effects. Multi-speaker separation goes a step further: once speech is isolated, it splits that speech into an individual stream for each person talking. Noise reduction only cleans up unwanted sound — it doesn’t identify or separate distinct speakers.

Is multi-speaker separation available via API?

Yes. You can integrate multi-speaker separation into your own applications through the AudioShake API. It’s not currently available via the AudioShake SDK, which is reserved for real-time, on-device products.

What are confidence scores?

Confidence scores measure how much to trust each moment of a result. Every job returns two — one for whether speech was assigned to the right speaker, one for whether overlapping voices were pulled apart cleanly — because audio can be separated well and labeled wrong, or labeled right and still carry bleed. Each is a value from 0 to 1, returned per 20 ms frame and as a whole-file score weighted by speech activity, so silence can’t inflate it. Instead of a single pass/fail, you see exactly where a recording was difficult, so you can process the confident material automatically and send a reviewer to the moments that need an ear.

What languages does multi-speaker separation support?

Multi-Speaker is an acoustic model, not a language model — it works on how voices sound, not what they’re saying — and it’s trained across a wide range of languages, so accented, code-switched, and non-English audio holds up rather than merely being tolerated.

Why can’t I use just a diarizer?

Diarization identifies who is speaking and when — but it can’t recover voices that overlap. AudioShake separates overlapping speakers into isolated tracks, improves speaker attribution, and provides confidence scores so you can automate quality control or focus review where it matters.

Get in touch.
Thank you! Your submission has been received!
Oops! Something went wrong while submitting the form.