Multi-Speaker 2.0: Keep Every Voice in the Conversation

AudioShake
September 1, 2026

One recording of several people goes in. A clean, labeled, per-person track comes out, even when people talk over each other.

Real conversations are full of interjections. People start before you finish, they laugh over the punchline, or they say “mm-hm” in the middle of your point to show support.

Today, almost every system that handles conversational audio deals with this in the same, incomplete way–using an approach called diarization to track one speaker versus another over the length of an audio file.

Tracking “who spoke when” is sufficient for many workflows, but breaks down in others:

  • Post-production: In unscripted TV, documentaries, and interviews, you can’t cut or edit one person’s line when it’s entangled with another voice.
  • Analysis: background audio is lost, hindering richer metadata extraction like audio events
  • Transcription & Captioning: When two people overlap, transcription has to assign the moment to one speaker, and whatever the other person said is lost. It’s one of the most common failure points in transcription pipelines.
  • Voice AI Training Data: Diarization labels a mixed file, so overlapped regions are contaminated no matter how good the labels are. Labs discard them, and end up with hours of conversations where nobody ever interrupts.

AudioShake Multi-Speaker keeps it all. One recording goes in, and clean, labeled per-speaker tracks come out — including the moments when two people are talking at once. We also retain background audio in its own track, allowing you to discard, edit, or extract metadata from it.

Loading cliptwo speakers · 53 s excerpt
Original mix
0:00 / 0:53
Switch stems while it plays

Check out more samples on Hugging Face

One mixed recording in, two labelled tracks out
Assignment confidence

Did this moment go to the right speaker? It falls where two voices land on top of each other, and anywhere a stretch of speech could plausibly belong to either person — the places a diarizer is most likely to hand a word to the wrong one. It also dips for a fraction of a second at every handover, which says nothing about the speech on either side, so the row below marks only sustained uncertainty inside a turn.

Separation confidence

How cleanly was this voice pulled out of the mix? It falls where the isolated stem is most likely to carry traces of someone else, so it tells you which passages will still sound crowded on their own.

Output
Confidence

What’s new in Multi-Speaker 2.0

Multi-Speaker 2.0 combines high-accuracy diarization with AudioShake’s industry-leading sound separation to achieve new milestones:

Capabilities:

  • One model, on whatever audio you have. A 1970s phone recording at 8 kHz, a meeting-room capture at 16 kHz, a studio master at 48 kHz — same model, same API call or AudioShake Studio upload, with no conversion step or having to choose a variant. Version 1 was built for professional audio: film, TV, radio. Version 2 covers the full range–from 8kHz to 48kHz, which means call center recordings, field audio, and archive tape are now in scope. This makes AudioShake Multi-Speaker 2 the world’s first Multi-Speaker separator designed for both low- and high-resolution.
  • No Limits on Language. Multi-Speaker is an acoustic model, not a language model — it works on how voices sound, not what they’re saying. It’s also trained across a wide range of languages, so accented, code-switched, and non-English audio holds up rather than merely being tolerated.
  • You get the labels, not just the audio. Alongside the separated tracks, every job returns a structured results file: who spoke when and how many people are in the recording. The timeline is aligned to the tracks themselves, so there’s no second pass with a standalone diarizer and nothing to reconcile afterward. In addition, multiple speakers can be active at the same moment, instead of one of them getting dropped. That’s what lets the output go straight into captioning, subtitling, dubbing, and edit prep, where you never want to lose an interjection or moment of overlap.
  • It tells you when not to trust it. Every job returns confidence scores in 20ms intervals, as well as for the whole file — so you can send a reviewer to the four seconds that need an ear instead of the whole hour. That means you can process the material with high confidence scores automatically, and route the rest.

Quality improvements:

  • Cleaner separation: Up to 32% less bleed and distortion between voices than Version 1.0.
  • Built for audio that isn’t clean. On noisy Multi-Speaker recordings, Multi-Speaker 2.0 removes 75% more of the background noise and interference than Version 1.0, and also improves the perceptual quality.
  • Nearly 4× fewer downstream transcription errors than the strongest open-source speech separator.
  • Overlap without a penalty. Transcribe our separated tracks from a conversation with up to 20% overlap of interfering speakers and you get the same accuracy as if each person had been recorded on their own microphone.
  • 30% fewer labeling errors than the leading open-source diarization system — and lower error on all 15 public test sets we ran, across meetings, phone calls, broadcast, and web video. On telephone audio, the hardest category for most systems, the gap is 38%.

Who it’s for

Labs building conversational AI. Real multi-party recordings turned into structured data. Each output comes as speaker-isolated tracks (or “stems”), with turn boundaries and quality scores attached. This lets you filter a corpus by confidence instead of by ear, which makes training on millions of hours feasible.

Teams running voice and media pipelines. Clean per-speaker audio feeding transcription, translation, dubbing, captioning, or analysis — on the same kind of phone-quality and far-field recordings those pipelines actually receive.

Post-production. Unscripted TV, documentaries, podcast and radio interviews all live with crosstalk and bleed. Multi-Speaker 2.0 gives editors a track per person, and the control that comes with it.