Why We Built The Refinery: Teaching Machines How People Really Talk

Think of all the conversations you have each day at work, at home, and on the go. We talk over each other, interrupt, finish each other's sentences–and do it all with music, traffic and room noise behind us. In fact, one analysis of natural conversation found nearly one interruption per minute.
Most of the tools built to understand speech are trained on something much tidier: clean, well-separated speech, often with one person talking at a time. But real conversations don’t work that way. It’s difficult to train a model to understand who said what during moments of overlap if the training data itself doesn’t reliably reflect who said what. Synthetic data can help fill the gap, but simulated overlap doesn’t necessarily capture the messy dynamics of real human conversation.
The data problem
To train on real conversational dynamics, a lab needs clean, speaker-tracked audio at scale. Historically, that has often meant collecting or manufacturing conversations specifically for AI training–an effective but painstaking process that can be hard to scale.
The problem isn't that real-world conversational data doesn't exist. It's that most of it isn't usable for training in its current form. A huge amount of real conversation already exists in company archives, licensed datasets, and data marketplaces. The catch is that these recordings primarily exist as finished mixes, with every voice and sound combined into one file. The conversations are there, but the individual speaker tracks are not accessible.
What The Refinery does
The Refinery takes raw, real-world audio and returns structured, training-ready data. It works directly from recorded audio, with no original stems or session files needed. The resulting stems are isolated from the real recording itself — never synthesized, hallucinated, or filled in.
- Speaker separation. Built on AudioShake's Multi-Speaker 2.0 technology, The Refinery can recover overlapping voices as separate, labeled audio tracks. Traditional diarization can mark that two people spoke at once; it cannot recover those voices as separate audio. Multi-Speaker 2.0 does both, preserving interruptions, backchannels and overlap for downstream training. In our LibriCSS evaluation, the separated tracks produced 4.1× fewer transcription errors than the tested open-source separation baseline. Read the full technical evaluation and methodology.
- Quality and confidence scoring. Separation alone isn't enough at corpus scale: you also need to know which outputs to trust. Every multi-speaker output includes confidence scores for speaker assignment and separation quality, so teams can filter and rank a dataset programmatically, keep the strongest material, discard weak stems, or route uncertain segments to human review instead of listening through thousands of hours themselves. See examples of confidence scores at work on Hugging Face.
- Dialogue, music and background separation. Leveraging the same models used by some of the world’s largest media companies, The Refinery can isolate speech from complex sound, and split dialogue, music and background sound from finished recordings. It can identify and remove copyrighted music from training data, or, for music AI applications, separate finished mixes into more than ten instrument and vocal stem types.
Who it's for
- Frontier and voice-AI labs can turn existing corpora into training data for ASR, diarization, speaker identification, TTS, and conversational AI.
- Content owners can unlock archives that couldn't easily be used for machine learning because the speakers were trapped inside finished recordings.
- Data providers and marketplaces can upgrade existing inventory into separated, quality-scored datasets.
Over the past year, we've deployed early versions of The Refinery privately with frontier AI labs, and processed 100M+ minutes of audio. Customers include multimodal training data company Luel and voice-AI company Rime, alongside some of the largest AI labs and data marketplaces in the world.
“AudioShake has helped us process thousands of hours of clean, speaker-separated data, enabling the world’s leading labs to build better models.” — William Namgyal, CEO & Co-founder, Luel
Built on years of separation research
The Refinery grows out of the same separation research behind our work in music, film, broadcast and AI. Customers like ESPN, Universal Music Group, Warner Bros. Post Production Creative Services, and NFL Films use AudioShake models to isolate speakers, dialogue, music and instruments for editing, distribution and compliance. The Refinery applies that technology to a new problem: making the world's existing audio usable as structured data for AI. The goal isn't to manufacture a cleaner version of conversation, but to take the complexity of the world around us and make it useful for training.
Learn more about The Refinery.