More than a Demo: AudioShake Live Beats the Cloud

When AudioShake debuted Dialogue RT, its real-time dialogue isolation model, at the NAB show in April, they brought a Dell Pro Max with GB10 to run demos on the show floor. CEO and co-founder Jessica Powell and her team assumed that any major broadcaster running 40 live feeds would eventually implement their tech on giant rack hardware. They took it as a matter of fact that nobody was going to buy 40 small workstations and run them side by side.
Then, AudioShake’s engineers managed to configure 16 live broadcast feeds to run simultaneously on a single small form factor workstation. Even more impressive, the GB10 isolated the dialogue on each feed with a delay of a mere 11 milliseconds, fast enough for live broadcasts. To everyone’s surprise, the demo machine transformed from a prop to production hardware.
The art of audio deconstruction
AudioShake’s core business is source separation, technology that deconstructs a finished audio mix into its individual components. The company sells this infrastructure to enterprises: major labels use it to extract instruments and vocals, movie studios use it to isolate dialogue and effects, sports broadcasters to process live feeds, and big tech companies for AI data prep.
Live audio has long had a noise problem. The standard fix has been noise suppression, the technology that muffles a passing siren or background music on a video call. But noise suppression only attenuates the noise floor by turning down background sound without removing it.
AudioShake trains their models on millions of minutes of individual sounds so it learns the exact characteristics of components such as a bass, guitar, piano, crowd noise, or speech. At runtime, the model searches the audio image for its target, “coloring in” the pixels that match and blacking out the rest.
The output is two independent streams: a clean dialogue stem and a separate background stem. It eliminates the need to manually adjust thresholds during a live broadcast, reduces the number of mics needed, and feeds clean dialogue directly into captioning and translation systems.
This is not a generative technology. Some speech enhancement tools regenerate and approximate the voice, which can result in words no one said, a hard no-go in broadcast, news, and forensics, where ground truth matters. “We’re subtracting. We’re never adding any information,” Powell said.
11 milliseconds that the cloud can’t compute
Separating sound in real time is incredibly hard. Training teaches the model what to expect from individual components that make up a sound, but a model working on a live feed isn’t able to gather extensive context about the environment. The models aren’t the size of LLMs, but they are still large deep-learning models that AudioShake must shrink to hit strict latency targets.
AudioShake handles post-production in the cloud, but live broadcast operates with more stringent requirements. AudioShake puts the window for a live broadcast chain at roughly 9 to 15 milliseconds before lip-syncing problems become perceptible. Dialogue RT runs at 11 milliseconds, measured from the input signal to the isolated dialogue output.
“You do need beefier compute than what you’re traditionally getting in the cloud to run that and to hit those latency targets,” Powell said. “The live workflow can really only be done with the compute in a Dell workstation.”
“We found a way to get 16 feeds onto a single machine and realized, this workstation that we thought was a great demo machine, turns out you can do a lot more on a GB10, alone,” Powell said.
Powered by an NVIDIA GB10 Grace Blackwell Superchip and 128 GB of unified memory, the Dell Pro Max proved it could handle 16 concurrent live feeds, each staying within an 11-millisecond budget, on a machine small enough to sit next to a monitor and quiet enough to live in the control room.
Dialogue RT is now deployed in proofs of concept across a number of AudioShake’s customers. “We’re the first to have ultra-low latency, high-quality isolation that requires no additional context or speaker data,” Powell said. “Turn it on and it starts working immediately.”
What started as a show-floor demo is now doing the real work.

Next stop, IBC
AudioShake will rock IBC, the broadcast industry’s European trade show, with the European launch of Dialogue RT. At the show, AudioShake will demonstrate Dialogue RT running 16 live feeds simultaneously on a single Dell Pro Max with GB10. Find AudioShake at Dell’s booth 7.B47 at the RAI Convention Center in Amsterdam, September 11-14.