Isolating Dialogue in Live Broadcast at 11 ms with AudioShake on NVIDIA DGX Spark

How AudioShake closed the latency gap that kept AI dialogue isolation out of the live chain by leveraging Dell Pro Max with GB10 — and what changes once live audio becomes a controllable data stream rather than a mixed signal
A referee makes a call, but the crowd noise swallows the explanation. A caption track garbles at the moment the audience most needs it to be right. Each of those failures has the same cause: the speech arrived with the surrounding noise bleeding into it, and every process downstream — the on-air mix, the caption engine, the translation feed, the archive — inherited that mixture.
Broadcast teams have compensated for this since live broadcasting began. They place microphones to favor the voice, ride gain against the crowd, insert denoisers and tune their thresholds, and when a feed has to drive transcription, they build a second path for it: a bus mixed for the ASR engine rather than for air. Those compensations only hold as long as the venue behaves as expected. The moment something happens that nobody configured for, the team has to repeat the fix on every single feed — and on productions with a dozen feeds or more, that adds up fast.
Until now, isolating speech as its own signal has only existed with large buffers, or in post-production, where separation models can take as long as they need to process a large audio file. Live is significantly harder: the model has to work on the fly, and the isolated result has to land within the roughly 10–15 ms threshold audiences tolerate — an order of magnitude tighter than post-production models run in.
Two stems from one feed
AudioShake’s Dialogue RT crosses that threshold. It isolates dialogue from a live feed with 11 ms of end-to-end latency, measured from model input to isolated output, running on NVIDIA DGX Spark and NVIDIA Blackwell architecture GPUs. Dialogue RT makes its European debut at IBC 2026, where AudioShake is running it live on NVIDIA DGX Spark at Stand 14.F46 and on the Dell Pro Max with GB10 at Stand 7.B47.
Dialogue RT takes a live mixed feed — commentary, crowd, PA bleed, ambient sound — and returns two signals: a dialogue stem carrying the isolated speech, and a background stem carrying everything else.
Isolation is a marked departure from traditional approaches like noise suppression. Suppression attenuates unwanted sound inside a single mixed output, so nothing downstream can address speech and background separately — a caption engine, a dubbing workflow, and a loudness process all receive the same compromise.
Isolation hands the caption engine, the dubbing workflow, and the loudness process each the element it needs. It lets a broadcast engineer decide how much background belongs under the dialogue instead of accepting a single ratio set by a noise threshold.
Concurrency across your production workflow
A single live production generates many feeds. A stadium event carries commentary positions, pitch-side reporters, and a press-conference room; a rights holder distributing internationally adds a commentary track per language. Isolation delivered one machine per feed ties hardware count to feed count: isolating 16 feeds means 16 processing paths to rack, power, and monitor.
With NVIDIA, Dialogue RT can run up to 16 concurrent isolations on a single run. 16 feeds enter, and 16 dialogue stems and 16 background stems come out, while the device count stays at one.
- A single source of truth for audio. Instead of managing parallel feeds for production and transcription, teams can isolate dialogue directly from the main feed and route it downstream — eliminating duplicate workflows.
- Fewer microphones, less complexity. In many live productions, engineers deploy dozens of microphones to compensate for bleed between crowd, PA, and commentary. Dialogue RT reduces that dependency — enabling cleaner results with fewer inputs and simpler signal chains.
- Better captioning and ASR accuracy. Using Dialogue RT to feed isolated dialogue rather than a noisy mix directly improves transcription performance by removing background interference at the source. This also enables easier downstream dubbing, localization, and international distribution — all from a single live source.
- A more hands-off mix, with control where it matters. Broadcast engineers can set dialogue isolation on the primary feed and avoid constantly managing crowd noise, stadium PA, or unpredictable field conditions. Unlike denoising tools, which require ongoing threshold tuning, Dialogue RT adapts in real time to changing environments — while still giving engineers the ability to dial in or override as needed.
Coming to European broadcasters at IBC 2026
Dialogue RT is available to European broadcasters for the first time at IBC 2026, where AudioShake will demonstrate the next generation of the model. Dialogue RT is available now via the AudioShake SDK, for direct integration into existing broadcast infrastructure. Get started at dashboard.audioshake.ai or read the documentation at developer.audioshake.ai.
AudioShake is at Stand 14.F46 and Stand 7.B47 at IBC 2026 in Amsterdam, 11–14 September, running live demos of Dialogue RT.