WorldSonus: Open-Weight Real-Time Stereo Sound for AI World Models

Introduction
WorldSonus, from Noiz AI and the Hong Kong University of Science and Technology (HKUST) with MetaX and Shanghai Jiao Tong University, went public on October 7, 2026. It's an open-weight video-to-audio model that watches a video stream and generates 48 kHz stereo sound in 100 ms chunks while the video plays. It's built for AI world models and gameplay footage, which usually come out silent. The inference code is on GitHub and the weights are on Hugging Face, so you can run it today on your own CUDA GPU.
The license is CC BY-NC 4.0, so it's free for research and personal projects but not for commercial work. Training code and datasets aren't included.
Pro Tip
Write your prompt the way the official examples do: describe the sounds you want, then say what to leave out. The README's own example ends with "No speech." Video-to-audio models love to add murmuring voices or background music to anything with people in it, and one short negative line keeps the track clean.
What Shipped
From the paper (arXiv, October 6), the GitHub repo (public release October 7), and the model card:
- Causal streaming: it only looks at past and current frames, never ahead. Each 100 ms audio chunk is generated as soon as its three video frames arrive, and a bounded 5-second memory keeps long sessions stable.
- Speed: 41.2 ms to generate each 100 ms chunk on one NVIDIA H100, a real-time factor of 0.41. Over a 10-second session including cold start, the RTF is 0.466.
- Live prompt edits: the paper shows text instructions changing mid-stream (for example, adding an off-screen sound) without resetting the session.
- Camera-aligned stereo: it was trained on curated clean-stereo clips plus panoramic ambisonic recordings decoded into view-matched stereo, so a sound source on the left of the frame ends up in the left channel.
- Training data: 1,465 hours of audio (999 hours of paired stereo video-audio and 466 hours audio-only).
- Files: a 4.1 GB generator checkpoint (
worldsonus_150k.pt) and a 301 MB causal stereo codec. The model uses frozen DINOv3 and T5Gemma 2 encoders, which you download separately under their own licenses.
The Numbers
The authors compare WorldSonus against offline video-to-audio models that see the whole clip in advance (AudioX, ThinkSound, PrismAudio) and against V-AURA, a streaming model that outputs mono. Lower FAD means more realistic audio. Lower DeSync means better timing between sound and picture.
ModelModeVGGSound 5 s FADVGGSound 5 s DeSyncInteractive 30 s FADInteractive 30 s DeSync
WorldSonus
Causal, 100 ms chunks
1.73
0.686
2.03
0.867
PrismAudio
Offline
2.09
0.539
6.08
0.765
ThinkSound
Offline
2.94
0.481
7.00
0.901
AudioX
Offline
3.00
0.986
5.52
1.125
V-AURA
Streaming, mono
4.06
1.287
13.33
1.266
WorldSonus has the best audio realism in both tests, and the gap gets wider on 30-second gameplay and real-world clips, where the offline models fall apart. It doesn't win on timing. ThinkSound and PrismAudio stay tighter to the picture on short clips, which makes sense because they can see future frames and WorldSonus can't. These are the authors' own benchmarks. The "Interactive" set is a test set they built themselves.
Run It
You need Python 3.10+, a CUDA GPU, and FFmpeg. You also have to accept the license terms for the DINOv3 and T5Gemma 2 encoders on Hugging Face and log in before the encoder download will work.
git clone https://github.com/NoizAI/WorldSonus.git
cd WorldSonus
bash scripts/setup.sh .venv
source .venv/bin/activate
pip install -e '.[features]'
python scripts/download_model.py --output assets # generator + codec + stats
python scripts/download_encoders.py --output assets # DINOv3 + T5Gemma 2 (gated)
# Offline: video + sound description in, stereo 48 kHz WAV out
python -m worldsonus.infer \
--video walk.mp4 \
--prompt "Boots crunch on wet gravel, a shallow river splashes nearby, wind moves through pine trees. No speech. No music." \
--output walk.wavTo get a 10-second window, add --start 5 --seconds 10. Leave out --prompt for video-only conditioning. For live playback, the README streams raw PCM straight into FFplay:
python -m worldsonus.infer --video walk.mp4 \
--prompt "A train passes beside a river." \
--stream --pcm-stdout \
| ffplay -f f32le -ar 48000 -ac 2 -i - -nodisp -autoexitTo put the generated track back on your clip:
ffmpeg -i walk.mp4 -i walk.wav -map 0:v -map 1:a -c:v copy -c:a aac -b:a 256k -shortest walk-with-sound.mp4Here are two prompts to try on world-model or gameplay captures. These are mine, not from the authors:
Heavy rain on a tin roof, water dripping from gutters, a distant thunder roll. No speech. No music.Footsteps echo across a stone courtyard, a fountain trickles on the left, pigeons flutter overhead. No speech.The first run compiles CUDA kernels, so expect a slow start. The authors say this is a research inference release and doesn't promise any particular live latency.
Why It Matters
Open world models like Matrix-Game and Genie-style interactive video can now render explorable scenes in real time, but they're still silent. Most open video-to-audio models need the whole clip before they generate anything, which doesn't work for a stream. WorldSonus is the first open-weight release we've seen that combines real-time causal generation, stereo that follows the camera, and prompts you can change mid-stream. That makes it a practical way to add sound to AI-generated gameplay, walkthroughs, and silent video drafts, as long as you stay within the non-commercial license.
Conclusion
WorldSonus is an open-weight (CC BY-NC 4.0) video-to-audio model from Noiz AI and HKUST. It streams 48 kHz stereo sound in 100 ms chunks at 0.41 RTF on one H100, and you can change the text prompt while it runs. It beats offline baselines on audio realism (FAD 1.73 on VGGSound and 2.03 on 30-second interactive clips), but trails them a little on sync for short clips. You can run it locally today from GitHub and Hugging Face.
Primary: WorldSonus paper on arXiv (Oct 6, 2026), WorldSonus code on GitHub (public Oct 7, 2026), WorldSonus weights on Hugging Face, project page and demos.
—Anabel ♡
