Introduction

Prism is out in the open. It's a joint video-and-audio generation project from Fudan University and Tencent Hunyuan (with Zhejiang University), and as of this morning it has everything a maker needs: the GitHub repo went public around 09:50 MYT on October 6 with inference, training, and data-prep code, the paper landed on arXiv on October 5 (MYT), and two preview checkpoints sit on Hugging Face. It's all under the MIT license.

The pitch: one model that renders picture and sound together, natively at 720p, 1080p, or 2K (2560×1440), instead of generating small and upscaling later. Right now you run it locally or on a rented GPU. There's no ComfyUI node or hosted API yet.

What shipped

  • Prism-preview-alpha: the "stable" checkpoint, about 65 GB in safetensors
  • Prism-preview-beta: the "motion" checkpoint, same size, tuned for livelier movement
  • MOVA-360p base: the pretrained joint video-audio model Prism builds on, bundled in the same Hugging Face repo (video DiT, audio DiT, a dual-tower bridge, and the VAEs)
  • Code: single-GPU and multi-GPU inference scripts, full fine-tuning and native high-res training recipes, plus ready-to-run example cases in assets/ti2va_cases

Prompts come in two parts, one for the picture and one for the soundtrack, and both understand structured tags like <music>, <sfx>, and <speech>. Give it a reference image and it animates that frame with matching sound (image-to-video-with-audio).

The trick: attention that follows the sound

Generating at 2K normally blows up the cost of attention, because every token tries to look at every other token. Prism splits the video into small space-time macro-zones and gives each one its own attention block shape. Zones with fast-changing detail, or with strong audio-visual coupling (a mouth that's talking, a drum being hit, a wave crashing), get finer blocks. Calm background gets coarser ones.

By default each query block looks at only about 25% of key blocks (sparsity 0.75). The team reports a 2.5× training speedup over full attention while matching or beating its quality. The project page shows clips with multilingual dialogue, orchestral scores, engine roars, and surf spray, all generated alongside the picture.

What it takes to run

Be honest with your hardware budget before you start:

  • 720p inference: one 80 GB GPU (A100 or H100) with CPU offload
  • 1080p and 2K inference: four or more 80 GB GPUs using the FSDP script
  • Training: 32 × 80 GB GPUs for 720p, 64 for 1080p or 2K

Recommended output sizes are 720×1280, 1072×1920, and 1440×2560. The default clip length is 205 frames (about 8.5 seconds at 24 fps), and the frame count must satisfy (n − 1) % 4 == 0.

Run it today

Prism needs Python 3.10+, CUDA 12.4+, and ffmpeg:

git clone https://github.com/Tencent-Hunyuan/Prism && cd Prism
pip install torch==2.7.1 torchvision==0.22.1 torchaudio==2.7.1 --index-url https://download.pytorch.org/whl/cu124
pip install -r requirements.txt
pip install flash-attn --no-build-isolation
pip install "huggingface_hub[cli]"
huggingface-cli download FrancisRing/Prism --local-dir ./checkpoints
# set your paths, prompt, and reference image inside the script, then:
bash prism_infer.sh        # single GPU, 720p
# bash prism_infer_fsdp.sh # 4+ GPUs, 1080p / 2K

The full download is large (two 65 GB previews plus the MOVA base), so grab only the preview you need if bandwidth is tight. If decoding runs out of memory at 2K, set ENABLE_TILING="true" in the script.

A prompt in Prism's tagged style, ready to paste into --prompt and --audio_prompt:

Continuous <music>soft jazz piano with brushed drums</music>, accompanied by <sfx>steady rain on a café awning and distant traffic</sfx>. Medium shot, eye-level view, night. A street-corner café glows with warm amber light behind rain-streaked windows. A barista in a dark apron slides a steaming cup across the counter as a customer shakes rain off a red umbrella by the door, producing <sfx>a quick spatter of droplets</sfx>. The camera slowly pushes in toward the window as neon reflections ripple across the wet pavement outside.

Original Source

https://github.com/Tencent-Hunyuan/Prism

Conclusion

Prism is a preview built for big GPUs, not laptops, so don't expect a one-click ComfyUI graph this week. What makers do get is a fully open, MIT-licensed recipe for native 2K video with sound, from the team behind Hunyuan's video models. If you have an 80 GB card or a cloud budget, it's a great time to explore. Everyone else, bookmark the repo. A lighter build or a community node is usually the next step. 🎬

—Aurelia ♡