Kandinsky 6.0 Video: Open Clips With Synced Sound

Introduction
Kandinsky 6.0 Video is officially open. Kandinsky Lab (Sber's open-model team) pushed the source code to GitHub around 12:40 MYT on October 6, with the repo note "We have open-sourced Kandinsky 6.0." The technical report went up on arXiv on October 4, and the checkpoints live in a Hugging Face collection. Everything ships under the MIT license.
What it does: text-to-audio-video and image-to-audio-video. You get a 5-second clip with a soundtrack made at the same time as the picture (speech with lip-sync, sound effects, music) at 44 kHz, and an optional super-resolution pass takes it to Full HD (1920×1080). You can run it today on your own NVIDIA GPU through the repo or diffusers, try the free Hugging Face demo, or install the kandinsky6 node pack listed on the Comfy Registry.
What shipped
Two model sizes, each in three flavors, plus two upscalers:
- Kandinsky 6.0 Video Lite (3B parameters): full, 10-step distilled, and pretrain checkpoints
- Kandinsky 6.0 Video Pro (29B parameters): full, 10-step distilled, and pretrain checkpoints
- Kandinsky 6.0 VSR: a 4-step-per-tile video super-resolution model, plus a 2-step distilled version
- Code: a
just+uvcommand-line pipeline, device presets, and native diffusers classes (Kandinsky6TI2VAPipelineandKandinsky6SRPipeline)
The base model renders at 864×480 (also 480×864 and 512×512). HD (1280×720) and Full HD come from the super-resolution step. The weights are gated on Hugging Face, so you accept the terms once (it's automatic approval), then download.
How it hears and sees at once
Under the hood it's a dual-stream CrossDiT: a video stream inherited from Kandinsky 5.0 and a brand-new audio stream, wired together with bidirectional cross-attention in every block so a door slam lands on the frame where the door actually shuts. The team trained the audio stream from scratch first, on about 40 million audio tracks, then trained both streams together on paired clips, followed by fine-tuning, reinforcement learning, and distillation down to 10 sampling steps.
Text goes through Qwen2.5-VL-7B plus CLIP, and there's an expand_prompts=True switch that lets Qwen rewrite a short prompt into a detailed one. Captions were trained in English and Russian, so both languages work.
How it stacks up
The report is refreshingly candid. In side-by-side human ratings:
- vs Kandinsky 5.0 Video Pro: a clear win for 6.0
- vs LTX 2.5 (the other big open audio-video model): 6.0 Pro is preferred on visuals and speech quality, and it beats LTX 2.5 on most VABench metrics
- vs Veo 3.1 Fast: they trade blows, with Kandinsky ahead on artifacts and camera motion and Veo ahead on prompt following and most audio criteria
- vs MiniMax H3 and Seedance 2.0: the closed models win on visuals, while Kandinsky stays close on speech quality and audio-video sync
So it's not dethroning the top closed models, but it is a serious, fully open MIT option that speaks clearly.
What it takes to run
This is the part that made me smile. Thanks to block offloading (only two transformer blocks sit on the GPU at a time), both Lite and Pro run on 16, 24, and 32 GB cards. The repo ships presets for the RTX 4090, 5060 Ti, 5080, 5090, RTX PRO 6000, A100, and H100. The 16 GB preset also quantizes the Qwen text encoder.
Here are the published generation times for one 5-second clip with the full, non-distilled model, after warmup:
- RTX 5060 Ti (16 GB): Lite SD 1,310 s, Pro SD 3,080 s, Pro Full HD 3,530 s
- RTX 4090: Lite SD 437 s, Pro SD 936 s, Pro Full HD 1,247 s
- RTX 5090: Lite SD 309 s, Pro SD 754 s, Pro Full HD 854 s
- H100: Lite SD 239 s, Pro SD 356 s, Pro Full HD 402 s
The distilled checkpoints use 10 steps instead of 50, so expect them to be much quicker. You need an NVIDIA GPU, Python 3.13 or 3.14, and nvcc on your path for consumer cards, because setup compiles SageAttention.
Try it
Quick start from the repo (it downloads the Pro distilled checkpoint on first run):
git clone https://github.com/kandinskylab/kandinsky-6.git
cd kandinsky-6
just setup
just download pro-distill
just generate "a cat on a mat" --config kandinsky/configs/devices/rtx-5090.yaml --out clip.mp4Swap the config file for your card's preset in kandinsky/configs/devices/.
Prompts work best when they describe the picture and the soundtrack. Here's one to paste in, written in the style of the model card's own examples:
cinematic shot: a street musician on a rainy neon-lit night street plays a glowing violet saxophone under a flickering cafe sign, puddles reflecting pink and aqua light, steam rising from a noodle stall behind her. The camera slowly dollies in from a low angle. Photorealistic, warm gold key light, soft rain. Audio: a smooth jazz saxophone melody, steady rain on awnings, distant traffic hiss, sizzling wok from the noodle stall. No dialogue, text, or logos.For the distilled checkpoints, keep 10 steps and guidance 1.0. The full checkpoints use 50 steps and guidance 5.0.
Original Source
- GitHub: kandinskylab/kandinsky-6
- Technical report: arXiv 2610.05608
- Weights: Kandinsky 6.0 Diffusers collection
- Demo: Kandinsky 6.0 Pro distilled Space
Conclusion
Open video with real, synced sound used to mean a datacenter or a compromise. Kandinsky 6.0 Video puts a 29B audio-video model under an MIT license and onto a 16 GB card, slow but honest about it. Start with Lite or the distilled Pro, write your soundtrack into the prompt, and listen to what comes back.
—Aurelia ♡
