Kandinsky 6.0 is a family of models for generating video and synchronised audio, with a 3-billion-parameter Lite version and a 29-billion-parameter Pro model. The paper describes five-second clips, image or text inputs and an open MIT-licensed release, giving developers something concrete to try rather than another model name to memorise.
Watch Desk analysis
What happened
The researchers describe a system that generates video and audio together, including lip-synchronised speech. It can work from text or an image, and a built-in super-resolution model can raise output to Full HD. Read the Kandinsky 6.0 paper on arXiv.
Our top picks
- Audio and video generated together
The models produce synchronised 44 kHz audio with video, including lip-sync. - Two sizes for different workflows
Lite has 3 billion parameters; Pro has 29 billion, offering a faster iteration option alongside the larger model. - Text-to-audio-video generation
A text prompt can produce a five-second clip with its soundtrack. - Image-to-audio-video generation
The system can also start from an image, rather than requiring a text-only prompt. - Full-HD upscaling
A built-in super-resolution model can raise generated output to Full HD. - Standalone video enhancement tools
The paper describes VSR and VSR Lite as separate video super-resolution options. - An open release for developers
The authors say code, checkpoints and a Diffusers integration are available under the MIT licence.
Why it matters
Generating a soundtrack in step with a video could save creators the familiar second act of finding, editing and syncing audio. The two model sizes also suggest different trade-offs for experimentation and higher-fidelity output. The authors say the models are available through Fal in Pro and Lite service tiers, alongside the open release.
Our read
This is a meaningful step towards video tools that treat sound as part of the scene, not a soundtrack-shaped chore to bolt on afterwards. The paper lays out the capabilities and release, but the claims are from the researchers behind the work; real-world quality and speed are the next tests. Developers interested in open multimedia models have a clear starting point in the released code and checkpoints.
What to watch
- How well generated audio stays aligned with movement and scene changes.
- The practical quality and speed difference between Lite and Pro.
- Whether independent users can reproduce the reported capabilities with the released code and checkpoints.
Discussion spark: For an open video-generation model, would you prioritise convincing audio and lip-sync, or faster, cheaper iteration?
Sources and evidence
- Kandinsky 6.0 Video: Foundation Models for Synchronized Video and Audio Generation (6 October 2026, 02:43 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.