Meituan has released LongCat-Video, a 13.6-billion-parameter model for generating and continuing video, alongside an updated open-source framework for audio-driven avatars. The practical draw is a single release covering text-to-video, image-to-video and video continuation, with a stated 720p, 30fps output target.
Watch Desk analysis
What happened
The LongCat-Video project page describes a coarse-to-fine generation approach and Block Sparse Attention. Meituan also released LongCat-Video-Avatar-1.5, an updated framework for audio-driven human video generation that uses the Whisper-Large audio encoder. Explore the LongCat-Video project.
Why it matters
Video generation is moving beyond turning a prompt into a clip: continuing existing footage and generating from an image offer more flexible starting points for creators. The accompanying avatar framework adds an audio-driven route, so this is a broader toolkit than a single text-to-video demo.
The project page describes the model and its intended capabilities, but does not provide comparative quality or speed results here. A 720p, 30fps specification tells us the target format, not whether the output will look convincing or how quickly it can be produced. The pixels still have to earn their keep.
Our read
This is a substantial open-source release, and its combination of video tasks with an avatar framework makes it worth a closer look. Developers can inspect the project and assess whether it fits their work; the next useful evidence will be examples and evaluations that show how the model performs beyond its headline specifications.
What to watch
- Whether Meituan publishes model weights, usage instructions and further deployment details.
- How well generation quality holds across text, image and continuation tasks.
- Whether LongCat-Video-Avatar-1.5 produces convincing results with audio in practice.
Discussion spark: For open video-generation tools, which matters more to you: support for several creation workflows, or independently tested quality and speed?
Sources and evidence
- meituan-longcat/LongCat-Video: Foundational Video Generation Model and Avatar Framework (4 October 2026, 07:08 UTC)
Watch Desk is operated by WittyWires as an independent cross-cutting AI news tracker. It does not speak for the organisations or people it covers.