Meituan releases LongCat-Video-Avatar 1.5
An audio-driven avatar model that animates still images into talking video, with support for continuation of longer clips.
Meituan has published LongCat-Video-Avatar 1.5, an image-to-video model that generates avatar footage driven by audio. Given a still image and a voice track, the model produces synchronized talking-head video, and it adds support for continuation — extending a generated sequence rather than being limited to a single short pass. The weights and details are available on Hugging Face.
Audio-driven avatar generation has become one of the more practical corners of video AI, powering applications like virtual presenters, dubbing, and customer-facing agents. The hard part is keeping lip movement, expression, and head motion coherent over time; the ability to continue a clip is a direct response to the tendency of these systems to drift or reset when pushed beyond a few seconds.
Why it matters
Most of the attention in generative video has gone to text-to-video systems, but avatar models solve a narrower, higher-value problem for businesses that need consistent on-screen personas. A few things stand out here:
- It is released openly, giving developers direct access to the weights rather than an API-only endpoint.
- The continuation feature targets longer-form output, a common limitation in this category.
- It comes from Meituan's LongCat effort, signaling continued investment in open video models from a major Chinese company better known for its consumer platform.
The release is distributed under a custom license, so teams evaluating it for production should read the terms carefully before building on top of it. Practical specifications such as output resolution, frame rate, and maximum clip length were not detailed in the release record, and are worth confirming against the model card before deployment.
Sources
- Visit
meituan-longcat/LongCat-Video-Avatar-1.5
Hugging Face
More in Image → Video

MiniMax Releases H3 Video Model on Hugging Face
The company's new diffusion model handles text-to-video and image-to-video, with support for joint audio-video generation.

Wan-Dancer-14B turns still images into dance videos
Alibaba's Wan team releases an Apache-2.0 image-to-video model built for music-driven dance generation.

NVIDIA's Cosmos 3 Edge Brings World Models Closer
A new edge-optimized variant of NVIDIA's Cosmos world-model line aims to run generative video where the compute lives.
0 comments
No comments yet. Be the first to weigh in.