Alibaba Releases Wan2.2, a 14B MoE Video Model
The new open-source diffusion model from the team behind Qwen uses a Mixture-of-Experts architecture to animate still images.

Alibaba's Qwen team has introduced a new contender in the open-source video generation space with the release of Wan2.2-I2V-A14B. This 14-billion parameter diffusion model is designed to animate still images, turning them into short video clips. It's available under a permissive Apache 2.0 license, encouraging broad adoption and experimentation.
What sets Wan2.2 apart is its use of a Mixture-of-Experts (MoE) architecture. While MoE has become a popular technique for efficiently scaling large language models, its application in video diffusion is less common. This approach allows the model to selectively activate parts of its network for a given task, potentially offering a more computationally efficient path to high-quality video generation at a large scale.
The model has been released in the popular Diffusers format, making it readily accessible for developers and researchers. This simplifies integration into existing pipelines and encourages experimentation within the community. The full model weights and usage instructions are available on the project's Hugging Face repository.
The arrival of Wan2.2 signifies a notable architectural experiment from a major AI lab. As the community pushes the boundaries of open video generation, exploring techniques like MoE could prove crucial for creating more powerful and efficient models that can compete with leading closed-source alternatives.
Sources
- Visit
Wan-AI/Wan2.2-I2V-A14B-Diffusers
Hugging Face
More in Image → Video

MiniMax Releases H3 Video Model on Hugging Face
The company's new diffusion model handles text-to-video and image-to-video, with support for joint audio-video generation.

Wan-Dancer-14B turns still images into dance videos
Alibaba's Wan team releases an Apache-2.0 image-to-video model built for music-driven dance generation.

NVIDIA's Cosmos 3 Edge Brings World Models Closer
A new edge-optimized variant of NVIDIA's Cosmos world-model line aims to run generative video where the compute lives.
0 comments
No comments yet. Be the first to weigh in.