Zhipu AI Releases SCAIL-2 for Character Animation
The new open-source diffusion model from the company's research arm generates video clips from a single character image and a sequence of poses.

Chinese AI firm Zhipu AI, through its research arm zai, has released SCAIL-2, an open-source model designed for a specific and challenging task: character animation. The new diffusion model can take a single static image of a character and bring it to life as a short video clip, following a user-provided sequence of poses.
The model works by conditioning its video generation on two key inputs: the reference image of the character and a control video representing the desired motion, typically as a skeletal pose estimation. This method gives creators granular control over the final animation, allowing them to precisely direct the character's movements rather than relying on a simple text prompt.
Why It Matters
While many recent open video models focus on general-purpose text-to-video generation, SCAIL-2 provides a specialized tool for animators, game developers, and creative technologists. By focusing on pose-driven control, it opens up new workflows for creating character-centric content with a high degree of consistency and directorial input.
Released under the permissive MIT license, SCAIL-2 allows for broad adoption and commercial use, encouraging developers to integrate it into new applications and build upon the core technology. The model and usage instructions are available on the official zai organization page on Hugging Face.
Sources
- Visit
zai-org/SCAIL-2
Hugging Face
More in Image → Video

Viggle Releases Viggle-Animate for Character Swaps
The open image-to-video model targets character replacement and video editing, distilled from MiniMax-H3.
Ant Research releases 4DAnyone for 4D human video
The new open model turns a single input into multiview video and reconstructs humans for novel-view synthesis.

MiniMax Releases H3 Video Model on Hugging Face
The company's new diffusion model handles text-to-video and image-to-video, with support for joint audio-video generation.
0 comments
No comments yet. Be the first to weigh in.