MiniMax Releases H3 Video Model on Hugging Face
The company's new diffusion model handles text-to-video and image-to-video, with support for joint audio-video generation.

MiniMax has published MiniMax-H3, a diffusion-based video generation model, on Hugging Face. The release supports both text-to-video and image-to-video workflows, positioning it against a growing field of open video generators.
What sets H3 apart, according to its model page, is support for joint audio-video generation — producing a soundtrack alongside the visuals rather than leaving audio as a separate post-processing step. That capability is still relatively rare in openly available video models, most of which output silent clips.
Why it matters
Open video generation has moved quickly over the past year, but audio remains a persistent gap. A model that generates sound and picture together lowers the friction for creators who would otherwise need to stitch in music or effects manually.
A few caveats are worth noting:
- MiniMax has not published parameter counts, resolution, frame rate, or maximum clip duration in the release record.
- The model ships under a non-standard "other" license, so teams should review the terms before commercial use.
As with any newly posted checkpoint, real-world quality and the practical limits of the audio-video feature will become clearer once the community begins testing it. For now, the weights and documentation are live on Hugging Face.
Sources
- Visit
MiniMaxAI/MiniMax-H3
Hugging Face
More in Text → Video
LingBot-Video puts a 30B MoE behind embodied AI video
A DiT-based mixture-of-experts model activates just 3B parameters per step and ships under an Apache 2.0 license.

NVIDIA's Cosmos 3 Edge Brings World Models Closer
A new edge-optimized variant of NVIDIA's Cosmos world-model line aims to run generative video where the compute lives.

JD.com Enters Open-Source AI Video with JoyAI-Echo
The Chinese e-commerce giant has released a new model capable of generating long-form, multi-shot videos with synchronized audio from text prompts.
0 comments
No comments yet. Be the first to weigh in.