Baidu Releases NAVA for Text-to-Video with Audio
The new model from the Chinese tech giant uses a Multimodal Diffusion Transformer to generate synchronized audio and video from text or image prompts.

Baidu has released the weights for NAVA, a new generative model capable of producing video complete with synchronized audio from a variety of inputs. NAVA, which stands for Native Audio-Video Animation, can take either a text prompt or a combination of text and an image to generate short video clips. The model and examples are available on its Hugging Face repository.
Under the hood, NAVA employs a sophisticated architecture known as a Multimodal Diffusion Transformer (MMDiT). This design allows the model to process and integrate different data types—like text and image features—within the same transformer blocks, creating a more cohesive understanding of the prompt. The model is built upon Baidu's own Wan2.2 video foundation model, extending its capabilities into multimodal generation.
A More Efficient Method
Instead of traditional diffusion methods, NAVA is trained using a flow-matching technique. This is a more recent approach to training generative models that can lead to more efficient training and faster inference times, as it learns the direct path from noise to a final, coherent output. This choice of technique points to a growing trend toward more computationally efficient generative architectures.
The release of NAVA adds another significant open-weights model to the competitive text-to-video landscape. Its ability to generate audio natively alongside video is a key differentiator, as audio is often a separate, post-processing step for other models. While the model is publicly available, it uses a custom license, so developers and researchers should review the terms before incorporating it into their work.
Sources
- Visit
baidu/NAVA
Hugging Face
More in Text → Video

MiniMax Releases H3 Video Model on Hugging Face
The company's new diffusion model handles text-to-video and image-to-video, with support for joint audio-video generation.
Lightricks Releases LTX-2.5 Video Model
The latest LTX generator handles text-to-video, image-to-video, and audio-video pipelines in one open release.

Alibaba's Wan2.2-Animate-2 14B lands under Apache 2.0
A permissively licensed 14B video model aimed at character animation joins the growing Wan family.
0 comments
No comments yet. Be the first to weigh in.