Prism Brings Joint Video-Audio Generation to Diffusion
A new MIT-licensed video diffusion transformer pairs high-resolution image-to-video output with synchronized audio and sparse attention.

A new open model called Prism has arrived on Hugging Face, positioning itself as a high-resolution video diffusion transformer that generates video and audio together rather than treating sound as a separate pass. Released under a permissive MIT license, it targets the image-to-video task — turning still frames into motion — with an architecture built around sparse attention to keep the computation tractable at higher resolutions.
The pairing of video and audio synthesis in a single model is the detail worth watching. Most open video generators focus on the visual stream alone, leaving creators to source or synthesize sound separately. By folding audio into the generation process, Prism aims for output where motion and sound are produced jointly, which can matter for coherence in scenes where the two need to align.
Why it matters
The open video generation space has moved quickly, but several things still set a release apart:
- A genuinely permissive MIT license, which lowers the barrier for research and commercial experimentation.
- Joint video-audio generation, still uncommon among open releases.
- Sparse attention, an efficiency choice aimed at making high-resolution output more practical.
For now, the record leaves key specifics — parameter count, output resolution, frame rate, and clip length — unstated, so practitioners will want to verify capabilities directly against the model page. As an initial release, Prism is best read as an early signal of where open video models are heading: toward integrated, multimodal output rather than video alone.
Sources
- Visit
FrancisRing/Prism
Hugging Face
More from FrancisRing
All FrancisRing releases →StableAvatar Brings Open Source Talking Heads to Life
A new diffusion-based model from developer FrancisRing animates still images into talking avatars using only an audio track.
More in Image → Video
All Image → Video →
Viggle Releases Viggle-Animate for Character Swaps
The open image-to-video model targets character replacement and video editing, distilled from MiniMax-H3.
Ant Research releases 4DAnyone for 4D human video
The new open model turns a single input into multiview video and reconstructs humans for novel-view synthesis.

MiniMax Releases H3 Video Model on Hugging Face
The company's new diffusion model handles text-to-video and image-to-video, with support for joint audio-video generation.
0 comments
No comments yet. Be the first to weigh in.