# Prism Brings Joint Video-Audio Generation to Diffusion

> A new MIT-licensed video diffusion transformer pairs high-resolution image-to-video output with synchronized audio and sparse attention.

Published by The Open Weights on Sep 24, 2026. Canonical: https://theopenweights.com/news/prism-acat

## Key facts

- Company: FrancisRing
- Model: Prism
- Version: 1
- Category: Image → Video
- Modalities: Image → Video
- License: MIT (Open weights, commercial use allowed)
- Significance: notable
- Published: 2026-09-24
- Last verified: 2026-10-10
- Hugging Face: https://huggingface.co/FrancisRing/Prism
- Canonical URL: https://theopenweights.com/news/prism-acat

A new open model called **Prism** has arrived on Hugging Face, positioning itself as a high-resolution video diffusion transformer that generates video and audio together rather than treating sound as a separate pass. Released under a permissive MIT license, it targets the image-to-video task — turning still frames into motion — with an architecture built around sparse attention to keep the computation tractable at higher resolutions.

The pairing of video and audio synthesis in a single model is the detail worth watching. Most open video generators focus on the visual stream alone, leaving creators to source or synthesize sound separately. By folding audio into the generation process, Prism aims for output where motion and sound are produced jointly, which can matter for coherence in scenes where the two need to align.

## Why it matters

The open video generation space has moved quickly, but several things still set a release apart:

- A genuinely permissive **MIT license**, which lowers the barrier for research and commercial experimentation.
- **Joint video-audio generation**, still uncommon among open releases.
- **Sparse attention**, an efficiency choice aimed at making high-resolution output more practical.

For now, the record leaves key specifics — parameter count, output resolution, frame rate, and clip length — unstated, so practitioners will want to verify capabilities directly against the [model page](https://huggingface.co/FrancisRing/Prism). As an initial release, Prism is best read as an early signal of where open video models are heading: toward integrated, multimodal output rather than video alone.

## Get the model

- [Hugging Face](https://huggingface.co/FrancisRing/Prism)

## Sources

- [FrancisRing/Prism](https://huggingface.co/FrancisRing/Prism) — Hugging Face, Sep 24, 2026

---
Source: The Open Weights (https://theopenweights.com/). Aggregated and written by Claude, curated by humans. Cite as: "Prism Brings Joint Video-Audio Generation to Diffusion", The Open Weights, Sep 24, 2026, https://theopenweights.com/news/prism-acat