The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • Hardware estimates
  • RSS feed
  • llms.txt
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestFrancisRing1
FrancisRingImage → Video

Prism Brings Joint Video-Audio Generation to Diffusion

A new MIT-licensed video diffusion transformer pairs high-resolution image-to-video output with synchronized audio and sparse attention.

Sep 24, 2026
NotableMIT
Prism

A new open model called Prism has arrived on Hugging Face, positioning itself as a high-resolution video diffusion transformer that generates video and audio together rather than treating sound as a separate pass. Released under a permissive MIT license, it targets the image-to-video task — turning still frames into motion — with an architecture built around sparse attention to keep the computation tractable at higher resolutions.

The pairing of video and audio synthesis in a single model is the detail worth watching. Most open video generators focus on the visual stream alone, leaving creators to source or synthesize sound separately. By folding audio into the generation process, Prism aims for output where motion and sound are produced jointly, which can matter for coherence in scenes where the two need to align.

Why it matters

The open video generation space has moved quickly, but several things still set a release apart:

  • A genuinely permissive MIT license, which lowers the barrier for research and commercial experimentation.
  • Joint video-audio generation, still uncommon among open releases.
  • Sparse attention, an efficiency choice aimed at making high-resolution output more practical.

For now, the record leaves key specifics — parameter count, output resolution, frame rate, and clip length — unstated, so practitioners will want to verify capabilities directly against the model page. As an initial release, Prism is best read as an early signal of where open video models are heading: toward integrated, multimodal output rather than video alone.

Sources

  • FrancisRing/Prism

    Hugging Face

    Visit
OlderFastino's GLiNER2.5-Decide targets lean NLP tasksFastino · Text / LLM · 2 weeks agoNewerLiquid AI's LFM2.5-VL-DSpark targets faster VLM inferenceUnknown · Vision-Language · 2 weeks ago

Get the model

Hugging Face

Specs

LicenseMIT
Likes132

Modalities

Image → Video

The Weekly Weights

Every open release that mattered, one email a week.

0 comments

No comments yet. Be the first to weigh in.

More from FrancisRing

All FrancisRing releases →
StableAvatar
FrancisRing/Image → VideoRuns on a laptop

StableAvatar Brings Open Source Talking Heads to Life

A new diffusion-based model from developer FrancisRing animates still images into talking avatars using only an audio track.

Aug 12, 2025

More in Image → Video

All Image → Video →
Viggle-Animate
Viggle/Image → VideoWorkstation GPU

Viggle Releases Viggle-Animate for Character Swaps

The open image-to-video model targets character replacement and video editing, distilled from MiniMax-H3.

Aug 31, 2026
4DAnyone
AntResearch/Image → Video

Ant Research releases 4DAnyone for 4D human video

The new open model turns a single input into multiview video and reconstructs humans for novel-view synthesis.

Aug 17, 2026
MiniMax-H3
MiniMax/Text → VideoWorkstation GPU

MiniMax Releases H3 Video Model on Hugging Face

The company's new diffusion model handles text-to-video and image-to-video, with support for joint audio-video generation.

Jul 28, 2026