The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestNVIDIA1.0
NVIDIAAny-to-Any

NVIDIA's Audio-Visual Flamingo Fuses Sound and Sight

A fully open multimodal model aims to reason jointly across audio, images, and long-form video.

Jul 16, 2026
NotableOther

NVIDIA has introduced Audio-Visual Flamingo, an open audio-visual language model designed to reason jointly over audio, images, and long-form video. Rather than treating sound as an afterthought bolted onto a vision system, the model is built to bring the two streams together, according to the research paper on Hugging Face.

Most widely used multimodal systems lean heavily on the visual channel, describing what appears on screen while ignoring what can be heard. That gap matters for real-world footage, where dialogue, music, and ambient sound often carry as much meaning as the picture. Audio-Visual Flamingo targets that shortfall by treating audio as a first-class input alongside frames.

Why it matters

The emphasis on "long and complex videos" is notable. Extended clips test a model's ability to track events over time and stitch together audio and visual cues that unfold minutes apart — a harder problem than answering questions about a single image.

  • Joint reasoning across audio, images, and video in one model
  • A focus on long-form, complex footage rather than short clips
  • A fully open release, lowering the barrier for researchers to build on the work

The openness is the practical draw here. By publishing the approach rather than gating it behind an API, NVIDIA gives the research community a foundation to probe audio-visual understanding directly. As an initial 1.0 release, it establishes a baseline the team and others can iterate on.

Sources

  • Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

    HF Papers

    Visit

Get the model

HF Papers

Specs

LicenseOTHER

Modalities

Any-to-AnyVision-Language

0 comments

No comments yet. Be the first to weigh in.

More in Any-to-Any

Inkling Small
Thinkingmachines/Vision-Language

Thinking Machines Debuts Inkling Small, a Compact Multimodal MoE

The Apache-2.0 model brings mixture-of-experts efficiency to image, audio, and text tasks in a smaller footprint.

Jul 27, 2026
A.X-K2 Raon Speech 21B-A3B
KRAFTON/Any-to-Any

KRAFTON releases A.X-K2 Raon speech MoE model

The game maker's new open model blends text-to-speech and speech recognition in a single 21B mixture-of-experts system with just 3B active parameters.

Jul 27, 2026
Mage-VL
Microsoft/Vision-Language

Microsoft's Mage-VL Streams Video Natively

A codec-native multimodal foundation model aims to understand live video and vision-language input in real time.

Jul 26, 2026