The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestNVIDIA1.0
NVIDIAAny-to-Any

NVIDIA's Audio-Visual Flamingo Fuses Sound and Sight

A fully open multimodal model aims to reason jointly across audio, images, and long-form video.

Jul 16, 2026
NotableOther

NVIDIA has introduced Audio-Visual Flamingo, an open audio-visual language model designed to reason jointly over audio, images, and long-form video. Rather than treating sound as an afterthought bolted onto a vision system, the model is built to bring the two streams together, according to the research paper on Hugging Face.

Most widely used multimodal systems lean heavily on the visual channel, describing what appears on screen while ignoring what can be heard. That gap matters for real-world footage, where dialogue, music, and ambient sound often carry as much meaning as the picture. Audio-Visual Flamingo targets that shortfall by treating audio as a first-class input alongside frames.

Why it matters

The emphasis on "long and complex videos" is notable. Extended clips test a model's ability to track events over time and stitch together audio and visual cues that unfold minutes apart — a harder problem than answering questions about a single image.

  • Joint reasoning across audio, images, and video in one model
  • A focus on long-form, complex footage rather than short clips
  • A fully open release, lowering the barrier for researchers to build on the work

The openness is the practical draw here. By publishing the approach rather than gating it behind an API, NVIDIA gives the research community a foundation to probe audio-visual understanding directly. As an initial 1.0 release, it establishes a baseline the team and others can iterate on.

Sources

  • Audio-Visual Flamingo: Open Audio-Visual Intelligence for Long and Complex Videos

    HF Papers

    Visit

Get the model

HF Papers

Specs

LicenseOTHER

Modalities

Any-to-AnyVision-Language

0 comments

No comments yet. Be the first to weigh in.

More in Any-to-Any

StepFun/Text → Speech

StepFun's StepAudio 3 Realtime targets live voice AI

The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Sep 11, 2026
SenseTime/Any-to-Any

SenseTime's SenseNova-U1.5 Unifies Vision Tasks

An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

Sep 9, 2026
SenseNova-U1.5-8B-MoT
SenseTime/Any-to-Any

SenseTime Releases SenseNova U1.5 8B Any-to-Any Model

The new 8B multimodal model handles text, images, and image editing within a single native architecture.

Aug 19, 2026