The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestMicrosoft1.0
MicrosoftVision-Language

Microsoft's Mage-VL Streams Video Natively

A codec-native multimodal foundation model aims to understand live video and vision-language input in real time.

Jul 26, 2026
NotableOther
Mage-VL

Microsoft has released Mage-VL, a multimodal foundation model designed to process video and vision-language input as it streams rather than after it is fully decoded. The company describes it as a "codec-native" system, meaning it is built to work directly with compressed video formats — a design choice aimed squarely at real-time understanding. Details are available on the model's Hugging Face page.

Why codec-native matters

Most vision-language models expect fully decoded frames, which adds latency and compute overhead when applied to live or long-form video. By operating closer to the encoded stream, Mage-VL is positioned to reduce that overhead and keep pace with continuous input — the kind of workload that matters for live captioning, monitoring, and interactive assistants.

The release covers both general vision-language tasks and streaming video understanding, placing it among a growing class of models that treat video as a first-class modality rather than a sequence of still images.

Key points from the release:

  • Primary focus on real-time video and vision-language understanding
  • A codec-native architecture intended to cut decoding latency
  • Distributed under a custom ("other") license, so teams should review terms before deployment

Microsoft has not published parameter counts, context length, or benchmark figures alongside this initial 1.0 release, so independent evaluation will be needed to gauge how it compares with existing streaming and video-language systems. For now, its arrival signals continued momentum toward models that can keep up with video in motion, not just static frames.

Sources

  • Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

    HF Papers

    Visit
  • microsoft/Mage-VL

    Hugging Face

    Visit

Get the model

Hugging FaceHF Papers

Specs

Size9.5 GB
PrecisionBF16
ArchitectureMageVLForConditionalGeneration
LicenseOTHER
Downloads27.4K
Likes405

Modalities

Vision-LanguageAny-to-Any

0 comments

No comments yet. Be the first to weigh in.

More in Vision-Language

Agnes-3.0-Flash
Agnes AI/Vision-Language

Agnes-3.0-Flash arrives as a multimodal reasoning model

The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.

Sep 11, 2026
SenseTime/Any-to-Any

SenseTime's SenseNova-U1.5 Unifies Vision Tasks

An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

Sep 9, 2026
inclusionAI/Vision-Language

LLaDA-UI Brings Diffusion Decoding to GUI Agents

inclusionAI's 16.7B MoE vision-language model uses block-wise diffusion to drive graphical interface tasks.

Sep 8, 2026