The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestMicrosoft1.0
MicrosoftVision-Language

Microsoft's Mage-VL Streams Video Natively

A codec-native multimodal foundation model aims to understand live video and vision-language input in real time.

Jul 26, 2026
NotableOther
Mage-VL

Microsoft has released Mage-VL, a multimodal foundation model designed to process video and vision-language input as it streams rather than after it is fully decoded. The company describes it as a "codec-native" system, meaning it is built to work directly with compressed video formats — a design choice aimed squarely at real-time understanding. Details are available on the model's Hugging Face page.

Why codec-native matters

Most vision-language models expect fully decoded frames, which adds latency and compute overhead when applied to live or long-form video. By operating closer to the encoded stream, Mage-VL is positioned to reduce that overhead and keep pace with continuous input — the kind of workload that matters for live captioning, monitoring, and interactive assistants.

The release covers both general vision-language tasks and streaming video understanding, placing it among a growing class of models that treat video as a first-class modality rather than a sequence of still images.

Key points from the release:

  • Primary focus on real-time video and vision-language understanding
  • A codec-native architecture intended to cut decoding latency
  • Distributed under a custom ("other") license, so teams should review terms before deployment

Microsoft has not published parameter counts, context length, or benchmark figures alongside this initial 1.0 release, so independent evaluation will be needed to gauge how it compares with existing streaming and video-language systems. For now, its arrival signals continued momentum toward models that can keep up with video in motion, not just static frames.

Sources

  • microsoft/Mage-VL

    Hugging Face

    Visit
  • Mage-VL: An Efficient Codec-Native Streaming Multimodal Foundation Model

    HF Papers

    Visit

Get the model

Hugging FaceHF Papers

Specs

Size9.5 GB
PrecisionBF16
ArchitectureMageVLForConditionalGeneration
LicenseOTHER
Downloads431.5K
Likes220

Modalities

Any-to-AnyVision-Language

0 comments

No comments yet. Be the first to weigh in.

More in Vision-Language

Inkling Small
Thinkingmachines/Vision-Language

Thinking Machines Debuts Inkling Small, a Compact Multimodal MoE

The Apache-2.0 model brings mixture-of-experts efficiency to image, audio, and text tasks in a smaller footprint.

Jul 27, 2026
Apertus v1.5 70B
Swiss Ai/Text / LLM

Apertus v1.5 70B arrives with an Apache-2.0 license

Switzerland's open-model effort ships a 70-billion-parameter, multilingual and multimodal system that anyone can use, modify, and deploy.

Jul 24, 2026
Fara1.5-27B
Microsoft/Vision-Language

Microsoft's Fara1.5-27B targets computer-use agents

A 27B-parameter vision-language model built to drive browsers and desktop apps like a human operator.

Jul 17, 2026