The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestOpenMOSSMOSS-VL
OpenMOSSVision-Language

OpenMOSS Debuts MOSS-VL for Real-Time Vision Interaction

A new open vision-language model family uses gated cross-attention to enable streaming, low-latency multimodal exchanges.

Aug 14, 2026
NotableOther

OpenMOSS has released MOSS-VL, a vision-language model family aimed at real-time, streaming multimodal interaction. According to the group's technical report, the model pairs visual and text understanding with an architecture built for low-latency exchanges rather than one-shot question answering.

The defining choice is a gated cross-attention mechanism, which the team uses to fuse visual signals into the language stream as they arrive. That design is what allows the system to process input continuously and respond while a scene or conversation is still unfolding, instead of waiting for a complete input before generating a reply.

Why it matters

Most open vision-language systems are optimized for static images and full-prompt inference. A model built around streaming interaction points toward more responsive assistants, live video understanding, and agent-style workflows where timing is as important as accuracy.

  • Modalities: vision-language plus text generation
  • Core technique: gated cross-attention for real-time fusion
  • Focus: streaming, low-latency interaction over batch inference

As an initial release, MOSS-VL leaves some open questions—parameter counts, context length, and licensing terms are not fully specified in the record. Still, it extends OpenMOSS's ongoing effort to ship open multimodal systems, and the streaming emphasis distinguishes it from the crowded field of image-first VLMs. The full details are laid out in the report on Hugging Face.

Sources

  • MOSS-VL Technical Report

    HF Papers

    Visit

Get the model

HF Papers

Specs

LicenseOTHER

Modalities

Text / LLMVision-Language

0 comments

No comments yet. Be the first to weigh in.

More in Vision-Language

LFM2.5-VL-3B
LiquidAI/Vision-Language

Liquid AI's LFM2.5-VL-3B targets on-device vision

The 3-billion-parameter vision-language model is tuned for faster multimodal work on edge hardware.

Aug 12, 2026
North Micro Vision Instruct
Cohere/Vision-Language

Cohere Labs releases compact North Micro Vision model

A small multilingual vision-language model built for instruction following arrives under a research-only license.

Aug 10, 2026
Muse Glimmer 30B
Meta AI/Vision-Language

Meta's Muse Glimmer 30B Targets Local Agentic Coding

A 30-billion-parameter multimodal model built to run locally, released under Apache 2.0 with an eye on agentic coding workflows.

Aug 10, 2026