The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestZhipu AI4.5V
Zhipu AIVision-Language

Zhipu AI Releases Open Vision Model GLM-4.5V

The new Mixture-of-Experts model offers strong multimodal reasoning capabilities under a permissive MIT license.

Aug 10, 2025
Major releaseMIT
GLM-4.5V

Chinese AI firm Zhipu AI has released GLM-4.5V, a new open-source vision-language model (VLM). The model, which uses a Mixture-of-Experts (MoE) architecture, is designed for sophisticated tasks that require understanding and reasoning about both text and images simultaneously.

According to the release notes, GLM-4.5V is built upon the company's GLM-4.5-Air-Base model. The key advancement is its capacity for what Zhipu AI describes as strong multimodal reasoning. This makes it suitable for complex applications like detailed image analysis, visual question answering, and generating text grounded in visual information. The model weights and code are available now on Hugging Face.

Why it matters

The release is significant for two main reasons. First, it adds a powerful, openly accessible VLM to the ecosystem, a domain where proprietary models have often dominated. Second, its release under the permissive MIT license removes significant barriers for both commercial and research applications, allowing developers to freely build upon and integrate the technology.

The MoE architecture also suggests an efficient design, capable of activating only the necessary expert sub-networks during inference. This can lead to faster performance and lower computational costs compared to dense models of a similar capability level, making advanced multimodal AI more accessible to a wider range of developers and organizations.

Sources

  • zai-org/GLM-4.5V

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Active params12B active
Size215.4 GB
PrecisionBF16
ArchitectureGlm4vMoeForConditionalGeneration
LicenseMIT
Downloads45.9K
Likes722

Modalities

Vision-LanguageReasoning

0 comments

No comments yet. Be the first to weigh in.

More in Vision-Language

Agnes-3.0-Flash
Agnes AI/Vision-Language

Agnes-3.0-Flash arrives as a multimodal reasoning model

The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.

Sep 11, 2026
SenseTime/Any-to-Any

SenseTime's SenseNova-U1.5 Unifies Vision Tasks

An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

Sep 9, 2026
inclusionAI/Vision-Language

LLaDA-UI Brings Diffusion Decoding to GUI Agents

inclusionAI's 16.7B MoE vision-language model uses block-wise diffusion to drive graphical interface tasks.

Sep 8, 2026