The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestMicrosoftZero-4B
MicrosoftVision-Language

Microsoft previews GELab-Zero-4B, a compact GUI agent

The 4-billion-parameter vision-language model targets on-screen and mobile automation, built atop Qwen3-VL.

Jun 30, 2026
UpdateOther
GELab-Zero-4B (preview)

Microsoft has quietly posted a preview of GELab-Zero-4B, a compact vision-language model aimed at graphical user interface and mobile automation. According to the model's Hugging Face page, the roughly 4-billion-parameter model is built on Alibaba's Qwen3-VL foundation and framed as an agent that can perceive and act on screens.

GUI agents are a fast-growing niche in applied AI: rather than just describing an image, these models are trained to interpret app layouts, buttons, and text fields, then plan and execute the taps or clicks needed to complete a task. The vision component is what lets the model "see" an interface the way a user would, which is essential when there's no clean API to work against.

Why it matters

  • At 4B parameters, GELab-Zero-4B sits in a size range that can plausibly run closer to the edge, which matters for mobile control where latency and privacy are concerns.
  • Building on Qwen3-VL rather than a proprietary base signals Microsoft's continued use of open weights as a starting point for specialized agents.
  • The "preview" and "Zero" labeling suggests this is an early, experimental checkpoint rather than a finished product.

Microsoft has released the model under an "other" license, and key details such as context length and formal benchmarks aren't specified in the record. As a preview, it's best read as a research artifact and a signal of where Microsoft's GUI-agent work is heading rather than a production-ready release.

Sources

  • microsoft/GELab-Zero-4B-preview-Sico-Evolution

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Parameters4B
Size8.9 GB
PrecisionBF16
ArchitectureQwen3VLForConditionalGeneration
LicenseOTHER
Downloads452
Likes56

Modalities

Vision-Language

0 comments

No comments yet. Be the first to weigh in.

More in Vision-Language

Inkling Small
Thinkingmachines/Vision-Language

Thinking Machines Debuts Inkling Small, a Compact Multimodal MoE

The Apache-2.0 model brings mixture-of-experts efficiency to image, audio, and text tasks in a smaller footprint.

Jul 27, 2026
Mage-VL
Microsoft/Vision-Language

Microsoft's Mage-VL Streams Video Natively

A codec-native multimodal foundation model aims to understand live video and vision-language input in real time.

Jul 26, 2026
Apertus v1.5 70B
Swiss Ai/Text / LLM

Apertus v1.5 70B arrives with an Apache-2.0 license

Switzerland's open-model effort ships a 70-billion-parameter, multilingual and multimodal system that anyone can use, modify, and deploy.

Jul 24, 2026