The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestQwen · Alibaba1.0-4B
Qwen · AlibabaVision-Language

Qwen-Drive 1.0 targets autonomous driving with a 4B VLM

Alibaba's Qwen team brings its vision-language stack to the road with a compact model built for perception and motion planning.

Aug 27, 2026
UpdateOther
Qwen-Drive-1.0-4B

Alibaba's Qwen group has released Qwen-Drive-1.0-4B, a vision-language model aimed squarely at autonomous driving. At roughly 4 billion parameters, it is designed to handle both perception — interpreting what a vehicle's cameras see — and motion planning, the harder task of deciding what the car should do next.

The move signals Qwen's push into a vertical domain rather than a general-purpose chatbot. Most of the Qwen lineup so far has focused on broad language and multimodal reasoning; Qwen-Drive narrows that lens to a single, safety-critical application where visual understanding and spatial planning have to work together.

Why it matters

Autonomous driving has become a proving ground for vision-language models, which can reason about scenes in natural language rather than relying solely on specialized detection pipelines. A few points stand out about this release:

  • It is a dense 4B model, small enough to be practical for on-vehicle or edge deployment scenarios.
  • It combines perception and planning in a single VLM, reflecting the industry's interest in end-to-end approaches.
  • It ships under a custom license, so teams should review the terms before commercial use.

This is the first entry in the Qwen-Drive family, and the details available at launch are modest — the record lists no published context length or benchmark figures. As with earlier Qwen releases, the real test will be how the model performs against established driving benchmarks and whether Qwen expands the family with larger variants. For now, its arrival on Hugging Face gives researchers a compact, open starting point to experiment with driving-focused multimodal models.

Sources

  • Qwen/Qwen-Drive-1.0-4B

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Parameters4B
Size9.1 GB
PrecisionBF16
ArchitectureQwenDriveForPlanning
LicenseOTHER
Downloads985
Likes60

Modalities

Vision-Language

0 comments

No comments yet. Be the first to weigh in.

More in Vision-Language

OpenBMB/Embeddings

NeoMME: a single-tower multilingual multimodal encoder

A new open encoder aims to make document retrieval faster by treating text and images natively in one model.

Sep 3, 2026
DeepSeek-V4-Flash-Vision-Exp
DeepSeek/Vision-Language

DeepSeek adds vision to its V4 Flash line

An experimental, MIT-licensed vision-language model brings image understanding to DeepSeek's fast V4 Flash architecture.

Aug 31, 2026
H company/Embeddings

H Company's NeoMME rethinks visual document retrieval

A single-tower multimodal encoder aims to make multilingual document search cheaper to fine-tune and run.

Aug 30, 2026