The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestBaidu1.0
BaiduVision-Language

Baidu Releases Qianfan-OCR for Document Intelligence

The new vision-language model from the Chinese tech giant is designed for complex, multilingual optical character recognition and layout analysis.

Mar 18, 2026
NotableOther
Qianfan-OCR

Chinese technology company Baidu has released Qianfan-OCR, a new vision-language model specialized for optical character recognition and document understanding. The model is aimed at developers who need to extract text and structural information from complex documents across multiple languages.

As a document intelligence model, Qianfan-OCR is designed to go beyond simple text transcription. Its capabilities include recognizing tables, analyzing page layouts, and handling a wide variety of languages. This makes it suitable for digitizing complex materials like invoices, structured forms, and academic papers that mix text with other elements.

A New Tool for Digitization

The release adds a powerful new option to the growing ecosystem of open models for document processing. Baidu's entry provides a strong multilingual solution for enterprise and archival applications where documents often contain complex formatting. This is a critical task for businesses looking to automate data entry and researchers digitizing large volumes of text.

The model weights and usage instructions are available on the Hugging Face Hub. Potential users should note that it is released under a custom End User License Agreement, which may place restrictions on certain use cases compared to more permissive open-source licenses.

Sources

  • baidu/Qianfan-OCR

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Size9.5 GB
PrecisionBF16
ArchitectureQianfanOCRForConditionalGeneration
LicenseOTHER
Downloads253.9K
Likes1.2K

Modalities

Vision-Language

0 comments

No comments yet. Be the first to weigh in.

More in Vision-Language

Agnes-3.0-Flash
Agnes AI/Vision-Language

Agnes-3.0-Flash arrives as a multimodal reasoning model

The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.

Sep 11, 2026
SenseTime/Any-to-Any

SenseTime's SenseNova-U1.5 Unifies Vision Tasks

An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

Sep 9, 2026
inclusionAI/Vision-Language

LLaDA-UI Brings Diffusion Decoding to GUI Agents

inclusionAI's 16.7B MoE vision-language model uses block-wise diffusion to drive graphical interface tasks.

Sep 8, 2026