Baidu releases Unlimited-OCR under permissive MIT license
The Chinese tech giant's multilingual vision-language model targets text extraction across languages and document types.

Baidu has published Unlimited-OCR on Hugging Face, a multilingual vision-language model aimed at optical character recognition. The release lands under the MIT license, one of the most permissive options available, meaning developers can use, modify, and ship the model commercially with minimal restrictions.
Unlimited-OCR is built as a vision-language system rather than a traditional OCR pipeline. That approach has become increasingly common for document understanding, where models read images and produce structured or plain text while drawing on broader language reasoning to handle layout, context, and multiple scripts.
Why it matters
OCR remains one of the most practical and widely deployed AI tasks, underpinning everything from document digitization to data entry and accessibility tools. A few points stand out about this release:
- The MIT license lowers the barrier for commercial adoption and downstream fine-tuning.
- A multilingual focus suggests the model is intended to work across scripts rather than English-only text.
- Backing from Baidu, a major player in Chinese-language AI, adds weight to the multilingual claim.
Baidu has not yet detailed parameter counts, context length, or benchmark results in the available record, so independent evaluation will be the real test of how Unlimited-OCR compares to established open OCR systems. For now, the model is available to download and try directly from its Hugging Face repository.
Sources
- Visit
baidu/Unlimited-OCR
Hugging Face
More in Vision-Language
Agnes-3.0-Flash arrives as a multimodal reasoning model
The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.
SenseTime's SenseNova-U1.5 Unifies Vision Tasks
An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.
LLaDA-UI Brings Diffusion Decoding to GUI Agents
inclusionAI's 16.7B MoE vision-language model uses block-wise diffusion to drive graphical interface tasks.
0 comments
No comments yet. Be the first to weigh in.