TeleOCR brings bilingual document parsing to Qwen2.5-VL
A new vision-language model targets Chinese and English OCR and structured document extraction, released openly on Hugging Face.

TeleOCR, a new vision-language model focused on optical character recognition and document parsing, has arrived on Hugging Face. Built on top of Alibaba's Qwen2.5-VL foundation, it targets both Chinese and English text, positioning itself for the messy reality of real-world documents rather than clean, single-language scans.
Document understanding has become one of the more practical proving grounds for open vision-language models. Beyond simply reading characters, the task increasingly means preserving layout, tables, and reading order so that downstream systems can turn a scanned page into usable structured data. TeleOCR is pitched at exactly this class of work.
Why it matters
Bilingual OCR is harder than it looks. Chinese-English mixed documents — invoices, forms, reports, and contracts — routinely trip up models tuned for a single script. By fine-tuning a capable general VLM like Qwen2.5-VL for this purpose, TeleOCR aims to give developers an open alternative to closed OCR APIs.
A few things to keep in mind about this first release:
- It is an initial version (v1) built on the Qwen2.5-VL base.
- The license is listed as "other," so teams should review terms before commercial use.
- Parameter count, context length, and benchmark figures were not specified in the release record.
As with any early open release, the real test will be how TeleOCR performs against established OCR pipelines on complex layouts. For now, the model is available to download and evaluate directly from its Hugging Face repository.
Sources
- Visit
XingChen-AGI/TeleOCR
Hugging Face
More in Vision-Language
Xiaomi distills MiMo V2.6 into a 9B model
The new MiMo-V2.6-Distill-Qwen-9B targets agentic workloads, coding, and tool use in a size that fits on modest hardware.
Apple's LensVLM-9B targets long-context vision tasks
A 9-billion-parameter vision-language model built on Qwen3.5-9B leans on visual-text compression to stretch its usable context.

Xiaomi expands MiMo line with V2.6 multimodal models
The new Flash, Pro, and Distill variants add vision, audio, agentic behavior, and long-context handling to Xiaomi's open MiMo family.
0 comments
No comments yet. Be the first to weigh in.