StarDoc-AI Releases TeleOCR for Document Parsing
A vision-language model built on Qwen2.5-VL targets bilingual Chinese and English text extraction from documents.

StarDoc-AI has published TeleOCR, a vision-language model aimed squarely at document parsing and optical character recognition. Built on top of Alibaba's Qwen2.5-VL foundation, the model is designed to read and extract text from documents in both Chinese and English.
OCR remains one of the most practical applications of multimodal AI. While general-purpose vision-language models can describe images and answer questions about them, document parsing demands precise handling of dense text, tables, and layout — the kind of structured extraction that businesses rely on for digitizing records, invoices, and forms.
Why it matters
By fine-tuning an established open model rather than training from scratch, StarDoc-AI inherits Qwen2.5-VL's visual grounding while specializing it for a narrower, higher-value task. Bilingual support is notable here:
- Chinese-language document OCR is underserved by many Western-built models
- English coverage keeps the model useful for cross-border workflows
- A specialized OCR model can outperform generalists on layout-heavy pages
The release arrives under a custom license, and the record does not specify parameter count or context length. Teams evaluating document-processing pipelines will want to test TeleOCR against both dedicated OCR engines and the broader family of open vision-language models before committing. The weights and details are available on the model's Hugging Face page.
Sources
- Visit
StarDoc-AI/TeleOCR
Hugging Face
More in Vision-Language
Xiaomi distills MiMo V2.6 into a 9B model
The new MiMo-V2.6-Distill-Qwen-9B targets agentic workloads, coding, and tool use in a size that fits on modest hardware.
Apple's LensVLM-9B targets long-context vision tasks
A 9-billion-parameter vision-language model built on Qwen3.5-9B leans on visual-text compression to stretch its usable context.

Xiaomi expands MiMo line with V2.6 multimodal models
The new Flash, Pro, and Distill variants add vision, audio, agentic behavior, and long-context handling to Xiaomi's open MiMo family.
0 comments
No comments yet. Be the first to weigh in.