The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestStarDoc AI1
StarDoc AIVision-Language

StarDoc-AI Releases TeleOCR for Document Parsing

A vision-language model built on Qwen2.5-VL targets bilingual Chinese and English text extraction from documents.

Aug 14, 2026
NotableOther
TeleOCR

StarDoc-AI has published TeleOCR, a vision-language model aimed squarely at document parsing and optical character recognition. Built on top of Alibaba's Qwen2.5-VL foundation, the model is designed to read and extract text from documents in both Chinese and English.

OCR remains one of the most practical applications of multimodal AI. While general-purpose vision-language models can describe images and answer questions about them, document parsing demands precise handling of dense text, tables, and layout — the kind of structured extraction that businesses rely on for digitizing records, invoices, and forms.

Why it matters

By fine-tuning an established open model rather than training from scratch, StarDoc-AI inherits Qwen2.5-VL's visual grounding while specializing it for a narrower, higher-value task. Bilingual support is notable here:

  • Chinese-language document OCR is underserved by many Western-built models
  • English coverage keeps the model useful for cross-border workflows
  • A specialized OCR model can outperform generalists on layout-heavy pages

The release arrives under a custom license, and the record does not specify parameter count or context length. Teams evaluating document-processing pipelines will want to test TeleOCR against both dedicated OCR engines and the broader family of open vision-language models before committing. The weights and details are available on the model's Hugging Face page.

Sources

  • StarDoc-AI/TeleOCR

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Size2.8 GB
PrecisionBF16
ArchitectureQwen2_5_VLForConditionalGeneration
LicenseOTHER
Downloads32.1K
Likes244

Modalities

Vision-Language

0 comments

No comments yet. Be the first to weigh in.

More in Vision-Language

MiMo-V2.6
Xiaomi/Vision-Language

Xiaomi distills MiMo V2.6 into a 9B model

The new MiMo-V2.6-Distill-Qwen-9B targets agentic workloads, coding, and tool use in a size that fits on modest hardware.

Sep 21, 2026
LensVLM-9B
Apple/Vision-Language

Apple's LensVLM-9B targets long-context vision tasks

A 9-billion-parameter vision-language model built on Qwen3.5-9B leans on visual-text compression to stretch its usable context.

Sep 21, 2026
MiMo-V2.6 (Flash/Pro
Xiaomi/Any-to-Any

Xiaomi expands MiMo line with V2.6 multimodal models

The new Flash, Pro, and Distill variants add vision, audio, agentic behavior, and long-context handling to Xiaomi's open MiMo family.

Sep 21, 2026