The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestJinaaiv1
JinaaiVision-Language

Jina AI Releases Jina-OCR-v1 for Document Intelligence

The new vision-language model targets multilingual OCR and structured document understanding on a DeepSeek-VL backbone.

Sep 1, 2026
NotableOther
jina-ocr-v1

Jina AI has published jina-ocr-v1, a vision-language model aimed squarely at optical character recognition and document intelligence. Rather than positioning itself as a general-purpose multimodal assistant, the model is tuned for the practical task of turning images of pages into usable, structured text across multiple languages.

The release is built on a DeepSeek-VL backbone, borrowing a well-regarded open vision-language foundation and adapting it toward reading and parsing documents. That lineage suggests an emphasis on robust image encoding, which matters for OCR workloads where dense text, tables, and mixed layouts routinely trip up general models.

Why it matters

OCR remains one of the most commercially important corners of applied AI, powering everything from invoice processing to archival digitization. A dedicated open model in this space gives developers an alternative to closed cloud APIs and to repurposing large general VLMs that can be expensive and inconsistent on structured text.

  • Multilingual coverage, useful for documents that mix scripts or languages
  • A DeepSeek-VL foundation, tying it to an established open backbone
  • A document-intelligence focus rather than broad chat-style multimodality

Jina AI is best known for its embeddings and retrieval tooling, so an OCR-oriented VLM extends its footprint further into the document-processing pipeline. The model card on Hugging Face is the authoritative reference for licensing and usage details, which teams should review before deployment.

Sources

  • jinaai/jina-ocr-v1

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Size6.7 GB
PrecisionBF16
ArchitectureDeepseekOCRForCausalLM
LicenseOTHER
Downloads1.4K
Likes96

Modalities

Vision-Language

0 comments

No comments yet. Be the first to weigh in.

More in Vision-Language

Agnes-3.0-Flash
Agnes AI/Vision-Language

Agnes-3.0-Flash arrives as a multimodal reasoning model

The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.

Sep 11, 2026
SenseTime/Any-to-Any

SenseTime's SenseNova-U1.5 Unifies Vision Tasks

An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

Sep 9, 2026
inclusionAI/Vision-Language

LLaDA-UI Brings Diffusion Decoding to GUI Agents

inclusionAI's 16.7B MoE vision-language model uses block-wise diffusion to drive graphical interface tasks.

Sep 8, 2026