The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestNVIDIA2.0
NVIDIAVision-Language

NVIDIA's Nemotron-Parse 2.0 targets document OCR

A compact vision-language model built to turn scanned pages and complex layouts into structured, machine-readable text.

Jun 30, 2026
UpdateOther
Nemotron-Parse-2.0

NVIDIA has published Nemotron-Parse 2.0, a vision-language model aimed squarely at optical character recognition and document parsing. Rather than serving as a general chat assistant, the model is purpose-built to read images of documents and extract their text and structure.

Document parsing sits at the messy intersection of vision and language: forms, invoices, receipts, and scientific papers mix dense text with tables, columns, and figures that trip up naive OCR pipelines. A model tuned for this task is meant to preserve layout and reading order, not just dump characters.

Why it matters

Extracting clean, structured data from documents is one of the most common enterprise uses of AI, and it increasingly feeds retrieval and agent workflows that depend on accurate source text. Specialized parsers like this are attractive because they can be smaller and cheaper to run than routing every page through a large general-purpose model.

A few things to note from the release:

  • It is distributed as a vision-language model on Hugging Face under a custom ("other") license, so teams should review the terms before production use.
  • NVIDIA has not published parameter counts or context details in the record, so evaluation on your own documents remains the best guide.

As the first entry in the Nemotron-Parse line, version 2.0 signals NVIDIA's continued push to broaden its open-weight Nemotron family beyond text generation into practical document intelligence.

Sources

  • nvidia/NVIDIA-Nemotron-Parse-2.0

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Size3.6 GB
PrecisionFP32
ArchitectureNemotronParseForConditionalGeneration
LicenseOTHER
Downloads13.1K
Likes97

Modalities

Vision-Language

0 comments

No comments yet. Be the first to weigh in.

More in Vision-Language

OpenMOSS/Vision-Language

OpenMOSS Debuts MOSS-VL for Real-Time Vision Interaction

A new open vision-language model family uses gated cross-attention to enable streaming, low-latency multimodal exchanges.

Aug 14, 2026
LFM2.5-VL-3B
LiquidAI/Vision-Language

Liquid AI's LFM2.5-VL-3B targets on-device vision

The 3-billion-parameter vision-language model is tuned for faster multimodal work on edge hardware.

Aug 12, 2026
North Micro Vision Instruct
Cohere/Vision-Language

Cohere Labs releases compact North Micro Vision model

A small multilingual vision-language model built for instruction following arrives under a research-only license.

Aug 10, 2026