Zhipu AI Releases Multilingual GLM-OCR Vision Model
The new vision-language model from the creators of the GLM series is specialized for recognizing and extracting text from images across multiple languages.

Zhipu AI, the company behind the prominent GLM series of large language models, has released a new open-source model focused on a classic computer vision task: optical character recognition (OCR). The new model, called GLM-OCR, is a vision-language model (VLM) designed specifically to identify and extract text embedded in images.
The key feature of GLM-OCR is its multilingual capability. According to the project's official release page, the model is trained to handle text in Chinese, English, Korean, and Japanese, making it a potentially valuable tool for applications that need to process documents and images from across East Asia and the English-speaking world. You can find the model and usage instructions on its Hugging Face repository.
Why it matters
High-quality OCR is a foundational technology for digitizing documents, parsing user interfaces, and powering accessibility tools. While powerful OCR services are available through proprietary APIs, strong open-source alternatives empower developers to build applications with more privacy and control. GLM-OCR provides a new, specialized tool for this purpose, particularly for developers working with multilingual content.
While Zhipu AI has released the model weights, potential users should note the license. The model is available under a custom license that places limitations on its use for online services, a key distinction from more permissive licenses like Apache 2.0. This restricts its use in certain commercial applications, so developers should review the terms carefully before integrating it into their projects.
Sources
- Visit
zai-org/GLM-OCR
Hugging Face
More in Vision-Language
Agnes-3.0-Flash arrives as a multimodal reasoning model
The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.
SenseTime's SenseNova-U1.5 Unifies Vision Tasks
An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.
LLaDA-UI Brings Diffusion Decoding to GUI Agents
inclusionAI's 16.7B MoE vision-language model uses block-wise diffusion to drive graphical interface tasks.
0 comments
No comments yet. Be the first to weigh in.