H Company releases Holo4 for computer-use agents
The new vision-language model is built to drive generalist agents that operate real software interfaces.
H Company has released Holo4, a vision-language model aimed squarely at one of the hardest problems in applied AI: building agents that can actually operate a computer the way a person does. According to the company's launch post on Hugging Face, Holo4 is positioned as the engine for "generalist computer-use agents" — systems that perceive a screen and act on it.
Computer-use is a distinct discipline from general chat or coding. A model has to read a graphical interface, locate the right controls, and chain actions together reliably across many steps. That requires tight coupling between visual understanding and action planning, which is why Holo4 is framed as a vision-language model rather than a text-only system.
Why it matters
The race to make agents that can click, type, and navigate real applications has drawn in most major labs, and a credible open-weight entry changes the calculus for developers who want to build on top of this capability rather than rent it through a closed API.
- Holo4 targets generalist use, not a single app or workflow
- It is distributed through Hugging Face under a custom ("other") license
- The focus is on perception-plus-action rather than conversation alone
As always, the real test will be reliability on long, multi-step tasks, where small perception errors compound quickly. For now, Holo4 adds another serious option to the growing field of models built specifically to put agents to work inside everyday software.
Sources
More from OpenAI
All OpenAI releases →Reflection releases Beam, a 501B open-weight model
The startup's first frontier-scale model ships with downloadable weights and a mixture-of-experts design aimed at reasoning.
Thinking Machines ships Inkling, its first open model
The Mira Murati-founded lab makes its debut with an open-weights, reasoning-focused language model.
OpenAI Releases 21B Open-Weight MoE Model
The new `gpt-oss-20b` is an Apache 2.0-licensed Mixture-of-Experts model designed to run efficiently on consumer-grade hardware.
More in Vision-Language
All Vision-Language →
Liquid AI's d1-3B brings multimodal models to the edge
The new LFM2-based d1-3B is a compact vision-language model aimed at running decisions directly on-device.
Perplexity releases a 27B model for multimodal routing
The open-weight 'decider' model is designed to classify queries and route them inside Perplexity's stack.
Cloudflare's Clef brings structured decisions to open models
The new open-weight vision-language family outputs typed, structured results and arrives alongside a reinforcement-learning fine-tuning platform.
0 comments
No comments yet. Be the first to weigh in.