Tencent's UI-Mate-27B targets desktop automation
A 27B vision-language model built to see and operate graphical interfaces, tested on OSWorld and WindowsAgentArena.
Tencent has published UI-Mate-27B, a 27-billion-parameter vision-language model designed to act as a computer-use agent — reading what's on screen and carrying out multi-step tasks across desktop applications. Unlike a general-purpose chat model, it is tuned specifically for the loop of perceiving an interface, deciding on an action, and executing it.
The model is a dense (non-mixture-of-experts) VLM, placing it in the mid-size tier where capability and deployability meet. Tencent points to evaluation on two well-known agent benchmarks, OSWorld and WindowsAgentArena, which measure how reliably a system can complete real tasks inside actual operating-system environments rather than in simplified sandboxes.
Why it matters
GUI agents are one of the more demanding frontiers in applied AI: a model has to ground language in pixels, track state across steps, and recover from mistakes. Open weights in this category are still relatively scarce, so a 27B model aimed squarely at desktop control gives researchers and builders something concrete to test and fine-tune.
- Purpose-built for computer-use and GUI automation, not generic vision tasks
- 27B dense parameters, released under a custom ("other") license
- Benchmarked on OSWorld and WindowsAgentArena
As always with agent models, the practical questions — latency, reliability on unfamiliar apps, and license terms for commercial use — will determine how far UI-Mate-27B travels beyond the benchmark tables. The weights and details are available now on Hugging Face.
Sources
- Visit
tencent/UI-Mate-27B
Hugging Face
More in Vision-Language

Ornith 1.5 arrives as a 397B MoE multimodal model
The MIT-licensed release spans a 397B mixture-of-experts flagship plus 9B and 35B-A3B variants for lighter deployments.
OpenMOSS Debuts MOSS-VL for Real-Time Vision Interaction
A new open vision-language model family uses gated cross-attention to enable streaming, low-latency multimodal exchanges.

Liquid AI's LFM2.5-VL-3B targets on-device vision
The 3-billion-parameter vision-language model is tuned for faster multimodal work on edge hardware.
0 comments
No comments yet. Be the first to weigh in.