Microsoft's Mage-VL Streams Video Natively
A codec-native multimodal foundation model aims to understand live video and vision-language input in real time.
Company
Releases
A codec-native multimodal foundation model aims to understand live video and vision-language input in real time.
A BitNet-quantized speech recognition model trades GPU dependence for efficient CPU inference in English and Chinese.
A compact model handles both text-to-image generation and instruction-based edits at native resolution, under a permissive MIT license.
A 27B-parameter vision-language model built to drive browsers and desktop apps like a human operator.
The 4-billion-parameter vision-language model targets on-screen and mobile automation, built atop Qwen3-VL.
A compact Qwen3-derived model built to explore repositories, released under a permissive MIT license.
The new open-source automatic speech recognition model handles multilingual transcription and speaker identification out of the box.
The new 500-million-parameter model is designed for generating natural, long-form speech with very low latency for interactive applications.
The 7-billion-parameter model is designed to understand and interact with graphical user interfaces, building on Alibaba's open-source Qwen2.5-VL.
The new 1.5-billion-parameter text-to-speech model is designed to generate natural, multi-speaker audio for podcasts and other long-form content.