inclusionAI Adds Vision to Ling 3.0 Flash
The new Ling-3.0-flash-VL brings a mixture-of-experts vision-language model to inclusionAI's open lineup under a permissive MIT license.

inclusionAI has released Ling-3.0-flash-VL, a vision-language model that extends its Ling 3.0 Flash series into multimodal territory. Built on the bailing_moe_v3_vl architecture, the model pairs image understanding with text generation and ships under a permissive MIT license.
The "flash" branding signals a focus on speed and efficiency, and the model uses a mixture-of-experts (MoE) design — an approach that activates only a fraction of total parameters per token to keep inference costs down while preserving capacity. That makes it a natural fit for teams that want multimodal reasoning without the overhead of a dense frontier model.
Why it matters
MIT licensing is the headline for developers. It allows commercial use, modification, and redistribution with minimal restrictions, lowering the barrier for building products on top of the model. Combined with an efficiency-oriented MoE architecture, that positions Ling-3.0-flash-VL as a practical option for vision-language applications.
A few details worth noting:
- Modalities: vision-language plus text generation
- Architecture: bailing_moe_v3_vl, a mixture-of-experts design
- License: MIT, permissive for commercial and research use
inclusionAI has not published parameter counts or context-length figures in the release record, so buyers will want to check the model card for specifics before deploying. As the first vision-language entry in the Ling 3.0 Flash family, it broadens an already active open-weights lineup.
Sources
- Visit
inclusionAI/Ling-3.0-flash-VL
Hugging Face
More in Vision-Language
NeoMME: a single-tower multilingual multimodal encoder
A new open encoder aims to make document retrieval faster by treating text and images natively in one model.
DeepSeek adds vision to its V4 Flash line
An experimental, MIT-licensed vision-language model brings image understanding to DeepSeek's fast V4 Flash architecture.
H Company's NeoMME rethinks visual document retrieval
A single-tower multimodal encoder aims to make multilingual document search cheaper to fine-tune and run.
0 comments
No comments yet. Be the first to weigh in.