inclusionAI's Ming 2.0 Tackles Any-to-Any Multimodality
The new open-source Mixture-of-Experts model can process and generate content across text, images, and audio in any combination.
AI research group inclusionAI has released Ming-flash-omni 2.0, an ambitious open-source model designed to natively handle text, images, and audio. Released under a permissive MIT license, the model aims to provide a single, unified system for 'any-to-any' multimodal tasks.
Unlike many multimodal models that primarily link text and images, Ming 2.0 is built to process and generate content across all three modalities interchangeably. This could enable capabilities like generating an image from an audio clip, describing a picture with spoken words, or transcribing speech, all within one framework.
The model utilizes a Mixture-of-Experts (MoE) architecture, a design that can lead to more efficient computation by only activating relevant parts of the network for a given task. While specific details on its parameter count and training data are not yet public, the MoE approach suggests a focus on scalable performance.
This release represents another step forward for complex, open-source AI systems that can perceive and create in ways more analogous to human senses. Researchers and developers can explore the model's capabilities on its Hugging Face repository.
Sources
- Visit
inclusionAI/Ming-flash-omni-2.0
Hugging Face
More in Any-to-Any
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
SenseTime's SenseNova-U1.5 Unifies Vision Tasks
An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

SenseTime Releases SenseNova U1.5 8B Any-to-Any Model
The new 8B multimodal model handles text, images, and image editing within a single native architecture.
0 comments
No comments yet. Be the first to weigh in.