The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestinclusionAIflash-omni Preview
inclusionAIAny-to-Any

inclusionAI Debuts 'Any-to-Any' Multimodal MoE Model

The new Ming-flash-omni-Preview aims to handle any combination of data modalities using an efficient Mixture of Experts architecture.

Oct 14, 2025
NotableMIT
Ming-flash-omni-Preview

AI research group inclusionAI has released Ming-flash-omni-Preview, a new open-source model designed for true multimodal flexibility. Released under a permissive MIT license, the model pursues an "any-to-any" capability, meaning it's built to process and generate a wide combination of data types, not just text and images.

This approach, often called "omnimodal," represents a significant step beyond models that are limited to specific input-output pairs, like text-to-image or audio-to-text. An any-to-any system can theoretically accept a mix of inputs—say, an image, a line of text, and an audio clip—and generate a relevant output in a requested modality.

The model is built on a Mixture of Experts (MoE) architecture, a technique that improves computational efficiency by routing inputs to specialized subnetworks, or "experts," rather than engaging the entire model for every token. According to the release card, Ming-flash-omni is based on a previous model called Ling-flash-2.0.

As major labs pursue closed, highly capable omnimodal models, the release of an open alternative like Ming-flash-omni-Preview provides researchers and developers with a valuable tool for experimentation. While labeled as a preview, it offers a foundational component for building applications that require a more fluid and comprehensive understanding of diverse data streams.

Sources

  • inclusionAI/Ming-flash-omni-Preview

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Size208.8 GB
PrecisionBF16
ArchitectureBailingMM2NativeForConditionalGeneration
LicenseMIT
Downloads2.7K
Likes70

Modalities

Any-to-Any

0 comments

No comments yet. Be the first to weigh in.

More in Any-to-Any

StepFun/Text → Speech

StepFun's StepAudio 3 Realtime targets live voice AI

The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Sep 11, 2026
SenseTime/Any-to-Any

SenseTime's SenseNova-U1.5 Unifies Vision Tasks

An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

Sep 9, 2026
SenseNova-U1.5-8B-MoT
SenseTime/Any-to-Any

SenseTime Releases SenseNova U1.5 8B Any-to-Any Model

The new 8B multimodal model handles text, images, and image editing within a single native architecture.

Aug 19, 2026