The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestinclusionAI1.5
inclusionAIAny-to-Any

Ming-Lite-Omni 1.5 Brings Any-to-Any Modality to Open Source

The new MIT-licensed model from inclusionAI can process and generate a mix of text, images, audio, and video, pushing the boundaries of open multimodal AI.

Jul 15, 2025
NotableMIT
Ming-Lite-Omni 1.5

Startup inclusionAI has released Ming-Lite-Omni 1.5, a new open-source model designed to handle a wide array of data types simultaneously. Published under a permissive MIT license, the model aims to provide "any-to-any" omni-modal capabilities, a significant step forward for generalized AI research and development. The model and its components are available now on Hugging Face.

Unlike many multimodal models that operate on a fixed input-to-output path (like text-to-image), an omni-modal system is designed to fluidly process and generate content across various formats. Ming-Lite-Omni can reportedly understand and create content using text, images, audio, and video, allowing for more complex and integrated AI applications.

A Flexible Foundation for Multimodal AI

The model's true significance lies in its combination of advanced architecture and an unrestrictive license. This opens the door for developers and researchers to experiment with sophisticated multimodal tasks that have largely been the domain of closed, proprietary systems. Potential applications could include:

  • Generating a video with a descriptive soundtrack from a single text prompt.
  • Creating a detailed textual summary of an audio-visual recording.
  • Answering questions about a video by analyzing both its frames and its spoken audio.

While specific benchmarks have not been released, the "Lite" designation in its name suggests that Ming-Lite-Omni may be a more computationally accessible version of this complex technology. Its release provides a valuable new tool for building the next generation of AI that can see, hear, and communicate in multiple dimensions.

Sources

  • inclusionAI/Ming-Lite-Omni-1.5

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Size37.8 GB
PrecisionBF16
ArchitectureBailingMMNativeForConditionalGeneration
LicenseMIT
Downloads1.4K
Likes84

Modalities

Any-to-Any

0 comments

No comments yet. Be the first to weigh in.

More in Any-to-Any

StepFun/Text → Speech

StepFun's StepAudio 3 Realtime targets live voice AI

The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Sep 11, 2026
SenseTime/Any-to-Any

SenseTime's SenseNova-U1.5 Unifies Vision Tasks

An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

Sep 9, 2026
SenseNova-U1.5-8B-MoT
SenseTime/Any-to-Any

SenseTime Releases SenseNova U1.5 8B Any-to-Any Model

The new 8B multimodal model handles text, images, and image editing within a single native architecture.

Aug 19, 2026