The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestOpenBMB4.5
OpenBMBAny-to-Any

OpenBMB Releases 'Any-to-Any' Multimodal Model

The new MiniCPM-o 4.5 model from the open-source research group can process and generate interleaved combinations of images, text, and audio.

Feb 3, 2026
NotableOther
MiniCPM-o 4.5

The open-source AI community OpenBMB has released MiniCPM-o 4.5, a new model that significantly expands the possibilities for multimodal interaction. Unlike many models that process one type of input to produce a single type of output, MiniCPM-o is designed for "any-to-any" communication, capable of handling a mix of text, images, and audio in a single conversational flow.

This approach aims to create more natural and fluid interactions with AI. The model's "full-duplex" support suggests it can understand interleaved inputs—for example, a user could provide an image, ask a question in text, and follow up with a spoken clarification. In response, the model could generate its own combination of text, a new image, and synthesized speech.

Why It Matters

This release represents a move beyond simple, turn-based tasks like image captioning. It points toward AI systems that can participate in dynamic, multi-format conversations. By handling various data streams simultaneously, MiniCPM-o could power more sophisticated applications in areas like:

  • Interactive educational tools
  • Advanced accessibility software
  • Complex creative and design assistants

While technical details like parameter count were not specified in the release record, the model's architecture itself is the key development. Researchers can explore its capabilities directly, as it is available on Hugging Face. The model provides an open-source foundation for building the next generation of conversational AI agents.

Sources

  • openbmb/MiniCPM-o-4_5

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Size18.7 GB
PrecisionBF16
ArchitectureMiniCPMO
LicenseOTHER
Downloads693.8K
Likes1.5K

Modalities

Any-to-AnyVision-Language
2 versions — view changelog

0 comments

No comments yet. Be the first to weigh in.

More in Any-to-Any

StepFun/Text → Speech

StepFun's StepAudio 3 Realtime targets live voice AI

The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Sep 11, 2026
SenseTime/Any-to-Any

SenseTime's SenseNova-U1.5 Unifies Vision Tasks

An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

Sep 9, 2026
SenseNova-U1.5-8B-MoT
SenseTime/Any-to-Any

SenseTime Releases SenseNova U1.5 8B Any-to-Any Model

The new 8B multimodal model handles text, images, and image editing within a single native architecture.

Aug 19, 2026