Meituan Releases LongCat-Next 'Any-to-Any' AI Model
The Chinese tech company has released the weights for a unified model that can process and generate combinations of text, images, audio, and video.

Chinese technology company Meituan has released the weights for LongCat-Next, an ambitious 'any-to-any' multimodal model. Published under a permissive MIT license, the model marks a significant step towards more flexible and generalized AI systems that can operate across a wide spectrum of data types.
Unlike most multimodal models that handle specific input-output pairs, such as text-to-image or image-to-text, LongCat-Next is designed for true combinatorial flexibility. It can accept any mix of text, images, audio, and video as input and generate any combination of those modalities as output. For example, it could take an image and an audio clip as prompts and produce a descriptive paragraph and a short video in response.
A Unified Architecture
The model achieves this versatility through a unified framework. Instead of stitching together separate, specialized encoders and decoders for each data type, LongCat-Next uses a single, end-to-end trained network. This architecture relies on a shared vocabulary to represent and process information from different sources, enabling it to generate coherent, multimodal content from complex prompts.
The release of LongCat-Next on the Hugging Face Hub provides researchers and developers with a powerful tool for exploring the frontiers of multimodal AI. Its open-ended capabilities and permissive license encourage experimentation in creative content generation, data synthesis, and complex reasoning tasks that span multiple domains.
Sources
- Visit
meituan-longcat/LongCat-Next
Hugging Face
More in Any-to-Any
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
SenseTime's SenseNova-U1.5 Unifies Vision Tasks
An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

SenseTime Releases SenseNova U1.5 8B Any-to-Any Model
The new 8B multimodal model handles text, images, and image editing within a single native architecture.
0 comments
No comments yet. Be the first to weigh in.