Tencent's WeMM-Embedding-9B Unifies Text, Image and Video
The WeChat team releases a 9-billion-parameter multimodal embedding model that maps three modalities into one shared vector space.
Tencent has released WeMM-Embedding-9B, a multimodal embedding model from its WeChat team designed to place text, images, and video into a single shared representation space. At roughly 9 billion parameters, it is aimed squarely at retrieval and search tasks that need to compare content across different formats.
Embedding models are the quiet workhorses behind modern search, recommendation, and retrieval-augmented generation. What distinguishes WeMM-Embedding-9B is its multimodal scope: rather than handling text alone, it encodes images and video into the same vector space, so a text query can surface a relevant frame, or an image can retrieve related clips without an intermediate captioning step.
Why it matters
Most widely used open embedding models are text-first, with multimodal support bolted on or limited to still images. A model that treats video as a native modality is comparatively rare, and could be useful for:
- Cross-modal search, where queries and results span text, images, and video
- Content moderation and deduplication at scale
- Retrieval pipelines feeding multimodal assistants
The release lands under a custom license rather than a standard permissive one, so teams should review the terms on the model page before building on it. Tencent has not published detailed benchmark figures alongside the initial drop, so real-world evaluation will fall to early adopters comparing it against established multimodal retrieval baselines.
Sources
- Visit
tencent/WeMM-Embedding-9B
Hugging Face
More from Tencent
All Tencent releases →Tencent's T1 Targets Long-Horizon Terminal Work
A 122B mixture-of-experts model trained with reinforcement learning claims state-of-the-art results on Terminal-Bench.

Tencent Previews Hunyuan Hy4, an Apache MoE Model
The company's next-generation Hunyuan language model arrives as an early preview with a permissive license and a mixture-of-experts design.

Tencent's AuK Bundles Voice Cloning and Speech Editing
The new open-weights model handles zero-shot TTS alongside enhancement and separation, aiming to be a broad speech toolkit rather than a single-purpose voice engine.
More in Embeddings
All Embeddings →EmbeddingGemma 2 Expands to Any Modality
Google DeepMind's compact embedding model now spans text, vision, audio, and video in a single vector space.
NeoMME: a single-tower multilingual multimodal encoder
A new open encoder aims to make document retrieval faster by treating text and images natively in one model.
H Company's NeoMME rethinks visual document retrieval
A single-tower multimodal encoder aims to make multilingual document search cheaper to fine-tune and run.
0 comments
No comments yet. Be the first to weigh in.