# Tencent's WeMM-Embedding-9B Unifies Text, Image and Video

> The WeChat team releases a 9-billion-parameter multimodal embedding model that maps three modalities into one shared vector space.

Published by The Open Weights on Aug 25, 2026. Canonical: https://theopenweights.com/news/wemm-embedding-9b-nrln

## Key facts

- Company: Tencent
- Model: WeMM
- Version: 9B
- Category: Embeddings
- Modalities: Embeddings, Vision-Language
- License: Other (Open weights, commercial use allowed)
- Parameters: 9B
- Architecture: Qwen3_5ForConditionalGeneration
- Precision: BF16
- Weights size: 18.8 GB
- Significance: notable
- Published: 2026-08-25
- Last verified: 2026-09-02
- Hugging Face: https://huggingface.co/tencent/WeMM-Embedding-9B
- Canonical URL: https://theopenweights.com/news/wemm-embedding-9b-nrln

Tencent has released [WeMM-Embedding-9B](https://huggingface.co/tencent/WeMM-Embedding-9B), a multimodal embedding model from its WeChat team designed to place text, images, and video into a single shared representation space. At roughly 9 billion parameters, it is aimed squarely at retrieval and search tasks that need to compare content across different formats.

Embedding models are the quiet workhorses behind modern search, recommendation, and retrieval-augmented generation. What distinguishes WeMM-Embedding-9B is its multimodal scope: rather than handling text alone, it encodes images and video into the same vector space, so a text query can surface a relevant frame, or an image can retrieve related clips without an intermediate captioning step.

## Why it matters

Most widely used open embedding models are text-first, with multimodal support bolted on or limited to still images. A model that treats video as a native modality is comparatively rare, and could be useful for:

- Cross-modal search, where queries and results span text, images, and video
- Content moderation and deduplication at scale
- Retrieval pipelines feeding multimodal assistants

The release lands under a custom license rather than a standard permissive one, so teams should review the terms on the [model page](https://huggingface.co/tencent/WeMM-Embedding-9B) before building on it. Tencent has not published detailed benchmark figures alongside the initial drop, so real-world evaluation will fall to early adopters comparing it against established multimodal retrieval baselines.

## Get the model

- [Hugging Face](https://huggingface.co/tencent/WeMM-Embedding-9B)

## Sources

- [tencent/WeMM-Embedding-9B](https://huggingface.co/tencent/WeMM-Embedding-9B) — Hugging Face, Aug 25, 2026

---
Source: The Open Weights (https://theopenweights.com/). Aggregated and written by Claude, curated by humans. Cite as: "Tencent's WeMM-Embedding-9B Unifies Text, Image and Video", The Open Weights, Aug 25, 2026, https://theopenweights.com/news/wemm-embedding-9b-nrln