The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • Hardware estimates
  • RSS feed
  • llms.txt
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestTencent9B
TencentEmbeddings

Tencent's WeMM-Embedding-9B Unifies Text, Image and Video

The WeChat team releases a 9-billion-parameter multimodal embedding model that maps three modalities into one shared vector space.

Aug 25, 2026
NotableOther
WeMM

Tencent has released WeMM-Embedding-9B, a multimodal embedding model from its WeChat team designed to place text, images, and video into a single shared representation space. At roughly 9 billion parameters, it is aimed squarely at retrieval and search tasks that need to compare content across different formats.

Embedding models are the quiet workhorses behind modern search, recommendation, and retrieval-augmented generation. What distinguishes WeMM-Embedding-9B is its multimodal scope: rather than handling text alone, it encodes images and video into the same vector space, so a text query can surface a relevant frame, or an image can retrieve related clips without an intermediate captioning step.

Why it matters

Most widely used open embedding models are text-first, with multimodal support bolted on or limited to still images. A model that treats video as a native modality is comparatively rare, and could be useful for:

  • Cross-modal search, where queries and results span text, images, and video
  • Content moderation and deduplication at scale
  • Retrieval pipelines feeding multimodal assistants

The release lands under a custom license rather than a standard permissive one, so teams should review the terms on the model page before building on it. Tencent has not published detailed benchmark figures alongside the initial drop, so real-world evaluation will fall to early adopters comparing it against established multimodal retrieval baselines.

Sources

  • tencent/WeMM-Embedding-9B

    Hugging Face

    Visit
OlderNVIDIA's PhoneLLM Targets Voice Agents on the LinePipecat Ai · Text / LLM · 2 months agoNewerZhipu releases GLM-5.3-Flash under MIT licenseZhipu AI · Text / LLM · 2 months ago

Get the model

Hugging Face

Specs

Parameters9B
Size18.8 GB
PrecisionBF16
ArchitectureQwen3_5ForConditionalGeneration
LicenseOTHER
Downloads5.8K
Likes138

Can you run it?

Runs on a laptop — about 10.4 GB at FP8.

  • BF16 (as published)

    24 GB GPU (RTX 3090 / 4090) · 32 GB Mac

    19.8 GB
  • FP8

    12 GB GPU (RTX 3060 / 4070) · 16 GB Mac

    10.4 GB
Your machine

Based on Hugging Face weights metadata. Assumes an 8K context and ~1 GB runtime overhead; actual needs vary. How we estimate


Modalities

EmbeddingsVision-Language

The Weekly Weights

Every open release that mattered, one email a week.

0 comments

No comments yet. Be the first to weigh in.

More from Tencent

All Tencent releases →
Tencent/ReasoningWorkstation GPU

Tencent's T1 Targets Long-Horizon Terminal Work

A 122B mixture-of-experts model trained with reinforcement learning claims state-of-the-art results on Terminal-Bench.

Sep 9, 2026
Hunyuan Hy4 (preview)
Tencent/Text / LLMDatacenter

Tencent Previews Hunyuan Hy4, an Apache MoE Model

The company's next-generation Hunyuan language model arrives as an early preview with a permissive license and a mixture-of-experts design.

Aug 27, 2026
AuK
Tencent/Text → Speech

Tencent's AuK Bundles Voice Cloning and Speech Editing

The new open-weights model handles zero-shot TTS alongside enhancement and separation, aiming to be a broad speech toolkit rather than a single-purpose voice engine.

Aug 18, 2026

More in Embeddings

All Embeddings →
EmbeddingGemma 2
Google DeepMind/EmbeddingsRuns anywhere

EmbeddingGemma 2 Expands to Any Modality

Google DeepMind's compact embedding model now spans text, vision, audio, and video in a single vector space.

Sep 14, 2026
OpenBMB/Embeddings

NeoMME: a single-tower multilingual multimodal encoder

A new open encoder aims to make document retrieval faster by treating text and images natively in one model.

Sep 3, 2026
H company/Embeddings

H Company's NeoMME rethinks visual document retrieval

A single-tower multimodal encoder aims to make multilingual document search cheaper to fine-tune and run.

Aug 30, 2026