The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • Hardware estimates
  • RSS feed
  • llms.txt
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestZhipu AI5.3-Flash
Zhipu AIText / LLM

Zhipu releases GLM-5.3-Flash under MIT license

A speed-tuned member of the GLM-5.3 family arrives with open weights and mixture-of-experts design aimed at fast, low-cost inference.

Aug 25, 2026
NotableMIT
GLM-5.3-Flash

Zhipu AI has published GLM-5.3-Flash on Hugging Face, a speed-oriented variant of its GLM-5.3 line released under the permissive MIT license. The model ships as an open-weights download, letting developers run and modify it without the usage restrictions attached to many competing releases (Hugging Face).

The "Flash" designation signals a build tuned for lower latency and cheaper inference rather than maximum capability. It uses a mixture-of-experts (MoE) architecture, which activates only a subset of parameters per token — a common way to hold down compute cost while keeping a large total parameter pool. Zhipu positions the broader GLM-5.3 family as competitive with frontier closed models.

What's included

  • Open weights distributed under the MIT license
  • A mixture-of-experts design geared toward fast inference
  • Support for text, vision-language, and reasoning tasks

Why it matters

Fast, permissively licensed models are increasingly where open-source AI competes hardest, since teams often care as much about serving cost and deployment freedom as raw benchmark scores. An MIT-licensed Flash tier lets startups and researchers build on GLM without negotiating commercial terms, and adds to the growing roster of capable Chinese open-weight models challenging the assumption that frontier-adjacent quality requires a closed API. Buyers should confirm exact context length and parameter details on the model card, which the record leaves unspecified.

Sources

  • zai-org/GLM-5.3-Flash

    Hugging Face

    Visit
OlderTencent's WeMM-Embedding-9B Unifies Text, Image and VideoTencent · Embeddings · 2 months agoNewerIBM's Granite 4.2 Adds Reasoning to Open LLM LineIBM · Text / LLM · 2 months ago

Get the model

Hugging Face

Specs

Context window1.0M tokens
Size328.3 GB
PrecisionFP8
ArchitectureGlm5NextForConditionalGeneration
LicenseMIT
Downloads6.5M
Likes2.8K

Can you run it?

Needs a multi-GPU server — about 194 GB at 4-bit.

  • FP8 (as published)

    8×H100 80 GB node

    329.7 GB
  • 4-bit

    4×H100 80 GB · 512 GB Mac Studio

    194 GB
Your machine
GGUF builds
  • Mixture-of-experts: every expert must be loaded, so memory follows total parameters, not the active slice.
  • KV cache sized for 8,192 tokens of context; longer prompts need proportionally more.

Based on Hugging Face weights metadata. Assumes an 8K context and ~1 GB runtime overhead; actual needs vary. How we estimate


Modalities

Text / LLMVision-LanguageReasoning

The Weekly Weights

Every open release that mattered, one email a week.

0 comments

No comments yet. Be the first to weigh in.

More from Zhipu AI

All Zhipu AI releases →
Zhipu AI/Text / LLM

Zhipu's GLM-5.3 Targets Coding at a Fraction of the Cost

The open-weight MoE model from Zhipu AI aims to match frontier closed systems on coding tasks while undercutting them on price.

Aug 23, 2026
GLM-5.3
Zhipu AI/Text / LLMDatacenter

Zhipu AI Releases MIT-Licensed GLM-5.2 MoE Model

The new bilingual model from the Chinese AI firm uses a Mixture of Experts architecture and sparse attention under a fully permissive license.

Jun 17, 2026
SCAIL-2
Zhipu AI/Image → Video

Zhipu AI Releases SCAIL-2 for Character Animation

The new open-source diffusion model from the company's research arm generates video clips from a single character image and a sequence of poses.

Jun 9, 2026

More in Text / LLM

All Text / LLM →
Mistral AI/Text / LLMDatacenter

Mistral Large 4 arrives as a trillion-param MoE

The new flagship is a sparse mixture-of-experts model with roughly 49B active parameters per token.

Oct 6, 2026
NVIDIA/Text / LLM

Falcon-Emirati tunes an LLM for local dialect

TII's Falcon family gets a variant built around Emirati Arabic, aiming at culture and nuance rather than generic Gulf Arabic.

Oct 6, 2026
OpenAI/Text / LLMMulti-GPU server

Reflection releases Beam, a 501B open-weight model

The startup's first frontier-scale model ships with downloadable weights and a mixture-of-experts design aimed at reasoning.

Oct 5, 2026