Benchmarks
How open-weight models stack up — across reasoning, code, vision, and image generation. Cached from public sources and refreshed every 12 hours.
Artificial Analysis Intelligence Index — a blended measure of reasoning, knowledge, and coding across open-weight and proprietary language models.
Best value · intelligence vs. price
How much capability each open model delivers per dollar. Models on the frontier (top-left) lead on value.
Every board
Artificial Analysis Intelligence Index — a blended measure of reasoning, knowledge, and coding across open-weight and proprietary language models.
| # | Model | Intelligence |
|---|---|---|
| 1 | GLM-5.3 Zhipu AI | 44.9 |
| 2 | Kimi K3 Moonshot AI | 43.8 |
| 3 | GLM 5.3 Flash Zhipu AI | 41.9 |
| 4 | Qwen3.8 2.4T A95B Qwen · Alibaba | 40.0 |
| 5 | Qwen3.8-Flash-Next Qwen · Alibaba | 39.9 |
Human-preference Elo for text-to-image models from the Artificial Analysis Arena — open-weight and proprietary, ranked head-to-head.
| # | Model | Elo |
|---|---|---|
| 1 | FLUX.2 [dev] Black Forest Labs | 1,000 |
| 2 | FLUX.2 [dev] Turbo Fal | 999 |
| 3 | HiDream-O1-Image HiDream | 978 |
| 4 | HunyuanImage 3.0 Instruct Tencent | 966 |
| 5 | HunyuanImage 3.0 Tencent | 945 |
Human-preference Elo for instruction-based image editing models, ranked head-to-head in the Artificial Analysis Arena.
| # | Model | Elo |
|---|---|---|
| 1 | HunyuanImage 3.0 Instruct Tencent | 1,027 |
| 2 | FLUX.2 [klein] 9B Black Forest Labs | 1,014 |
| 3 | FLUX.2 [dev] Black Forest Labs | 1,000 |
| 4 | FLUX.2 [dev] Turbo Fal | 998 |
| 5 | FLUX.2 [klein] Base 9B Black Forest Labs | 973 |
Human-preference Elo for text-to-video models from the Artificial Analysis Arena — the fastest-moving open-vs-proprietary race in generative AI.
| # | Model | Elo |
|---|---|---|
| 1 | LTX-2.5 Fast Lightricks | 1,217 |
| 2 | LTX-2 Fast Lightricks | 1,125 |
| 3 | LTX-2.3 Fast Lightricks | 1,121 |
| 4 | Wan 2.2 A14B Qwen · Alibaba | 1,116 |
| 5 | HunyuanVideo-1.5 Tencent | 1,020 |
Human-preference Elo for image-to-video models, ranked head-to-head in the Artificial Analysis Arena.
| # | Model | Elo |
|---|---|---|
| 1 | LTX-2.5 Fast Lightricks | 1,214 |
| 2 | LTX-2 Fast Lightricks | 1,189 |
| 3 | LTX-2.3 Fast Lightricks | 1,147 |
| 4 | HunyuanVideo-1.5 Tencent | 1,133 |
| 5 | Wan 2.2 A14B Qwen · Alibaba | 1,105 |
Human-preference Elo for text-to-speech models from the Artificial Analysis Arena — open-weight and proprietary voices, ranked head-to-head.
| # | Model | Elo |
|---|---|---|
| 1 | Falcon 2 Murf AI | 1,155 |
| 2 | Kokoro 82M v1.0 Kokoro | 1,060 |
| 3 | Maya1 Maya Research | 1,040 |
| 4 | Higgs Audio V3 TTS Boson AI | 1,037 |
| 5 | Chatterbox Resemble AI | 1,019 |
Aggregate of contamination-resistant benchmarks — IFEval, BBH, MATH, GPQA, MUSR, MMLU-Pro — for open-weight language models. The final snapshot of Hugging Face's Open LLM Leaderboard.
| # | Model | Avg. |
|---|---|---|
| 1 | MaziyarPanahi/calme-3.2-instruct-78b | 52.1 |
| 2 | MaziyarPanahi/calme-3.1-instruct-78b | 51.3 |
| 3 | dfurman/CalmeRys-78B-Orpo-v0.1 | 51.2 |
| 4 | MaziyarPanahi/calme-2.4-rys-78b | 50.8 |
| 5 | huihui-ai/Qwen2.5-72B-Instruct-abliterated | 48.1 |
The most-downloaded open-weight models on Hugging Face over the last 30 days — the community's working set, straight from the Hub and refreshed every 12 hours.
| # | Model | Downloads |
|---|---|---|
| 1 | sentence-transformers/all-MiniLM-L6-v2 | 255.6M |
| 2 | cross-encoder/ms-marco-MiniLM-L6-v2 | 88.7M |
| 3 | BAAI/bge-small-en-v1.5 BAAI | 64.5M |
| 4 | google/electra-base-discriminator Google DeepMind | 56.4M |
| 5 | google-bert/bert-base-uncased | 47.6M |
The most-liked open-weight models on Hugging Face — a durable signal of community esteem spanning every modality, refreshed every 12 hours.
| # | Model | Likes |
|---|---|---|
| 1 | Qwen/Qwen3.8-27B Qwen · Alibaba | 15.5K |
| 2 | Black Forest Labs | 14.8K |
| 3 | deepseek-ai/DeepSeek-R1 DeepSeek | 14.3K |
| 4 | Moonshot AI | 11.4K |
| 5 | stabilityai/stable-diffusion-xl-base-1.0 Stability AI | 8.2K |