Agnes-3.0-Flash arrives as a multimodal reasoning model
The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.
Category · text
Open models tuned for step-by-step problem solving — math, logic, and multi-step planning — that show their work and trade extra compute for harder answers.
65 releases
The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.
A new mixture-of-experts model trained on verified tool interactions arrives as an early preview under an MIT license.
A compact foundation model targets math reasoning and agentic search with tool use, and its makers are releasing it fully open.
A 122B mixture-of-experts model trained with reinforcement learning claims state-of-the-art results on Terminal-Bench.
A compact 4-billion-parameter model built for tool use, coding, and multi-step reasoning arrives from TokenRhythm.
A new open 35B model is tuned for multi-step, co-working tasks rather than one-shot answers.
The flagship model uses a mixture-of-experts design that activates just 23 billion parameters per token, keeping inference costs in check.
The company's next-generation Hunyuan language model arrives as an early preview with a permissive license and a mixture-of-experts design.
The latest update to IBM's Apache 2.0 model family leans into structured reasoning while keeping its enterprise-friendly licensing.
A speed-tuned member of the GLM-5.3 family arrives with open weights and mixture-of-experts design aimed at fast, low-cost inference.
The open-weight MoE model from Zhipu AI aims to match frontier closed systems on coding tasks while undercutting them on price.
The MIT-licensed model activates just 3B parameters per token, borrowing the Qwen3.5 MoE design for text and reasoning tasks.
An MoE model built on Qwen3.5-35B-A3B aims at complex, multi-step tasks rather than one-shot answers.
The company's newest flagship targets reasoning and coding while keeping a permissive open-source license.
The Chinese lab pushes a higher-capability checkpoint of its V4 line to Hugging Face under a permissive MIT license.
A 30-billion-parameter multimodal model built to run locally, released under Apache 2.0 with an eye on agentic coding workflows.
Alibaba's Qwen team pushes its largest sparse model yet, activating 95 billion parameters per token from a 2.4-trillion-parameter pool.
The new dense model brings optional step-by-step thinking and function calling under a permissive Apache-2.0 license.
Motif Technologies debuts a mixture-of-experts language model built around grouped differential latent attention for long-context reasoning and code.
A permissively licensed reasoning model that pairs a mixture-of-experts design with ternary weights, aiming for efficiency.
A 30-billion-parameter mixture-of-experts model activates just 3 billion parameters per token, using a hybrid Mamba design to keep inference fast.
A 30B mixture-of-experts model with just 3B active parameters aims at fast, agentic coding workloads.
The MIT-licensed mixture-of-experts model returns in an updated build shipping with FP8 weights for cheaper inference.
The latest checkpoint in DeepSeek's V4 line leans into agentic workflows while keeping the permissive MIT license.
The new mixture-of-experts model activates 37B parameters per token and targets English, Korean, and Spanish reasoning tasks.
A new open mixture-of-experts model with 16B total parameters and just 3B active is tuned to run on AMD's own accelerator stack.
The Mira Murati-founded lab makes its debut with an open-weights, reasoning-focused language model.
A new mixture-of-experts model learns to reason through reinforcement learning alone, without human-annotated chains of thought.
A new 30B mixture-of-experts model from NVIDIA handles both listening and speaking within a single audio-text architecture.
The new open-weights family adds a mixture-of-experts design, encoder-free multimodal inputs, and an optional thinking mode.
The new Apache-2.0 mixture-of-experts model activates just 6B parameters per token, trading raw density for cheaper inference.
The new flagship arrives as a mixture-of-experts system with FP8 weights and open reasoning capabilities under a permissive license.
The new dense model ships in GGUF format under a permissive MIT license, aimed at local and self-hosted deployment.
A 75-billion-parameter mixture-of-experts reasoning model that activates just 9 billion parameters per token.
The 75B-parameter model activates just 9B per token and ships in NVIDIA's NVFP4 format for efficient inference.
The flagship of a new open model family arrives under a permissive MIT license, with reasoning among its stated strengths.
Alibaba's new MoE model acts as a language world model, generating the environments that agents act within.
InclusionAI's new mixture-of-experts model bets that agent-horizon scaling can rival far larger systems on long-running tasks.
The new bilingual model from the Chinese AI firm uses a Mixture of Experts architecture and sparse attention under a fully permissive license.
The open-weights multimodal model leans into coding and agentic tasks, extending Moonshot's Kimi line into a new scale bracket.
The new 3-billion-parameter model from the Chinese tech giant focuses on challenging benchmarks in mathematics, coding, and graduate-level questions.
The new open-weight model from MiniMax AI combines vision, coding, and reasoning using a Mixture-of-Experts architecture.
Moonshot AI's open-weights mixture-of-experts model reportedly outperformed Claude, GPT-5.5, and Gemini on a programming challenge.
The new 30-billion parameter Mixture-of-Experts model handles text and images while using only 3 billion active parameters for inference.
The new flagship model combines a Mixture-of-Experts architecture with a permissive MIT license, positioning it for wide commercial adoption.
The new Mixture of Experts model from the Beijing-based AI lab is optimized for fast, efficient conversational AI and carries a fully permissive license.
The new 30-billion-parameter Mixture-of-Experts model handles any combination of modalities with just 3 billion active parameters.
The new Qwen3.6-35B-A3B from Alibaba's Qwen team combines vision and language capabilities using an efficient sparse architecture.
The new conversational language model from the Chinese AI company uses a Mixture-of-Experts architecture and 8-bit weights, but is released under a restrictive custom license.
The new bilingual model from the Chinese AI firm features an efficient Mixture-of-Experts architecture and a fully permissive MIT license.
The new Mixture-of-Experts model from the Chinese AI company combines an advanced architecture with a fully permissive MIT license for commercial use.
The new 4-billion-parameter vision-language model is specialized for tasks in radiology, pathology, and complex clinical reasoning.
The new vision-language model from the Chinese AI firm uses a Mixture-of-Experts architecture and is now available on Hugging Face.
The new ERNIE 4.5 VL model brings advanced multimodal reasoning to the open-source community with an efficient Mixture-of-Experts architecture.
The new Mixture-of-Experts model is designed for complex tasks but arrives in a custom compressed format with a restrictive license.
The Shanghai-based AI startup has released a new Mixture-of-Experts model focused on complex reasoning, coding, and agentic tasks.
The new Mixture-of-Experts model is available under a permissive MIT license and is optimized for complex reasoning and coding tasks.
The new 30-billion-parameter Mixture-of-Experts model from Alibaba's Qwen team is designed to show its reasoning process for complex multimodal tasks.
The new DeepSeek-V3.1-Base is a massive 671-billion-parameter Mixture-of-Experts model designed for efficient, large-scale research and development.
The new Mixture-of-Experts model offers strong multimodal reasoning capabilities under a permissive MIT license.
The new `gpt-oss-20b` is an Apache 2.0-licensed Mixture-of-Experts model designed to run efficiently on consumer-grade hardware.
The new 117-billion-parameter `gpt-oss-120b` is a Mixture-of-Experts model focused on reasoning, released under a permissive Apache 2.0 license.
The new Mixture-of-Experts model combines massive scale with a fully permissive license, targeting complex reasoning and agentic applications.
The new Mixture-of-Experts model brings massive scale to the open-weights community, focusing on complex reasoning and coding tasks with a 128K context window.
The new GLM-4.1V-9B-Thinking model makes its vision and chain-of-thought reasoning capabilities available under a permissive MIT license.