# The Open Weights > The daily record of open-source AI. New open-weight model releases, leaderboards, and upcoming launches — aggregated from Hugging Face, arXiv, GitHub, lab blogs and Hacker News, deduplicated, and refreshed every 12 hours. The Open Weights (https://theopenweights.com/) tracks every open-weight AI model release: who made it, what it does, its license, parameter count, context window and benchmark standing, with links to the weights. Articles are written by Claude from primary sources and curated by humans; every article lists its sources. Content is free to read and cite — please attribute "The Open Weights" and link the canonical URL. ## Machine-readable - [Full text of every article](https://theopenweights.com/llms-full.txt): llms-full.txt, Markdown, newest first - Per-article Markdown: append `/index.md` to any article URL, e.g. https://theopenweights.com/news//index.md - [RSS](https://theopenweights.com/rss.xml): the 20 newest releases - [Sitemap](https://theopenweights.com/sitemap.xml) · [News sitemap](https://theopenweights.com/news-sitemap.xml) (last 48 hours) ## Key pages - [Latest releases](https://theopenweights.com/news): the feed of new open-source model releases - [New today](https://theopenweights.com/today): the day's most significant release - [Leaderboards](https://theopenweights.com/leaderboards): benchmark and human-preference rankings for open-weight models - [Models](https://theopenweights.com/models): every model family with a version changelog - [Trending](https://theopenweights.com/trending): models by 7-day download velocity - [Upcoming](https://theopenweights.com/upcoming): announced and rumored launches - [Companies](https://theopenweights.com/companies): the labs shipping open weights - [Categories](https://theopenweights.com/categories): releases by modality - [About & editorial policy](https://theopenweights.com/about) ## Categories - [Text / LLM](https://theopenweights.com/categories/llm-text): Open-weight large language models for chat, writing, and general reasoning — the foundation models you can download, fine-tune, and self-host instead of calling a closed API. - [Reasoning](https://theopenweights.com/categories/reasoning): Open models tuned for step-by-step problem solving — math, logic, and multi-step planning — that show their work and trade extra compute for harder answers. - [Code](https://theopenweights.com/categories/code): Open-weight coding models for autocomplete, refactoring, and agentic development — the engines behind self-hosted copilots and local code assistants. - [Vision-Language](https://theopenweights.com/categories/vlm): Open vision-language models that read images alongside text — for document understanding, screenshots, charts, and visual question answering you can run yourself. - [Text → Image](https://theopenweights.com/categories/text-to-image): Open diffusion and autoregressive image generators you can run on your own hardware — from photorealism to illustration, design, and concept art. - [Image Editing](https://theopenweights.com/categories/image-edit): Open models that edit existing images from a prompt — inpainting, object removal, style transfer, and instruction-based edits, all with weights you control. - [Text → Video](https://theopenweights.com/categories/text-to-video): Open text-to-video models that turn a prompt into motion — the fast-moving frontier of open-weight generative video you can self-host. - [Image → Video](https://theopenweights.com/categories/image-to-video): Open models that animate a still image into video — bringing photos and artwork to life with controllable motion, runnable on your own GPUs. - [Text → Speech](https://theopenweights.com/categories/tts-audio): Open text-to-speech and voice models for natural narration, voice cloning, and real-time speech — self-hostable alternatives to closed voice APIs. - [Speech → Text](https://theopenweights.com/categories/asr-stt): Open speech-recognition models that transcribe audio to text across languages — for captioning, voice interfaces, and pipelines you run on your own infrastructure. - [Music](https://theopenweights.com/categories/music): Open models that generate music and audio from text or melody — instrumentals, sound design, and full tracks, with weights you can download and tune. - [Embeddings](https://theopenweights.com/categories/embeddings): Open embedding models that turn text and images into vectors — the retrieval and semantic-search backbone of self-hosted RAG and recommendation systems. - [Text → 3D](https://theopenweights.com/categories/text-to-3d): Open models that generate 3D meshes and assets from text or images — for games, simulation, and design, with open weights you can build pipelines around. - [Any-to-Any](https://theopenweights.com/categories/multimodal-any): Open any-to-any models that take and produce text, images, audio, and more in a single network — the most general open-weight systems being released. ## Leaderboards - [LLM Intelligence](https://theopenweights.com/leaderboards/llm-intelligence): Artificial Analysis Intelligence Index — a blended measure of reasoning, knowledge, and coding across open-weight and proprietary language models. Source: Artificial Analysis. Captured 2026-10-07. Top: 1. GLM-5.3 (Zhipu AI) 44.8; 2. Kimi K3 (Moonshot AI) 43.6; 3. GLM 5.3 Flash (Zhipu AI) 41.8; 4. Qwen3.8 2.4T A95B (Qwen · Alibaba) 39.9; 5. Qwen3.8-Flash-Next (Qwen · Alibaba) 39.8 - [Text-to-Image](https://theopenweights.com/leaderboards/text-to-image): Human-preference Elo for text-to-image models from the Artificial Analysis Arena — open-weight and proprietary, ranked head-to-head. Source: Artificial Analysis Arena. Captured 2026-10-07. Top: 1. Qwen-Image-2.1 (Qwen · Alibaba) 1035; 2. FLUX.2 [dev] (Black Forest Labs) 1000; 3. FLUX.2 [dev] Turbo 998; 4. HunyuanImage 3.0 Instruct (Tencent) 996; 5. HunyuanImage 3.0 (Tencent) 983 - [Image Editing](https://theopenweights.com/leaderboards/image-editing): Human-preference Elo for instruction-based image editing models, ranked head-to-head in the Artificial Analysis Arena. Source: Artificial Analysis Arena. Captured 2026-10-07. Top: 1. Qwen-Image-2.1 (Qwen · Alibaba) 1074; 2. HunyuanImage 3.0 Instruct (Tencent) 1068; 3. HunyuanImage 3.0 (Tencent) 1029; 4. FLUX.2 [klein] 9B (Black Forest Labs) 1013; 5. FLUX.2 [dev] Turbo 1004 - [Text-to-Video](https://theopenweights.com/leaderboards/text-to-video): Human-preference Elo for text-to-video models from the Artificial Analysis Arena — the fastest-moving open-vs-proprietary race in generative AI. Source: Artificial Analysis Arena. Captured 2026-10-07. Top: 1. FLUX 3 (Black Forest Labs) 1091; 2. LTX-2.5 Fast (Lightricks) 937; 3. LTX-2.3 Fast (Lightricks) 857 - [Image-to-Video](https://theopenweights.com/leaderboards/image-to-video): Human-preference Elo for image-to-video models, ranked head-to-head in the Artificial Analysis Arena. Source: Artificial Analysis Arena. Captured 2026-10-07. Top: 1. LTX-2.5 Fast (Lightricks) 1215; 2. LTX-2 Fast (Lightricks) 1191; 3. LTX-2.3 Fast (Lightricks) 1150; 4. HunyuanVideo-1.5 (Tencent) 1134; 5. Wan 2.2 A14B (Qwen · Alibaba) 1107 - [Text-to-Speech](https://theopenweights.com/leaderboards/text-to-speech): Human-preference Elo for text-to-speech models from the Artificial Analysis Arena — open-weight and proprietary voices, ranked head-to-head. Source: Artificial Analysis Arena. Captured 2026-10-06. Top: 1. Kokoro 82M v1.0 1064; 2. Maya1 1049; 3. Higgs Audio V3 TTS 1038; 4. Chatterbox (Resemble AI) 1025; 5. Zonos-v0.1 (Zyphra) 1000 - [Open LLM](https://theopenweights.com/leaderboards/open-llm): Aggregate of contamination-resistant benchmarks — IFEval, BBH, MATH, GPQA, MUSR, MMLU-Pro — for open-weight language models. The final snapshot of Hugging Face's Open LLM Leaderboard. Source: Open LLM Leaderboard. Captured 2026-06-17. Top: 1. MaziyarPanahi/calme-3.2-instruct-78b 52.1; 2. MaziyarPanahi/calme-3.1-instruct-78b 51.3; 3. dfurman/CalmeRys-78B-Orpo-v0.1 51.2; 4. MaziyarPanahi/calme-2.4-rys-78b 50.8; 5. huihui-ai/Qwen2.5-72B-Instruct-abliterated 48.1 - [Most Downloaded](https://theopenweights.com/leaderboards/most-downloaded): The most-downloaded open-weight models on Hugging Face over the last 30 days — the community's working set, straight from the Hub and refreshed every 12 hours. Source: Hugging Face Hub. Captured 2026-10-07. Top: 1. sentence-transformers/all-MiniLM-L6-v2 234015516; 2. cross-encoder/ms-marco-MiniLM-L6-v2 83236540; 3. BAAI/bge-small-en-v1.5 (BAAI) 62278766; 4. sentence-transformers/paraphrase-multilingual-MiniLM-L12-v2 50332290; 5. google/electra-base-discriminator (Google DeepMind) 46039671 - [Community Favorites](https://theopenweights.com/leaderboards/most-liked): The most-liked open-weight models on Hugging Face — a durable signal of community esteem spanning every modality, refreshed every 12 hours. Source: Hugging Face Hub. Captured 2026-10-07. Top: 1. Qwen/Qwen3.8-27B (Qwen · Alibaba) 17163; 2. black-forest-labs/FLUX.1-dev (Black Forest Labs) 15415; 3. deepseek-ai/DeepSeek-R1 (DeepSeek) 14328; 4. moonshotai/Kimi-K3 (Moonshot AI) 11616; 5. stabilityai/stable-diffusion-xl-base-1.0 (Stability AI) 8279 ## Companies - [Qwen · Alibaba](https://theopenweights.com/companies/alibaba-qwen): 37 models - [NVIDIA](https://theopenweights.com/companies/nvidia): 29 models - [Tencent](https://theopenweights.com/companies/tencent): 18 models - [inclusionAI](https://theopenweights.com/companies/inclusion-ai): 15 models - [Google DeepMind](https://theopenweights.com/companies/google): 14 models - [Zhipu AI](https://theopenweights.com/companies/zhipu-ai): 13 models - [Microsoft](https://theopenweights.com/companies/microsoft): 11 models - [Unknown](https://theopenweights.com/companies/unknown): 10 models - [Baidu](https://theopenweights.com/companies/baidu): 9 models - [OpenMOSS](https://theopenweights.com/companies/openmoss): 9 models - [OpenBMB](https://theopenweights.com/companies/openbmb): 8 models - [DeepSeek](https://theopenweights.com/companies/deepseek): 8 models - [Meituan](https://theopenweights.com/companies/meituan): 7 models - [Xiaomi](https://theopenweights.com/companies/xiaomi): 6 models - [Black Forest Labs](https://theopenweights.com/companies/black-forest-labs): 6 models - [Mistral AI](https://theopenweights.com/companies/mistral): 6 models - [LiquidAI](https://theopenweights.com/companies/liquidai): 5 models - [ByteDance](https://theopenweights.com/companies/bytedance): 4 models - [MiniMax](https://theopenweights.com/companies/minimax): 4 models - [Cohere](https://theopenweights.com/companies/cohere): 4 models - [Moonshot AI](https://theopenweights.com/companies/moonshot-ai): 4 models - [StepFun](https://theopenweights.com/companies/stepfun): 4 models - [SenseTime](https://theopenweights.com/companies/sensetime): 4 models - [Bosonai](https://theopenweights.com/companies/bosonai): 4 models - [Resemble AI](https://theopenweights.com/companies/resemble-ai): 3 models - [IBM](https://theopenweights.com/companies/ibm): 3 models - [KRAFTON](https://theopenweights.com/companies/krafton): 3 models - [nineninesix](https://theopenweights.com/companies/nineninesix): 3 models - [Poolside](https://theopenweights.com/companies/poolside): 3 models - [Cactus Compute](https://theopenweights.com/companies/cactus-compute): 3 models - [Skywork](https://theopenweights.com/companies/skywork): 2 models - [Maya Research](https://theopenweights.com/companies/maya-research): 2 models - [Supertone](https://theopenweights.com/companies/supertone): 2 models - [JD](https://theopenweights.com/companies/jd): 2 models - [Motif Technologies](https://theopenweights.com/companies/motif-technologies): 2 models - [Krea](https://theopenweights.com/companies/krea): 2 models - [Meta AI](https://theopenweights.com/companies/meta): 2 models - [moondream](https://theopenweights.com/companies/moondream): 2 models - [YatharthS](https://theopenweights.com/companies/yatharths): 2 models - [Datalab To](https://theopenweights.com/companies/datalab-to): 2 models - [Soul AILab](https://theopenweights.com/companies/soul-ailab): 2 models - [HumeAI](https://theopenweights.com/companies/humeai): 2 models - [ekwek](https://theopenweights.com/companies/ekwek): 2 models - [robbyant](https://theopenweights.com/companies/robbyant): 2 models - [BAAI](https://theopenweights.com/companies/baai): 2 models - [Ai Sage](https://theopenweights.com/companies/ai-sage): 2 models - [Prism Ml](https://theopenweights.com/companies/prism-ml): 2 models - [Thinkingmachines](https://theopenweights.com/companies/thinkingmachines): 2 models - [Internlm](https://theopenweights.com/companies/internlm): 2 models - [Edge0](https://theopenweights.com/companies/edge0): 2 models - [IFM](https://theopenweights.com/companies/ifm): 2 models - [Nex Agi](https://theopenweights.com/companies/nex-agi): 2 models - [Yandex](https://theopenweights.com/companies/yandex): 2 models - [Fastino](https://theopenweights.com/companies/fastino): 2 models - [H company](https://theopenweights.com/companies/h-company): 2 models - [AIDC-AI](https://theopenweights.com/companies/aidc-ai): 1 model - [HiDream.ai](https://theopenweights.com/companies/hidream-ai): 1 model - [FrancisRing](https://theopenweights.com/companies/francisring): 1 model - [Zyphra](https://theopenweights.com/companies/zyphra): 1 model - [T-Tech](https://theopenweights.com/companies/t-tech): 1 model - [FreedomIntelligence](https://theopenweights.com/companies/freedomintelligence): 1 model - [RaphaelLiu](https://theopenweights.com/companies/raphaelliu): 1 model - [MisoLabs](https://theopenweights.com/companies/misolabs): 1 model - [zhifeixie](https://theopenweights.com/companies/zhifeixie): 1 model - [Kyutai](https://theopenweights.com/companies/kyutai): 1 model - [rednote-hilab](https://theopenweights.com/companies/rednote-hilab): 1 model - [NexaAI](https://theopenweights.com/companies/nexaai): 1 model - [EPFL VITA](https://theopenweights.com/companies/epfl-vita): 1 model - [HKUSTAudio](https://theopenweights.com/companies/hkustaudio): 1 model - [chetwinlow1](https://theopenweights.com/companies/chetwinlow1): 1 model - [Allen Institute for AI](https://theopenweights.com/companies/allen-ai): 1 model - [Kuaishou](https://theopenweights.com/companies/kuaishou): 1 model - [k2-fsa](https://theopenweights.com/companies/k2-fsa): 1 model - [GAIR](https://theopenweights.com/companies/gair): 1 model - [neuphonic](https://theopenweights.com/companies/neuphonic): 1 model - [Lightricks](https://theopenweights.com/companies/lightricks): 1 model - [OpenAI](https://theopenweights.com/companies/openai): 1 model - [WeiboAI](https://theopenweights.com/companies/weiboai): 1 model - [Vibevoice](https://theopenweights.com/companies/vibevoice): 1 model - [Quark Vision](https://theopenweights.com/companies/quark-vision): 1 model - [Aratako](https://theopenweights.com/companies/aratako): 1 model - [Alpha-VLLM](https://theopenweights.com/companies/alpha-vllm): 1 model - [LightOn](https://theopenweights.com/companies/lightonai): 1 model - [FlashLabs](https://theopenweights.com/companies/flashlabs): 1 model - [Nanbeige](https://theopenweights.com/companies/nanbeige): 1 model - [Ideogram Ai](https://theopenweights.com/companies/ideogram-ai): 1 model - [Aoi Ot](https://theopenweights.com/companies/aoi-ot): 1 model - [Huaichang](https://theopenweights.com/companies/huaichang): 1 model - [Kugelaudio](https://theopenweights.com/companies/kugelaudio): 1 model - [Fishaudio](https://theopenweights.com/companies/fishaudio): 1 model - [Nari Labs](https://theopenweights.com/companies/nari-labs): 1 model - [Boogu](https://theopenweights.com/companies/boogu): 1 model - [Stability AI](https://theopenweights.com/companies/stability-ai): 1 model - [Owensong](https://theopenweights.com/companies/owensong): 1 model - [InternScience](https://theopenweights.com/companies/internscience): 1 model - [Upstage](https://theopenweights.com/companies/upstage): 1 model - [ATH MaaS](https://theopenweights.com/companies/ath-maas): 1 model - [MuScriptor](https://theopenweights.com/companies/muscriptor): 1 model - [Deepreinforce Ai](https://theopenweights.com/companies/deepreinforce-ai): 1 model - [Open Gigaai](https://theopenweights.com/companies/open-gigaai): 1 model - [Kwaipilot](https://theopenweights.com/companies/kwaipilot): 1 model - [Swiss Ai](https://theopenweights.com/companies/swiss-ai): 1 model - [Amd](https://theopenweights.com/companies/amd): 1 model - [LGAI EXAONE](https://theopenweights.com/companies/lgai-exaone): 1 model - [Nyralabs](https://theopenweights.com/companies/nyralabs): 1 model - [Skt](https://theopenweights.com/companies/skt): 1 model - [Deepgrove](https://theopenweights.com/companies/deepgrove): 1 model - [Lodestones](https://theopenweights.com/companies/lodestones): 1 model - [Dots Studio](https://theopenweights.com/companies/dots-studio): 1 model - [Shunyalabs](https://theopenweights.com/companies/shunyalabs): 1 model - [IndexTeam](https://theopenweights.com/companies/indexteam): 1 model - [Superwhisper](https://theopenweights.com/companies/superwhisper): 1 model - [Audio8](https://theopenweights.com/companies/audio8): 1 model - [AntResearch](https://theopenweights.com/companies/antresearch): 1 model - [Apodex](https://theopenweights.com/companies/apodex): 1 model - [Thomsonreuters](https://theopenweights.com/companies/thomsonreuters): 1 model - [BreezeBlue](https://theopenweights.com/companies/breezeblue): 1 model - [Pipecat Ai](https://theopenweights.com/companies/pipecat-ai): 1 model - [Fla Hub](https://theopenweights.com/companies/fla-hub): 1 model - [XHToken](https://theopenweights.com/companies/xhtoken): 1 model - [Jinaai](https://theopenweights.com/companies/jinaai): 1 model - [Mrfakename](https://theopenweights.com/companies/mrfakename): 1 model - [Agnes AI](https://theopenweights.com/companies/agnes-ai): 1 model - [Thesysdev](https://theopenweights.com/companies/thesysdev): 1 model - [Netease Youdao](https://theopenweights.com/companies/netease-youdao): 1 model - [TaichuAI](https://theopenweights.com/companies/taichuai): 1 model - [Viggle](https://theopenweights.com/companies/viggle): 1 model - [FermionResearch](https://theopenweights.com/companies/fermionresearch): 1 model - [Ukisai](https://theopenweights.com/companies/ukisai): 1 model - [TokenRhythm](https://theopenweights.com/companies/tokenrhythm): 1 model - [Convaiinnovations](https://theopenweights.com/companies/convaiinnovations): 1 model - [Apple](https://theopenweights.com/companies/apple): 1 model - [StarDoc AI](https://theopenweights.com/companies/stardoc-ai): 1 model - [XingChen AGI](https://theopenweights.com/companies/xingchen-agi): 1 model - [Pyannote](https://theopenweights.com/companies/pyannote): 1 model - [Cloudflare](https://theopenweights.com/companies/cloudflare): 1 model - [Aleph Alpha](https://theopenweights.com/companies/aleph-alpha): 1 model - [Perplexity Ai](https://theopenweights.com/companies/perplexity-ai): 1 model - [Reflection AI](https://theopenweights.com/companies/reflection-ai): 1 model - [Autotrust](https://theopenweights.com/companies/autotrust): 1 model - [Soofi Project](https://theopenweights.com/companies/soofi-project): 0 models - [Ornith Ai](https://theopenweights.com/companies/ornith-ai): 0 models - [M A P](https://theopenweights.com/companies/m-a-p): 0 models ## Releases (474, newest first) - [Mistral Large 4 arrives as a trillion-param MoE](https://theopenweights.com/news/mistral-large-4-uu8c) — Mistral AI · Text / LLM · Other · Oct 6, 2026: Mistral AI previews Mistral Large 4, a 1T-parameter mixture-of-experts model activating about 49B parameters per token. - [Falcon-Emirati tunes an LLM for local dialect](https://theopenweights.com/news/falcon-emirati-5wji) — NVIDIA · Text / LLM · Other · Oct 6, 2026: TII releases Falcon-Emirati, a Falcon LLM specialized for Emirati Arabic dialect and cultural context. - [Reflection releases Beam, a 501B open-weight model](https://theopenweights.com/news/beam-np3t) — OpenAI · Text / LLM · Other · Oct 5, 2026: Reflection's Beam is a 501B-parameter mixture-of-experts model released as open weights, with a focus on text and reasoning tasks. - [Reflection AI debuts Beam, a 501B open model](https://theopenweights.com/news/beam-oedh) — Reflection AI · Text / LLM · Other · Oct 5, 2026: Reflection AI has released Beam, a 501B-parameter open-weight language model built for text generation and reasoning. - [Kandinsky 6.0 Generates Video and Audio Together](https://theopenweights.com/news/kandinsky-6-0-video-yhja) — Unknown · Text → Video · Other · Oct 3, 2026: Kandinsky 6.0 Video is a text-to-video family that generates synchronized audio and video, shipping in 3B Lite and 29B Pro variants. - [Allen Institute Open-Sources AstaBrief Report Model](https://theopenweights.com/news/astabrief-1ktv) — Allen Institute for AI · Text / LLM · Apache 2.0 · Oct 2, 2026: AI2 has open-sourced AstaBrief, the fast report-generation model that powers its Asta research assistant, under Apache 2.0. - [Aleph Alpha releases Kolibri, a sovereign reasoning model](https://theopenweights.com/news/kolibri-1-v2uw) — Aleph Alpha · Reasoning · Other · Oct 2, 2026: Aleph Alpha's Kolibri-1 is an open-weight reasoning MoE model built for German and English, pitched as a sovereign European option. - [Perplexity releases a 27B model for multimodal routing](https://theopenweights.com/news/pplx-decider-v1-27b-8u0o) — Perplexity Ai · Vision-Language · Apache 2.0 · Oct 1, 2026: Perplexity published pplx-decider-v1-27b, a 27B open-weight vision-language model for query classification and routing, under Apache 2.0. - [Cloudflare's Clef brings structured decisions to open models](https://theopenweights.com/news/clef-wde9) — Cloudflare · Vision-Language · Other · Oct 1, 2026: Cloudflare releases Clef, open-weight VLMs that turn image-text input into typed structured output, plus an RL fine-tuning platform. - [Cactus Compute's Whistle brings speech-to-text to the edge](https://theopenweights.com/news/whistle-gdzg) — Cactus Compute · Speech → Text · Other · Sep 30, 2026: Cactus Compute released Whistle, an on-device speech-to-text model optimized for edge and WebAssembly deployment, supporting English, German, and French. - [JEV-27B-VL Pairs Vision-Language With Calibrated Odds](https://theopenweights.com/news/jev-27b-vl-2it8) — Autotrust · Vision-Language · Other · Sep 30, 2026: JEV-27B-VL is a 27B vision-language model that emits typed decisions alongside calibrated probability estimates. - [AREX-2 arrives as an open deep-research agent model](https://theopenweights.com/news/arex-2-o62j) — BAAI · Reasoning · Other · Sep 29, 2026: AREX-2 is an open deep-research agent model built for tool use, long context and self-improvement, now on Hugging Face. - [Phonon-2 brings on-device ASR to Apple Silicon](https://theopenweights.com/news/phonon-2-prj5) — FermionResearch · Speech → Text · Other · Sep 28, 2026: Phonon-2 is a ~0.6B English ASR model, quantized from Parakeet to run on-device on Apple Silicon. - [H Company's Holo4 Takes On Computer-Use Agents](https://theopenweights.com/news/holo4-11m0) — H company · Vision-Language · Other · Sep 28, 2026: H Company's Holo4 is a vision-language model designed to power generalist computer-use agents that navigate real software interfaces. - [Liquid AI's LFM2.5-VL-DSpark targets faster VLM inference](https://theopenweights.com/news/lfm2-5-vl-dspark-vobr) — Unknown · Vision-Language · Other · Sep 24, 2026: Liquid AI releases LFM2.5-VL-DSpark, a vision-language model optimized for accelerated inference. - [Fastino's GLiNER2.5-Decide targets lean NLP tasks](https://theopenweights.com/news/gliner2-5-decide-enf3) — Fastino · Text / LLM · Other · Sep 23, 2026: Fastino releases GLiNER2.5-Decide, a compact model for NER, intent, sentiment and topic classification, now on Hugging Face. - [Black Forest Labs brings FLUX to robotics](https://theopenweights.com/news/flux-3-action-mx81) — Black Forest Labs · Any-to-Any · Other · Sep 22, 2026: Black Forest Labs released FLUX 3 Action Base, a FLUX-derived world-action model for robotics, distributed via LeRobot on Hugging Face. - [Xiaomi distills MiMo V2.6 into a 9B model](https://theopenweights.com/news/mimo-v2-6-distill-qwen-9b-zef5) — Xiaomi · Vision-Language · Other · Sep 21, 2026: Xiaomi releases MiMo-V2.6-Distill-Qwen-9B, a 9B model distilled from MiMo V2.6 for agentic, coding, and tool-use tasks. - [Apple's LensVLM-9B targets long-context vision tasks](https://theopenweights.com/news/lensvlm-9b-8tr2) — Apple · Vision-Language · Other · Sep 21, 2026: Apple's LensVLM-9B is a 9B vision-language model on Qwen3.5-9B, using visual-text compression for long-context multimodal work. - [Xiaomi expands MiMo line with V2.6 multimodal models](https://theopenweights.com/news/distill-40w8) — Xiaomi · Any-to-Any · Other · Sep 21, 2026: Xiaomi's MiMo V2.6 arrives as a family of multimodal RL-trained models spanning Flash, Pro, and Distill variants with vision, audio, and agent capabilities. - [Xiaomi's MiMo V2.6-Pro-RL Targets Agentic Multimodal Work](https://theopenweights.com/news/mimo-v2-6-pro-rl-eglk) — Xiaomi · Vision-Language · Other · Sep 21, 2026: Xiaomi releases MiMo-V2.6-Pro-RL, a reinforcement-learning-tuned multimodal model spanning vision, audio, video, and long-context reasoning. - [Audio8-ASR-Infinite brings streaming bilingual speech recognition](https://theopenweights.com/news/audio8-asr-infinite-2rwh) — Edge0 · Speech → Text · Other · Sep 21, 2026: Audio8-ASR-Infinite is a streaming, real-time speech-recognition model supporting Chinese and English, now on Hugging Face. - [Moondream shrinks Parakeet ASR for CPUs](https://theopenweights.com/news/parakeet-redux-b5ng) — moondream · Speech → Text · Other · Sep 18, 2026: Moondream's Parakeet Redux is a ternary-quantized speech recognition model tuned for CPU and Apple Silicon, covering English, German and French. - [Convai's Laya targets fast AI decisions and guardrails](https://theopenweights.com/news/laya-system-one-pdfx) — Convaiinnovations · Text / LLM · Other · Sep 18, 2026: Convai Innovations released Laya, a System-One decision and classification model built for routing, scoring and guardrails. - [inclusionAI's Ming-Image targets graphic design](https://theopenweights.com/news/ming-image-0-1-design-pqrh) — inclusionAI · Text → Image · MIT · Sep 17, 2026: inclusionAI releases Ming-Image-0.1-Design, an MIT-licensed text-to-image model built for graphic design with strong text rendering and RGBA output. - [Microsoft Releases FrogNano-4B Under MIT License](https://theopenweights.com/news/frognano-4b-wi64) — Microsoft · Text / LLM · MIT · Sep 17, 2026: Microsoft's FrogNano-4B is a compact 4B-parameter, MIT-licensed text model now available on Hugging Face. - [Ternary-Bonsai-2 packs a 27B model into 2-bit form](https://theopenweights.com/news/ternary-bonsai-2-27b-y04w) — Prism Ml · Text / LLM · Other · Sep 16, 2026: Ternary-Bonsai-2-27B is a 2-bit, 27B-parameter model with hybrid attention built for local inference on CUDA and Metal hardware. - [Swift-1.5 Trims Reasoning Tokens on a 27B Qwen Model](https://theopenweights.com/news/swift-1-5-qwen3-8-27b-dg9s) — Ukisai · Reasoning · Other · Sep 16, 2026: Swift-1.5-Qwen3.8-27B is a token-efficient reasoning tune claiming ~58% fewer thinking tokens and about 2x speed. - [Xing4.0 arrives as a 29B MoE with 4B active params](https://theopenweights.com/news/xing4-0-29b-a4b-dpmk) — XingChen AGI · Text / LLM · Apache 2.0 · Sep 16, 2026: Xing4.0-29B-A4B is a 29B-parameter MoE text model with 4B active params, released under Apache-2.0. - [Cactus Needle 3: tiny on-device tool-calling models](https://theopenweights.com/news/cactus-needle-3-n6y8) — Cactus Compute · Text / LLM · Apache 2.0 · Sep 16, 2026: Cactus Needle 3 packs on-device tool-calling into 8–29MB models under Apache 2.0, targeting automation tasks without the cloud. - [Nari Labs Ships Qwen3-Based TTS and ASR Models](https://theopenweights.com/news/nari-qwen3-tts-qwen3-asr-pkbt) — Unknown · Text → Speech · Other · Sep 14, 2026: Nari Labs releases Qwen3-based TTS and ASR models aimed at high accuracy, low latency and cost, topping Coval voice benchmarks. - [EmbeddingGemma 2 Expands to Any Modality](https://theopenweights.com/news/embeddinggemma-2-g4n1) — Google DeepMind · Embeddings · Gemma · Sep 14, 2026: Google DeepMind releases EmbeddingGemma 2, a sub-1B multimodal embedding model covering text, images, audio, and video under the Gemma license. - [Qwen-Image-2.1 Adds RGBA to Image Generation](https://theopenweights.com/news/qwen-image-2-1-haxk) — Qwen · Alibaba · Text → Image · Other · Sep 14, 2026: Qwen-Image-2.1 is an open image generation and editing model from Alibaba's Qwen team, now with RGBA transparency support. - [Yandex releases 80B MoE base model under Apache 2.0](https://theopenweights.com/news/aliceai-foundation-80b-a3b-h8qu) — Yandex · Text / LLM · Apache 2.0 · Sep 12, 2026: Yandex's AliceAI-Foundation-80B-A3B-Base is an Apache-2.0 MoE foundation model with 3B active params and Russian/English support. - [StepFun's StepAudio 3 Realtime targets live voice AI](https://theopenweights.com/news/stepaudio-3-realtime-54sk) — StepFun · Text → Speech · Other · Sep 11, 2026: StepFun releases StepAudio 3 Realtime, an audio-language foundation model for realtime spoken interaction with a listen-converse-think-act loop. - [Agnes-3.0-Flash arrives as a multimodal reasoning model](https://theopenweights.com/news/agnes-3-0-flash-z86d) — Agnes AI · Vision-Language · Other · Sep 11, 2026: Agnes-3.0-Flash is a multimodal reasoning model with hybrid attention and long-context support, now on Hugging Face. - [InternLM's Atria Dawn Preview Targets Agentic Tasks](https://theopenweights.com/news/atria-dawn-tjby) — Internlm · Reasoning · MIT · Sep 11, 2026: InternLM releases Atria Dawn Preview, a MoE agentic reasoning model trained on verified tool interactions, under MIT. - [ZGCM-1 arrives as a fully open 7B reasoning model](https://theopenweights.com/news/zgcm-1-dqzh) — Unknown · Reasoning · Other · Sep 10, 2026: ZGCM-1 is a fully open 7B foundation model built for math reasoning and agentic search with tool use. - [StepFun's StepAudio 3 Music Plans Before It Plays](https://theopenweights.com/news/stepaudio-3-music-pxsm) — StepFun · Music · Other · Sep 10, 2026: StepFun's StepAudio 3 Music generates long-form tracks using explicit musical planning and text-controlled generation. - [StepFun's StepAudio 3 Gen Unifies TTS and Music](https://theopenweights.com/news/stepaudio-3-gen-7rfm) — StepFun · Text → Speech · Other · Sep 10, 2026: StepFun's StepAudio 3 Gen is a discrete autoregressive model unifying text-to-speech, voice design, sound effects, and music in one system. - [Yandex releases a sparse T5 model with 0.6B active params](https://theopenweights.com/news/aliceai-t5-35b-a0-6b-sfeo) — Yandex · Text / LLM · Other · Sep 10, 2026: Yandex's AliceAI-T5-35B-A0.6B is an encoder-decoder MoE model with 35B total and 0.6B active parameters, trained with a UL2 objective. - [NetEase Youdao debuts Confucius4-R2T2 streaming ASR](https://theopenweights.com/news/confucius4-r2t2-gsok) — Netease Youdao · Speech → Text · Other · Sep 10, 2026: NetEase Youdao's Confucius4-R2T2 is a streaming, low-latency multilingual ASR model with vLLM support, released on Hugging Face. - [OpenMOSS Unveils YuE2-3B Music Generation Model](https://theopenweights.com/news/yue2-3b-m3dl) — Mrfakename · Music · CC BY-NC 4.0 · Sep 10, 2026: OpenMOSS releases YuE2-3B, a 3B music model with symbolic planning and agentic editing, supporting Chinese and English under a non-commercial license. - [Tencent's T1 Targets Long-Horizon Terminal Work](https://theopenweights.com/news/t1-terminal-agent-qet4) — Tencent · Reasoning · Other · Sep 9, 2026: Tencent's T1 is a 122B MoE model trained with RL for long-horizon terminal tasks, reporting SOTA on Terminal-Bench. - [SenseTime's SenseNova-U1.5 Unifies Vision Tasks](https://theopenweights.com/news/sensenova-u1-5-s9kc) — SenseTime · Any-to-Any · Other · Sep 9, 2026: SenseTime's SenseNova-U1.5 is an 8B encoder-free, VAE-free unified multimodal model for understanding, reasoning, and image generation. - [OpenMOSS releases YuE2-3B for music generation](https://theopenweights.com/news/yue2-3b-e3s3) — M A P · Music · CC BY-NC 4.0 · Sep 9, 2026: OpenMOSS's YuE2-3B is a 3B music generation model with symbolic planning and agentic editing, released under CC-BY-NC-4.0 for English and Chinese. - [LLaDA-UI Brings Diffusion Decoding to GUI Agents](https://theopenweights.com/news/llada-ui-t7r4) — inclusionAI · Vision-Language · Other · Sep 8, 2026: inclusionAI unveils LLaDA-UI, a 16.7B MoE diffusion vision-language GUI agent using block-parallel decoding. - [Edge0's 35B MoE Aims for SSD-Backed Edge Inference](https://theopenweights.com/news/edge0-35b-a3b-vzm8) — Edge0 · Text / LLM · Apache 2.0 · Sep 8, 2026: Edge0-35B-A3B-preview is an Apache-2.0 MoE with 3B active params, tuned for SSD-offloaded edge inference via trained routing prediction. - [Nex-N2.5-Pro arrives as an Apache-2.0 MoE vision model](https://theopenweights.com/news/nex-n2-5-pro-466u) — Nex Agi · Vision-Language · Apache 2.0 · Sep 8, 2026: Nex-N2.5-Pro is an Apache-2.0 mixture-of-experts vision-language model built on a qwen3_5_moe architecture. - [Nex AGI releases Nex-N2.5-mini, an open MoE multimodal model](https://theopenweights.com/news/nex-n2-5-mini-xwbr) — Nex Agi · Text / LLM · Apache 2.0 · Sep 8, 2026: Nex AGI's Nex-N2.5-mini is an Apache-2.0 mixture-of-experts model that handles both text and vision tasks. - [OUI-1: A Gemma-based diffusion model for generative UI](https://theopenweights.com/news/oui-1-5lcn) — Thesysdev · Text / LLM · Gemma · Sep 7, 2026: OUI-1 is a Gemma-based diffusion language model for generative UI tasks, released under the Gemma license on Hugging Face. - [OpenBMB's MiniCPM5-2B targets on-device AI](https://theopenweights.com/news/minicpm5-2b-30zi) — OpenBMB · Text / LLM · Other · Sep 6, 2026: OpenBMB releases MiniCPM5-2B, a 2B-parameter on-device LLM with long-context and tool-calling support. - [NeoHorse-1-9B targets agentic coding tasks](https://theopenweights.com/news/neohorse-1-9b-523c) — TokenRhythm · Text / LLM · Other · Sep 5, 2026: NeoHorse-1-9B is a 9B dense model from inclusionAI focused on agentic tool-use and coding, built on Qwen3.5. - [NeoHorse-1-4B tunes Qwen3.5 for agentic work](https://theopenweights.com/news/neohorse-1-4b-2c8n) — TokenRhythm · Text / LLM · Other · Sep 5, 2026: NeoHorse-1-4B is a 4B agentic model fine-tuned from Qwen3.5-4B for tool use, coding, and reasoning. - [inclusionAI Adds Vision to Ling 3.0 Flash](https://theopenweights.com/news/ling-3-0-flash-vl-eb2x) — inclusionAI · Vision-Language · MIT · Sep 4, 2026: inclusionAI releases Ling-3.0-flash-VL, an MIT-licensed MoE vision-language model built on the bailing_moe_v3_vl architecture. - [Taichu 5.0 Brings a 9B Vision-Language Model to Hugging Face](https://theopenweights.com/news/zdtaichu5-0-9b-39lz) — TaichuAI · Vision-Language · Other · Sep 4, 2026: SenseTime's ZDTaichu5.0-9B is a 9B vision-language model with spatial reasoning, video understanding, and agent skills. - [Occamy-1.0 targets long-horizon agent work at 35B](https://theopenweights.com/news/occamy-1-0-b5gt) — Unknown · Text / LLM · Other · Sep 3, 2026: Occamy-1.0 is an open 35B dense model built for long-horizon, multi-step agent tasks and reasoning-heavy collaboration. - [NeoMME: a single-tower multilingual multimodal encoder](https://theopenweights.com/news/neomme-mzbr) — OpenBMB · Embeddings · Other · Sep 3, 2026: NeoMME is a single-tower, multimodal-native and multilingual encoder built for efficient document retrieval and embeddings. - [inclusionAI Tunes Ling-3.0-flash for Finance](https://theopenweights.com/news/ling-3-0-flash-fin-rxsn) — inclusionAI · Text / LLM · Other · Sep 3, 2026: inclusionAI releases Ling-3.0-flash-Fin, a mixture-of-experts model tuned for financial research and agentic workflows. - [LLaDA-Image: A Fully Open 6B Image Generator](https://theopenweights.com/news/llada-image-4mh6) — inclusionAI · Text → Image · Other · Sep 2, 2026: inclusionAI's LLaDA-Image is a 6B open image generator with a frozen VLM, full training recipe, and a distilled few-step variant. - [Microsoft's VibeVoice ASR brings streaming speech-to-text](https://theopenweights.com/news/vibevoice-asr-streaming-7b-102r) — Microsoft · Speech → Text · Other · Sep 2, 2026: Microsoft released VibeVoice-ASR-Streaming-7B, a 7B multilingual speech-to-text model built for real-time transcription. - [RWKV7-G1j arrives as a 13.3B attention-free model](https://theopenweights.com/news/rwkv7-g1j-13-3b-qmte) — Fla Hub · Text / LLM · Apache 2.0 · Sep 2, 2026: RWKV7-G1j is a 13.3B-parameter recurrent, attention-free multilingual language model released on Hugging Face under Apache 2.0. - [IFM releases K2-Horizon, a 375B open-weight MoE](https://theopenweights.com/news/k2-horizon-375b-a23b-jg0y) — IFM · Text / LLM · Other · Sep 1, 2026: IFM's K2-Horizon is a 375B-parameter open-weight mixture-of-experts model that activates 23B parameters per token, with a focus on reasoning. - [IFM releases K2-Horizon-7B with open pretraining data](https://theopenweights.com/news/k2-horizon-7b-ofj8) — IFM · Text / LLM · Other · Sep 1, 2026: IFM's K2-Horizon-7B is a dense 7B open-weight LLM released with its pretraining datasets on Hugging Face. - [IFM's K2-Horizon MoVA Ships as a 36B MoE Model](https://theopenweights.com/news/k2-horizon-mova-36b-a4b-dunh) — IFM · Text / LLM · Other · Sep 1, 2026: IFM releases K2-Horizon-MoVA-36B-A4B, an open-weight 36B mixture-of-experts text model that activates 4B parameters per token. - [NVIDIA's Nemotron-3 Brings Streaming Speaker Diarization](https://theopenweights.com/news/nemotron-3-diarization-wjxi) — NVIDIA · Speech → Text · Other · Sep 1, 2026: NVIDIA releases Nemotron-3-Diarization, a streaming speaker diarization model built on the Sortformer architecture. - [Jina AI Releases Jina-OCR-v1 for Document Intelligence](https://theopenweights.com/news/jina-ocr-v1-mgfi) — Jinaai · Vision-Language · Other · Sep 1, 2026: Jina AI's jina-ocr-v1 is a multilingual OCR and document-intelligence VLM built on a DeepSeek-VL backbone, now available on Hugging Face. - [NVIDIA's Kumo Takes Aim at Tabular Prediction](https://theopenweights.com/news/kumo-tabular-eie8) — NVIDIA · Any-to-Any · Other · Sep 1, 2026: NVIDIA releases Kumo Tabular, a foundation model built for structured prediction, claiming a stronger accuracy-efficiency balance. - [Viggle Releases Viggle-Animate for Character Swaps](https://theopenweights.com/news/viggle-animate-p4z6) — Viggle · Image → Video · Other · Aug 31, 2026: Viggle-Animate is an open image-to-video model for character replacement and video editing, distilled from MiniMax-H3. - [DeepSeek adds vision to its V4 Flash line](https://theopenweights.com/news/deepseek-v4-flash-vision-exp-2l4q) — DeepSeek · Vision-Language · MIT · Aug 31, 2026: DeepSeek releases V4-Flash-Vision-Exp, an experimental MoE vision-language model under an MIT license. - [H Company's NeoMME rethinks visual document retrieval](https://theopenweights.com/news/neomme-f44g) — H company · Embeddings · Other · Aug 30, 2026: H Company's NeoMME is a single-tower multimodal, multilingual encoder for visual document retrieval using compact late-interaction embeddings. - [Tencent Previews Hunyuan Hy4, an Apache MoE Model](https://theopenweights.com/news/hunyuan-hy4-aazc) — Tencent · Text / LLM · Apache 2.0 · Aug 27, 2026: Tencent has published Hy4-preview, an early look at its v4 Hunyuan MoE language model, released under Apache 2.0. - [Qwen-Drive 1.0 targets autonomous driving with a 4B VLM](https://theopenweights.com/news/qwen-drive-1-0-4b-9prw) — Qwen · Alibaba · Vision-Language · Other · Aug 27, 2026: Qwen releases Qwen-Drive 1.0, a 4B vision-language model for autonomous driving perception and motion planning. - [Breeze-TTS-2 Brings Open Voice Cloning to English](https://theopenweights.com/news/breeze-tts-2-dicn) — BreezeBlue · Text → Speech · Other · Aug 25, 2026: BreezeBlue releases Breeze-TTS-2, an open English text-to-speech model with voice cloning and directional control. - [IBM's Granite 4.2 Adds Reasoning to Open LLM Line](https://theopenweights.com/news/granite-4-2-k1r1) — IBM · Text / LLM · Apache 2.0 · Aug 25, 2026: IBM has released Granite 4.2, an update to its open, Apache 2.0 licensed LLM family with a new focus on reasoning-oriented workloads. - [Zhipu releases GLM-5.3-Flash under MIT license](https://theopenweights.com/news/glm-5-3-flash-82hx) — Zhipu AI · Text / LLM · MIT · Aug 25, 2026: Zhipu AI's GLM-5.3-Flash is a fast MoE variant of GLM-5.3, released with open MIT weights and multimodal, reasoning capabilities. - [Tencent's WeMM-Embedding-9B Unifies Text, Image and Video](https://theopenweights.com/news/wemm-embedding-9b-nrln) — Tencent · Embeddings · Other · Aug 25, 2026: Tencent's WeMM-Embedding-9B is a 9B multimodal embedding model aligning text, image, and video in a shared space. - [NVIDIA's PhoneLLM Targets Voice Agents on the Line](https://theopenweights.com/news/phonellm-alpha-1-mp8w) — Pipecat Ai · Text / LLM · Other · Aug 24, 2026: PhoneLLM alpha-1 is a Nemotron-H MoE model built for voice agents, with tool-use and function-calling aimed at phone applications. - [XHToken releases Spark-X2.5, a 4B open LLM](https://theopenweights.com/news/spark-x2-5-4b-abea) — XHToken · Text / LLM · Apache 2.0 · Aug 24, 2026: XHToken's Spark-X2.5-4B is an Apache-2.0, instruction-tuned 4B text model now available on Hugging Face. - [Zhipu's GLM-5.3 Targets Coding at a Fraction of the Cost](https://theopenweights.com/news/glm-5-3-4zfa) — Zhipu AI · Text / LLM · MIT · Aug 23, 2026: Zhipu AI's open-weight GLM-5.3 is a mixture-of-experts model pitched as competitive on coding at roughly a fifth the cost of closed rivals. - [Liquid AI's LFM2.5-DSpark targets faster inference](https://theopenweights.com/news/lfm2-5-dspark-suu4) — Unknown · Text / LLM · Other · Aug 20, 2026: Liquid AI's LFM2.5-DSpark is a text model tuned for speed, claiming up to 3.2x faster inference. - [Liquid AI's LFM2.5-DSpark targets faster inference](https://theopenweights.com/news/lfm2-5-dspark-wrx1) — NVIDIA · Text / LLM · Other · Aug 20, 2026: Liquid AI's LFM2.5-DSpark is an efficient text model promising up to 3.2x faster inference for on-device and cost-sensitive use. - [SenseTime Releases SenseNova U1.5 8B Any-to-Any Model](https://theopenweights.com/news/sensenova-u1-5-8b-mot-9429) — SenseTime · Any-to-Any · Other · Aug 19, 2026: SenseTime's SenseNova U1.5 8B is a native any-to-any multimodal model with built-in image generation and editing. - [Audio8 debuts a compact 0.1B preview TTS model](https://theopenweights.com/news/audio8-tts-0-1b-z891) — Audio8 · Text → Speech · Other · Aug 19, 2026: Audio8-TTS-Preview-0.1b is a 0.1B-parameter text-to-speech model with zero-shot voice cloning, released as an early preview on Hugging Face. - [Thomson Reuters enters the model race with Thomson-1.0](https://theopenweights.com/news/thomson-1-0-small-i9ut) — Thomsonreuters · Text / LLM · Other · Aug 18, 2026: Thomson Reuters debuts Thomson-1.0-Small, an MoE language model trained on its proprietary data assets, on Hugging Face. - [Tencent's AuK Bundles Voice Cloning and Speech Editing](https://theopenweights.com/news/auk-r1aw) — Tencent · Text → Speech · Other · Aug 18, 2026: Tencent releases AuK, a zero-shot text-to-speech model with voice cloning, speech editing, enhancement and separation. - [Ornith 1.5 arrives as a 397B MoE multimodal model](https://theopenweights.com/news/ornith-1-5-397b-lskb) — Ornith Ai · Text / LLM · MIT · Aug 18, 2026: Ornith 1.5 launches as a 397B MoE vision-language model under MIT, with 9B and 35B-A3B siblings for smaller footprints. - [Ornith 1.5 Brings a Lean 35B MoE to Open Weights](https://theopenweights.com/news/ornith-1-5-35b-a3b-w6je) — Ornith Ai · Text / LLM · MIT · Aug 18, 2026: Ornith-1.5-35B-A3B is an MIT-licensed 35B mixture-of-experts text model with 3B active parameters, built on the Qwen3.5 MoE architecture. - [Ant Research releases 4DAnyone for 4D human video](https://theopenweights.com/news/4danyone-bi56) — AntResearch · Image → Video · Other · Aug 17, 2026: Ant Group's 4DAnyone brings multiview video generation and 4D human reconstruction to Hugging Face under a custom license. - [Apodex 1.1 mini targets long-horizon agentic work](https://theopenweights.com/news/apodex-1-1-mini-nb79) — Apodex · Text / LLM · Other · Aug 17, 2026: Apodex 1.1 mini is an agentic mixture-of-experts model built on Qwen3.5-35B-A3B, tuned for long-horizon complex work. - [OpenMOSS Debuts MOSS-VL for Real-Time Vision Interaction](https://theopenweights.com/news/moss-vl-98so) — OpenMOSS · Vision-Language · Other · Aug 14, 2026: OpenMOSS introduces MOSS-VL, a vision-language model family designed for real-time streaming interaction using gated cross-attention. - [Fastino releases GLiNER 2.5 Multi for extraction](https://theopenweights.com/news/gliner-2-5-multi-9k1l) — Fastino · Text / LLM · Other · Aug 14, 2026: Fastino's GLiNER 2.5 Multi handles multilingual NER, relation extraction, classification and JSON extraction in one open model. - [TeleOCR brings bilingual document parsing to Qwen2.5-VL](https://theopenweights.com/news/teleocr-w084) — XingChen AGI · Vision-Language · Other · Aug 14, 2026: TeleOCR is an open document-parsing and OCR vision-language model built on Qwen2.5-VL for Chinese and English text. - [StarDoc-AI Releases TeleOCR for Document Parsing](https://theopenweights.com/news/teleocr-huln) — StarDoc AI · Vision-Language · Other · Aug 14, 2026: StarDoc-AI's TeleOCR is a Qwen2.5-VL-based vision-language model for parsing Chinese and English documents. - [Tencent's UI-Mate-27B targets desktop automation](https://theopenweights.com/news/ui-mate-27b-6n3c) — Tencent · Vision-Language · Other · Aug 14, 2026: Tencent releases UI-Mate-27B, a 27B vision-language GUI agent for desktop automation, on Hugging Face. - [DeepSeek Releases V4-Pro-0813 With Open Weights](https://theopenweights.com/news/deepseek-v4-pro-0813-7o26) — DeepSeek · Text / LLM · MIT · Aug 13, 2026: DeepSeek has published DeepSeek-V4-Pro-0813, a mixture-of-experts text and reasoning model, with open weights under an MIT license. - [DeepSeek Releases V4-Pro, an MIT-Licensed MoE Model](https://theopenweights.com/news/deepseek-v4-pro-0813-leus) — DeepSeek · Text / LLM · MIT · Aug 13, 2026: DeepSeek's V4-Pro is a mixture-of-experts model tuned for reasoning and code, released under the permissive MIT license. - [Superwhisper's s1-mini polishes raw speech-to-text output](https://theopenweights.com/news/s1-mini-6bye) — Superwhisper · Speech → Text · Other · Aug 12, 2026: Superwhisper releases s1-mini, a Qwen3-based model for normalizing, punctuating and truecasing ASR output. - [Liquid AI's LFM2.5-VL-3B targets on-device vision](https://theopenweights.com/news/lfm2-5-vl-3b-5uex) — LiquidAI · Vision-Language · Other · Aug 12, 2026: Liquid AI released LFM2.5-VL-3B, a 3B-parameter vision-language model built for faster on-device multimodal tasks. - [NVIDIA Opens Magpie TTS for Multilingual Voice Agents](https://theopenweights.com/news/nvidia-magpie-tts-2enb) — NVIDIA · Text → Speech · Other · Aug 10, 2026: NVIDIA releases Magpie TTS, an open-weights multilingual text-to-speech model built for low-latency voice agents and full deployment control. - [NVIDIA Opens Magpie TTS for Multilingual Voice Agents](https://theopenweights.com/news/magpie-tts-5vid) — NVIDIA · Text → Speech · Other · Aug 10, 2026: NVIDIA releases Magpie TTS, an open-weights multilingual text-to-speech model built for low-latency voice agents. - [NVIDIA opens Magpie TTS for multilingual voice agents](https://theopenweights.com/news/magpie-tts-multilingual-o1ie) — NVIDIA · Text → Speech · Other · Aug 10, 2026: NVIDIA has released Magpie TTS, an open-weights, low-latency multilingual text-to-speech model built for real-time voice agents. - [Cohere Labs releases compact North Micro Vision model](https://theopenweights.com/news/north-micro-vision-ppiy) — Cohere · Vision-Language · CC BY-NC 4.0 · Aug 10, 2026: Cohere Labs debuts North Micro Vision Instruct, a compact multilingual vision-language model released under CC-BY-NC-4.0. - [Vak Conformer targets speech recognition in six Indic languages](https://theopenweights.com/news/vak-conformer-wmpt) — Shunyalabs · Speech → Text · Other · Aug 10, 2026: Shunya Labs' Vak Conformer is a Conformer-based speech recognition model covering six Indic languages, now on Hugging Face. - [Bilibili's IndexTTS-2.5 Refines Zero-Shot Voice Cloning](https://theopenweights.com/news/indextts-2-5-s7an) — IndexTeam · Text → Speech · Other · Aug 10, 2026: Bilibili releases IndexTTS-2.5, a zero-shot voice-cloning TTS model with emotion control and cross-lingual support for zh, en, and ja. - [inclusionAI Releases Ling-3.0-tiny, an MIT-Licensed MoE](https://theopenweights.com/news/ling-3-0-tiny-gl7r) — inclusionAI · Text / LLM · MIT · Aug 10, 2026: inclusionAI's Ling-3.0-tiny is a hybrid mixture-of-experts text model released on Hugging Face under the permissive MIT license. - [Meta's Muse Glimmer 30B Targets Local Agentic Coding](https://theopenweights.com/news/muse-glimmer-30b-3d3h) — Meta AI · Vision-Language · Apache 2.0 · Aug 10, 2026: Meta releases Muse Glimmer 30B, an Apache-2.0 multimodal coding model built for local, agentic use. - [dots3-note preview brings audio and vision to one model](https://theopenweights.com/news/dots3-note-k0oo) — Dots Studio · Any-to-Any · Other · Aug 9, 2026: dots-studio releases dots3-note preview, a multimodal long-context agentic model with audio and vision support, on Hugging Face. - [Qwen releases 2.4T-parameter open MoE with 95B active](https://theopenweights.com/news/qwen3-8-2-4t-a95b-ohbw) — Qwen · Alibaba · Text / LLM · Other · Aug 8, 2026: Qwen3.8-2.4T-A95B is Qwen's largest open Mixture-of-Experts model, with 2.4T total and 95B active parameters per token. - [IBM's Granite 4.2 30B adds reasoning and tool calls](https://theopenweights.com/news/granite-4-2-30b-9jv3) — IBM · Text / LLM · Apache 2.0 · Aug 7, 2026: IBM releases Granite 4.2 30B, a dense LLM with reasoning and tool-calling support, under Apache-2.0. - [MiniMax Opens Its Music3 Text-to-Music Model](https://theopenweights.com/news/minimax-music3-cu1a) — MiniMax · Music · Other · Aug 7, 2026: MiniMax has released Music3, an open text-to-music generation model, on Hugging Face. - [Motif 3 brings a sparse MoE approach to reasoning](https://theopenweights.com/news/motif-3-a5d4) — Motif Technologies · Text / LLM · Other · Aug 7, 2026: Motif Technologies released Motif 3, a sparse MoE language model using grouped differential latent attention for reasoning and long context. - [MameLoshnLM Brings Yiddish Into the Open LLM Era](https://theopenweights.com/news/mameloshnlm-tr65) — OpenMOSS · Text / LLM · Other · Aug 5, 2026: OpenMOSS's MameLoshnLM is an open 8B model built for Yiddish, released alongside a dedicated evaluation benchmark. - [Qwen3.8-27B Brings Vision to a Dense Model](https://theopenweights.com/news/qwen3-8-27b-zi22) — Qwen · Alibaba · Vision-Language · Apache 2.0 · Aug 5, 2026: Qwen3.8-27B is a dense 27B-parameter vision-language model with image-text-to-text and chat support, released under Apache-2.0. - [Deepgrove's Maple Preview bets on ternary-weight MoE](https://theopenweights.com/news/maple-5v71) — Deepgrove · Reasoning · MIT · Aug 4, 2026: Deepgrove's Maple Preview is a ternary-weight MoE reasoning model released under an MIT license on Hugging Face. - [NVIDIA's Nemotron 3.5 Lightning trims MoE for speed](https://theopenweights.com/news/nemotron-3-5-lightning-30b-a3b-i2lm) — NVIDIA · Text / LLM · Other · Aug 4, 2026: NVIDIA releases Nemotron 3.5 Lightning, a 30B MoE with 3B active params, hybrid Mamba architecture, and NVFP4 quantization for efficient reasoning. - [Mistral debuts Shieldstral, a 3B safety model](https://theopenweights.com/news/shieldstral-9bzq) — Mistral AI · Vision-Language · Apache 2.0 · Aug 4, 2026: Mistral AI releases Shieldstral, a 3B open-weights multimodal moderation model for text and image safety, under Apache 2.0. - [inclusionAI ships Ling-3.0-flash, an MIT-licensed MoE model](https://theopenweights.com/news/ling-3-0-flash-ta67) — inclusionAI · Text / LLM · MIT · Aug 2, 2026: inclusionAI released Ling-3.0-flash, an MIT-licensed hybrid mixture-of-experts text model in the Bailing series, now available on Hugging Face. - [Liquid AI ships LFM2.5, a 2.6B on-device model](https://theopenweights.com/news/lfm2-5-2-6b-x6qp) — LiquidAI · Embeddings · Other · Aug 1, 2026: Liquid AI's LFM2.5-2.6B is a compact multilingual language model released in GGUF format for efficient on-device inference. - [NVIDIA's Nemotron 3.5 Lightning Blends Mamba and MoE](https://theopenweights.com/news/nemotron-3-5-lightning-30b-a3b-ui67) — NVIDIA · Text / LLM · Other · Aug 1, 2026: NVIDIA releases Nemotron 3.5 Lightning, a hybrid Mamba-MoE model with 3B active params tuned for agentic coding. - [Meituan Ships a Lighter, Sparser LongCat-Flash](https://theopenweights.com/news/longcat-flash-lite-sparse-guvc) — Meituan · Text / LLM · MIT · Jul 31, 2026: Meituan releases LongCat-Flash-Lite-Sparse, a lighter MoE language model available on Hugging Face under an MIT license. - [DeepSeek Ships V4-Flash, a 304B MoE Tuned for Agents](https://theopenweights.com/news/deepseek-v4-flash-0731-7oaz) — DeepSeek · Text / LLM · MIT · Jul 31, 2026: DeepSeek's V4-Flash-0731 is a 304B-parameter mixture-of-experts model with stronger agentic capabilities, released under MIT. - [DeepSeek Refreshes V4-Flash With New 0731 Checkpoint](https://theopenweights.com/news/deepseek-v4-flash-0731-s1rd) — DeepSeek · Text / LLM · MIT · Jul 31, 2026: DeepSeek released a refreshed 0731 checkpoint of its V4-Flash MoE language model under an MIT license with FP8/8-bit weights. - [Kroma: An MIT-Licensed Text-to-Image Model for ComfyUI](https://theopenweights.com/news/kroma-1hlg) — Lodestones · Text → Image · MIT · Jul 31, 2026: Lodestones releases Kroma 1.0, an MIT-licensed text-to-image model derived from the Krea 2 lineage and tuned for ComfyUI. - [NVIDIA's Nemotron VoiceChat 11B Targets Spoken AI](https://theopenweights.com/news/nemotron-voicechat-11b-i4zj) — NVIDIA · Text → Speech · Other · Jul 29, 2026: NVIDIA releases Nemotron VoiceChat 11B, an English voice conversation model built on Nemotron Nano 9B v2, on Hugging Face. - [Needle2 Packs an Agentic LLM Into 14MB](https://theopenweights.com/news/needle2-j0jc) — Cactus Compute · Text / LLM · Apache 2.0 · Jul 29, 2026: Cactus Compute's Needle2 is a 14MB agentic LLM with tool/function calling, built for phones, wearables, and robots. - [LG AI Research debuts K-EXAONE 2.0, a 750B MoE model](https://theopenweights.com/news/k-exaone-2-0-750b-a37b-rhz2) — LGAI EXAONE · Text / LLM · Other · Jul 29, 2026: LG AI Research released K-EXAONE 2.0, a 750B-parameter MoE model with 37B active params, spanning English, Korean, and Spanish. - [Liquid AI's LFM2.5-2.6B targets on-device agents](https://theopenweights.com/news/lfm2-5-2-6b-dps4) — LiquidAI · Embeddings · Other · Jul 28, 2026: Liquid AI releases LFM2.5-2.6B, a compact multilingual LLM designed to power local agents and edge deployments. - [MiniMax Releases H3 Video Model on Hugging Face](https://theopenweights.com/news/minimax-h3-5ssh) — MiniMax · Text → Video · Other · Jul 28, 2026: MiniMax has released H3, a diffusion model for text- and image-to-video generation that also supports synchronized audio-video output. - [SK Telecom Releases A.X-K2 Multilingual LLM](https://theopenweights.com/news/a-x-k2-oei0) — Skt · Text / LLM · Apache 2.0 · Jul 28, 2026: SK Telecom's A.X-K2 is a multilingual text model spanning EN, KO, ZH, JA, and ES, released on Hugging Face under Apache-2.0. - [SenseTime Debuts SenseNova U1.5 8B Multimodal Preview](https://theopenweights.com/news/sensenova-u1-5-8b-mot-n5qz) — SenseTime · Any-to-Any · Other · Jul 28, 2026: SenseTime's SenseNova U1.5 8B MoT is an any-to-any multimodal preview with 4K image generation and editing. - [Audio8 debuts a 0.6B multilingual zero-shot TTS preview](https://theopenweights.com/news/audio8-tts-0-6b-rtx1) — Audio8 · Text → Speech · Other · Jul 28, 2026: Audio8-TTS-Preview-0.6b is a 600M-parameter multilingual text-to-speech model with zero-shot voice cloning, now on Hugging Face. - [Thinking Machines Debuts Inkling Small, a Compact Multimodal MoE](https://theopenweights.com/news/inkling-small-den7) — Thinkingmachines · Vision-Language · Apache 2.0 · Jul 27, 2026: Thinking Machines released Inkling Small, an Apache-2.0 multimodal mixture-of-experts model handling image, audio, and text. - [Liquid AI's LFM2.5 encoder targets fast CPU inference](https://theopenweights.com/news/lfm2-5-encoder-230m-75r2) — LiquidAI · Embeddings · Other · Jul 27, 2026: Liquid AI's LFM2.5-Encoder-230M is a compact bidirectional encoder for fast, long-context embeddings on CPU, covering English and German. - [Liquid AI ships a 350M encoder built for CPUs](https://theopenweights.com/news/lfm2-5-encoder-350m-t23r) — LiquidAI · Embeddings · Apache 2.0 · Jul 27, 2026: Liquid AI's LFM2.5-Encoder-350M is a compact bidirectional encoder for fast long-context inference on CPU, under Apache 2.0. - [KRAFTON releases A.X-K2 Raon speech MoE model](https://theopenweights.com/news/a-x-k2-raon-speech-21b-a3b-9rct) — KRAFTON · Any-to-Any · Other · Jul 27, 2026: KRAFTON's A.X-K2 Raon Speech is a 21B/3B-active MoE model unifying TTS and ASR in one any-to-any architecture. - [Microsoft's Mage-VL Streams Video Natively](https://theopenweights.com/news/mage-vl-jlzt) — Microsoft · Vision-Language · Other · Jul 26, 2026: Microsoft releases Mage-VL, a codec-native streaming multimodal model built for real-time video and vision-language understanding. - [Apertus v1.5 70B arrives with an Apache-2.0 license](https://theopenweights.com/news/apertus-v1-5-70b-kmu1) — Swiss Ai · Text / LLM · Apache 2.0 · Jul 24, 2026: Swiss AI's Apertus v1.5 70B is a multilingual, multimodal open model released under a permissive Apache-2.0 license. - [Microsoft's VibeVoice ASR Goes BitNet for CPU Speech](https://theopenweights.com/news/vibevoice-asr-bitnet-pz17) — Microsoft · Speech → Text · Other · Jul 24, 2026: Microsoft releases VibeVoice ASR BitNet, a multilingual speech-to-text model quantized for efficient CPU inference in English and Chinese. - [AMD's Instella-MoE Brings Reasoning to ROCm Hardware](https://theopenweights.com/news/instella-moe-16b-a3b-think-ygvm) — Amd · Reasoning · Other · Jul 23, 2026: AMD releases Instella-MoE 16B A3B Think, an open reasoning model with 3B active params, optimized for ROCm. - [Kwaipilot Releases KAT-Coder V2.5 Dev, an Agentic MoE Coder](https://theopenweights.com/news/kat-coder-v2-5-dev-tv94) — Kwaipilot · Code · Other · Jul 23, 2026: Kwaipilot's KAT-Coder V2.5 Dev is an agentic-coding mixture-of-experts model built on Qwen3.5, now available on Hugging Face. - [Lightricks Releases LTX-2.5 Video Model](https://theopenweights.com/news/ltx-2-5-t9ub) — Lightricks · Image → Video · Other · Jul 23, 2026: Lightricks' LTX-2.5 open video model supports text-to-video and image-to-video generation on Hugging Face. - [Upstage's Solar Open2 arrives as a 250B MoE model](https://theopenweights.com/news/solar-open2-250b-fhuj) — Upstage · Text / LLM · Other · Jul 22, 2026: Upstage releases Solar Open2, a 250B mixture-of-experts LLM for English and Korean, published openly on Hugging Face. - [Microsoft's Mage-Flow packs image editing into 4B](https://theopenweights.com/news/mage-flow-c5wr) — Microsoft · Text → Image · MIT · Jul 21, 2026: Microsoft's Mage-Flow is a 4B open model for text-to-image generation and instruction-based editing, released under MIT. - [Nanbeige 4.2 arrives as a compact 3B bilingual model](https://theopenweights.com/news/nanbeige4-2-3b-8a6o) — Nanbeige · Text / LLM · Other · Jul 21, 2026: Nanbeige releases 4.2-3B, a 3-billion-parameter bilingual EN/ZH text model now available on Hugging Face. - [Motif Technologies debuts Motif 3 Beta, an MoE model](https://theopenweights.com/news/motif-3-ybaz) — Motif Technologies · Text / LLM · Other · Jul 20, 2026: Motif Technologies has published Motif 3 Beta, a preview mixture-of-experts LLM aimed at long-context, multilingual tasks. - [Microsoft's Fara1.5-27B targets computer-use agents](https://theopenweights.com/news/fara1-5-27b-ni3f) — Microsoft · Vision-Language · Other · Jul 17, 2026: Microsoft releases Fara1.5-27B, a 27B vision-language agent for browser and desktop automation, on Hugging Face. - [NVIDIA's Audio-Visual Flamingo Fuses Sound and Sight](https://theopenweights.com/news/audio-visual-flamingo-ghkj) — NVIDIA · Any-to-Any · Other · Jul 16, 2026: NVIDIA's Audio-Visual Flamingo is an open audio-visual LLM built for joint reasoning over sound, images, and long videos. - [German consortium releases open 30B model Soofi S](https://theopenweights.com/news/soofi-s-izq4) — Unknown · Text / LLM · Other · Jul 16, 2026: Soofi S is an open 30B dense model from a German AI consortium, reported to top benchmarks in both English and German. - [NVIDIA's Nemotron-3-Embed 8B tops RTEB retrieval test](https://theopenweights.com/news/nemotron-3-embed-8b-qagw) — NVIDIA · Embeddings · Other · Jul 16, 2026: NVIDIA's Nemotron-3-Embed 8B text-embedding model ranked #1 overall on the RTEB retrieval benchmark, aiming at agentic search. - [InternLM Previews 397B Vision-Language Model](https://theopenweights.com/news/intern-s2-397b-2934) — Internlm · Vision-Language · Apache 2.0 · Jul 16, 2026: InternLM has released Intern-S2-Preview, a 397B-parameter vision-language model, on Hugging Face under Apache-2.0. - [Mistral's Shieldstral brings safety checks to images](https://theopenweights.com/news/shieldstral-1-0-3b-5n2j) — Mistral AI · Vision-Language · Apache 2.0 · Jul 16, 2026: Mistral released Shieldstral 1.0 3B, an open-weights multimodal safety classifier under Apache 2.0. - [inclusionAI ships LLaDA2.2-flash diffusion LLM](https://theopenweights.com/news/llada2-2-flash-p1lk) — inclusionAI · Text / LLM · Apache 2.0 · Jul 16, 2026: inclusionAI's LLaDA2.2-flash is an Apache-2.0, diffusion-based mixture-of-experts language model now on Hugging Face. - [Thinking Machines ships Inkling, its first open model](https://theopenweights.com/news/inkling-ttu9) — OpenAI · Any-to-Any · Other · Jul 15, 2026: Thinking Machines has released Inkling, its first open-weights language model, focused on text and reasoning. - [Thinking Machines Lab debuts Inkling, its first open model](https://theopenweights.com/news/inkling-wh8c) — Thinkingmachines · Any-to-Any · Apache 2.0 · Jul 15, 2026: Inkling is Thinking Machines Lab's first open-weights model: an Apache 2.0 mixture-of-experts system handling image and audio inputs. - [CrisperWhisper 2.0 Large targets verbatim transcription](https://theopenweights.com/news/crisperwhisper-2-0-large-gpmf) — Nyralabs · Speech → Text · Other · Jul 15, 2026: CrisperWhisper 2.0 Large is a Whisper-based verbatim ASR model with disfluency handling and word-level timestamps for English and German. - [Alibaba's Wan2.2-Animate-2 14B lands under Apache 2.0](https://theopenweights.com/news/wan2-2-animate-2-14b-gotn) — Qwen · Alibaba · Image → Video · Apache 2.0 · Jul 14, 2026: Wan2.2-Animate-2 is a 14B, Apache-licensed image-to-video and text-to-video model from Alibaba's Wan team, now on Hugging Face. - [NVIDIA's Nemotron 3 Embed tops the RTEB leaderboard](https://theopenweights.com/news/nemotron-3-embed-1b-dsbz) — NVIDIA · Embeddings · Other · Jul 14, 2026: NVIDIA's Nemotron 3 Embed 1B, a text embedding model, ranks #1 overall on the RTEB retrieval benchmark. - [OpenMOSS Debuts MOSS-VL-Realtime for Live Video](https://theopenweights.com/news/moss-vl-realtime-jevs) — OpenMOSS · Vision-Language · Other · Jul 14, 2026: OpenMOSS releases MOSS-VL-Realtime, a streaming vision-language model built for real-time video and image understanding. - [SberDevices releases GigaAM Multilingual ASR model](https://theopenweights.com/news/gigaam-multilingual-hulu) — Ai Sage · Speech → Text · MIT · Jul 14, 2026: SberDevices' GigaAM Multilingual is an open, MIT-licensed ASR model covering Russian, English, and Kazakh speech. - [Boogu-Image-0.1 Brings Unified Multimodal to Open Source](https://theopenweights.com/news/boogu-image-0-1-h0m4) — Unknown · Any-to-Any · Apache 2.0 · Jul 13, 2026: Boogu-Image-0.1 is an Apache-2.0 unified multimodal model for bilingual text-to-image generation and instruction-based image editing. - [inclusionAI's Ring-Zero Scales Zero-RL to a Trillion Parameters](https://theopenweights.com/news/ring-zero-ru07) — inclusionAI · Reasoning · Apache 2.0 · Jul 13, 2026: inclusionAI's Ring-Zero is a trillion-parameter MoE model that elicits emergent chain-of-thought reasoning via zero-RL, with no human-labeled data. - [Poolside releases Laguna-S-2.1 coding model](https://theopenweights.com/news/laguna-s-2-1-3lt9) — Poolside · Code · Other · Jul 13, 2026: Poolside published Laguna-S-2.1, a code-focused language model, on Hugging Face under the OpenMDW license. - [Alibaba's OvisOCR2 turns page images into Markdown](https://theopenweights.com/news/ovisocr2-8zer) — ATH MaaS · Vision-Language · Apache 2.0 · Jul 13, 2026: Alibaba's OvisOCR2 is a 0.8B end-to-end document parser that converts page images to Markdown, covering text, tables and formulas. - [Qwen Enters Music Generation With Qwen-Music](https://theopenweights.com/news/qwen-music-66sf) — Qwen · Alibaba · Music · Qwen · Jul 12, 2026: Qwen-Music generates high-fidelity songs with vocals from text, lyrics and musical prompts, marking Alibaba's Qwen team's first music model. - [Wan-Dancer-14B turns still images into dance videos](https://theopenweights.com/news/wan-dancer-14b-40fg) — Qwen · Alibaba · Image → Video · Apache 2.0 · Jul 10, 2026: Wan-Dancer-14B is a 14B image-to-video model from Alibaba's Wan team for music-driven dance generation, released under Apache 2.0. - [LingBot-Video puts a 30B MoE behind embodied AI video](https://theopenweights.com/news/lingbot-video-30b-a3b-h8nh) — robbyant · Text → Video · Apache 2.0 · Jul 8, 2026: LingBot-Video 30B-A3B is a DiT-based mixture-of-experts text-to-video model built for embodied intelligence, released under Apache 2.0. - [NVIDIA's Audex Unifies Audio Understanding and Speech](https://theopenweights.com/news/nemotron-labs-audex-30b-a3b-qo1w) — NVIDIA · Any-to-Any · Other · Jul 6, 2026: NVIDIA released Nemotron-Labs-Audex-30B-A3B, a mixture-of-experts model that unifies audio understanding and speech generation with 3B active parameters. - [GigaChat 3.5 arrives as a 432B mixture-of-experts model](https://theopenweights.com/news/gigachat3-5-432b-a28b-67tq) — Ai Sage · Text / LLM · Other · Jul 5, 2026: GigaChat 3.5 is a 432B mixture-of-experts instruct model with 28B active parameters, multilingual coverage, and hybrid attention. - [Bonsai-27B Brings 1-Bit Quantization to Local Inference](https://theopenweights.com/news/bonsai-27b-lf5s) — Prism Ml · Text / LLM · Other · Jul 4, 2026: Bonsai-27B is a 27B parameter LLM using 1-bit/ternary quantization and hybrid attention, distributed in GGUF for on-device inference. - [Tencent releases Hunyuan Hy3 under Apache 2.0](https://theopenweights.com/news/hunyuan-hy3-ome6) — Tencent · Text / LLM · Apache 2.0 · Jul 2, 2026: Tencent's Hunyuan Hy3 is a mixture-of-experts conversational LLM released on Hugging Face under the permissive Apache-2.0 license. - [NVIDIA's Cosmos 3 Edge Brings World Models Closer](https://theopenweights.com/news/cosmos-3-edge-rufq) — NVIDIA · Text → Video · Other · Jul 1, 2026: NVIDIA releases Cosmos 3 Edge, a world-model variant tuned for edge deployment with text-to-video and image-to-video generation. - [NVIDIA distills Qwen-Image for few-step generation](https://theopenweights.com/news/qwen-image-flash-cm77) — NVIDIA · Text → Image · Other · Jul 1, 2026: NVIDIA released Qwen-Image-Flash, a DMD2-distilled text-to-image model that cuts inference to just a few steps for faster generation. - [Google DeepMind's Gemma 4 Goes Multimodal and MoE](https://theopenweights.com/news/gemma-4-l0lj) — Google DeepMind · Any-to-Any · Gemma · Jul 1, 2026: Gemma 4 brings mixture-of-experts and encoder-free multimodal models with a thinking mode and long context, under the Gemma license. - [Mistral's Leanstral 1.5 puts 119B in a lean MoE](https://theopenweights.com/news/leanstral-1-5-119b-a6b-fzf5) — Mistral AI · Text / LLM · Apache 2.0 · Jul 1, 2026: Mistral AI releases Leanstral 1.5, a 119B-parameter MoE model with 6B active per token, under Apache-2.0. - [German Consortium Debuts Soofi S, an Open 30B MoE Model](https://theopenweights.com/news/soofi-s-yo8r) — Soofi Project · Text / LLM · Other · Jul 1, 2026: Soofi S Base is an open 30B-parameter Mamba-2 MoE model from a German consortium, reportedly leading English and German benchmarks. - [Liquid AI's LFM2.5-230M targets phones and robots](https://theopenweights.com/news/lfm2-5-230m-2ck6) — NVIDIA · Text / LLM · Other · Jul 1, 2026: Liquid AI released LFM2.5-230M, a 230M-parameter text model tuned to run locally on phones, single-board computers, and robots. - [Liquid AI's LFM2.5-230M targets phones and robots](https://theopenweights.com/news/lfm2-5-230m-wazo) — Mistral AI · Text / LLM · Other · Jul 1, 2026: Liquid AI's LFM2.5-230M is a 230M-parameter text model optimized for phones, Raspberry Pi and robots. - [Liquid AI's LFM2.5 230M targets phones and robots](https://theopenweights.com/news/lfm2-5-230m-rz5x) — IBM · Text / LLM · Other · Jul 1, 2026: Liquid AI released LFM2.5 230M, a compact text model designed to run on phones, single-board computers, and robots. - [GigaAI Releases Giga-World-1 Under Apache 2.0](https://theopenweights.com/news/giga-world-1-4egn) — Open Gigaai · Image → Video · Apache 2.0 · Jul 1, 2026: GigaAI's Giga-World-1 is an Apache-2.0 image-to-video world model built for physically grounded generation and robot policy learning. - [NVIDIA's Nemotron-Parse 2.0 targets document OCR](https://theopenweights.com/news/nemotron-parse-2-0-a7pt) — NVIDIA · Vision-Language · Other · Jun 30, 2026: NVIDIA released Nemotron-Parse 2.0, a vision-language model for OCR and document parsing, on Hugging Face. - [Microsoft previews GELab-Zero-4B, a compact GUI agent](https://theopenweights.com/news/gelab-zero-4b-f2ea) — Microsoft · Vision-Language · Other · Jun 30, 2026: Microsoft's GELab-Zero-4B is a preview 4B vision-language model for GUI and mobile agent tasks, built on Qwen3-VL. - [MuScriptor Large Turns Real Music Into MIDI](https://theopenweights.com/news/muscriptor-large-82l7) — MuScriptor · Music · CC BY-NC 4.0 · Jun 30, 2026: MuScriptor Large is an open audio-to-MIDI model for multi-instrument transcription of real music mixes, released under CC BY-NC 4.0. - [Meituan releases LongCat-2.0 language model](https://theopenweights.com/news/longcat-2-0-m7er) — Meituan · Text / LLM · Other · Jun 30, 2026: Meituan has published LongCat-2.0, a text-generation language model now available on Hugging Face under a custom license. - [Google DeepMind Releases TabFM for Tabular Data](https://theopenweights.com/news/tabfm-1-0-0-psmt) — Google DeepMind · Any-to-Any · Other · Jun 29, 2026: Google DeepMind's TabFM 1.0.0 is a tabular foundation model for zero-shot in-context classification and regression, now on Hugging Face. - [SenseTime's SenseNova-Vision-7B-MoT Goes Any-to-Any](https://theopenweights.com/news/sensenova-vision-7b-mot-lhsr) — SenseTime · Any-to-Any · Other · Jun 29, 2026: SenseTime releases SenseNova-Vision-7B-MoT, a 7B any-to-any multimodal model spanning vision-language, generation, editing and perception. - [DeepSeek Releases V4-Pro, an MIT-Licensed MoE Model](https://theopenweights.com/news/deepseek-v4-pro-d50q) — DeepSeek · Text / LLM · MIT · Jun 27, 2026: DeepSeek has published DeepSeek-V4-Pro, an MIT-licensed mixture-of-experts model with FP8 weights and reasoning support, on Hugging Face. - [DeepSeek Releases V4-Flash for Low-Latency Inference](https://theopenweights.com/news/deepseek-v4-flash-qi3s) — DeepSeek · Text / LLM · MIT · Jun 27, 2026: DeepSeek-V4-Flash is a faster, lighter mixture-of-experts variant tuned for lower-latency text inference, released under MIT. - [DeepReinforce Releases Ornith 1.0, a 35B Reasoning Model](https://theopenweights.com/news/ornith-1-0-35b-5yql) — Deepreinforce Ai · Text / LLM · MIT · Jun 25, 2026: DeepReinforce AI's Ornith 1.0 is a 35B dense text and reasoning model released in GGUF format under an MIT license. - [Liquid AI's LFM2.5-230M targets on-device language tasks](https://theopenweights.com/news/lfm2-5-230m-us0z) — LiquidAI · Text / LLM · Other · Jun 24, 2026: Liquid AI released LFM2.5-230M, a 230M-parameter multilingual text model designed for fast, low-resource inference on edge devices. - [NVIDIA's Nemotron 3 Puzzle Runs Big on a Lean Budget](https://theopenweights.com/news/nemotron-labs-3-puzzle-75b-a9b-46io) — NVIDIA · Text / LLM · Other · Jun 24, 2026: NVIDIA released Nemotron-Labs-3-Puzzle-75B-A9B, a sparse MoE reasoning model with 75B total and 9B active parameters, on Hugging Face. - [NVIDIA's Nemotron-3 Puzzle Brings a Lean MoE to Reasoning](https://theopenweights.com/news/nemotron-3-puzzle-75b-a9b-ko5i) — NVIDIA · Text / LLM · Other · Jun 24, 2026: NVIDIA releases Nemotron-3 Puzzle, a 75B latent-MoE model with 9B active params, multi-token prediction, and NVFP4 quantization. - [DeepReinforce debuts Ornith-1.0, a 397B MoE model](https://theopenweights.com/news/ornith-1-0-397b-ten1) — Deepreinforce Ai · Text / LLM · MIT · Jun 23, 2026: DeepReinforce released Ornith-1.0-397B, an MIT-licensed mixture-of-experts model with text and reasoning capabilities. - [Tencent's Moebius packs inpainting into 0.2B params](https://theopenweights.com/news/moebius-lelh) — Tencent · Image Editing · Other · Jun 22, 2026: Moebius is a 0.2B image inpainting model from Tencent that claims quality near 10B-scale systems, small enough to run client-side. - [Qwen's AgentWorld Simulates Worlds for AI Agents](https://theopenweights.com/news/qwen-agentworld-35b-a3b-ryrn) — Qwen · Alibaba · Text / LLM · Qwen · Jun 22, 2026: Qwen-AgentWorld-35B-A3B is a Mixture-of-Experts language world model with 35B total and 3B active params, aimed at simulating agentic environments. - [Baidu's PP-OCRv6 packs 50-language OCR into tiny models](https://theopenweights.com/news/pp-ocrv6-mnzg) — Baidu · Vision-Language · Apache 2.0 · Jun 22, 2026: Baidu releases PP-OCRv6, a 50-language OCR model family ranging from 1.5M to 34.5M parameters under Apache 2.0. - [Gepard 1.0 brings bilingual TTS with voice cloning](https://theopenweights.com/news/gepard-1-0-13dp) — nineninesix · Text → Speech · Apache 2.0 · Jun 22, 2026: Gepard 1.0 is an open autoregressive TTS model with voice cloning for English and Spanish, released under Apache 2.0. - [Agents-A1: A 35B MoE Built for Agentic Scaling](https://theopenweights.com/news/agents-a1-ezpm) — InternScience · Reasoning · Apache 2.0 · Jun 22, 2026: Agents-A1 is a 35B-parameter MoE agentic model from inclusionAI, released under Apache 2.0 and claiming trillion-scale performance via agent-horizon scaling. - [DeepReinforce's Ornith-1.0-9B Targets Agentic Coding](https://theopenweights.com/news/ornith-1-0-9b-o2ja) — Deepreinforce Ai · Text / LLM · MIT · Jun 21, 2026: DeepReinforce releases Ornith-1.0-9B, a compact MIT-licensed model tuned for agentic coding, now available on Hugging Face. - [Ornith-1.0-35B brings a mid-size MoE to agentic coding](https://theopenweights.com/news/ornith-1-0-35b-y50v) — Deepreinforce Ai · Text / LLM · MIT · Jun 21, 2026: Ornith-1.0-35B is a 35B mixture-of-experts model for agentic coding, released under a permissive MIT license on Hugging Face. - [Poolside releases Laguna XS 2.1 code model](https://theopenweights.com/news/laguna-xs-2-1-s9l9) — Poolside · Code · Other · Jun 20, 2026: Poolside's Laguna XS 2.1, a compact code-focused LLM, is now available on Hugging Face under an open model license. - [LIFT: A Qwen3.5-Based VLM for PDF-to-JSON Extraction](https://theopenweights.com/news/lift-828w) — Datalab To · Vision-Language · OpenRAIL-M · Jun 19, 2026: Datalab releases LIFT 1.0, a Qwen3.5-based vision-language model for extracting structured JSON from PDFs and documents. - [Baidu releases Unlimited-OCR under permissive MIT license](https://theopenweights.com/news/unlimited-ocr-5bqq) — Baidu · Vision-Language · MIT · Jun 19, 2026: Baidu's Unlimited-OCR is a multilingual vision-language OCR model released on Hugging Face under the permissive MIT license. - [Cohere releases Apache-licensed Arabic speech model](https://theopenweights.com/news/cohere-transcribe-arabic-55ph) — Cohere · Speech → Text · Apache 2.0 · Jun 18, 2026: Cohere Labs has released an Arabic-focused speech recognition model supporting Arabic and English under Apache 2.0. - [Krea 2 Arrives as Open-Weights Text-to-Image Model](https://theopenweights.com/news/krea-2-g1jz) — Krea · Text → Image · Other · Jun 18, 2026: Krea 2 launches as an open-weights text-to-image diffusion model, offered in raw and turbo flavors on Hugging Face. - [Krea 2 Arrives as a 12B Open-Weights Image Model](https://theopenweights.com/news/krea-2-qw7c) — Krea · Text → Image · Other · Jun 18, 2026: Krea 2 is a 12B open-weights text-to-image model, released alongside a faster Turbo variant for quicker generation. - [Zhipu AI Releases MIT-Licensed GLM-5.2 MoE Model](https://theopenweights.com/news/glm-5-2-bs32) — Zhipu AI · Text / LLM · MIT · Jun 17, 2026: Zhipu AI has released GLM-5.2, a new bilingual Mixture of Experts language model with sparse attention, available under the permissive MIT license. - [Resemble AI's Inflect-Nano-v1 puts TTS on local hardware](https://theopenweights.com/news/inflect-nano-v1-1hlt) — Owensong · Text → Speech · Apache 2.0 · Jun 16, 2026: Resemble AI releases Inflect-Nano-v1, a sub-1B experimental English text-to-speech model under Apache 2.0 for local use. - [Boogu-Image 0.1 Edit arrives with Apache license](https://theopenweights.com/news/boogu-image-0-1-edit-kso3) — Boogu · Image Editing · Apache 2.0 · Jun 16, 2026: Boogu-Image 0.1 Edit is an Apache-licensed open-weight diffusion model for image editing, with ComfyUI support out of the gate. - [Poolside Releases Laguna-M.1, an Open MoE Model](https://theopenweights.com/news/laguna-m-1-ldp1) — Poolside · Text / LLM · Apache 2.0 · Jun 15, 2026: Poolside debuts Laguna-M.1, an Apache-2.0 mixture-of-experts LLM for text and code with vLLM and SGLang support. - [Microsoft's FastContext is a 4B sub-agent for code](https://theopenweights.com/news/fastcontext-1-0-4b-9m7y) — Microsoft · Text / LLM · MIT · Jun 14, 2026: Microsoft releases FastContext 1.0 4B SFT, a Qwen3-4B fine-tune built as a repository-exploration sub-agent, under an MIT license. - [Moonshot AI releases Kimi K3, a 2.8T-parameter MoE model](https://theopenweights.com/news/kimi-k3-xyt3) — Moonshot AI · Text / LLM · Other · Jun 13, 2026: Moonshot AI has released Kimi K3, a ~2.8T-parameter mixture-of-experts multimodal model with open weights on Hugging Face. - [Weibo AI Releases VibeThinker-3B, a Compact Reasoning Model](https://theopenweights.com/news/vibethinker-3b-3am6) — WeiboAI · Reasoning · Other · Jun 12, 2026: Weibo AI has released VibeThinker-3B, a compact 3B model specializing in complex reasoning tasks like math, code generation, and GPQA benchmarks. - [Moonshot AI Releases Kimi, a Multimodal Coding Model](https://theopenweights.com/news/kimi-k2-7-code-jmbg) — Moonshot AI · Code · Other · Jun 11, 2026: Moonshot AI has released Kimi-K2.7-Code, a Mixture-of-Experts model designed for coding tasks that can also process visual inputs like images. - [Zyphra Releases Open-Source Zonos 2 TTS Model](https://theopenweights.com/news/zonos-2-u7e2) — Zyphra · Text → Speech · Apache 2.0 · Jun 11, 2026: Zyphra has released Zonos 2, an open-weight text-to-speech model available for commercial use under the permissive Apache 2.0 license. - [Google Releases Open-Source DiffusionGemma 26B Model](https://theopenweights.com/news/diffusiongemma-26b-a4b-h4ke) — Google DeepMind · Text / LLM · Apache 2.0 · Jun 9, 2026: Google has released DiffusionGemma 26B, an open-source, instruction-tuned Mixture-of-Experts model that uses a novel diffusion architecture for text. - [Zhipu AI Releases SCAIL-2 for Character Animation](https://theopenweights.com/news/scail-2-xvt5) — Zhipu AI · Image → Video · MIT · Jun 9, 2026: Zhipu AI's research arm has released SCAIL-2, an open-source, MIT-licensed diffusion model for pose-driven character animation from a single image. - [PaddleOCR's PP-OCRv6 Adds a Medium Detection Model](https://theopenweights.com/news/pp-ocrv6-medium-detection-ktee) — Baidu · Vision-Language · Apache 2.0 · Jun 9, 2026: Baidu releases PP-OCRv6 Medium Detection, an Apache-2.0 text-line detection model under 1B params, on Hugging Face. - [Cohere Releases North-Mini-Code, an Open MoE Model](https://theopenweights.com/news/north-mini-code-1-0-wul0) — Cohere · Code · Apache 2.0 · Jun 5, 2026: Cohere has released North-Mini-Code 1.0, an Apache 2.0-licensed Mixture-of-Experts model specialized for code generation and agentic chat applications. - [Boson AI releases Higgs TTS v3, a 4B speech model](https://theopenweights.com/news/higgs-tts-v3-4b-txct) — Bosonai · Text → Speech · Other · Jun 4, 2026: Boson AI's Higgs TTS v3 is a 4B-parameter multilingual text-to-speech model offering expressive control and voice cloning. - [Higgs TTS 3 lands as a 4B multilingual speech model](https://theopenweights.com/news/higgs-tts-3-4b-losw) — Bosonai · Text → Speech · Other · Jun 4, 2026: Higgs TTS 3 4B is an expressive, controllable multilingual text-to-speech model, now available on Hugging Face. - [Boson AI's Higgs Audio v3 Offers Expressive, Multilingual TTS](https://theopenweights.com/news/higgs-audio-v3-tts-4b-2be3) — Bosonai · Text → Speech · Other · Jun 4, 2026: Boson AI has released Higgs Audio v3, a 4-billion-parameter text-to-speech model for expressive, controllable, and multilingual audio generation. - [Ideogram 4.0 arrives as an open-weight image model](https://theopenweights.com/news/ideogram-4-0-dxje) — Stability AI · Text → Image · Other · Jun 3, 2026: Ideogram 4.0 is a 9.3B-parameter open-weight text-to-image model, released publicly via GitHub by its OSS team. - [Ideogram 4.0 arrives as an open-weight image model](https://theopenweights.com/news/ideogram-4-0-bw1q) — Black Forest Labs · Text → Image · Other · Jun 3, 2026: Ideogram 4.0, a 9.3B open-weight text-to-image model, has been released with code on GitHub. - [MiniMax Releases M3, a Multimodal MoE Model](https://theopenweights.com/news/minimax-m3-1lji) — MiniMax · Vision-Language · Other · Jun 2, 2026: MiniMax AI has released MiniMax-M3, an open-weight, multimodal Mixture-of-Experts model designed for vision, coding, and agentic reasoning tasks. - [JD.com Enters Open-Source AI Video with JoyAI-Echo](https://theopenweights.com/news/joyai-echo-sgij) — JD · Text → Video · Other · Jun 2, 2026: JD.com has released JoyAI-Echo, an open-source text-to-video model based on LTX-Video research that generates long-form, multi-shot videos with audio. - [Ideogram 4.0: A 9.3B Open-Weight Text-to-Image Model](https://theopenweights.com/news/ideogram-4-0-abj7) — Ideogram Ai · Text → Image · Other · May 30, 2026: Ideogram has released Ideogram 4.0, a 9.3B parameter open-weight text-to-image model using a Diffusion Transformer (DiT) architecture. - [Baidu Releases NAVA for Text-to-Video with Audio](https://theopenweights.com/news/nava-xei6) — Baidu · Text → Video · Other · May 29, 2026: Baidu has released NAVA, a model that generates synchronized video and audio from text and image prompts using a modern flow-matching technique. - [Stability AI's Demon brings real-time music diffusion to local GPUs](https://theopenweights.com/news/demon-y899) — Stability AI · Music · Other · May 27, 2026: Demon is an open-source real-time music diffusion engine that runs at 25Hz on a single local GPU. - [MOSS-TTS Aims for More Robust Speech Synthesis](https://theopenweights.com/news/moss-tts-v1-5-4pw8) — OpenMOSS · Text → Speech · Other · May 25, 2026: OpenMOSS has released MOSS-TTS v1.5, a multilingual text-to-speech model using a novel decoding method to improve the robustness of generated audio. - [Google Releases Gemma 4 12B Multimodal Model](https://theopenweights.com/news/gemma-4-12b-noh4) — Google DeepMind · Any-to-Any · Apache 2.0 · May 23, 2026: Google DeepMind has released Gemma 4 12B, a new open 'any-to-any' multimodal model designed for flexible input and output across different data types. - [Google Releases Gemma 4, a 12B 'Any-to-Any' Model](https://theopenweights.com/news/gemma-4-12b-lou5) — Google DeepMind · Any-to-Any · Gemma · May 23, 2026: Google DeepMind has released Gemma 4 12B Instruct, a new open-weight model with a unified 'any-to-any' architecture for flexible multimodal reasoning. - [NVIDIA Releases Cosmos3 Image-to-Video World Model](https://theopenweights.com/news/cosmos3-super-image2video-ojli) — NVIDIA · Image → Video · Other · May 21, 2026: NVIDIA has released Cosmos3 Super Image2Video, a new generative model designed to create video from a single still image, part of its world-model research. - [Meituan releases LongCat-Video-Avatar 1.5](https://theopenweights.com/news/longcat-video-avatar-1-5-ptbe) — Meituan · Image → Video · Other · May 21, 2026: Meituan's LongCat-Video-Avatar 1.5 turns a still image and audio into talking-head video, with support for clip continuation. - [OpenBMB's MiniCPM5-1B targets on-device AI](https://theopenweights.com/news/minicpm5-1b-qv4h) — OpenBMB · Text / LLM · Other · May 21, 2026: OpenBMB releases MiniCPM5-1B, a 1B-parameter on-device LLM with long-context and tool-calling support. - [KRAFTON releases 1B zero-shot voice-cloning TTS](https://theopenweights.com/news/raon-opentts-1b-mr3n) — KRAFTON · Text → Speech · Other · May 21, 2026: KRAFTON's Raon-OpenTTS-1B is a 1B-parameter zero-shot English text-to-speech model built on a flow-matching diffusion transformer. - [MisoLabs Debuts MisoTTS, an Open Voice Model](https://theopenweights.com/news/misotts-te3d) — MisoLabs · Text → Speech · Other · May 21, 2026: MisoLabs has released MisoTTS, a new open-source text-to-speech model that uses a decoder-only architecture inspired by the Llama family of LLMs. - [Mega-ASR Improves on Qwen for Speech Recognition](https://theopenweights.com/news/mega-asr-mvn4) — zhifeixie · Speech → Text · Apache 2.0 · May 19, 2026: Mega-ASR is a new 1.7B-parameter speech recognition model released under an Apache 2.0 license. It's a fine-tune of Qwen3-ASR for English and Chinese. - [OpenMOSS Releases Transcribe-Diarize ASR Model](https://theopenweights.com/news/moss-transcribe-diarize-xcc6) — OpenMOSS · Speech → Text · Other · May 19, 2026: OpenMOSS releases MOSS Transcribe-Diarize, an open ASR model handling long-form audio with speaker diarization and timestamped output. - [NVIDIA Releases SANA, a Camera-Controllable Video Model](https://theopenweights.com/news/sana-wm-bidirectional-3rui) — NVIDIA · Image → Video · Apache 2.0 · May 18, 2026: NVIDIA has released SANA-WM, an open-source world model for camera-controllable video generation using a novel bidirectional diffusion technique. - [NVIDIA Releases Nemotron-3.5 Streaming ASR Model](https://theopenweights.com/news/nemotron-3-5-asr-streaming-0-6b-jn8s) — NVIDIA · Speech → Text · Other · May 15, 2026: NVIDIA has released Nemotron 3.5 ASR Streaming, a 600M-parameter model designed for real-time, multilingual speech-to-text transcription. - [ByteDance Releases Lance, a Unified Generative AI Model](https://theopenweights.com/news/lance-8jr6) — ByteDance · Any-to-Any · Apache 2.0 · May 15, 2026: ByteDance has released Lance, a 3B parameter open-source model for unified image and video generation, editing, and understanding via a single architecture. - [SenseTime Releases 8B 'Any-to-Any' Infographic Model](https://theopenweights.com/news/sensenova-u1-8b-mot-infographic-e0f6) — SenseTime · Any-to-Any · Other · May 14, 2026: SenseTime has released SenseNova U1, an 8B 'any-to-any' multimodal model. It specializes in generating and editing complex infographics within a conversation. - [GLiGuard: A Sub-1B Model for Faster LLM Guardrails](https://theopenweights.com/news/gliguard-ylbc) — OpenBMB · Text / LLM · Apache 2.0 · May 12, 2026: GLiGuard is an Apache-2.0 small language model for LLM safety moderation, with claimed 16x faster guardrail enforcement. - [Needle: A 26M-Parameter Model Built for Tool Calling](https://theopenweights.com/news/needle-gvr1) — OpenBMB · Code · Apache 2.0 · May 12, 2026: Needle is a 26M-parameter model distilled from Gemini for tool calling, small enough to run on local hardware under Apache 2.0. - [Lightricks Releases LoRA for AI Lip-Dubbing](https://theopenweights.com/news/ltx-2-3-22b-ic-lora-lipdub-azlx) — Lightricks · Image → Video · Other · May 11, 2026: Lightricks has released an IC-LoRA, a specialized adapter for its LTX-2.3 video model, designed to enable high-quality, identity-preserving lip-dubbing. - [Tencent Releases 1.8B Model for Multilingual Translation](https://theopenweights.com/news/hunyuan-mt2-1-8b-pb3x) — Tencent · Text / LLM · Other · May 11, 2026: Tencent has released Hunyuan-MT2, a 1.8B parameter dense model focused on high-quality multilingual machine translation, available with a custom license. - [Supertone Releases On-Device Multilingual TTS Model](https://theopenweights.com/news/supertonic-3-47h3) — Supertone · Text → Speech · Other · May 6, 2026: Supertone has released Supertonic 3, a multilingual text-to-speech model optimized for on-device use with ONNX support and a non-commercial license. - [Liquid AI ships a 350M multilingual embedding model](https://theopenweights.com/news/lfm2-5-embedding-350m-v6x6) — LiquidAI · Embeddings · Other · May 5, 2026: Liquid AI's LFM2.5-Embedding-350M is a compact multilingual text embedding model built for on-device retrieval and search. - [Kimi K2.6 tops closed models in coding test](https://theopenweights.com/news/kimi-k2-6-9pup) — Moonshot AI · Vision-Language · Other · May 3, 2026: Moonshot AI's open-weights MoE model Kimi K2.6 is reported to beat leading closed models in a coding challenge. - [NVIDIA Releases PiD for High-Quality Image Upscaling](https://theopenweights.com/news/nvidia-pid-pixel-diffusion-decoder-38mh) — NVIDIA · Image Editing · Other · Apr 28, 2026: NVIDIA has released PiD, a pixel-diffusion VAE decoder for super-resolution. The component is designed to work with Stability AI's Z-Image base model. - [NVIDIA Releases Efficient Nemotron-3 Multimodal MoE](https://theopenweights.com/news/nemotron-3-nano-omni-30b-a3b-reasoning-cbno) — NVIDIA · Any-to-Any · Other · Apr 24, 2026: NVIDIA has released Nemotron-3 Nano Omni, a 30B parameter Mixture-of-Experts model for multimodal reasoning that uses only 3B active parameters. - [Google Releases Gemma 4 Multimodal Open Model](https://theopenweights.com/news/gemma-4-26b-a4b-moe-z8tk) — Google DeepMind · Any-to-Any · Apache 2.0 · Apr 23, 2026: Google DeepMind has released Gemma 4, a 26B parameter multimodal model. It uses a sparse mixture-of-experts architecture with 4B active parameters. - [Google Releases Multimodal Gemma 4 31B Model](https://theopenweights.com/news/gemma-4-31b-yue0) — Google DeepMind · Any-to-Any · Apache 2.0 · Apr 23, 2026: Google DeepMind has released Gemma 4 31B, a new instruction-tuned, multimodal 'any-to-any' model with an open Apache 2.0 license for commercial use. - [Google Releases 4B Multimodal Gemma 4 Assistant](https://theopenweights.com/news/gemma-4-e4b-assistant-u21f) — Google DeepMind · Any-to-Any · Apache 2.0 · Apr 23, 2026: Google has released Gemma 4 E4B-it, a 4-billion-parameter, instruction-tuned multimodal model designed for flexible 'any-to-any' assistant tasks. - [Google Releases 2B Multimodal Gemma 4 Assistant Model](https://theopenweights.com/news/gemma-4-e2b-assistant-xoc9) — Google DeepMind · Any-to-Any · Apache 2.0 · Apr 23, 2026: Google has released Gemma 4 E2B-it, a 2-billion-parameter, instruction-tuned multimodal assistant model under the permissive Apache 2.0 license. - [Xiaomi Releases MiMo Model for Speech Recognition](https://theopenweights.com/news/mimo-v2-5-asr-n0ry) — Xiaomi · Speech → Text · MIT · Apr 23, 2026: Xiaomi has released MiMo-V2.5-ASR, a new open-source model for automatic speech recognition in Mandarin Chinese, English, and Cantonese. - [LLaDA2.0-Uni: A Unified MoE for Vision Tasks](https://theopenweights.com/news/llada2-0-uni-4qj9) — inclusionAI · Any-to-Any · Apache 2.0 · Apr 22, 2026: LLaDA2.0-Uni is a new Apache 2.0 licensed Mixture-of-Experts model for unified image understanding, generation, and editing, built on a diffusion framework. - [DeepSeek Releases V4-Pro, an Open MoE Contender](https://theopenweights.com/news/deepseek-v4-pro-y3sa) — DeepSeek · Text / LLM · MIT · Apr 22, 2026: DeepSeek has released V4-Pro, a new Mixture-of-Experts (MoE) language model. Its fully permissive MIT license makes it notable for commercial use. - [DeepSeek Releases V4-Flash, a Fast MIT-Licensed MoE Model](https://theopenweights.com/news/deepseek-v4-flash-m3ve) — DeepSeek · Text / LLM · MIT · Apr 22, 2026: DeepSeek has released DeepSeek-V4-Flash, a Mixture-of-Experts model optimized for speed with FP8 weights and a permissive MIT license for commercial use. - [SenseTime Releases 8B Any-to-Any Multimodal Model](https://theopenweights.com/news/sensenova-u1-8b-mot-k2dz) — SenseTime · Any-to-Any · Other · Apr 22, 2026: SenseTime has released SenseNova-U1-8B-MoT, an 8B multimodal model designed for unified image understanding, generation, editing, and interleaved output. - [Alibaba's Qwen Releases Open 27B Vision Model](https://theopenweights.com/news/qwen3-6-27b-v8co) — Qwen · Alibaba · Vision-Language · Apache 2.0 · Apr 21, 2026: Alibaba's Qwen team has released Qwen3.6-27B, a new 27-billion-parameter dense vision-language model available under the permissive Apache 2.0 license. - [NVIDIA Releases Nemotron-3-Nano Omni-Modal MoE](https://theopenweights.com/news/nemotron-3-nano-omni-30b-a3b-reasoning-bpwf) — NVIDIA · Any-to-Any · Other · Apr 20, 2026: NVIDIA has released Nemotron-3-Nano-Omni, a 30B parameter Mixture-of-Experts model designed for efficient, 'any-to-any' multimodal reasoning. - [Resemble AI Releases Dramabox Voice Cloning TTS Model](https://theopenweights.com/news/dramabox-tts-od1y) — Resemble AI · Text → Speech · Other · Apr 17, 2026: Resemble AI has released Dramabox, a text-to-speech model featuring one-shot voice cloning capabilities built on a diffusion-transformer architecture. - [IBM Releases 2B Granite Model for Multilingual Speech](https://theopenweights.com/news/granite-speech-4-1-2b-ce4g) — IBM · Speech → Text · Apache 2.0 · Apr 16, 2026: IBM has released Granite Speech 4.1, a 2-billion-parameter open-source model for multilingual automatic speech recognition. - [Qwen Releases 35B Multimodal Mixture-of-Experts Model](https://theopenweights.com/news/qwen3-6-35b-a3b-ut3q) — Qwen · Alibaba · Vision-Language · Apache 2.0 · Apr 15, 2026: Alibaba's Qwen team has released Qwen3.6-35B-A3B, a 35B-parameter multimodal model using a Mixture-of-Experts architecture for greater efficiency. - [Motif Releases 2B Open-Source Text-to-Video Model](https://theopenweights.com/news/motif-video-2b-z7yy) — Motif Technologies · Text → Video · Apache 2.0 · Apr 14, 2026: Motif Technologies has released Motif-Video-2B, a 2-billion-parameter open-source model for text-to-video and image-to-video generation. - [Moonshot AI Releases Kimi-K2.6 Multimodal Model](https://theopenweights.com/news/kimi-k2-6-ravo) — Moonshot AI · Vision-Language · Other · Apr 14, 2026: Moonshot AI has released Kimi-K2.6, a new vision-language model. The weights are available for research purposes under a non-commercial license. - [OpenBMB Releases MiniCPM-V for On-Device Vision](https://theopenweights.com/news/minicpm-v-4-6-ghgd) — OpenBMB · Vision-Language · Apache 2.0 · Apr 13, 2026: OpenBMB has released MiniCPM-V-4.6, a lightweight vision-language model that combines Llama 3 and SigLIP for on-device and mobile applications. - [NVIDIA's Nemotron TwoTower mixes diffusion and Mamba](https://theopenweights.com/news/nemotron-twotower-30b-a3b-d2oa) — NVIDIA · Text / LLM · Other · Apr 11, 2026: NVIDIA releases Nemotron TwoTower 30B-A3B, a hybrid diffusion/Mamba MoE base model with 30B total and 3B active parameters. - [NVIDIA's Nemotron TwoTower is a MoE experiment](https://theopenweights.com/news/nemotron-twotower-30b-a3b-1pls) — NVIDIA · Text / LLM · Other · Apr 11, 2026: NVIDIA released Nemotron TwoTower 30B-A3B, an experimental MoE base model with 3B active params, on Hugging Face. - [MiniMax Releases M2.7, an MoE Model with FP8 Weights](https://theopenweights.com/news/minimax-m2-7-r9py) — MiniMax · Text / LLM · Other · Apr 9, 2026: MiniMax has released MiniMax-M2.7, a conversational language model featuring a Mixture-of-Experts architecture and efficient FP8 weights for research use. - [Baidu Releases 8B Text-to-Image Model ERNIE-Image](https://theopenweights.com/news/ernie-image-9iz4) — Baidu · Text → Image · Apache 2.0 · Apr 7, 2026: Baidu has released ERNIE-Image, an 8-billion-parameter text-to-image model under a permissive Apache 2.0 license for commercial use and research. - [Black Forest Labs Releases Open FLUX.2 Image Decoder](https://theopenweights.com/news/flux-2-small-decoder-42ux) — Black Forest Labs · Text → Image · Apache 2.0 · Apr 6, 2026: Black Forest Labs has released the FLUX.2 small decoder, a key component of a new transformer-based architecture for text-to-image and image editing. - [Zhipu AI Releases Open-Source GLM-5.1 MoE Model](https://theopenweights.com/news/glm-5-1-ufei) — Zhipu AI · Text / LLM · MIT · Apr 3, 2026: Zhipu AI has released GLM-5.1, a bilingual (EN/ZH) Mixture-of-Experts LLM with a permissive MIT license and a new attention mechanism. - [OpenBMB Releases VoxCPM2 for Expressive TTS](https://theopenweights.com/news/voxcpm2-1326) — OpenBMB · Text → Speech · Other · Apr 3, 2026: OpenBMB has released VoxCPM2, a diffusion-based text-to-speech model capable of multilingual synthesis and zero-shot voice cloning from short audio clips. - [MOSS-TTS-Nano Delivers Multilingual Speech at 100M Params](https://theopenweights.com/news/moss-tts-nano-100m-63iz) — OpenMOSS · Text → Speech · Other · Apr 2, 2026: OpenMOSS-Team releases MOSS-TTS-Nano, a compact 100M parameter text-to-speech model supporting English, Mandarin, Cantonese, and mixed-language text. - [Tencent Releases 2B Vision Model for Robotics](https://theopenweights.com/news/hy-embodied-0-5-bhzk) — Tencent · Vision-Language · Other · Apr 2, 2026: Tencent has released HY-Embodied 0.5, a 2-billion-parameter vision-language model with an end-to-end architecture for multi-object tracking. - [Tencent Releases HY-OmniWeaving for Multi-Image Video](https://theopenweights.com/news/hy-omniweaving-yu2r) — Tencent · Image → Video · Other · Mar 31, 2026: Tencent has released HY-OmniWeaving, an open-source model that generates video by combining multiple images and text prompts based on HunyuanVideo-1.5. - [JD.com Releases Open-Source Bilingual Image Editor](https://theopenweights.com/news/joyai-image-edit-xznj) — JD · Image Editing · Apache 2.0 · Mar 31, 2026: JD.com has released JoyAI-Image-Edit, an Apache 2.0 licensed model for text-guided image editing that supports prompts in both English and Chinese. - [KRAFTON Releases 9B Bilingual Speech Model](https://theopenweights.com/news/raon-speech-9b-aymv) — KRAFTON · Any-to-Any · CC BY-NC 4.0 · Mar 30, 2026: KRAFTON has released Raon-Speech-9B, a 9-billion-parameter model for English and Korean speech recognition and synthesis under a non-commercial license. - [OmniVoice TTS Offers Zero-Shot Multilingual Voice Cloning](https://theopenweights.com/news/omnivoice-62u6) — k2-fsa · Text → Speech · Other · Mar 30, 2026: OmniVoice is a new zero-shot text-to-speech model capable of multilingual voice cloning and style control using a single three-second audio prompt. - [HKUST Releases Audio-Omni, a Unified Audio Model](https://theopenweights.com/news/audio-omni-a60k) — HKUSTAudio · Any-to-Any · CC BY-NC 4.0 · Mar 27, 2026: Researchers at HKUST have released Audio-Omni, a versatile diffusion model for any-to-any audio generation, conversion, and editing tasks. - [Meituan Releases LongCat-Next 'Any-to-Any' AI Model](https://theopenweights.com/news/longcat-next-x6er) — Meituan · Any-to-Any · MIT · Mar 25, 2026: Chinese tech company Meituan has released LongCat-Next, a multimodal AI model designed to handle any combination of text, image, audio, and video inputs and outputs. - [Cohere Releases Top-Ranked Multilingual Transcription Model](https://theopenweights.com/news/cohere-transcribe-03-2026-vg9w) — Cohere · Speech → Text · Other · Mar 24, 2026: Cohere has released a new multilingual transcription model that has taken the top spot on the Hugging Face Open ASR Leaderboard for speech recognition. - [Irodori-TTS v2 Offers Open Japanese Speech Synthesis](https://theopenweights.com/news/irodori-tts-500m-v2-twfp) — Aratako · Text → Speech · MIT · Mar 23, 2026: Researcher Aratako has released Irodori-TTS-500M v2, an open-source Japanese text-to-speech model based on the VITS architecture with an MIT license. - [GAIR Releases daVinci-MagiHuman for Video Generation](https://theopenweights.com/news/davinci-magihuman-5hmp) — GAIR · Image → Video · Other · Mar 21, 2026: GAIR releases daVinci-MagiHuman, an open-source model that generates video with audio from text, image, and audio inputs. - [Baidu Releases Qianfan-OCR for Document Intelligence](https://theopenweights.com/news/qianfan-ocr-8hbc) — Baidu · Vision-Language · Other · Mar 18, 2026: Baidu has released Qianfan-OCR, a new vision-language model for multilingual document intelligence tasks, including layout analysis and text extraction. - [Needle: A 26M-Param Model Built for On-Device Tool Calls](https://theopenweights.com/news/needle-yfob) — Cactus Compute · Code · MIT · Mar 16, 2026: Needle is a 26M-parameter encoder-decoder distilled for tool and function calling, MIT-licensed and aimed at on-device edge deployment. - [Google Releases Gemma 4, a 26B Vision-Language Model](https://theopenweights.com/news/gemma-4-26b-a4b-rzlm) — Google DeepMind · Any-to-Any · Apache 2.0 · Mar 11, 2026: Google DeepMind has released Gemma 4 26B, an open-source, instruction-tuned vision-language model featuring a Mixture-of-Experts (MoE) architecture. - [Google Releases Multimodal Gemma 4 31B Model](https://theopenweights.com/news/gemma-4-31b-xj24) — Google DeepMind · Any-to-Any · Gemma · Mar 11, 2026: Google DeepMind has released Gemma 4 31B IT, a new 31-billion-parameter open-weights model capable of processing both text and image inputs. - [Black Forest Labs Releases 9B FLUX.2 klein Image Model](https://theopenweights.com/news/flux-2-klein-9b-mlgd) — Black Forest Labs · Text → Image · Other · Mar 9, 2026: Black Forest Labs has released FLUX.2 klein, a 9-billion-parameter open-weight model for high-quality image generation and editing, based on a distilled architecture. - [Fish Audio's S2-Pro Brings Expressive TTS to Open Source](https://theopenweights.com/news/fish-audio-s2-pro-xc7o) — Fishaudio · Text → Speech · Other · Mar 9, 2026: Fish Audio has released S2-Pro, a multilingual, instruction-following text-to-speech model capable of zero-shot voice cloning from brief audio clips. - [Lightricks LTX-2.3 Generates Video and Audio Together](https://theopenweights.com/news/ltx-2-3-fvz7) — Lightricks · Image → Video · Other · Mar 4, 2026: Lightricks has released LTX-2.3, a new model that generates video and audio streams jointly from text, image, or audio inputs. - [NVIDIA's New 3B VLM Pinpoints Objects in Images](https://theopenweights.com/news/locateanything-3b-kowy) — NVIDIA · Vision-Language · Other · Mar 2, 2026: NVIDIA has released LocateAnything-3B, a 3-billion-parameter Vision-Language Model designed for precise object detection and visual grounding. - [Google Releases Compact Gemma 4 E2B Multimodal Model](https://theopenweights.com/news/gemma-4-e2b-3990) — Google DeepMind · Any-to-Any · Gemma · Mar 2, 2026: Google DeepMind has released Gemma 4 E2B, a new 2-billion-parameter multimodal model capable of processing both image and text inputs. - [Google's Gemma 4 Arrives with Any-to-Any Multimodal Skills](https://theopenweights.com/news/gemma-4-e2b-a40j) — Google DeepMind · Any-to-Any · Gemma · Mar 2, 2026: Google DeepMind has released Gemma 4 E2B IT, a 2B parameter, instruction-tuned multimodal model capable of processing text, vision, and audio. - [Google Releases Gemma 4 E4B, a 4B Multimodal Model](https://theopenweights.com/news/gemma-4-e4b-7zmz) — Google DeepMind · Any-to-Any · Gemma · Mar 2, 2026: Google DeepMind has released Gemma 4 E4B, a new 4-billion-parameter open model capable of processing both image and text inputs to generate text. - [Google's Gemma 4 Debuts with Any-to-Any Multimodality](https://theopenweights.com/news/gemma-4-e4b-edpx) — Google DeepMind · Any-to-Any · Apache 2.0 · Mar 2, 2026: Google has released Gemma 4 E4B Instruct, a compact 4B parameter model with 'any-to-any' multimodal capabilities for handling diverse inputs and outputs. - [Xiaomi Releases Bilingual Image Editing Model FireRed 1.1](https://theopenweights.com/news/firered-image-edit-1-1-1y5n) — Xiaomi · Image Editing · Apache 2.0 · Mar 2, 2026: Xiaomi's FireRedTeam has released FireRed Image Edit 1.1, an open-source model for instruction-based image editing in English and Chinese. - [Alibaba's Qwen Releases Compact 0.8B Vision Model](https://theopenweights.com/news/qwen3-5-0-8b-s4x3) — Qwen · Alibaba · Vision-Language · Apache 2.0 · Feb 28, 2026: Alibaba's Qwen team has released Qwen3.5-0.8B, a compact vision-language model with 800 million parameters available under an Apache 2.0 license. - [IBM Releases 1B Granite Model for Multilingual Speech](https://theopenweights.com/news/granite-4-0-1b-speech-xfxz) — IBM · Speech → Text · Apache 2.0 · Feb 27, 2026: IBM has released Granite 4.0 1B Speech, a 1-billion-parameter open-source model for multilingual automatic speech recognition under an Apache 2.0 license. - [Alibaba's Qwen team releases 4B vision-language model](https://theopenweights.com/news/qwen3-5-4b-xcb0) — Qwen · Alibaba · Vision-Language · Apache 2.0 · Feb 27, 2026: Alibaba's Qwen team has released Qwen3.5-4B, a 4-billion-parameter vision-language model available under the permissive Apache 2.0 license. - [Qwen Releases 9B Multimodal Model in New 3.5 Series](https://theopenweights.com/news/qwen3-5-9b-uah1) — Qwen · Alibaba · Vision-Language · Apache 2.0 · Feb 27, 2026: Alibaba's Qwen team has released Qwen3.5-9B, a 9-billion parameter vision-language model licensed under Apache 2.0 for commercial and research use. - [Moonshine: Open STT Models Aim to Beat Whisper](https://theopenweights.com/news/moonshine-stt-475j) — Resemble AI · Speech → Text · MIT · Feb 24, 2026: Resemble AI's Moonshine is an open-weight, MIT-licensed STT model family claiming to outperform Whisper Large v3 on accuracy. - [Qwen Releases Flagship 122B Multimodal MoE Model](https://theopenweights.com/news/qwen3-5-122b-a10b-zglq) — Qwen · Alibaba · Vision-Language · Apache 2.0 · Feb 24, 2026: Alibaba's Qwen team has released Qwen3.5-122B-A10B, a 122B parameter Mixture-of-Experts model for vision and text, under an Apache 2.0 license. - [Qwen Releases 27B Vision Model with Long Context](https://theopenweights.com/news/qwen3-5-27b-t3hz) — Qwen · Alibaba · Vision-Language · Apache 2.0 · Feb 24, 2026: Alibaba's Qwen has released Qwen3.5-27B, a 27B parameter vision-language model with a 131K context window, available under the Apache 2.0 license. - [Qwen Releases Efficient 35B Multimodal MoE Model](https://theopenweights.com/news/qwen3-5-35b-a3b-vs1s) — Qwen · Alibaba · Vision-Language · Apache 2.0 · Feb 24, 2026: Alibaba's Qwen team has released a 35B-parameter Mixture of Experts model with vision capabilities and a permissive Apache 2.0 license. - [Qwen releases flagship 397B multimodal MoE](https://theopenweights.com/news/qwen3-5-397b-a17b-d2xo) — Qwen · Alibaba · Vision-Language · Apache 2.0 · Feb 16, 2026: Alibaba's Qwen project has released Qwen3.5-397B-A17B, a massive 397B parameter multimodal Mixture-of-Experts model available under an Apache 2.0 license. - [Hume AI Releases 3B Multilingual Text-to-Speech Model](https://theopenweights.com/news/tada-3b-ml-msos) — HumeAI · Text → Speech · Other · Feb 16, 2026: Hume AI has released Tada-3B-ML, a 3 billion parameter text-to-speech model supporting over 10 languages with a focus on expressive vocal control. - [Kani-TTS-2 Offers New Open-Source Voice Generation](https://theopenweights.com/news/kani-tts-2-english-4urg) — nineninesix · Text → Speech · Apache 2.0 · Feb 12, 2026: Independent researcher 'nineninesix' has released Kani-TTS-2, a new open-source English text-to-speech model available on Hugging Face. - [MiniMax Releases M2.5 Mixture-of-Experts Model](https://theopenweights.com/news/minimax-m2-5-dvuu) — MiniMax · Text / LLM · Other · Feb 12, 2026: Chinese AI company MiniMax has released MiniMax-M2.5, a Mixture-of-Experts language model featuring efficient FP8 weights and a non-commercial license. - [Zhipu AI Releases Open-Source GLM-5 MoE Model](https://theopenweights.com/news/glm-5-w6db) — Zhipu AI · Text / LLM · MIT · Feb 11, 2026: Zhipu AI has released GLM-5, a Mixture-of-Experts language model featuring sparse attention and a permissive MIT license for open-source development. - [inclusionAI's Ming 2.0 Tackles Any-to-Any Multimodality](https://theopenweights.com/news/ming-flash-omni-2-0-bg44) — inclusionAI · Any-to-Any · MIT · Feb 10, 2026: inclusionAI has released Ming-flash-omni 2.0, an open-source, any-to-any multimodal Mixture-of-Experts model for text, image, and audio processing. - [Nanbeige Releases 3B Chinese-Enhanced Language Model](https://theopenweights.com/news/nanbeige4-1-3b-thkp) — Nanbeige · Text / LLM · Other · Feb 10, 2026: Nanbeige has released Nanbeige4.1-3B, a 3-billion-parameter, Llama-based model trained on a 3.5T token bilingual dataset for enhanced Chinese performance. - [MOSS-TTS: A New Multilingual Text-to-Speech Model](https://theopenweights.com/news/moss-tts-j0m7) — OpenMOSS · Text → Speech · Other · Feb 6, 2026: The OpenMOSS Team has released MOSS-TTS, a multilingual text-to-speech model supporting Chinese, English, and Japanese with a novel architecture. - [Soul-AILab Releases Zero-Shot Singing Voice Model](https://theopenweights.com/news/soulx-singer-sp75) — Soul AILab · Music · Apache 2.0 · Feb 6, 2026: Soul-AILab has released SoulX-Singer, an open-source, zero-shot model for singing voice synthesis that supports both English and Chinese. - [OpenBMB Releases 'Any-to-Any' Multimodal Model](https://theopenweights.com/news/minicpm-o-4-5-xb5m) — OpenBMB · Any-to-Any · Other · Feb 3, 2026: OpenBMB has released MiniCPM-o 4.5, an open-source model for 'any-to-any' multimodal conversation across text, images, and audio. - [MiniCPM-o 4.5 Offers 'Any-to-Any' Multimodal AI](https://theopenweights.com/news/minicpm-o-4-5-mqtq) — OpenBMB · Any-to-Any · Apache 2.0 · Feb 2, 2026: OpenBMB has released MiniCPM-o 4.5, a full-duplex, 'any-to-any' multimodal model in the GGUF format for efficient local inference. - [Qwen Releases Coder-Next, A New Open MoE Coding Model](https://theopenweights.com/news/qwen3-coder-next-v7g5) — Qwen · Alibaba · Code · Apache 2.0 · Jan 30, 2026: Qwen, Alibaba's AI team, has released Qwen3-Coder-Next, an open-source Mixture-of-Experts (MoE) model for code generation under an Apache 2.0 license. - [Zhipu AI Releases Multilingual GLM-OCR Vision Model](https://theopenweights.com/news/glm-ocr-ax40) — Zhipu AI · Vision-Language · Other · Jan 30, 2026: Zhipu AI has released GLM-OCR, a new open vision-language model specialized for multilingual optical character recognition from images. - [Black Forest Labs Releases FLUX.2 Klein 9B](https://theopenweights.com/news/flux-2-klein-9b-p13j) — Black Forest Labs · Text → Image · Other · Jan 28, 2026: Black Forest Labs has released FLUX.2 Klein 9B, a 9B-parameter open-weight base model for text-to-image generation. - [OpenMOSS Releases MOVA, a 720p Multimodal Video Generator](https://theopenweights.com/news/mova-720p-q4hu) — OpenMOSS · Any-to-Any · Other · Jan 28, 2026: OpenMOSS has released MOVA 720p, a new open model that generates high-definition video with audio from both text and image inputs. - [Baidu Releases Open VLM for Advanced Document OCR](https://theopenweights.com/news/paddleocr-vl-1-5-hx3x) — Baidu · Vision-Language · Apache 2.0 · Jan 28, 2026: Baidu has released PaddleOCR-VL-1.5, an open vision-language model designed for complex document parsing, including tables, formulas, and charts. - [OpenMOSS Releases MOVA for Joint Video and Audio Gen](https://theopenweights.com/news/mova-360p-a2ai) — OpenMOSS · Image → Video · Other · Jan 28, 2026: OpenMOSS has released MOVA-360p, an open model for generating video and audio jointly from either text or image prompts at 360p resolution. - [Qwen Releases 0.6B Model for Audio-Text Alignment](https://theopenweights.com/news/qwen3-forcedaligner-0-6b-79f5) — Qwen · Alibaba · Speech → Text · Apache 2.0 · Jan 28, 2026: Alibaba's Qwen team has released Qwen3 ForcedAligner 0.6B, a compact open-source model for synchronizing audio recordings with existing text transcripts. - [Qwen3 Family Expands into Speech Recognition](https://theopenweights.com/news/qwen3-asr-1-7b-r2hk) — Qwen · Alibaba · Speech → Text · Apache 2.0 · Jan 28, 2026: Alibaba's Qwen team has released Qwen3-ASR-1.7B, a 1.7B-parameter open-source model for automatic speech recognition under the Apache 2.0 license. - [Qwen open-sources compact model for speech recognition](https://theopenweights.com/news/qwen3-asr-0-6b-wjhc) — Qwen · Alibaba · Speech → Text · Apache 2.0 · Jan 28, 2026: Alibaba's Qwen team has released Qwen3-ASR-0.6B, a compact 600M parameter model for automatic speech recognition, available under an Apache 2.0 license. - [DeepSeek-OCR-2 Tackles Multilingual Document AI](https://theopenweights.com/news/deepseek-ocr-2-lnoi) — DeepSeek · Vision-Language · Other · Jan 27, 2026: DeepSeek has released DeepSeek-OCR-2, a new open vision-language model designed for advanced, multilingual document understanding and text extraction. - [Lingbot-World Animates Images with Camera Control](https://theopenweights.com/news/lingbot-world-cam-a2dz) — robbyant · Image → Video · Apache 2.0 · Jan 26, 2026: Lingbot World Base Cam is a new open-source world model that generates camera-controllable video clips from a single static image. - [Alibaba's Qwen Team Releases Z-Image Diffusion Model](https://theopenweights.com/news/z-image-46ux) — Qwen · Alibaba · Text → Image · Apache 2.0 · Jan 23, 2026: Alibaba's Qwen team has released Z-Image, a new open-source text-to-image diffusion model available under the permissive Apache 2.0 license. - [LuxTTS Delivers Lightweight, Open-Source Speech Synthesis](https://theopenweights.com/news/luxtts-abdc) — YatharthS · Text → Speech · Apache 2.0 · Jan 22, 2026: LuxTTS, a new lightweight English text-to-speech model from the OpenMOSS community, has been released under an Apache 2.0 license and optimized for ONNX. - [Mistral Enters Speech AI with Voxtral Mini Model](https://theopenweights.com/news/voxtral-mini-4b-realtime-cb0s) — Mistral AI · Speech → Text · Apache 2.0 · Jan 21, 2026: Mistral AI has released Voxtral Mini, a 4B parameter open-source model for real-time, multilingual speech-to-text, marking its entry into audio AI. - [Microsoft Releases VibeVoice for Speech Transcription](https://theopenweights.com/news/vibevoice-asr-umgc) — Microsoft · Speech → Text · Other · Jan 21, 2026: Microsoft has released VibeVoice-ASR, a new speech-to-text model capable of transcription and speaker diarization in multiple languages for research use. - [Qwen Releases Open-Source Voice Cloning Model](https://theopenweights.com/news/qwen3-tts-0-6b-qp5z) — Qwen · Alibaba · Text → Speech · Apache 2.0 · Jan 21, 2026: Alibaba's Qwen team has released Qwen3-TTS, a 0.6B parameter open-source model for multilingual text-to-speech and voice cloning under an Apache 2.0 license. - [Qwen Releases a Compact Custom-Voice TTS Model](https://theopenweights.com/news/qwen3-tts-12hz-0-6b-customvoice-89qs) — Qwen · Alibaba · Text → Speech · Qwen · Jan 21, 2026: Qwen has released a 0.6B parameter text-to-speech model, Qwen3-TTS, which supports multilingual audio generation and custom voice cloning. - [Qwen Releases Open 1.7B Custom Voice Synthesis Model](https://theopenweights.com/news/qwen3-tts-1-7b-customvoice-ft6t) — Qwen · Alibaba · Text → Speech · Apache 2.0 · Jan 21, 2026: Alibaba's Qwen team has released Qwen3-TTS, a 1.7B parameter text-to-speech model that supports custom voice cloning from short audio samples. - [Qwen Unveils Open Model for Custom Voice Synthesis](https://theopenweights.com/news/qwen3-tts-12hz-1-7b-voicedesign-jyno) — Qwen · Alibaba · Text → Speech · Apache 2.0 · Jan 21, 2026: Alibaba's Qwen team has released Qwen3-TTS, a 1.7B-parameter open-source model for creating custom, multilingual voices from audio prompts. - [Zhipu AI Releases GLM-4.7-Flash MoE Model](https://theopenweights.com/news/glm-4-7-flash-x2iw) — Zhipu AI · Text / LLM · MIT · Jan 19, 2026: Zhipu AI has released GLM-4.7-Flash, a new lightweight Mixture-of-Experts language model with an MIT license, designed for fast inference. - [LightOn Releases OCR-2, a 1B Document AI Model](https://theopenweights.com/news/lightonocr-2-1b-rre8) — LightOn · Vision-Language · Other · Jan 16, 2026: LightOn has released LightOnOCR-2, a 1-billion-parameter vision model for document understanding. Based on Mistral, it extracts text from PDFs and forms. - [Black Forest Labs Releases 9B FLUX.2 Image Model](https://theopenweights.com/news/flux-2-klein-9b-4dhi) — Black Forest Labs · Text → Image · Other · Jan 14, 2026: Black Forest Labs has released FLUX.2-klein-base, a 9-billion-parameter text-to-image model optimized for efficiency with a unique transformer-based architecture. - [Soprano TTS Model Leverages Qwen3 Architecture](https://theopenweights.com/news/soprano-1-1-80m-z3jz) — ekwek · Text → Speech · Apache 2.0 · Jan 14, 2026: OpenMOSS has released Soprano-1.1-80M, a compact text-to-speech model based on the Qwen3 architecture and licensed under Apache 2.0 for open use. - [Black Forest Labs Releases Open-Source FLUX.2 Klein 4B](https://theopenweights.com/news/flux-2-klein-4b-co9g) — Black Forest Labs · Text → Image · Apache 2.0 · Jan 14, 2026: Black Forest Labs has released FLUX.2 klein, a 4B parameter distilled text-to-image and editing model with an open Apache 2.0 license. - [FLUX.2 Klein: A Compact 4B Open-Source Image Model](https://theopenweights.com/news/flux-2-klein-4b-rwg5) — Black Forest Labs · Text → Image · Apache 2.0 · Jan 14, 2026: Black Forest Labs has released FLUX.2 klein, a compact 4B-parameter open-source model for text-to-image generation based on the FLUX architecture. - [Black Forest Labs Releases 9B FLUX.2 Image Model](https://theopenweights.com/news/flux-2-klein-9b-zas5) — Black Forest Labs · Text → Image · Other · Jan 14, 2026: Black Forest Labs has released FLUX.2 klein, a new 9B parameter open-source text-to-image model based on the Diffusion Transformer (DiT) architecture. - [Black Forest Labs Releases New FLUX.2 Image Model](https://theopenweights.com/news/flux-2-klein-9b-k2lg) — Black Forest Labs · Text → Image · Other · Jan 14, 2026: Black Forest Labs has released FLUX.2 klein, a 9B parameter text-to-image model featuring a new architecture for fast, high-quality image generation. - [Hume AI Releases TADA 1B for Expressive Speech](https://theopenweights.com/news/tada-1b-6ir5) — HumeAI · Text → Speech · Llama Community · Jan 12, 2026: Hume AI has released TADA 1B, a 1B-parameter text-to-speech model built on Llama 3.2 that aims to generate more expressive and realistic audio. - [Google Releases TranslateGemma for Open Translation](https://theopenweights.com/news/translategemma-4b-8rdp) — Google DeepMind · Text / LLM · Gemma · Jan 12, 2026: Google has released TranslateGemma, a 4-billion-parameter model based on its Gemma architecture that is instruction-tuned for text translation. - [OpenMOSS Releases KugelAudio for European Languages](https://theopenweights.com/news/kugelaudio-0-open-f3dx) — Kugelaudio · Text → Speech · Other · Jan 11, 2026: OpenMOSS has released KugelAudio-0-open, a new text-to-speech model for European languages using a hybrid diffusion and autoregressive approach. - [Zhipu AI Releases Open, Bilingual GLM-Image Model](https://theopenweights.com/news/glm-image-vx9z) — Zhipu AI · Text → Image · MIT · Jan 8, 2026: Zhipu AI has released GLM-Image, an open-source text-to-image diffusion model notable for its bilingual proficiency in both Chinese and English. - [Google's MedGemma brings open vision AI to medicine](https://theopenweights.com/news/medgemma-1-5-4b-addh) — Google DeepMind · Vision-Language · Gemma · Jan 7, 2026: Google has released MedGemma 1.5, a 4B parameter open vision-language model fine-tuned for medical imaging and clinical reasoning tasks. - [Supertone Open-Sources Supertonic 2 Voice Model](https://theopenweights.com/news/supertonic-2-om5d) — Supertone · Text → Speech · OpenRAIL-M · Jan 6, 2026: AI audio company Supertone has released Supertonic 2, a multilingual, open-source text-to-speech model optimized for deployment with ONNX. - [Lightricks Releases LTX-2 Multimodal Video Generator](https://theopenweights.com/news/ltx-2-8nv7) — Lightricks · Image → Video · Other · Jan 3, 2026: Lightricks has released LTX-2, a versatile video synthesis model that accepts text, image, and audio prompts under a Responsible AI license. - [Moonshot AI Releases Kimi K2.5 Multimodal Model](https://theopenweights.com/news/kimi-k2-5-1iqs) — Moonshot AI · Vision-Language · Other · Jan 1, 2026: Moonshot AI has released Kimi K2.5, a new multimodal Mixture-of-Experts model with advanced vision-language and reasoning capabilities. - [Qwen Releases Bilingual Open-Source Image Model](https://theopenweights.com/news/qwen-image-2512-u4sd) — Qwen · Alibaba · Text → Image · Apache 2.0 · Dec 30, 2025: Alibaba's Qwen team has released Qwen-Image 2512, a new open-source text-to-image model with strong support for English and Chinese prompts. - [Qwen's Fun-Audio-Chat: An Open Speech-to-Speech LLM](https://theopenweights.com/news/fun-audio-8b-j2op) — Qwen · Alibaba · Any-to-Any · Apache 2.0 · Dec 23, 2025: Alibaba's Qwen team has released Fun-Audio-Chat-8B, an 8-billion-parameter audio-language model capable of speech-to-speech chat in English and Chinese. - [MiniMax Debuts M2.1, an MoE Model Optimized with FP8](https://theopenweights.com/news/minimax-m2-1-cvvd) — MiniMax · Text / LLM · Other · Dec 20, 2025: Chinese AI company MiniMax has released MiniMax-M2.1, a Mixture of Experts language model that utilizes FP8 weights for improved efficiency. - [Google Releases MedASR for Medical Transcription](https://theopenweights.com/news/medasr-gnc6) — Google DeepMind · Speech → Text · Other · Dec 18, 2025: Google DeepMind has released MedASR, an automatic speech recognition model specialized for transcribing medical dictation, particularly in radiology. - [MiraTTS Brings Qwen2 to Bilingual Speech Synthesis](https://theopenweights.com/news/miratts-fmsw) — YatharthS · Text → Speech · CC BY-NC 4.0 · Dec 17, 2025: MiraTTS is a new open-source text-to-speech model based on Qwen2. It supports English and Chinese and can be exported to ONNX for efficient inference. - [Soprano-80M: A Tiny TTS Model Based on Qwen3](https://theopenweights.com/news/soprano-80m-3zyt) — ekwek · Text → Speech · Apache 2.0 · Dec 17, 2025: Developer 'ekwek' has released Soprano-80M, a compact 80M-parameter text-to-speech model based on the Qwen3 architecture and licensed under Apache 2.0. - [Qwen Releases Open, Bilingual Image Editing Model](https://theopenweights.com/news/qwen-image-edit-2511-pnl9) — Qwen · Alibaba · Image Editing · Apache 2.0 · Dec 17, 2025: Alibaba's Qwen team has released Qwen-Image-Edit, an open-source diffusion model for instruction-based image editing in English and Chinese. - [NVIDIA Releases Streaming Speech-to-Text Model](https://theopenweights.com/news/nemotron-speech-streaming-en-0-6b-t8gq) — NVIDIA · Speech → Text · Other · Dec 17, 2025: NVIDIA has released Nemotron Speech Streaming EN, a 600M parameter model for low-latency English speech recognition for real-time applications. - [Qwen Releases Compact ASR Model for Streaming Audio](https://theopenweights.com/news/fun-asr-nano-2512-i9r7) — Qwen · Alibaba · Speech → Text · Other · Dec 15, 2025: Alibaba's Qwen has released Fun-ASR-Nano-2512, a compact model for real-time, multilingual speech-to-text with advanced diarization and hotword features. - [PersonaLive Model Animates Portraits in Real Time](https://theopenweights.com/news/personalive-9tsf) — Huaichang · Image → Video · Apache 2.0 · Dec 13, 2025: OpenBMB has released PersonaLive, an open-source model that uses diffusion techniques to create real-time portrait animations from a single photo. - [Tencent's HY-WorldPlay Creates 3D Scenes from One Image](https://theopenweights.com/news/hy-worldplay-h58y) — Tencent · Image → Video · Other · Dec 12, 2025: Tencent has released HY-WorldPlay, an AI world model that generates video and reconstructs 3D scenes from a single image. Released for research use. - [Alibaba Releases CosyVoice 3 for Expressive TTS](https://theopenweights.com/news/fun-cosyvoice3-0-5b-ee4a) — Qwen · Alibaba · Text → Speech · Other · Dec 11, 2025: Alibaba's FunAudioLLM team has released CosyVoice 3, a 500M-parameter text-to-speech model with multilingual support and zero-shot voice cloning capabilities. - [Zhipu AI Releases GLM-TTS for Zero-Shot Voice Cloning](https://theopenweights.com/news/glm-tts-9alc) — Zhipu AI · Text → Speech · Other · Dec 10, 2025: Zhipu AI has released GLM-TTS, a zero-shot voice cloning model for Chinese and English that uses flow matching and reinforcement learning techniques. - [Zhipu AI Releases Compact Bilingual Speech Model](https://theopenweights.com/news/glm-asr-nano-2512-jz26) — Zhipu AI · Speech → Text · MIT · Dec 9, 2025: Zhipu AI has released GLM-ASR-Nano, a compact model for automatic speech recognition in English and Chinese. It is available under a permissive MIT license. - [Zhipu AI Releases Fast, Open Vision Model GLM-4.6V-Flash](https://theopenweights.com/news/glm-4-6v-flash-9rt7) — Zhipu AI · Vision-Language · MIT · Dec 7, 2025: Zhipu AI has released GLM-4.6V-Flash, a new open-source vision-language model under the MIT license, designed for fast multimodal applications. - [VoxCPM 1.5 Brings Open-Source Voice Cloning](https://theopenweights.com/news/voxcpm-1-5-hqs2) — OpenBMB · Text → Speech · Apache 2.0 · Dec 5, 2025: OpenBMB released VoxCPM 1.5, a 500M parameter open-source TTS model with zero-shot voice cloning capabilities for English and Chinese. - [Meituan Releases Open, Bilingual Image Editing Model](https://theopenweights.com/news/longcat-image-edit-4ns6) — Meituan · Image Editing · Apache 2.0 · Dec 5, 2025: Meituan has released LongCat-Image-Edit, an open-source model for instruction-based image editing that supports both English and Chinese prompts. - [Baidu's Live-Avatar Animates Photos With Audio](https://theopenweights.com/news/live-avatar-xlen) — Quark Vision · Image → Video · Apache 2.0 · Dec 4, 2025: Baidu has released Live-Avatar, a 14B parameter open-source model that generates talking head videos from a single image and an audio track. - [Microsoft Releases VibeVoice for Real-Time AI Speech](https://theopenweights.com/news/vibevoice-realtime-0-5b-ktrl) — Microsoft · Text → Speech · Other · Dec 4, 2025: Microsoft has released VibeVoice-Realtime-0.5B, a model for low-latency, streaming text-to-speech built on Alibaba's Qwen2.5-0.5B foundation. - [Resemble AI Releases Chatterbox Turbo for Open TTS](https://theopenweights.com/news/chatterbox-turbo-gg2x) — Resemble AI · Text → Speech · MIT · Dec 2, 2025: Resemble AI has released Chatterbox Turbo, a faster open-source text-to-speech model for English that supports voice cloning and speech generation. - [DeepSeek-V3.2 Arrives With FP8 Weights, MIT License](https://theopenweights.com/news/deepseek-v3-2-8ghv) — DeepSeek · Text / LLM · MIT · Dec 1, 2025: DeepSeek has released DeepSeek-V3.2, a Mixture-of-Experts LLM with efficient FP8 weights and a permissive MIT license for broad commercial use. - [FlashLabs Releases Chroma-4B, an Any-to-Any Model](https://theopenweights.com/news/chroma-4b-ancy) — FlashLabs · Any-to-Any · Apache 2.0 · Nov 28, 2025: FlashLabs has released Chroma-4B, a 4B-parameter open model capable of processing text, images, and audio in any combination. - [Alibaba Releases Z-Image-Turbo, A Fast Open Image Model](https://theopenweights.com/news/z-image-turbo-kk9g) — Qwen · Alibaba · Text → Image · Apache 2.0 · Nov 25, 2025: Alibaba's Qwen team has released Z-Image-Turbo, a fast text-to-image model that generates 1024px images in a single step under an Apache 2.0 license. - [Black Forest Labs Releases Open-Source FLUX.2 Image Model](https://theopenweights.com/news/flux-2-dev-x9yr) — Black Forest Labs · Text → Image · Other · Nov 22, 2025: Black Forest Labs has released FLUX.2-dev, a new open-weight text-to-image and image editing model, as a developer preview for community testing. - [Tencent Releases HunyuanVideo 1.5 Generation Model](https://theopenweights.com/news/hunyuanvideo-1-5-1w06) — Tencent · Text → Video · Other · Nov 18, 2025: Tencent has released HunyuanVideo 1.5, an open-weights model for generating video from text and images under a custom license. - [Tencent Releases 1B Parameter HunyuanOCR Model](https://theopenweights.com/news/hunyuanocr-bksf) — Tencent · Vision-Language · Other · Nov 18, 2025: Tencent has released HunyuanOCR, a 1-billion-parameter vision-language model for end-to-end optical character recognition under an Apache 2.0 license. - [Mistral AI Releases Voxtral, an Open-Source TTS Model](https://theopenweights.com/news/voxtral-4b-tts-ton3) — Mistral AI · Text → Speech · Apache 2.0 · Nov 17, 2025: Mistral AI has released Voxtral 4B TTS, a 4-billion-parameter, multilingual text-to-speech model under a permissive Apache 2.0 license. - [Nari Labs Releases Dia2-2B, an Open Voice Cloning Model](https://theopenweights.com/news/dia2-2b-qszq) — Nari Labs · Text → Speech · Apache 2.0 · Nov 15, 2025: Nari Labs has released Dia2-2B, a 2-billion-parameter open-source text-to-speech model capable of high-quality, zero-shot voice cloning. - [Meta releases SAM 3 for image and video segmentation](https://theopenweights.com/news/sam-3-9gib) — Meta AI · Vision-Language · Other · Nov 7, 2025: Meta has released SAM 3, its newest Segment Anything Model for generating object masks across both images and video. - [Baidu Releases Open Vision-Language MoE Model](https://theopenweights.com/news/ernie-4-5-vl-28b-a3b-thinking-86bx) — Baidu · Vision-Language · Apache 2.0 · Nov 7, 2025: Baidu has released ERNIE 4.5 VL, a 28B-parameter open vision-language model using a Mixture-of-Experts architecture for complex reasoning tasks. - [Moonshot AI Releases Kimi-K2 Reasoning Model](https://theopenweights.com/news/kimi-k2-thinking-nc9g) — Moonshot AI · Reasoning · Other · Nov 4, 2025: Moonshot AI has released Kimi-K2-Thinking, a Mixture-of-Experts model for complex reasoning, distributed in a unique compressed format with a custom license. - [BAAI Releases Emu3.5, an 'Any-to-Any' Multimodal Model](https://theopenweights.com/news/emu3-5-z7ca) — BAAI · Any-to-Any · Apache 2.0 · Oct 31, 2025: The Allen Institute for AI has released Emu3.5, an open-source multimodal model designed for unified 'any-to-any' generation and understanding. - [Microsoft Releases Fara-7B Vision Agent Model](https://theopenweights.com/news/fara-7b-vccl) — Microsoft · Vision-Language · MIT · Oct 30, 2025: Microsoft has released Fara-7B, a 7B vision-language model based on Qwen2.5-VL, designed as an agent for interacting with computer GUIs. - [SoulX-Podcast 1.7B Offers Open Multi-Speaker TTS](https://theopenweights.com/news/soulx-podcast-1-7b-m4wq) — Soul AILab · Text → Speech · Apache 2.0 · Oct 27, 2025: OpenMOSS has released SoulX-Podcast 1.7B, an open-source model for generating conversational, multi-speaker text-to-speech audio in English and Chinese. - [Meituan Releases Open-Source LongCat-Video Model](https://theopenweights.com/news/longcat-video-myk7) — Meituan · Text → Video · MIT · Oct 24, 2025: Meituan has released LongCat-Video, an open-source model under the MIT license for generating video from text, images, or existing video clips. - [Meituan Debuts LongCat-Flash-Omni, an Any-to-Any AI Model](https://theopenweights.com/news/longcat-flash-omni-wk4l) — Meituan · Any-to-Any · MIT · Oct 23, 2025: Meituan has released LongCat-Flash-Omni, an open-source, any-to-any multimodal AI model that uses a Mixture-of-Experts (MoE) architecture. - [NVIDIA Releases Real-Time Speaker Diarization Model](https://theopenweights.com/news/streaming-sortformer-diarization-4spk-v2-1-3i55) — NVIDIA · Speech → Text · Other · Oct 22, 2025: NVIDIA has released a new Sortformer model for real-time speaker diarization, capable of distinguishing up to four speakers in a continuous audio stream. - [MiniMax Releases M2, an Open-Weight MoE for Agents](https://theopenweights.com/news/minimax-m2-uf5z) — MiniMax · Text / LLM · Other · Oct 22, 2025: Shanghai-based MiniMax has released MiniMax-M2, a new open-weight Mixture-of-Experts LLM designed for strong performance in reasoning, coding, and agentic tasks. - [Datalab Releases Chandra, a New OCR Vision Model](https://theopenweights.com/news/chandra-ocr-4f3h) — Datalab To · Vision-Language · OpenRAIL-M · Oct 21, 2025: Datalab has released Chandra, an open-source vision-language model based on Qwen2-VL, specialized for high-performance OCR and document parsing tasks. - [Kling Releases UniVideo for Generation and Understanding](https://theopenweights.com/news/univideo-5yrf) — Kuaishou · Any-to-Any · Apache 2.0 · Oct 18, 2025: Kling has released UniVideo, an open-source model that can both generate and understand video content. It's based on Qwen2.5-VL-7B and is licensed under Apache 2.0. - [Maya Research Releases Maya1, an Expressive TTS Model](https://theopenweights.com/news/maya1-5yhf) — Maya Research · Text → Speech · Apache 2.0 · Oct 18, 2025: Maya Research has released Maya1, a new open-source, Llama-based model for generating expressive text-to-speech audio under an Apache 2.0 license. - [DeepSeek-OCR Tackles Document Parsing with Vision AI](https://theopenweights.com/news/deepseek-ocr-hn77) — DeepSeek · Vision-Language · MIT · Oct 17, 2025: DeepSeek has released DeepSeek-OCR, an open-source vision-language model designed for efficient document parsing using optical context compression. - [Baidu Releases PaddleOCR-VL for Document AI](https://theopenweights.com/news/paddleocr-vl-qu3n) — Baidu · Vision-Language · Other · Oct 16, 2025: Baidu has released PaddleOCR-VL, a new open-source vision-language model based on ERNIE 4.5 for advanced document understanding and OCR. - [NVIDIA's Parakeet ASR Tackles Multi-Speaker Audio](https://theopenweights.com/news/multitalker-parakeet-streaming-0-6b-j759) — NVIDIA · Speech → Text · Other · Oct 15, 2025: NVIDIA has released Multitalker Parakeet Streaming 0.6B, an ASR model for real-time transcription and speaker diarization in multi-person conversations. - [inclusionAI Debuts 'Any-to-Any' Multimodal MoE Model](https://theopenweights.com/news/ming-flash-omni-ruk4) — inclusionAI · Any-to-Any · MIT · Oct 14, 2025: inclusionAI has released Ming-flash-omni-Preview, a new open-source, any-to-any multimodal model using a Mixture of Experts (MoE) architecture. - [Alibaba Releases Qwen3-VL, an 8B Open-Source Vision Model](https://theopenweights.com/news/qwen3-vl-8b-15l9) — Qwen · Alibaba · Vision-Language · Apache 2.0 · Oct 11, 2025: Alibaba's Qwen team has released Qwen3-VL-8B-Instruct, an 8-billion-parameter open-source vision-language model tuned for following instructions. - [Google Releases Compact FunctionGemma Model](https://theopenweights.com/news/functiongemma-270m-bxym) — Google DeepMind · Text / LLM · Gemma · Oct 8, 2025: Google has released FunctionGemma, a 270M-parameter model from the Gemma family, specifically optimized for function calling and integrating with external tools. - [EPFL Releases SVI for Streaming Image-to-Video](https://theopenweights.com/news/svi-dcx0) — EPFL VITA · Image → Video · MIT · Oct 8, 2025: EPFL's VITA lab has released SVI, an open-source model designed to generate long, coherent videos from a single image using a streaming, chunk-based approach. - [Krea Releases Open-Source Real-Time Video Model](https://theopenweights.com/news/krea-realtime-video-nl47) — Krea · Text → Video · Apache 2.0 · Oct 8, 2025: Krea has released an open-source, 14B-parameter model for real-time text-to-video and video-to-video generation under an Apache 2.0 license. - [inclusionAI Releases Ming-UniVision MoE Multimodal Model](https://theopenweights.com/news/ming-univision-16b-a3b-d0un) — inclusionAI · Any-to-Any · Apache 2.0 · Sep 30, 2025: inclusionAI has released Ming-UniVision, a 16B parameter open-source multimodal model using a Mixture-of-Experts architecture with 3B active parameters. - [Qwen Releases 30B MoE Vision Model, Qwen3-VL](https://theopenweights.com/news/qwen3-vl-30b-a3b-8y0y) — Qwen · Alibaba · Vision-Language · Apache 2.0 · Sep 30, 2025: Alibaba's Qwen team has released Qwen3-VL, a 30B Mixture-of-Experts vision-language model with 3B active parameters, available under an Apache 2.0 license. - [Kani TTS 370M Offers Compact Multilingual Speech](https://theopenweights.com/news/kani-tts-370m-6fuf) — nineninesix · Text → Speech · Other · Sep 30, 2025: Kani TTS 370M is a new, compact text-to-speech model for multilingual audio generation, built on the LFM2 architecture and available on Hugging Face. - [Ovi Syncs Audio and Video in New Open-Source Model](https://theopenweights.com/news/ovi-ulgk) — chetwinlow1 · Image → Video · Apache 2.0 · Sep 30, 2025: Ovi is a new 5B-parameter open-source model that generates video from a single image, complete with synchronized audio, under an Apache 2.0 license. - [Zhipu AI Releases Open-Weight MoE Model GLM-4.6](https://theopenweights.com/news/glm-4-6-zatb) — Zhipu AI · Text / LLM · MIT · Sep 29, 2025: Zhipu AI has released GLM-4.6, a Mixture-of-Experts (MoE) language model with open weights available under the permissive MIT license. - [Ming-UniAudio Brings MoE to Unified Audio AI](https://theopenweights.com/news/ming-uniaudio-16b-a3b-z4hb) — inclusionAI · Any-to-Any · Apache 2.0 · Sep 29, 2025: inclusionAI has released Ming-UniAudio, a 16B parameter Mixture-of-Experts model for unified audio tasks like speech recognition, translation, and TTS. - [ByteDance Releases Lynx for Identity-Preserving Video](https://theopenweights.com/news/lynx-cl2f) — ByteDance · Image → Video · Apache 2.0 · Sep 26, 2025: ByteDance has released Lynx, an open-source model for generating personalized, identity-preserving video clips from a single source image. - [Tencent Debuts HunyuanImage 3.0 with MoE Design](https://theopenweights.com/news/hunyuanimage-3-0-26kv) — Tencent · Text → Image · Other · Sep 25, 2025: Tencent has released HunyuanImage 3.0, a new text-to-image model using a Mixture-of-Experts architecture for improved instruction following. - [Tencent Releases HunyuanImage 3.0 Text-to-Image Model](https://theopenweights.com/news/hunyuanimage-3-0-g8fp) — Tencent · Text → Image · Other · Sep 25, 2025: Tencent has released the weights for HunyuanImage 3.0, a new text-to-image model that uses a Mixture-of-Experts (MoE) architecture for generation. - [Qwen Releases Open-Source Instruction-Based Image Editor](https://theopenweights.com/news/qwen-image-edit-2509-208u) — Qwen · Alibaba · Image Editing · Apache 2.0 · Sep 22, 2025: Alibaba's Qwen team has released Qwen-Image-Edit, a new open-source model that enables fine-grained image modification through text instructions. - [Qwen3-Omni Arrives With Any-to-Any Multimodality](https://theopenweights.com/news/qwen3-omni-30b-a3b-nwcf) — Qwen · Alibaba · Any-to-Any · Other · Sep 20, 2025: Alibaba's Qwen team has released Qwen3-Omni, a 30B parameter MoE model capable of any-to-any multimodal tasks involving text, vision, and audio. - [Xiaomi's MiMo-Audio 7B Tackles Complex Speech Tasks](https://theopenweights.com/news/mimo-audio-7b-udzb) — Xiaomi · Any-to-Any · MIT · Sep 18, 2025: Xiaomi has released MiMo-Audio-7B-Instruct, a 7-billion parameter, MIT-licensed model for versatile audio-text tasks like TTS and ASR. - [OpenBMB Releases VoxCPM for Open Voice Synthesis](https://theopenweights.com/news/voxcpm-0-5b-32pa) — OpenBMB · Text → Speech · Apache 2.0 · Sep 16, 2025: OpenBMB has released VoxCPM-0.5B, a 500M-parameter model for open text-to-speech and zero-shot voice cloning, available under an Apache 2.0 license. - [Qwen Releases 'Thinking' Multimodal MoE Model](https://theopenweights.com/news/qwen3-omni-30b-a3b-thinking-1f1p) — Qwen · Alibaba · Any-to-Any · Other · Sep 15, 2025: Alibaba's Qwen team has released Qwen3-Omni, a 30B parameter Mixture-of-Experts model that outputs its chain-of-thought process for omni-modal reasoning tasks. - [Qwen Releases 30B Model for Audio Captioning](https://theopenweights.com/news/qwen3-omni-30b-a3b-captioner-hnhq) — Qwen · Alibaba · Any-to-Any · Other · Sep 15, 2025: Alibaba's Qwen team has released a 30B parameter omni-modal model specialized for multilingual audio captioning, using an efficient MoE architecture. - [Neuphonic Releases NeuTTS Air for On-Device AI Speech](https://theopenweights.com/news/neutts-air-g1ab) — neuphonic · Text → Speech · Apache 2.0 · Sep 15, 2025: Neuphonic has released NeuTTS Air, an Apache 2.0 licensed text-to-speech model with voice cloning capabilities, optimized for on-device use. - [Moondream 3 Arrives in Preview Release](https://theopenweights.com/news/moondream-3-jbiz) — moondream · Vision-Language · Other · Sep 11, 2025: The preview for Moondream 3, a new 4-billion-parameter open-source vision-language model, has been released for community evaluation. - [ByteDance Releases HuMo for Human Video Generation](https://theopenweights.com/news/humo-f7h7) — ByteDance · Image → Video · Apache 2.0 · Sep 10, 2025: ByteDance has released HuMo, an open-source model for generating high-quality videos of humans from text descriptions or still images. - [Alibaba's Wan2.2 Adds Control to Open Video](https://theopenweights.com/news/wan2-2-vace-fun-a14b-pog2) — Qwen · Alibaba · Text → Video · Apache 2.0 · Sep 10, 2025: Alibaba has released Wan2.2-VACE-Fun-A14B, a 14B parameter open-source text-to-video model designed for highly controllable video synthesis. - [Qwen Releases 80B Mixture-of-Experts Model](https://theopenweights.com/news/qwen3-next-80b-a3b-dsac) — Qwen · Alibaba · Text / LLM · Apache 2.0 · Sep 9, 2025: Qwen has released Qwen3-Next-80B, an instruction-tuned Mixture-of-Experts model with 80B total parameters and 3B active parameters per token. - [Lumina-DiMOO: A Diffusion Model for Any-to-Any AI](https://theopenweights.com/news/lumina-dimoo-956h) — Alpha-VLLM · Any-to-Any · Apache 2.0 · Sep 9, 2025: Lumina-DiMOO is a new open-source model using a diffusion architecture for "any-to-any" multimodal tasks, capable of both understanding and generation. - [Tencent SRPO Fine-Tunes SDXL with Preference Optimization](https://theopenweights.com/news/srpo-ow7r) — Tencent · Text → Image · Other · Sep 8, 2025: Tencent has released SRPO, a fine-tuned Stable Diffusion XL model that uses a new preference optimization method to improve image quality and prompt alignment. - [Tencent Releases HunyuanImage 2.1 for Bilingual AI Art](https://theopenweights.com/news/hunyuanimage-2-1-fu6q) — Tencent · Text → Image · Other · Sep 5, 2025: Tencent has released HunyuanImage 2.1, a text-to-image model specializing in bilingual (Chinese/English) prompts and high-resolution 1024x1024 output. - [Microsoft Releases VibeVoice, a 7B Podcast TTS Model](https://theopenweights.com/news/vibevoice-7b-rrx3) — Vibevoice · Text → Speech · MIT · Sep 4, 2025: Microsoft has released VibeVoice-7B, an open-source model for generating long-form, multi-speaker, podcast-style audio in English and Chinese. - [Microsoft Releases VibeVoice, a Podcast-Ready TTS Model](https://theopenweights.com/news/vibevoice-large-ern9) — Aoi Ot · Text → Speech · MIT · Sep 4, 2025: Microsoft has released VibeVoice Large, an open-source text-to-speech model designed for generating long-form, multi-speaker audio for podcasts in EN/ZH. - [StepFun Releases Step-Audio 2 mini, a Unified Audio AI](https://theopenweights.com/news/step-audio-2-mini-dqvi) — StepFun · Any-to-Any · Apache 2.0 · Aug 28, 2025: StepFun has released Step-Audio 2 mini, a compact model for both speech-to-text and text-to-speech, available under an Apache 2.0 license. - [Tencent's Voyager Model Turns Images into 3D Worlds](https://theopenweights.com/news/hunyuanworld-voyager-qrnp) — Tencent · Image → Video · Other · Aug 27, 2025: Tencent releases HunyuanWorld-Voyager, a new world model that generates 3D-consistent videos from a single image, allowing for scene exploration. - [Microsoft Releases VibeVoice for Long-Form Audio](https://theopenweights.com/news/vibevoice-1-5b-mnfp) — Microsoft · Text → Speech · MIT · Aug 25, 2025: Microsoft has released VibeVoice-1.5B, an open-source text-to-speech model for generating long-form, multi-speaker audio in English and Chinese. - [Alibaba Releases 14B Model for Audio-Driven Video](https://theopenweights.com/news/wan2-2-s2v-14b-ih9s) — Qwen · Alibaba · Image → Video · Apache 2.0 · Aug 25, 2025: Alibaba's Qwen team has released Wan2.2-S2V-14B, a 14-billion parameter open model that animates still images based on an audio speech track. - [OpenBMB Releases Compact Multimodal Model MiniCPM-V 4.5](https://theopenweights.com/news/minicpm-v-4-5-id6s) — OpenBMB · Vision-Language · Other · Aug 24, 2025: OpenBMB has released MiniCPM-V 4.5, a compact vision-language model with advanced capabilities in OCR, multi-image processing, and video understanding. - [DeepSeek Releases 671B MoE Model Under MIT License](https://theopenweights.com/news/deepseek-v3-1-n5kg) — DeepSeek · Text / LLM · MIT · Aug 19, 2025: DeepSeek has released DeepSeek-V3.1-Base, a 671B parameter Mixture-of-Experts model with an MIT license, aimed at large-scale AI development. - [Qwen Releases Open Model for Image Editing](https://theopenweights.com/news/qwen-image-edit-aq8z) — Qwen · Alibaba · Image Editing · Apache 2.0 · Aug 17, 2025: Alibaba's Qwen team has released Qwen-Image-Edit, an open-source model for instruction-based image editing that supports English and Chinese prompts. - [NexaAI Releases OmniNeural-4B for On-Device AI](https://theopenweights.com/news/omnineural-4b-bi7v) — NexaAI · Any-to-Any · CC BY 4.0 · Aug 15, 2025: NexaAI has released OmniNeural-4B, a 4B parameter 'any-to-any' multimodal model optimized for on-device performance on Android and Snapdragon NPUs. - [Tencent Releases Controllable Game Video Model](https://theopenweights.com/news/hunyuan-gamecraft-1-0-niqn) — Tencent · Image → Video · Other · Aug 13, 2025: Tencent has released Hunyuan-GameCraft 1.0, an open image-to-video model designed to generate controllable, interactive video resembling game footage. - [StableAvatar Brings Open Source Talking Heads to Life](https://theopenweights.com/news/stableavatar-tl2x) — FrancisRing · Image → Video · MIT · Aug 12, 2025: StableAvatar is a new, MIT-licensed video diffusion model that generates talking-head animations from a single image and an audio file. - [Zhipu AI Releases Open Vision Model GLM-4.5V](https://theopenweights.com/news/glm-4-5v-6tcq) — Zhipu AI · Vision-Language · MIT · Aug 10, 2025: Zhipu AI has released GLM-4.5V, an open-source vision-language model using a Mixture-of-Experts architecture for advanced multimodal reasoning. - [Skywork Releases Open 'World Model' for Playable Video](https://theopenweights.com/news/matrix-game-2-0-ubi1) — Skywork · Image → Video · MIT · Aug 8, 2025: Skywork has released Matrix-Game 2.0, a 1.3B parameter open-source model that generates interactive, controllable video from a single static image. - [Google Releases Gemma 3 270M for On-Device AI](https://theopenweights.com/news/gemma-3-270m-2gln) — Google DeepMind · Text / LLM · Gemma · Aug 5, 2025: Google has released Gemma 3 270M, a 270-million-parameter text model with a 32K context window, optimized for on-device use and fine-tuning. - [OpenAI Releases 21B Open-Weight MoE Model](https://theopenweights.com/news/gpt-oss-20b-irah) — OpenAI · Reasoning · Apache 2.0 · Aug 4, 2025: OpenAI has released gpt-oss-20b, a ~21B parameter Mixture-of-Experts reasoning model under an Apache 2.0 license, designed for consumer hardware. - [OpenAI Releases Its First Open-Source MoE Model](https://theopenweights.com/news/gpt-oss-120b-35je) — OpenAI · Reasoning · Apache 2.0 · Aug 4, 2025: OpenAI has released gpt-oss-120b, its first open-weight model. The 117B parameter MoE is licensed under Apache 2.0 and excels at reasoning tasks. - [NVIDIA Releases Canary 1B v2 Multilingual Speech Model](https://theopenweights.com/news/canary-1b-v2-9jie) — NVIDIA · Speech → Text · CC BY 4.0 · Aug 4, 2025: NVIDIA has released Canary 1B v2, a 1B-parameter open model for multilingual speech recognition and speech-to-text translation. - [NVIDIA Releases 600M Parakeet for Speech Recognition](https://theopenweights.com/news/parakeet-tdt-0-6b-v3-p9bm) — NVIDIA · Speech → Text · CC BY 4.0 · Aug 4, 2025: NVIDIA has released Parakeet TDT 0.6B, a new 600-million-parameter multilingual model for automatic speech recognition, available under a CC-BY-4.0 license. - [Qwen releases open model for text-in-image generation](https://theopenweights.com/news/qwen-image-o5sj) — Qwen · Alibaba · Text → Image · Apache 2.0 · Aug 2, 2025: Qwen-Image is a new open-source text-to-image model from Alibaba specializing in generating images with legible text in English and Chinese. - [Qwen Releases Compact 30B MoE for Coding Agents](https://theopenweights.com/news/qwen3-coder-30b-a3b-lsuz) — Qwen · Alibaba · Code · Apache 2.0 · Jul 31, 2025: Qwen has released a 30B parameter MoE model with 3B active parameters, designed for efficient, agent-driven code generation under an Apache 2.0 license. - [New VLM `dots.ocr` Takes on Complex Documents](https://theopenweights.com/news/dots-ocr-ix9f) — rednote-hilab · Vision-Language · Other · Jul 30, 2025: Rednote-hilab has released dots.ocr, a 3B-parameter vision-language model specialized for parsing complex document layouts, tables, and formulas. - [Skywork Releases UniPic, a Unified 1.5B Vision Model](https://theopenweights.com/news/skywork-unipic-1-5b-hk15) — Skywork · Any-to-Any · Other · Jul 29, 2025: Skywork has released UniPic-1.5B, a unified autoregressive model that handles image understanding, generation, and editing in a single 1.5B parameter framework. - [Alibaba Releases Wan2.2, a 14B MoE Video Model](https://theopenweights.com/news/wan2-2-i2v-a14b-fek6) — Qwen · Alibaba · Image → Video · Apache 2.0 · Jul 28, 2025: Alibaba's Qwen team has released Wan2.2-I2V-A14B, a 14B parameter open-source model that uses a Mixture of Experts architecture for image-to-video generation. - [Tencent Releases Wan2.2, a 14B MoE Video Model](https://theopenweights.com/news/wan2-2-t2v-a14b-ww8a) — Tencent · Text → Video · Apache 2.0 · Jul 28, 2025: Tencent has released Wan2.2, a 14-billion parameter text-to-video model using a Mixture-of-Experts architecture, now available on Hugging Face. - [Qwen Releases Wan2.2, a 5B Open-Source Video Model](https://theopenweights.com/news/wan2-2-ti2v-5b-szwm) — Qwen · Alibaba · Text → Video · Apache 2.0 · Jul 28, 2025: Alibaba's Qwen team has released Wan2.2-TI2V-5B, a 5-billion-parameter open model for generating video from text prompts or images, licensed under Apache 2.0. - [Qwen Unveils Wan2.2, a 14B Open Text-to-Video Model](https://theopenweights.com/news/wan2-2-t2v-a14b-s7h0) — Qwen · Alibaba · Text → Video · Apache 2.0 · Jul 24, 2025: Alibaba's Qwen team has released Wan2.2-T2V-A14B, a 14B open-source Mixture-of-Experts model for text-to-video generation under an Apache 2.0 license. - [Qwen Releases Wan2.2, a 14B Image-to-Video Model](https://theopenweights.com/news/wan2-2-i2v-a14b-ic5j) — Qwen · Alibaba · Image → Video · Apache 2.0 · Jul 24, 2025: Alibaba's Qwen team has released Wan2.2-I2V-A14B, a 14B parameter, Apache 2.0-licensed, open-source model for generating video from a still image. - [Qwen Releases 480B Open-Source Model for Code Agents](https://theopenweights.com/news/qwen3-coder-480b-a35b-74pv) — Qwen · Alibaba · Code · Apache 2.0 · Jul 22, 2025: Alibaba's Qwen team has released Qwen3-Coder, a 480B parameter Mixture-of-Experts model for agentic coding with a 262K context length and an Apache-2.0 license. - [Z.ai Releases 355B Parameter GLM-4.5 Under MIT License](https://theopenweights.com/news/glm-4-5-6wo9) — Zhipu AI · Text / LLM · MIT · Jul 20, 2025: Z.ai has released GLM-4.5, a 355B-parameter Mixture-of-Experts model featuring a 128k context window and a permissive MIT license for commercial use. - [Qwen Releases Wan 2.2, a 5B Open Video AI Model](https://theopenweights.com/news/wan2-2-ti2v-5b-v52o) — Qwen · Alibaba · Text → Video · Apache 2.0 · Jul 18, 2025: Alibaba's Qwen team has released Wan2.2-TI2V-5B, a 5-billion-parameter open-source model for generating video from text and image inputs. - [HiDream.ai Releases 17B Open Image Editing Model](https://theopenweights.com/news/hidream-e1-1-5yvz) — HiDream.ai · Image Editing · MIT · Jul 16, 2025: HiDream.ai has released HiDream-E1.1, a 17-billion-parameter image editing model that follows text instructions. It is available under a permissive MIT license. - [Ming-Lite-Omni 1.5 Brings Any-to-Any Modality to Open Source](https://theopenweights.com/news/ming-lite-omni-1-5-oiyw) — inclusionAI · Any-to-Any · MIT · Jul 15, 2025: inclusionAI has released Ming-Lite-Omni 1.5, an MIT-licensed omni-modal model capable of handling any combination of text, image, audio, and video. - [Pusa V1: A New Open Model for Image-to-Video Animation](https://theopenweights.com/news/pusa-v1-k9bo) — RaphaelLiu · Image → Video · Apache 2.0 · Jul 14, 2025: Pusa V1 is a new 14B parameter open-source video model based on Wan2.1. It specializes in image-to-video with text prompts and frame control. - [T-Tech Releases T-one for Russian Speech Recognition](https://theopenweights.com/news/t-one-8kxp) — T-Tech · Speech → Text · Other · Jul 14, 2025: T-Tech has released T-one, a streaming speech recognition model trained on 30,000 hours of Russian audio and optimized for telephony. - [Moonshot AI Releases Trillion-Parameter Kimi-K2 Model](https://theopenweights.com/news/kimi-k2-mdgg) — Moonshot AI · Vision-Language · Other · Jul 11, 2025: Moonshot AI has released Kimi-K2-Instruct, a trillion-parameter Mixture-of-Experts model with a 128K context window for advanced reasoning and coding. - [Black Forest Labs Releases FLUX.1 Krea Image Model](https://theopenweights.com/news/flux-1-krea-dev-iy5y) — Black Forest Labs · Text → Image · Other · Jul 7, 2025: Black Forest Labs has released FLUX.1 Krea [dev], a 12-billion-parameter text-to-image model tuned for improved aesthetics and prompt adherence. - [ByteDance Releases Tar-7B for 'Any-to-Any' Multimodality](https://theopenweights.com/news/tar-7b-ry87) — ByteDance · Any-to-Any · Apache 2.0 · Jul 2, 2025: ByteDance has released Tar-7B, a 7B parameter open-source model designed for unified "any-to-any" multimodal understanding and generation, built on Qwen2.5. - [Boson AI Releases Higgs Audio v2 for Expressive TTS](https://theopenweights.com/news/higgs-audio-v2-3b-k9px) — Bosonai · Text → Speech · Other · Jul 1, 2025: Boson AI has released Higgs Audio v2, a 3B parameter model for expressive, multilingual text-to-speech, available under an Apache 2.0 license. - [Kyutai Releases 1.6B Bilingual TTS Model](https://theopenweights.com/news/fr-b8oz) — Kyutai · Text → Speech · CC BY 4.0 · Jun 30, 2025: Kyutai has released a 1.6B parameter text-to-speech model for English and French, built on the Moshi stack and available under a CC-BY-4.0 license. - [Zhipu AI Open-Sources 9B Vision Model with 'Thinking' Mode](https://theopenweights.com/news/glm-4-1v-9b-thinking-4w99) — Zhipu AI · Vision-Language · MIT · Jun 28, 2025: Zhipu AI has released GLM-4.1V-9B-Thinking, a 9-billion-parameter vision-language model with a chain-of-thought reasoning mode, available via an MIT license. - [Ovis-U1-3B Unifies Image Understanding and Generation](https://theopenweights.com/news/ovis-u1-3b-ycl0) — AIDC-AI · Any-to-Any · Apache 2.0 · Jun 28, 2025: AIDC-AI has released Ovis-U1-3B, a 3-billion-parameter open model that unifies image understanding, generation, and editing in a single architecture. - [NVIDIA Fuses LLM and ASR in Canary-Qwen 2.5B Model](https://theopenweights.com/news/canary-qwen-2-5b-roef) — NVIDIA · Speech → Text · Other · Jun 26, 2025: NVIDIA has released Canary-Qwen 2.5B, a speech-to-text model that pairs a specialized speech encoder with a general-purpose Qwen LLM decoder. - [Veena TTS Model Targets Indian Languages with Llama Base](https://theopenweights.com/news/veena-uvs8) — Maya Research · Text → Speech · Other · Jun 24, 2025: Maya Research released Veena, a 3B parameter text-to-speech model built on a Llama architecture to support high-quality synthesis for Hindi and English. - [Janus-4o-7B Adds Image Generation to 7B Multimodal AI](https://theopenweights.com/news/janus-4o-7b-fxgh) — FreedomIntelligence · Any-to-Any · Other · Jun 23, 2025: FreedomIntelligence has released Janus-4o-7B, a 7-billion-parameter multimodal model capable of text-to-image generation and image editing. - [pyannote ships community-1 diarization pipeline](https://theopenweights.com/news/pyannote-speaker-diarization-community-1-8iyn) — Pyannote · Speech → Text · Other · Apr 15, 2025: pyannote released speaker-diarization-community-1, an open-source pipeline for identifying who spoke when in audio recordings. - [Illustrious-XL v0.1 Arrives as Anime Image Base](https://theopenweights.com/news/illustrious-xl-v0-1-dpo4) — Unknown · Text → Image · Other · Sep 25, 2024: Illustrious-XL v0.1 is an SDXL-based anime text-to-image base model built to anchor downstream community fine-tunes. - [FLUX.1 Dev Becomes the Open-Weight Base to Beat](https://theopenweights.com/news/flux-1-dev-zetg) — Black Forest Labs · Text → Image · Other · Aug 2, 2024: Black Forest Labs' FLUX.1 Dev, a 12B rectified-flow text-to-image model, has become a widely adopted open-weight base for image generation.