The Open Weights
LatestModelsLeaderboardsCompanies
Subscribe
The Open Weights

The daily record of open-source AI. New model releases, leaderboards, and what's coming next — written for people who ship.

Refreshed every 12 hours

Discover

  • Latest releases
  • New today
  • Trending models

Browse

  • All models
  • Companies
  • Categories
  • Leaderboards

About

  • About
  • Editorial policy
  • RSS feed
  • Newsletter

© 2026 The Open Weights. An independent publication.

PrivacyTermsSMSAggregated by Claude · curated by humans.
LatestNVIDIA1
NVIDIAAny-to-Any

NVIDIA's Audex Unifies Audio Understanding and Speech

A new 30B mixture-of-experts model from NVIDIA handles both listening and speaking within a single audio-text architecture.

Jul 6, 2026
NotableOther
Nemotron-Labs-Audex-30B-A3B

NVIDIA has released Nemotron-Labs-Audex-30B-A3B, a mixture-of-experts model that folds audio understanding and speech generation into a single system. According to the model's Hugging Face page, it is designed as a unified audio-text architecture — able to both interpret incoming audio and produce speech, rather than splitting those jobs across separate specialized models.

The model uses a mixture-of-experts design with roughly 30 billion total parameters but only about 3 billion active per token, the arrangement its "30B-A3B" name signals. That approach keeps inference costs closer to a small dense model while giving the network a larger pool of specialized capacity to draw from, a pattern NVIDIA and others have leaned on across recent Nemotron releases.

Why it matters

Most production audio stacks still chain together distinct components — a speech recognizer, a language model, and a separate text-to-speech engine. A single model spanning both comprehension and generation could simplify those pipelines and reduce the latency and error accumulation that comes from stitching parts together.

  • Unified audio-text handling for both understanding and speech synthesis
  • Mixture-of-experts efficiency: ~30B total, ~3B active parameters
  • Reasoning listed among its capabilities alongside audio and TTS

The release ships under a custom NVIDIA license rather than a standard open-source one, so teams will want to check the terms before building on it. As an initial version, Audex sets a baseline for a family that NVIDIA appears positioned to iterate on.

Sources

  • nvidia/Nemotron-Labs-Audex-30B-A3B

    Hugging Face

    Visit

Get the model

Hugging Face

Specs

Parameters30B · MoE
Architecturenemotron_labs_audex
LicenseOTHER
Downloads512
Likes176

Modalities

Any-to-AnyText → SpeechReasoning

0 comments

No comments yet. Be the first to weigh in.

More in Any-to-Any

StepFun/Text → Speech

StepFun's StepAudio 3 Realtime targets live voice AI

The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.

Sep 11, 2026
SenseTime/Any-to-Any

SenseTime's SenseNova-U1.5 Unifies Vision Tasks

An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

Sep 9, 2026
SenseNova-U1.5-8B-MoT
SenseTime/Any-to-Any

SenseTime Releases SenseNova U1.5 8B Any-to-Any Model

The new 8B multimodal model handles text, images, and image editing within a single native architecture.

Aug 19, 2026