NVIDIA's Audex Unifies Audio Understanding and Speech
A new 30B mixture-of-experts model from NVIDIA handles both listening and speaking within a single audio-text architecture.

NVIDIA has released Nemotron-Labs-Audex-30B-A3B, a mixture-of-experts model that folds audio understanding and speech generation into a single system. According to the model's Hugging Face page, it is designed as a unified audio-text architecture — able to both interpret incoming audio and produce speech, rather than splitting those jobs across separate specialized models.
The model uses a mixture-of-experts design with roughly 30 billion total parameters but only about 3 billion active per token, the arrangement its "30B-A3B" name signals. That approach keeps inference costs closer to a small dense model while giving the network a larger pool of specialized capacity to draw from, a pattern NVIDIA and others have leaned on across recent Nemotron releases.
Why it matters
Most production audio stacks still chain together distinct components — a speech recognizer, a language model, and a separate text-to-speech engine. A single model spanning both comprehension and generation could simplify those pipelines and reduce the latency and error accumulation that comes from stitching parts together.
- Unified audio-text handling for both understanding and speech synthesis
- Mixture-of-experts efficiency: ~30B total, ~3B active parameters
- Reasoning listed among its capabilities alongside audio and TTS
The release ships under a custom NVIDIA license rather than a standard open-source one, so teams will want to check the terms before building on it. As an initial version, Audex sets a baseline for a family that NVIDIA appears positioned to iterate on.
Sources
- Visit
nvidia/Nemotron-Labs-Audex-30B-A3B
Hugging Face
More in Any-to-Any

Thinking Machines Debuts Inkling Small, a Compact Multimodal MoE
The Apache-2.0 model brings mixture-of-experts efficiency to image, audio, and text tasks in a smaller footprint.

KRAFTON releases A.X-K2 Raon speech MoE model
The game maker's new open model blends text-to-speech and speech recognition in a single 21B mixture-of-experts system with just 3B active parameters.

Microsoft's Mage-VL Streams Video Natively
A codec-native multimodal foundation model aims to understand live video and vision-language input in real time.
0 comments
No comments yet. Be the first to weigh in.