NVIDIA's Audex Unifies Audio Understanding and Speech
A new 30B mixture-of-experts model from NVIDIA handles both listening and speaking within a single audio-text architecture.

NVIDIA has released Nemotron-Labs-Audex-30B-A3B, a mixture-of-experts model that folds audio understanding and speech generation into a single system. According to the model's Hugging Face page, it is designed as a unified audio-text architecture — able to both interpret incoming audio and produce speech, rather than splitting those jobs across separate specialized models.
The model uses a mixture-of-experts design with roughly 30 billion total parameters but only about 3 billion active per token, the arrangement its "30B-A3B" name signals. That approach keeps inference costs closer to a small dense model while giving the network a larger pool of specialized capacity to draw from, a pattern NVIDIA and others have leaned on across recent Nemotron releases.
Why it matters
Most production audio stacks still chain together distinct components — a speech recognizer, a language model, and a separate text-to-speech engine. A single model spanning both comprehension and generation could simplify those pipelines and reduce the latency and error accumulation that comes from stitching parts together.
- Unified audio-text handling for both understanding and speech synthesis
- Mixture-of-experts efficiency: ~30B total, ~3B active parameters
- Reasoning listed among its capabilities alongside audio and TTS
The release ships under a custom NVIDIA license rather than a standard open-source one, so teams will want to check the terms before building on it. As an initial version, Audex sets a baseline for a family that NVIDIA appears positioned to iterate on.
Sources
- Visit
nvidia/Nemotron-Labs-Audex-30B-A3B
Hugging Face
More in Any-to-Any
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
SenseTime's SenseNova-U1.5 Unifies Vision Tasks
An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

SenseTime Releases SenseNova U1.5 8B Any-to-Any Model
The new 8B multimodal model handles text, images, and image editing within a single native architecture.
0 comments
No comments yet. Be the first to weigh in.