NVIDIA Releases Nemotron-3-Nano Omni-Modal MoE
The new 30-billion-parameter Mixture-of-Experts model handles any combination of modalities with just 3 billion active parameters.
NVIDIA has introduced Nemotron-3-Nano-Omni, a powerful new model designed for sophisticated multimodal tasks. The model employs a Mixture-of-Experts (MoE) architecture, a technique that activates only a fraction of its total parameters for any given task, leading to significant computational savings.
With a total of 30 billion parameters, Nemotron-3-Nano-Omni uses just 3 billion active parameters during inference. This efficient design allows it to deliver the performance of a much larger model without the corresponding computational overhead, making advanced reasoning more accessible. The model is available with BF16 weights, a common format for balancing performance and precision.
Any-to-Any Reasoning
The model's key feature is its "omni-modal" capability, allowing it to process and generate information across different formats seamlessly. It can handle what the company calls "any-to-any" tasks, meaning it can ingest a mix of inputs—such as text, images, and video—and produce a mix of outputs in response. This flexibility is critical for complex applications that require understanding context across multiple data types.
Nemotron-3-Nano-Omni represents a notable step forward for efficient and versatile AI. While it is governed by a custom NVIDIA license rather than a permissive open-source one, its availability on the Hugging Face Hub enables researchers and developers to experiment with its unique multimodal reasoning capabilities.
Sources
- Visit
nvidia/Nemotron-3-Nano-Omni-30B-A3B-Reasoning-BF16
Hugging Face
More in Any-to-Any
StepFun's StepAudio 3 Realtime targets live voice AI
The audio-language foundation model builds a listen-converse-think-act loop aimed at natural, low-latency spoken interaction.
SenseTime's SenseNova-U1.5 Unifies Vision Tasks
An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.

SenseTime Releases SenseNova U1.5 8B Any-to-Any Model
The new 8B multimodal model handles text, images, and image editing within a single native architecture.
0 comments
No comments yet. Be the first to weigh in.