dots3-note preview brings audio and vision to one model
An early build of a multimodal, long-context agentic system arrives on Hugging Face with support for both images and sound.

dots-studio has published an early preview of dots3-note, a multimodal model designed to handle images, audio, and text within a single long-context, agentic framework. The release is available now on Hugging Face.
The model is described as "multimodal-any," meaning it aims to accept a range of input types rather than specializing in a single one. Vision-language capabilities sit alongside audio understanding, positioning dots3-note as a general-purpose system for tasks that blend perception with reasoning.
What we know
- Multimodal support spanning vision and audio, plus text
- A long-context, agentic design intended for multi-step tasks
- Distributed under a non-standard "other" license
- Published as a preview, so specifications may still shift
Key technical details remain unpublished at this stage. The record lists no confirmed parameter count, context length, or benchmark results, and the license is marked simply as "other" — worth checking carefully before any commercial use.
Why it matters: Combining audio and vision in one openly available agentic model is still relatively uncommon, and preview releases like this offer early hands-on access for developers experimenting with cross-modal workflows. As with any preview, expect the details to firm up as the family matures.
Sources
- Visit
dots-studio/dots3-note-prev
Hugging Face
More in Any-to-Any
Cloudflare's Clef brings structured decisions to open models
The new open-weight vision-language family outputs typed, structured results and arrives alongside a reinforcement-learning fine-tuning platform.
Black Forest Labs brings FLUX to robotics
The image-model maker's FLUX 3 Action Base extends its generative stack into world-action modeling, released through the LeRobot ecosystem.

Xiaomi expands MiMo line with V2.6 multimodal models
The new Flash, Pro, and Distill variants add vision, audio, agentic behavior, and long-context handling to Xiaomi's open MiMo family.
0 comments
No comments yet. Be the first to weigh in.