Thinking Machines Debuts Inkling Small, a Compact Multimodal MoE
The Apache-2.0 model brings mixture-of-experts efficiency to image, audio, and text tasks in a smaller footprint.

Thinking Machines has released Inkling Small, a compact variant of its Inkling multimodal model family. Published on Hugging Face under an Apache-2.0 license, the model is built as a mixture-of-experts (MoE) system designed to process image, audio, and text inputs.
The "Small" designation signals a lighter-weight member of the lineup, aimed at teams that want multimodal capability without the compute demands of a full-scale flagship. MoE architectures activate only a subset of parameters per token, which typically lets developers run larger effective models at a fraction of the inference cost — a practical fit for a smaller, more deployable release.
Why it matters
Open multimodal models that reach beyond image-and-text into audio remain relatively rare, and a permissive Apache-2.0 license makes Inkling Small easy to adopt commercially. A few things stand out:
- Broad input support across image, audio, and text in a single model
- MoE efficiency, favoring lower active compute at inference time
- Apache-2.0 licensing, removing most barriers to commercial use
Thinking Machines has not published detailed parameter counts, context length, or benchmark figures alongside the release, so the model's precise capabilities and how it stacks up against peers remain to be seen. For now, the availability of an openly licensed, any-modality MoE is a meaningful addition to the growing catalog of accessible multimodal systems. Interested developers can find weights and any accompanying documentation on the model's Hugging Face page.
Sources
- Visit
thinkingmachines/Inkling-Small
Hugging Face
More in Vision-Language
Agnes-3.0-Flash arrives as a multimodal reasoning model
The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.
SenseTime's SenseNova-U1.5 Unifies Vision Tasks
An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.
LLaDA-UI Brings Diffusion Decoding to GUI Agents
inclusionAI's 16.7B MoE vision-language model uses block-wise diffusion to drive graphical interface tasks.
0 comments
No comments yet. Be the first to weigh in.