Thinking Machines Debuts Inkling Small, a Compact Multimodal MoE
The Apache-2.0 model brings mixture-of-experts efficiency to image, audio, and text tasks in a smaller footprint.

Thinking Machines has released Inkling Small, a compact variant of its Inkling multimodal model family. Published on Hugging Face under an Apache-2.0 license, the model is built as a mixture-of-experts (MoE) system designed to process image, audio, and text inputs.
The "Small" designation signals a lighter-weight member of the lineup, aimed at teams that want multimodal capability without the compute demands of a full-scale flagship. MoE architectures activate only a subset of parameters per token, which typically lets developers run larger effective models at a fraction of the inference cost — a practical fit for a smaller, more deployable release.
Why it matters
Open multimodal models that reach beyond image-and-text into audio remain relatively rare, and a permissive Apache-2.0 license makes Inkling Small easy to adopt commercially. A few things stand out:
- Broad input support across image, audio, and text in a single model
- MoE efficiency, favoring lower active compute at inference time
- Apache-2.0 licensing, removing most barriers to commercial use
Thinking Machines has not published detailed parameter counts, context length, or benchmark figures alongside the release, so the model's precise capabilities and how it stacks up against peers remain to be seen. For now, the availability of an openly licensed, any-modality MoE is a meaningful addition to the growing catalog of accessible multimodal systems. Interested developers can find weights and any accompanying documentation on the model's Hugging Face page.
Sources
- Visit
thinkingmachines/Inkling-Small
Hugging Face
More in Vision-Language

Microsoft's Mage-VL Streams Video Natively
A codec-native multimodal foundation model aims to understand live video and vision-language input in real time.
Apertus v1.5 70B arrives with an Apache-2.0 license
Switzerland's open-model effort ships a 70-billion-parameter, multilingual and multimodal system that anyone can use, modify, and deploy.
Microsoft's Fara1.5-27B targets computer-use agents
A 27B-parameter vision-language model built to drive browsers and desktop apps like a human operator.
0 comments
No comments yet. Be the first to weigh in.