Thinking Machines Lab debuts Inkling, its first open model
The lab's inaugural open-weights release is a mixture-of-experts system that takes image and audio inputs, shipped under a permissive Apache 2.0 license.

Thinking Machines Lab has released Inkling, described as its first open-weights model and a mixture-of-experts (MoE) multimodal system that accepts image and audio inputs alongside text. The model is published under the permissive Apache 2.0 license, according to the lab's announcement.
The MoE design is notable because it activates only a subset of the model's parameters for any given input, a routing approach that lets teams scale total capacity while keeping the compute cost of each forward pass in check. Pairing that architecture with native image and audio handling puts Inkling in the growing category of open models built to reason across more than one input type.
Why it matters
An Apache 2.0 license is one of the least restrictive terms an AI lab can choose, which means developers can study, fine-tune, and deploy Inkling commercially without the usage caveats attached to many so-called open releases.
- Open weights: the model parameters are available for download and self-hosting.
- Multimodal inputs: image and audio are supported in addition to text.
- MoE architecture: sparse expert routing aimed at efficiency at scale.
As an initial release, Inkling establishes a baseline rather than iterating on a prior version. The full picture — including parameter counts, context length, and benchmark results — will come into focus as the community begins testing the weights against comparable open multimodal systems.
Sources
More in Any-to-Any

Thinking Machines Debuts Inkling Small, a Compact Multimodal MoE
The Apache-2.0 model brings mixture-of-experts efficiency to image, audio, and text tasks in a smaller footprint.

KRAFTON releases A.X-K2 Raon speech MoE model
The game maker's new open model blends text-to-speech and speech recognition in a single 21B mixture-of-experts system with just 3B active parameters.

Microsoft's Mage-VL Streams Video Natively
A codec-native multimodal foundation model aims to understand live video and vision-language input in real time.
0 comments
No comments yet. Be the first to weigh in.