InternLM Previews 397B Vision-Language Model
The Intern-S2 preview arrives as a very large multimodal system under a permissive Apache-2.0 license.

InternLM has published an early look at its next-generation multimodal system, Intern-S2-Preview-397B, on Hugging Face. As the name suggests, this is a preview release rather than a finished model, offering a first glimpse at where the InternLM team is heading with its S2 line.
The headline figure is scale. At 397 billion parameters, Intern-S2 sits firmly in the upper tier of openly available models, and it is billed as a vision-language model capable of processing both images and text. Notably, the record lists it as a dense architecture rather than a mixture-of-experts design, which is unusual at this size and implies substantial hardware requirements to run.
Why it matters
Open releases at this parameter count remain rare, and the choice of an Apache-2.0 license is significant. That permissive terms allow commercial use, redistribution, and fine-tuning without the usage restrictions attached to many other large models.
- Modality: vision-language (image and text)
- Scale: 397B parameters, dense
- License: Apache-2.0
- Status: preview release
As a preview, key details such as context length and benchmark results are not yet spelled out, and prospective users should treat it as a work in progress. Still, the release signals continued momentum from InternLM in pushing large, openly licensed multimodal systems into the community's hands. Full details are on the model's Hugging Face page.
Sources
- Visit
internlm/Intern-S2-Preview-397B
Hugging Face
More in Vision-Language

Thinking Machines Debuts Inkling Small, a Compact Multimodal MoE
The Apache-2.0 model brings mixture-of-experts efficiency to image, audio, and text tasks in a smaller footprint.

Microsoft's Mage-VL Streams Video Natively
A codec-native multimodal foundation model aims to understand live video and vision-language input in real time.
Apertus v1.5 70B arrives with an Apache-2.0 license
Switzerland's open-model effort ships a 70-billion-parameter, multilingual and multimodal system that anyone can use, modify, and deploy.
0 comments
No comments yet. Be the first to weigh in.