InternLM Previews 397B Vision-Language Model
The Intern-S2 preview arrives as a very large multimodal system under a permissive Apache-2.0 license.

InternLM has published an early look at its next-generation multimodal system, Intern-S2-Preview-397B, on Hugging Face. As the name suggests, this is a preview release rather than a finished model, offering a first glimpse at where the InternLM team is heading with its S2 line.
The headline figure is scale. At 397 billion parameters, Intern-S2 sits firmly in the upper tier of openly available models, and it is billed as a vision-language model capable of processing both images and text. Notably, the record lists it as a dense architecture rather than a mixture-of-experts design, which is unusual at this size and implies substantial hardware requirements to run.
Why it matters
Open releases at this parameter count remain rare, and the choice of an Apache-2.0 license is significant. That permissive terms allow commercial use, redistribution, and fine-tuning without the usage restrictions attached to many other large models.
- Modality: vision-language (image and text)
- Scale: 397B parameters, dense
- License: Apache-2.0
- Status: preview release
As a preview, key details such as context length and benchmark results are not yet spelled out, and prospective users should treat it as a work in progress. Still, the release signals continued momentum from InternLM in pushing large, openly licensed multimodal systems into the community's hands. Full details are on the model's Hugging Face page.
Sources
- Visit
internlm/Intern-S2-Preview-397B
Hugging Face
More in Vision-Language
Agnes-3.0-Flash arrives as a multimodal reasoning model
The new release pairs vision-language understanding with a hybrid-attention design aimed at long-context reasoning.
SenseTime's SenseNova-U1.5 Unifies Vision Tasks
An 8B model drops the usual encoder and VAE in favor of a single native architecture spanning understanding, reasoning, and image generation.
LLaDA-UI Brings Diffusion Decoding to GUI Agents
inclusionAI's 16.7B MoE vision-language model uses block-wise diffusion to drive graphical interface tasks.
0 comments
No comments yet. Be the first to weigh in.