Qwen Releases Open Model for Image Editing
The new open-source model from Alibaba lets users edit images with simple text commands in both English and Chinese.
The Qwen team at Alibaba has released Qwen-Image-Edit, a new open-source model designed for sophisticated, instruction-based image editing. Released under the permissive Apache 2.0 license, the model adds a powerful new modality to the growing Qwen family of open AI tools.
Unlike text-to-image generation models that create visuals from scratch, Qwen-Image-Edit specializes in modifying existing images. Users can provide a source image and a text prompt to perform specific edits, such as changing an object's color, altering a background, or adding new elements. A key feature is its bilingual capability, allowing it to understand prompts in both English and Chinese.
Why It Matters
The release provides a powerful open alternative to the proprietary, AI-powered editing features found in tools like Adobe Photoshop or Canva. By making this technology accessible, Qwen enables developers to integrate advanced, in-painting and editing capabilities directly into their own applications, from creative suites to e-commerce platforms, without relying on a closed API.
The model is available for download and experimentation on the Hugging Face Hub. Example use cases highlighted by the team include:
- Object Modification: Change the color of a car or add stripes to a shirt.
- Background Editing: Replace a city background with a natural landscape.
- Style Transformation: Convert a photograph into the style of a watercolor painting.
Sources
- Visit
Qwen/Qwen-Image-Edit
Hugging Face
More in Image Editing

Microsoft's Mage-Flow packs image editing into 4B
A compact model handles both text-to-image generation and instruction-based edits at native resolution, under a permissive MIT license.
Boogu-Image-0.1 Brings Unified Multimodal to Open Source
A new Apache-licensed model family folds bilingual text-to-image generation and instruction editing into one system.

SenseTime's SenseNova-Vision-7B-MoT Goes Any-to-Any
A single 7B model from SenseTime folds vision-language understanding, image generation, editing, and perception into one system.
0 comments
No comments yet. Be the first to weigh in.