Qwen-Image-2.1: Compact, Efficient, and Unified Image Creation
The 7B visual parameter model unifies text-to-image generation, editing, and native RGBA transparency support in a single architecture.
- Uses 32 Single-Stream DiT layers with mixed-granularity attention and KV cache reuse for static editing context
- Supports up to 10 reference images simultaneously for multi-subject compositions, virtual try-ons, and room staging
- Enables local editing via color-coded circles, painted annotations, or separate input masks while preserving facial and product fidelity
- Generates and edits transparent layers directly, with the ability to extract subjects from standard RGB photos into RGBA layers
- Supports panorama, infographic, and storyboard generation alongside improved typography rendering
Designers and developers gain a lightweight open-source model capable of complex multi-image composition, native transparency, and targeted local editing.
Sources
Read this as text
Back to the AI news