Qwen-Image-3.0 Release Notes

Qwen has launched Qwen-Image-3.0, a third-generation foundational image generation model designed to transition image generation from being merely aesthetic to being a practical productivity tool. The model centers on the concept of "Real" (实), which is implemented through three primary technical pillars: Rich Content, Authentic Details, and Deep Knowledge.

Rich Content: Complex Layouts and High-Token Input

Qwen-Image-3.0 can render extremely complex, information-dense visual layouts by increasing the acceptable instruction length to 4.5k tokens. This allows the model to handle both horizontal expansion and vertical depth in image composition.

Horizontal Expansion and Spatial Control

The model demonstrates high semantic juxtaposition and spatial control, enabling it to lay out multiple distinct concepts in a single image without mutual interference. An example of this capability is the generation of a 3x3 grid in a single pass, where each cell contains a complex infographic (e.g., mathematical symbols, biology explainers, and medical diagrams) based on a prompt of 3.7k tokens.

Vertical Depth and Logical Nesting

Beyond parallel elements, the model supports semantic deconstruction and logical nesting. It can render multiple nested interfaces layer by layer within a single image, such as a VSCode interface containing a Qwen Chat interface, which in turn contains a WeChat interface and a coffee poster, while preserving the authentic style of each UI layer.

Authentic Details: Micro-Level Precision

Qwen-Image-3.0 achieves high rendering precision for micro-level details, focusing on legibility and photographic realism.

Precise Text Rendering

The model can render text as small as 10px, making it legible in dense environments. Key capabilities include:

  • Academic Papers: Accurate rendering of LaTeX formulas, subscripts, superscripts, Greek letters, and theorem numbering.
  • Realistic Documents: Generation of dense text in newspaper formats, simulating the authentic look of a physical newspaper.
  • Handwritten Annotations: Ability to overlay realistic red handwritten notes, including underlines and arrows, simulating a student's class notes.

Texture and Restoration

The model produces lifelike skin textures, including pores and hair strands. In editing tasks, it can restore damaged traditional paintings by maintaining original artistic styles, ink-wash gradients, and brushwork while removing mold spots and damage.

Deep Knowledge: Multilingual and World Awareness

Qwen-Image-3.0 leverages extensive world knowledge to render a broad variety of styles and interfaces.

Multilingual and UI Support

The model natively renders 12 languages and supports over 100 artistic styles. It can simulate mainstream user interfaces for web pages, games, and livestreams.

Knowledge-Driven Generation

The model can transform simple photographs into professional research figures by adding taxonomic information, morphological annotations, and scale bars. Additionally, it can connect to the internet to retrieve current information, such as generating a weather forecast image for a specific date and location, or incorporating specific IP figures into a scene.

Sources