Qwen-Image-2512 release notes / what's new

Qwen-Image-2512 is a foundational text-to-image model update that reduces the typical "AI-generated" aesthetic in favor of higher realism and finer detail. It specifically improves the rendering of human subjects, natural textures, and multimodal text compositions compared to the base Qwen-Image model released in August.

Model Performance and Benchmarking

Qwen-Image-2512 is positioned as the strongest open-source text-to-image model based on blind evaluations. According to the Qwen team, the model underwent over 10,000 rounds of blind evaluations on AI Arena, where it remained highly competitive against closed-source alternatives.

Enhanced Human Realism

Qwen-Image-2512 substantially refines the depiction of people by adding richer facial details and better environmental context. Key improvements include:

  • Age Accuracy: The model precisely captures age-specific cues, such as wrinkles in elderly subjects, avoiding the artificial smoothness often seen in earlier AI generations.
  • Fine Detail: Precision in rendering individual hair strands replaces the blurred textures found in the August release.
  • Semantic Adherence: The model more accurately follows complex postural instructions, such as specific body leaning or expressions.
  • Environmental Integration: Background objects (e.g., dormitory furniture or convention banners) are rendered with greater clarity and better integration with the subject.

Finer Natural Detail and Textures

The update extends high-fidelity rendering to landscapes, wildlife, and natural elements:

  • Nature and Landscapes: Improvements are evident in the rendering of water flow, foliage, and waterfall mist, with richer gradations of green and more realistic atmospheric effects like fog and spray.
  • Animal Fur: The model achieves high precision in animal textures, such as the distinct layering of guard hairs and soft undercoats in golden retrievers, or the coarse, dense coats of argali sheep.

Improved Text Rendering and Layout

Qwen-Image-2512 enhances the accuracy and layout of textual elements within images, enabling more complex multimodal compositions. Demonstrated capabilities include:

  • Structured Documents: The ability to generate professional-grade PPT slides with specific timelines, labels, and directional arrows.
  • Infographics: The creation of industrial technical infographics featuring distinct layout blocks, specific icons (e.g., rusted gears, conical flasks), and precise textual labels with logical markers (e.g., checkmarks and red crosses).
  • Complex Grids: The generation of multi-panel educational posters (such as a 3x4 grid) where each cell maintains a consistent narrative theme while remaining a distinct, high-quality photographic scene.

Sources