Qwen-Image Release: Native Text Rendering and Precise Image Editing
TL;DR
Qwen has released Qwen-Image, a 20B MMDiT image foundation model designed to solve the long-standing challenge of high-fidelity text rendering in AI-generated images. The model excels at multi-line layouts, bilingual text (English and Chinese), and precise image editing, outperforming existing state-of-the-art models across multiple generation and editing benchmarks.
Native Text Rendering Capabilities
Qwen-Image provides superior text rendering across diverse languages and layouts, supporting both alphabetic languages like English and logographic languages like Chinese. The model is capable of handling complex semantic and spatial requirements, including:
- Multi-line and Paragraph Rendering: The model can generate long paragraphs of text, including handwritten styles on surfaces like glass plates, and maintain accuracy even when the text occupies a small fraction (less than one-tenth) of the total image area.
- Complex Layouts: Qwen-Image can automatically arrange text and icons into structured formats, such as elegant infographics with multiple submodules, movie posters with hierarchical text (title, subtitle, cast, director), and corporate-style PPT slides.
- Bilingual Support: The model can switch between English and Chinese within a single image, maintaining high fidelity in both languages.
- Calligraphy and Stylization: The model supports specific artistic styles, such as traditional Chinese calligraphy for couplets and horizontal scrolls.
Image Editing and General Generation
Beyond text rendering, Qwen-Image is a versatile foundation model for general visual content creation:
- Precise Image Editing: Through an enhanced multi-task training paradigm, the model supports a wide range of editing operations, including style transfer, adding or deleting objects, enhancing details, editing existing text, and adjusting character poses while preserving semantic meaning and visual realism.
- General Generation: The model supports various artistic styles, ranging from photorealistic scenes and anime styles to impressionistic paintings and minimalist designs.
Benchmark Performance
Qwen-Image achieves state-of-the-art (SOTA) performance across a broad suite of public benchmarks, demonstrating its strength in both generation and editing:
- General Image Generation: Evaluated on GenEval, DPG, and OneIG-Bench.
- Image Editing: Evaluated on GEdit, ImgEdit, and GSO.
- Text Rendering: Specifically excels on LongText-Bench, ChineseWord, and TextCraft, where it significantly outperforms previous SOTA models, particularly in Chinese text generation.