Source-linked AI summary

Qwen-Image-2.0 Technical Report

Bing Zhao, Chenfei Wu, Deqing Li, Hao Meng, Jiahao Li, Jie Zhang, Jingren Zhou, Junyang Lin, Kaiyuan Gao, Kuan Cao, Kun Yan, Liang Peng, Lihan Jiang, Niantong Li, Ningyuan Tang, Shengming Yin, Tianhe Wu, Xiao Xu, Xiaoyue Chen, Xihua Wang, Yan Shu, Yanran Zhang, Yi Wang, Yilei Chen, Ying Ba, Yixian Xu, Yujia Wu, Yuxiang Chen, Zecheng Tang, Zekai Zhang, Zhendong Wang, Zihao Liu, Zikai Zhou, An Yang, Chen Cheng, Chenxu Lv, Dayiheng Liu, Fan Zhou, Hantian Xiong, Hongzhu Shi, Hu Wei, Huihong Zhao, Ivy Liu, Jianwei Zhang, Jiawei Zhang, Kai Chen, Kang He, Levon Xue, Lin Qu, Linhan Tang, Luwen Feng, Minggang Wu, Minmin Sun, Na Ni, Rui Men, Shuai Bai, Sishou Zheng, Tao Lan, Tianqi Zhang, Tingkun Wen, Wei Wang, Weixu Qiao, Weiyi Lu, Wenmeng Zhou, Xiaodong Deng, Xiaoxiao Xu, Xinlei Fang, Xionghui Chen, Yanan Wang, Yang Fan, Yichang Zhang, Yixuan Xu, Yu Wu, Zhiyuan Ma, Zhizhi Cai

arXiv:2605.10730v1cs.CV

TL;DR

Existing image models remain limited in long-text rendering, multilingual typography, high-resolution photorealism, and complex workflows. Qwen-Image-2.0 unifies generation and editing through multimodal conditioning and a diffusion-transformer architecture, with qualitative comparisons showing successful text rendering where competing models fail.

  • Problem

    Existing image models struggle with ultra-long text rendering, multilingual typography, and high-resolution photorealism in text-dense and compositionally complex workflows.

  • Method

    Qwen-Image-2.0 unifies generation and editing using Qwen3-VL conditioning, a VAE, and an MMDiT denoising backbone supported by staged data curation.

  • Results

    Qualitative Chinese text comparisons show Qwen-Image-2.0 successfully fulfills the specified text-rendering prompt while competing models produce omissions, errors, or unrelated content.

  • Takeaways & Limitations

    Qwen-Image-2.0 provides a unified foundation for practical text-to-image generation and instruction-based image editing across demanding visual-generation scenarios.

Abstract

from arXiv · show

We present Qwen-Image-2.0, an omni-capable image generation foundation model that unifies high-fidelity generation and precise image editing within a single framework. Despite recent progress, existing models still struggle with ultra-long text rendering, multilingual typography, high-resolution photorealism, robust instruction following, and efficient deployment, especially in text-rich and compositionally complex scenarios. Qwen-Image-2.0 addresses these challenges by coupling Qwen3-VL as the condition encoder with a Multimodal Diffusion Transformer for joint condition-target modeling, supported by large-scale data curation and a customized multi-stage training pipeline. This enables strong multimodal understanding while preserving flexible generation and editing capabilities. The model supports instructions of up to 1K tokens for generating text-rich content such as slides, posters, infographics, and comics, while significantly improving multilingual text fidelity and typography. It also enhances photorealistic generation with richer details, more realistic textures, and coherent lighting, and follows complex prompts more reliably across diverse styles. Extensive human evaluations show that Qwen-Image-2.0 substantially outperforms previous Qwen-Image models in both generation and editing, marking a step toward more general, reliable, and practical image generation foundation models.

1 Introduction

Qwen-Image-2.0 is introduced as a unified image generation foundation model designed to overcome fragile ultra-long text rendering, underdeveloped multilingual typography, and the difficulty of combining photorealism, generation, and editing in one system. Its architecture, data infrastructure, and progressive training recipe support text-rich outputs, high-resolution photorealism, broad stylistic expression, precise instruction following, unified editing, and faster inference.

  • Core challenges and approach: The model targets fragile ultra-long text rendering and underdeveloped multilingual typography, which cause glyph distortion, character omission, layout collapse, and limited language coverage.These bottlenecks restrict the usefulness of existing systems for text-dense applications such as slides, infographics, and posters.
  • Architecture and training: Qwen-Image-2.0 couples a Qwen3-VL encoder with a Multimodal Diffusion Transformer and uses progressive multi-stage training spanning pretraining, fine-tuning, and RLHF.A resolution curriculum scales from lower to higher resolutions to stabilize optimization and improve detail fidelity and high-resolution coherence.
  • Capabilities: Qwen-Image-2.0 supports prompts of up to 1K tokens and produces text-dense visual outputs with improved glyph fidelity, higher character accuracy, and more complex typography across languages.The claimed applications include slides, posters, and infographics.
  • Capabilities: With native 2K-resolution support, Qwen-Image-2.0 generates finer texture detail, more coherent lighting, and more realistic materials across portraits, natural scenes, and architectural imagery.The model also aims to maintain robust quality across diverse aesthetic settings and reduce quality fluctuation between artistic styles.
  • Capabilities: The model improves semantic understanding for complex, composition-heavy prompts while jointly optimizing architecture and training for faster inference without sacrificing visual quality.These properties are intended to support interactive creative workflows.
  • Core challenges and approach: Qwen-Image-2.0 unifies text-to-image generation and instruction-based image editing within a single architecture and training paradigm.The model is supported by fine-grained captioning and a multi-stage, multi-resolution data pipeline incorporating filtered corpora, editing pairs, synthetic data, and curated high-resolution data.

2 Data

Qwen-Image-2.0 uses a large-scale, diverse data pipeline spanning text-to-image generation and instruction-based editing, with fine-grained captions and progressively filtered multi-stage training data. An automated data flywheel further supports continuous, targeted model improvement while preserving data reliability and diversity.

  • Data construction: The data pipeline supports unified text-to-image generation and instruction-based editing through broad domain coverage, high-quality instructions, and reliable source-target consistency.T2I data spans realistic photography, graphic design, artistic content, and synthetic imagery, while editing data covers single-image and multi-image tasks.
  • Captioning: Four captioning schemes—General, Text, Knowledge, and Structured—represent detailed visual content, dense text and layout, world knowledge, and complex entities and relations.The schemes are tailored to different task types and image characteristics, including text-rich materials and diagrams.
  • Data flywheel: An automated data flywheel forms a closed loop for iterative capability enhancement across image generation and editing models.Its error-attribution mechanism enables targeted optimization, while vector retrieval enriches training-data diversity; manual intervention is limited to critical filtering.

3 Architecture

Qwen-Image-2.0 unifies text and image processing through a Qwen3-VL-conditioned MMDiT backbone, a VAE latent pathway, and shared positional modeling. Its architecture also strengthens high-resolution reconstruction and complex-prompt handling through high-compression VAE design and prompt enhancement.

  • Core architecture: The architecture couples Qwen3-VL as a condition encoder, a VAE for image latents, and an MMDiT backbone for joint text-image modeling.The unified framework supports T2I, TI2I, and interleaved multi-image inputs.
  • VAE design: The VAE uses a 16× compression ratio to reduce DiT training costs, while confronting the trade-off among compression ratio, reconstruction fidelity, and diffusability.Aggressive compression can create information bottlenecks, whereas more latent channels produce high-dimensional manifolds that are harder to diffuse.
  • VAE design: A residual autoencoder, 64 latent channels, text-rich training data, and semantic alignment loss improve reconstruction fidelity and latent-space diffusability.The f16c64 configuration preserves the same total channel bottleneck as the f8c16 baseline.
  • Core architecture: Text and image tokens share a transformer stream, with visual representations replaced by VAE latents and positional information encoded using MSRoPE.Qwen3-VL produces modality-aware representations before multimodal concatenation and processing.
  • Prompt enhancement: The Prompt Enhancement module combines supervised rewriting with generation-aware reinforcement learning to improve enhanced prompts for image generation and editing.Its training data uses degraded prompts, inverse reasoning traces, and fine-grained annotations spanning General, Portrait, Text, and Complex Text categories.

4 Training

Qwen-Image-2.0 is trained through a three-phase pipeline that progressively increases resolution and refines data composition, followed by RLHF alignment and few-step distillation. These procedures target semantic learning, visual detail, aesthetic quality, controllability, and efficient inference while preserving generation and editing capabilities.

  • Multistage training: Training uses pre-training, continual pre-training, and supervised fine-tuning, progressively adjusting resolution, filtering, and data composition from semantic representations to fine-grained visual details.The configurations are summarized in Table 2.
  • Pre-training: Pre-training runs for 700K steps at low resolution with a 9:1 T2I-to-TI2I mixture and a learning rate of 1 × 10−4.This stage primarily learns basic semantic representations and robust general-purpose visual representations.
  • Supervised fine-tuning: Supervised fine-tuning runs for approximately 10K steps with a learning rate of 1 × 10−5, using diverse categories, strict filtering, and manual curation to improve aesthetic quality.The procedure is intended to enhance fine-grained visual details while preserving world knowledge.
  • RLHF alignment: RLHF refines the base diffusion model with multi-dimensional reward signals and sample-efficient optimization, yielding consistent gains in perceptual quality and task-specific controllability across T2I and TI2I.Rewards cover aesthetics, image-text alignment, portrait quality, instruction following, and visual consistency; qualitative results report improved texture fidelity, realism, and consistency.
  • Few-step distillation: Distillation converts the multi-step Qwen-Image-2.0-Base teacher into a 4-NFE student whose outputs are visually comparable to the teacher’s 40-step results across diverse prompts and visual domains.The student preserves detailed appearance, coherent composition, and faithful semantic alignment while reducing function evaluations.

5 Benchmark and Qualitative Evaluation

Qwen-Image-2.0 performs strongly in blind user-preference benchmarking and qualitative evaluations, ranking highly on LMArena while improving text rendering, portrait realism, multilingual generation, slide generation, and image editing. Qualitative comparisons particularly highlight its accurate text rendering, photorealistic integration, and identity preservation.

  • Benchmark evaluation: Qwen-Image-2.0 ranks #9 globally and #1 among Chinese models on LMArena, with an ELO score of 1168 and performance above Nano Banana.LMArena uses anonymous same-prompt comparisons and an ELO-based, preference-oriented ranking system.
  • T2I qualitative evaluation: Qualitative T2I evaluation covers text rendering, portraits, multilingual text rendering, and slide generation across Figures 13–15 and 18–19.These evaluations assess generation across text-rich, portrait, multilingual, and presentation-oriented scenarios.
  • T2I qualitative evaluation: Compared with baselines, Qwen-Image-2.0 uniquely combines high-fidelity signboard text with photorealistic material textures and consistent natural lighting.Other models exhibit small-scale or erroneous text, omissions, hallucinated numbers, or flat signboard integration.
  • TI2I editing evaluation: In complex Chinese text editing, Qwen-Image-2.0 is the only model described as rendering classical Chinese poetry accurately and aesthetically.Baselines produce undersized text, duplicated poems, or character-level errors.
  • TI2I editing evaluation: Across single-image and multi-image editing, Qwen-Image-2.0 preserves object identity and fine-grained details while following compositional instructions.The identity-preservation task requires adding a carrot and tissue while transferring a hat and maintaining the cat’s expression and posture.

6 Conclusion

Qwen-Image-2.0 is presented as a versatile foundation model unifying text-to-image generation and instruction-based image editing. Its multimodal encoder, efficient MMDiT backbone, and high-compression VAE target major real-world image-generation challenges.

  • Unified capabilities: Qwen-Image-2.0 supports both text-to-image generation and instruction-based image editing within a single framework.The model is described as a versatile image generation foundation model.
  • Architecture: The framework combines a strong multimodal encoder, an efficient MMDiT backbone, and a high-compression VAE.These components are identified as the technical basis of Qwen-Image-2.0.
  • Targeted challenges: Qwen-Image-2.0 addresses long-text rendering, multilingual typography, high-resolution photorealism, and complex instruction following.These are presented as key challenges in real-world image generation.

7 Authors

The paper credits a large team of core contributors and additional contributors.

  • Core Contributors: Core contributors include Bing Zhao, Chenfei Wu, Deqing Li, Deqing Li, Hao Meng, Jiahao Li, Jie Zhang, Jingren Zhou, Junyang Lin, and many others.The listed core contributors span 35 named individuals.
  • Contributors: Additional contributors include An Yang, Chen Cheng, Chenxu Lv, Dayiheng Liu, Fan Zhou, Hantian Xiong, Hongzhu Shi, Hu Wei, and many others.The contributors list contains 45 entries, ending with an incomplete “Zhizhi” entry.
Loading 2605.10730v1…