Source-linked AI summary
Wan-Image: Pushing the Boundaries of Generative Visual Intelligence
Chaojie Mao, Chen-Wei Xie, Chongyang Zhong, Haoyou Deng, Jiaxing Zhao, Jie Xiao, Jinbo Xing, Jingfeng Zhang, Jingren Zhou, Jingyi Zhang, Jun Dan, Kai Zhu, Kang Zhao, Keyu Yan, Minghui Chen, Pandeng Li, Shuangle Chen, Tong Shen, Yu Liu, Yue Jiang, Yulin Pan, Yuxiang Tuo, Zeyinzi Jiang, Zhen Han, Ang Wang, Bang Zhang, Baole Ai, Bin Wen, Boang Feng, Feiwu Yu, Gang Wang, Haiming Zhao, He Kang, Jianjing Xiang, Jianyuan Zeng, Jinkai Wang, Junjie Zhou, Ke Sun, Linqian Wu, Pei Gong, Pingyu Wu, Ruiwen Wu, Tongtong Su, Wenmeng Zhou, Wenting Shen, Wenyuan Yu, Xianjun Xu, Xiaoming Huang, Xiejie Shen, Xin Xu, Yan Kou, Yangyu Lv, Yifan Zhai, Yitong Huang, Yun Zheng, Yuntao Hong, Zhe Zhang, Zhicheng Zhang
TL;DR
Existing visual generators often remain limited in rigorous design workflows that require controllability, typography, and identity preservation. Wan-Image addresses this gap with a unified multimodal architecture, scaled annotated data, and reinforcement learning, achieving highly competitive performance across professional generation tasks. It surpasses Seedream 5.0 Lite and GPT Image 1.5 overall and is comparable to Nano Banana Pro on challenging tasks.
Problem
Contemporary visual generators deliver high-fidelity images but remain constrained in rigorous workflows requiring precise control, complex typography, and identity preservation.
Method
Wan-Image unifies an MLLM-based Planner, a DiT-based Visualizer, a four-channel VAE, large-scale fine-grained data, and curated reinforcement learning.
Results
Wan-Image comprehensively surpasses Seedream 5.0 Lite and GPT Image 1.5, matches Nano Banana Pro in challenging text-to-image scenarios, and reaches around 80% pass rates in interactive editing and image-series generation.
Takeaways & Limitations
Wan-Image supports professional visual workflows through complex typography, hyper-diverse portraits, palette control, identity-preserving series, interactive editing, transparency, and efficient 4K generation.
Abstract
from arXiv · showhide
We present Wan-Image, a unified visual generation system explicitly engineered to paradigm-shift image generation models from casual synthesizers into professional-grade productivity tools. While contemporary diffusion models excel at aesthetic generation, they frequently encounter critical bottlenecks in rigorous design workflows that demand absolute controllability, complex typography rendering, and strict identity preservation. To address these challenges, Wan-Image features a natively unified multi-modal architecture by synergizing the cognitive capabilities of large language models with the high-fidelity pixel synthesis of diffusion transformers, which seamlessly translates highly nuanced user intents into precise visual outputs. It is fundamentally powered by large-scale multi-modal data scaling, a systematic fine-grained annotation engine, and curated reinforcement learning data to surpass basic instruction following and unlock expert-level professional capabilities. These include ultra-long complex text rendering, hyper-diverse portrait generation, palette-guided generation, multi-subject identity preservation, coherent sequential visual generation, precise multi-modal interactive editing, native alpha-channel generation, and high-efficiency 4K synthesis. Across diverse human evaluations, Wan-Image exceeds Seedream 5.0 Lite and GPT Image 1.5 in overall performance, reaching parity with Nano Banana Pro in challenging tasks. Ultimately, Wan-Image revolutionizes visual content creation across e-commerce, entertainment, education, and personal productivity, redefining the boundaries of professional visual synthesis.
1 Introduction
Wan-Image is designed to extend visual generation from casual creation into professional workflows requiring controllability, complex typography, identity preservation, and precise editing. It combines unified multimodal reasoning and pixel synthesis with scaled, finely annotated data to support a broad set of productivity capabilities and achieves highly competitive evaluation results.
- System design: Wan-Image combines an MLLM-based Planner for semantic reasoning with a DiT-based Visualizer and four-channel VAE for high-fidelity generation and transparency.A Prompt Enhancer captures nuanced intent, while an Image Refiner improves high-frequency detail and supports 4K output.
- Training foundation: Its data engine uses large-scale multimodal data, hierarchical fine-grained annotation, multi-stage training, multitask optimization, and curated reinforcement learning.The reinforcement learning data target aesthetic preferences, identity preservation, and complex instruction following.
- Professional capabilities: Wan-Image supports ultra-long complex typography, extreme aspect ratios up to 1:8, fine-grained portrait steering, and palette-guided generation.Palette control uses user-specified hex-code palettes with explicit color proportions.
- Professional capabilities: It also enables multi-subject identity preservation, coherent image series of up to 12 images, precise multimodal editing, native alpha-channel generation, and efficient 4K synthesis.These capabilities target consistency across subjects or image groups, localized edits, transparent backgrounds, and professional-resolution output.
- Evaluation: Wan-Image comprehensively surpasses Seedream 5.0 Lite and GPT Image 1.5, remains comparable to Nano Banana Pro on challenging text-to-image tasks, and reaches around 80% pass rates in editing and image-series generation.The reported challenging tasks include text rendering and photo realism.
2 Data
Wan-Image’s data pipeline combines understanding-oriented datasets, task-balanced taxonomy, multimodal retrieval, and fine-grained filtering to support diverse generative workflows. The resulting data emphasize coverage, quality, and real-world usability.
- Understanding Data: Understanding data combine general text-image resources, text-proxy data, and user-aligned queries for multimodal comprehension and generation support.Text-proxy data provide detailed visual prompts, while query annotations target practical interaction patterns.
- Task Coverage: The dataset covers text-to-image, image-to-image, image-series, and interleaved generation, including up to 9 input images and 12 output images.Series generation emphasizes visual consistency and logical thematic progression.
- Data Collection: A multimodal retrieval system links data retrieval, annotation, and cleaning in a closed-loop pipeline for acquiring category-relevant training data.The system addresses low acquisition efficiency and limited relevance in collecting key categories.
- Data Collection: Fine-grained image operators assess features, aesthetics, AI-generated content, low-level information, and overall image quality across complementary dimensions.These operators support more efficient filtering and higher-quality datasets.
- Data Collection: The coordinated filtering operators improve dataset clarity, realism, aesthetic quality, and overall usability.Representative operator distributions and visual examples are shown in Figure 10.
- Data Organization: A structured taxonomy controls sampling across semantic themes, attributes, styles, composition, and difficulty to improve coverage and mitigate distribution bias.Structured captions further support instruction following across varying prompt granularities.
3 Model
Wan-Image uses a unified architecture that connects multimodal understanding, planning, and diffusion-based visual synthesis. Its VAE and generation mechanisms target high-fidelity reconstruction, controllable editing, identity preservation, and multi-image workflows.
- Unified Architecture: A unified Transformer integrates a decoder-only understanding branch with a DiT generation branch trained under Rectified Flow.Dedicated Transformer experts share attention within each block for feature-space alignment.
- VAE: The VAE targets high-frequency reconstruction, alpha-channel modeling, and super-resolution for text, dense layouts, textures, fine boundaries, and transparent regions.Its architecture uses 16 × 16 overall compression with residual blocks across encoder and decoder stages.
- VAE: A three-stage curriculum progresses from 256 × 256 training to mixed resolutions and 4-channel GAN training, adapting toward 2K resolution.The final stage improves detail sharpness, transparent-boundary realism, and visual consistency.
- VAE Evaluation: 35.429 PSNR, 0.958 SSIM, and 0.010 LPIPS are achieved on 5,000 RGB images under 16×16 spatial downsampling.The comparison reports the proposed VAE’s best overall reconstruction performance.
- VAE Evaluation: The VAE preserves sharper character strokes and glyph boundaries in document regions than existing methods.Figure 13 provides a qualitative comparison of VAE models.
- Unified Architecture: Wan-Image combines an MLLM-based Planner with a DiT-based visual generator to transition from semantic understanding to visual synthesis.The Planner supports autonomous decision-making and task routing.
- Generation Planning: Think Mode and chain-of-thought-driven planning translate user intent into structured descriptions or instructions for generation tasks.Special tokens support richer outputs grounded in world knowledge.
- Prompt Enhancer: The prompt enhancer jointly supports T2I, I2I, T2S, and TI2S, separating preserved content from requested changes through common and difference fields.This explicit disentanglement supports reference-controlled editing.
4 Training
Wan-Image follows a progressive training pipeline that first establishes semantic understanding and then translates high-level intent into high-fidelity visual synthesis.
- Training Pipeline: Training begins with CT and multi-task SFT, followed by on-policy multi-teacher distillation to establish semantic priors.The understanding stage builds the cognitive foundation for later generation training.
- Training Pipeline: Generation training progresses from foundational PT to SFT for improved image quality.The pipeline connects semantic comprehension with visual synthesis.
- Training Pipeline: The overall paradigm translates high-level user intent into pixels through progressively integrated understanding and generation training.
4.1 Understanding Training
Understanding training strengthens Wan-Image’s multimodal reasoning and generation readiness through staged training, balanced text and multimodal data, and user-aligned query construction.
- Understanding Training: The understanding stage comprises CT, multi-task SFT, and on-policy multi-teacher distillation to improve multimodal understanding and real-world performance.
- Understanding Training: Pure-text CT data are mixed with multimodal CT data at a controlled ratio to preserve textual reasoning and comprehension capabilities.
- Understanding Training: Large-scale text-proxy data provide dense visual prompts that help the model adapt user inputs for generation-oriented tasks without degrading understanding.
- Understanding Training: A user-aligned query set uses annotated task, difficulty, reasoning, and visual-question attributes to construct a balanced set of practical interaction patterns.
4.2 Generation Training
Generation training progresses through PT, CT, and SFT, progressively increasing resolution and refining image quality and instruction following. The stages use multi-task data and increasingly targeted optimization to develop high-resolution generation capabilities.
- Training stages: PT, CT, and SFT form a three-stage generation training pipeline that progressively advances the model from basic to ultra-high-resolution image generation.Training resolution and data composition are adjusted across stages.
- Training stages: PT establishes foundational visual representations using 13.27T tokens over 713K steps at 192, 320, and 640 resolutions with a 7:3 T2I-to-I2I mixture.
- Training stages: CT shifts toward high-resolution understanding and generation, processing 8.85T tokens over 223K steps at resolutions from 512 to 2048.
- Multi-task training: The training recipe supports multiple tasks, including T2T, I2T, T2I, I2I, and interleaved generation, with CT adding T2S and TI2S data.
- Supervised fine-tuning: SFT focuses on aesthetic quality and instruction-following precision through 13K steps at 512–2048 resolution, using a reduced learning rate of 3 × 10−5.
- Optimization: All stages use AdamW with weight decay 0.02, gradient norm clipping at 0.5, and unconditional dropout of 0.05 for classifier-free guidance.
4.3 Reinforcement Learning with Human Feedback
After supervised fine-tuning, Wan-Image applies Cascade RL to align generation with human preferences and improve quality across domains. The framework combines multiple optimization methods with domain-wise training and achieves 15%–50% winning-rate gains over the SFT baseline.
- Framework: After SFT, the model is refined with preference data, reward models, and feedback-driven learning through a cascaded domain-wise reinforcement learning framework.
- Data construction: The reinforcement learning dataset uses prompts mutually exclusive from SFT data to prevent high rewards from memorized SFT responses.
- Reward modeling: A diverse reward benchmark is constructed to address preference overfitting and improve consistency between reward-model evaluation and reinforcement learning guidance.
- Cascade RL: Cascade RL combines DPO, ReFL, DenseGRPO, and ReFMA in a comprehensive reinforcement learning framework.
- Cascade RL: ReFMA uses anchor-based control across noise levels and reward-model gradients to guide policy optimization in a cascaded training strategy.
- Results: 15% to 50%: Cascade RL improves winning-rate scores over the SFT baseline across various scenarios.The reported visual improvements include photorealism, structural correctness, text rendering, and overall aesthetic quality.
4.4 Model Distillation
Model distillation targets inference-step reduction to lower latency for real-world applications. The section compares the teacher and distilled student models as part of this efficiency-focused process.
- Distillation objective: Inference-step distillation is designed to substantially reduce base-model latency for real-world efficiency requirements.
- Distillation objective: The original iterative generation scheme and classifier-free guidance create computational burden through a high number of function evaluations.
- Teacher–student comparison: The teacher model and distilled student model are compared visually to assess the distillation process.
5 Results
Wan-Image is evaluated across understanding and generative tasks using quantitative comparisons and qualitative visualizations. Results emphasize broad task versatility, specialized controllability, coherent sequences, and improvements from successive training stages and reinforcement learning.
- Evaluation: The evaluation combines quantitative analysis and qualitative visualizations across generative tasks and training stages.
- Evaluation: The model demonstrates superior capabilities across several key benchmarks in the reported radar-chart comparison.
- Task coverage: Wan-Image supports text-to-image, image-to-image, text-to-image-series, and text-image-to-image-series generation.
- Text-to-image: Text-to-image generation supports ultra-long text rendering, extreme aspect ratios, hyper-diverse portraits, palette-guided generation, and alpha-channel generation.
- Image-to-image: Image-to-image generation performs instruction-based editing and interactive manipulation while preserving layouts and identity consistency.
- Sequential generation: Series generation maintains logical progression, visual continuity, reference style, identity, and temporal consistency across sequences.
- Training evolution: The transition from PT to SFT improves instruction following and localized detail refinement.
- Reinforcement learning: Post-RL outputs are described as more natural and visually appealing, with improved global harmony compared with pre-RL outputs.
6 Conclusion
Wan-Image is presented as a unified, high-performance visual generation system for professional productivity workflows. It combines broad controllability and generation capabilities with competitive performance across diverse visual tasks.
- Wan-Image integrates an MLLM-based Planner with a DiT-based Visualizer to support professional visual generation workflows.The Planner provides semantic reasoning, while the Visualizer supports pixel synthesis.
- The system supports ultra-long complex typography, extreme aspect ratios up to 1:8, and high-efficiency 4K generation without compromising speed.
- Wan-Image provides fine-grained control through palette-guided generation, realistic portrait steering, and multi-subject identity preservation.
- It also supports coherent image series, precise multimodal interactive editing, and true alpha-channel generation for design integration.
- Extensive evaluations report highly competitive performance against state-of-the-art models for complex real-world design workflows.
7 Authors
This section lists the paper’s contributors and presents visual demonstrations and comparisons covering multiple generation tasks and training stages.
- Authors: The paper lists Chaojie Mao through Zhen Han as core contributors and Ang Wang through Zhicheng Zhang as additional contributors.
- Authors: Core and additional contributors are listed alphabetically by first name according to the contributor notes.
- Visual demonstrations: Figures demonstrate text-to-image, image-to-image, interactive editing, text-to-series, and text-and-image-to-series generation capabilities.
- Visual demonstrations: Additional demonstrations cover logical image series, extreme aspect ratios, alpha-channel generation, and ultra-long text rendering.
- Comparisons: The section includes visual comparisons across models, training stages, Cascade RL, and settings with or without PE.