Source-linked AI summary

Scaling Up to Excellence: Practicing Model Scaling for Photo-Realistic Image Restoration In the Wild

Fanghua Yu, Jinjin Gu, Zheyuan Li, Jinfan Hu, Xiangtao Kong, Xintao Wang, Jingwen He, Yu Qiao, Chao Dong

arXiv:2401.13627v2cs.CV

TL;DR

Image restoration needs stronger generative priors while preserving fidelity to degraded inputs. SUPIR scales the restoration model with SDXL, multimodal prompts, a 20-million-image dataset, negative-quality training, and restoration-guided sampling; experiments show strong perceptual quality and text-controlled restoration, alongside disadvantages on full-reference metrics.

  • Problem

    Improving image restoration requires stronger generative priors that raise perceptual quality and intelligence without compromising faithful recovery of the low-quality image.

  • Method

    SUPIR uses SDXL as a generative prior, a large adaptor and ZeroSFT connector, 20 million text-annotated high-quality images, multimodal prompts, negative-quality training, and restoration-guided sampling.

  • Results

    SUPIR achieves the best results on all reported non-reference metrics and supports image-content, quality, and semantic restoration control through textual prompts.

  • Takeaways & Limitations

    Model scaling, enriched image-text training, and controlled sampling expand restoration toward perceptually strong outputs that users can manipulate with text.

  • Takeaways & Limitations

    The method has disadvantages on full-reference metrics, despite achieving the best results on the reported non-reference metrics.

Abstract

from arXiv · show

We introduce SUPIR (Scaling-UP Image Restoration), a groundbreaking image restoration method that harnesses generative prior and the power of model scaling up. Leveraging multi-modal techniques and advanced generative prior, SUPIR marks a significant advance in intelligent and realistic image restoration. As a pivotal catalyst within SUPIR, model scaling dramatically enhances its capabilities and demonstrates new potential for image restoration. We collect a dataset comprising 20 million high-resolution, high-quality images for model training, each enriched with descriptive text annotations. SUPIR provides the capability to restore images guided by textual prompts, broadening its application scope and potential. Moreover, we introduce negative-quality prompts to further improve perceptual quality. We also develop a restoration-guided sampling method to suppress the fidelity issue encountered in generative-based restoration. Experiments demonstrate SUPIR's exceptional restoration effects and its novel capacity to manipulate restoration through textual prompts.

1. Introduction

SUPIR pursues large-scale intelligent image restoration by scaling the generative prior, training data, and supporting components. Its design combines efficient adaptation, text-enriched training, negative-quality guidance, and restoration-guided sampling to improve perceptual quality while preserving fidelity.

  • Model and adaptor scaling: SUPIR uses SDXL with 2.6 billion parameters as its generative prior and trains a large-scale adaptor with a ZeroSFT connector.The method also fine-tunes the image encoder for robustness to varied degradation and reduces computation by trimming ControlNet.
  • Training data: A dataset of over 20 million high-quality, high-resolution images with descriptive text annotations supports SUPIR’s scaled training.The training data is designed to provide a foundation for improving restoration quality and textual control.
  • Reported capabilities: SUPIR achieves the best visual quality across varied image-restoration tasks, particularly in complex and challenging real-world scenarios.The model also supports flexible restoration control through textual prompts, broadening how restoration can be specified.
  • Quality guidance: SUPIR integrates poor-quality samples and negative-quality prompts so sampling can guide outputs away from undesirable visual qualities.The approach uses prompts to improve perceptual effects while accounting for the model’s ability to interpret negative qualities.
  • Fidelity control: Restoration-guided sampling limits uncontrolled generation because excessive generative capacity can reduce fidelity to the low-quality input.The method selectively guides diffusion predictions toward the low-quality image during restoration.

2. Related Work

Image restoration has progressed from degradation-specific methods toward generative priors, but diffusion-based restoration remains constrained by the scale of its underlying generative models. Model scaling offers a route to stronger capabilities, although it introduces challenges involving design, data, computation, and other limitations.

  • Image Restoration: Traditional restoration methods target degradations such as super-resolution, denoising, and deblurring but often rely on specific degradation assumptions.Those assumptions limit generalization to other degradations and motivate blind image-restoration methods.
  • Generative Prior: Generative priors capture inherent image structures and support generation following the natural image distribution.Prior work uses GANs and other generative models, while this work focuses primarily on diffusion-based priors.
  • Generative Prior: Diffusion-based restoration methods remain constrained by the scale of their generative models, limiting further improvement in effectiveness.This limitation motivates exploring larger generative priors for image restoration.
  • Model Scaling: Model scaling has improved language, text-to-image, and image-segmentation models as their parameter counts and complexity increased.However, scaling is systematic and requires addressing model design, data collection, computing resources, and related limitations.

3. Method

SUPIR scales image restoration through a large SDXL generative prior, a dedicated adaptor, extensive annotated data, multimodal prompts, and restoration-guided sampling. These components improve low-quality image interpretation, controllability, perceptual quality, and fidelity during diffusion restoration.

  • Network Design: A degradation-robust encoder maps low-quality inputs into SDXL’s latent space while reducing the risk of interpreting degradation artifacts as image content.The encoder is fine-tuned with a fixed decoder against ground-truth images.
  • Training Data and Textual Modality: SUPIR trains on 20 million 1024×1024 high-quality, texture-rich images and supplements image information with textual descriptions and prompts.The system can generate descriptions from low-quality inputs and use them to guide restoration, including targeted completion of missing information.
  • Network Design: SUPIR uses SDXL with 2.6 billion parameters as a generative prior and adds a large-scale adaptor for image restoration.The adaptor includes a trimmed trainable encoder and a ZeroSFT connector to control SDXL efficiently.
  • Training Data and Textual Modality: Negative-quality prompts improve perceptual quality only when negative-quality samples are included during training.The authors add 100K SDXL-generated low-quality images so SUPIR can learn negative-quality concepts for classifier-free guidance.
  • Restoration-Guided Sampling: Restoration-guided sampling interpolates diffusion predictions with the low-quality latent to limit excessive generation and preserve fidelity.The guidance is strongest early in diffusion, when low-frequency image information is formed, and is controlled by τr.

4. Experiments

Experiments show that SUPIR restores challenging and real-world low-quality images with strong perceptual quality, while textual prompts and restoration-guided sampling provide controllable quality–fidelity trade-offs.

  • Comparison with Existing Methods: SUPIR accurately restores semantically correct textures and details under challenging degradation, unlike compared methods that produce incorrect details.
  • Comparison with Existing Methods: SUPIR achieves the best results on all non-reference metrics, although its visually stronger results do not necessarily lead to superior full-reference metrics.
  • Comparison with Existing Methods: On 60 real-world low-quality images, SUPIR achieves the best perceptual quality and significantly outperforms state-of-the-art methods in a 20-participant user study.
  • Controlling Restoration with Textual Prompts: SUPIR can restore or manipulate image content using textual prompts, but prompts become ineffective when they conflict with the low-quality input.
  • Ablation Study: The ZeroSFT connector preserves fidelity while retaining perceptual effects, outperforming zero convolution in full-reference performance.
  • Ablation Study: Training on large-scale high-quality data is important for SUPIR, with qualitative comparisons demonstrating its necessity.
  • Controlling Restoration with Textual Prompts: Negative-quality prompts improve perceptual quality, positive and negative prompts together yield the best perceptual results, and training without negative samples removes these gains.
  • Ablation Study: The restoration-guided sampling parameter τr controls the quality–fidelity trade-off: τr = 4 improves fidelity without compromising image quality and is selected by default.

5. Conclusion

SUPIR expands image restoration through model scaling, enriched training data, advanced design features, and controlled textual prompting, improving perceptual quality and restoration control.

  • SUPIR combines model scaling, dataset enrichment, and advanced design features to enhance perceptual quality and enable controlled textual prompts for image restoration.

A.1. Degradation-Robust Encoder

The degradation-robust encoder is placed before the adaptor to reduce the effects of degradation in low-quality inputs before diffusion processing.

  • The degradation-robust encoder reduces noise and blur from degraded low-quality inputs before their latent representations are decoded or supplied to the diffusion model.

A.2. LLaVA Annotation

SUPIR uses textual prompts during restoration, including automatically generated image descriptions and quality-oriented prompting, while negative prompts can introduce artifacts for semantically unclear inputs.

  • SUPIR combines LLaVA-generated image descriptions with a standardized positive quality prompt during textual restoration.
  • Negative prompts can cause artifacts when low-quality inputs lack clear semantics.

A.3. Limitations of Negative Prompt

Negative-quality prompts improve restored-image quality, but they can introduce artifacts when the restoration target lacks clear semantic definition. This limitation is attributed to misalignment between low-quality inputs and language concepts.

  • Negative-quality prompts substantially improve the image quality of restored images.
  • Negative prompts may introduce artifacts when the restoration target lacks clear semantic definition.
  • The issue likely stems from misalignment between low-quality inputs and language concepts.

A.4. Negative Samples Generation

SUPIR addresses the difficulty of using negative prompts by distilling negative concepts from SDXL and generating negative samples through image-to-image transformation. Direct text-to-image sampling is avoided because it produces meaningless outputs.

  • SUPIR distills negative concepts from the SDXL model to help the fine-tuned model comprehend negative-quality prompts.
  • Direct text-to-image sampling of negative samples often produces meaningless outputs.
  • Negative samples are generated from dataset training images using an image-to-image process with strength setting 0.5.

B. More Visual Results

Additional visual results compare SUPIR with other methods, examine metric–human assessment mismatches, and show the effects of negative samples and textual prompts. The examples highlight restoration of textures and details, controllable restoration, and artifacts caused by omitting negative samples during training.

  • Metric and human assessment: Full-reference metrics can disagree with human evaluation when SUPIR produces high-fidelity textures.
  • Negative samples: Using negative-quality prompts without negative samples in training may cause artifacts, whereas including negative samples enables further quality enhancement through CFG.
  • Annotations: LLaVA accurately predicts most image content even from low-quality inputs, supporting the annotation process used in the restoration pipeline.
  • Negative sample generation: Direct noise-to-image sampling yields meaningless negative samples, while synthetic generation from high-quality images provides the alternative shown in the pipeline.
  • Qualitative comparisons: SUPIR restores textures and details accurately under challenging degradation in qualitative comparisons with DiffBIR and PASD.
  • Textual prompt control: SUPIR supports controllable restoration through positive and negative textual prompts.
Loading 2401.13627v2…