Source-linked AI summary

Null-text Inversion for Editing Real Images using Guided Diffusion Models

Ron Mokady, Amir Hertz, Kfir Aberman, Yael Pritch, Daniel Cohen-Or

arXiv:2211.09794v1cs.CV

TL;DR

Editing real images with text-guided diffusion requires inversion that preserves both image fidelity and editing capability. The paper combines pivotal inversion with null-text optimization, achieving high-fidelity reconstruction and enabling Prompt-to-Prompt editing without tuning model weights.

  • Problem

    Text-guided diffusion editing of real images lacks an inversion method that faithfully reconstructs the image while preserving meaningful editing capabilities.

  • Method

    The method uses a DDIM-derived pivotal trajectory and optimizes only per-timestamp unconditional null-text embeddings for inversion.

  • Results

    The approach achieves near-perfect, high-fidelity reconstruction while retaining Prompt-to-Prompt editing on real images without tuning model weights.

  • Takeaways & Limitations

    A single inversion supports multiple intuitive local and global text edits while preserving the original image identity.

  • Takeaways & Limitations

    Inverting one image takes approximately one minute on a GPU, which is insufficient for real-time applications.

Abstract

from arXiv · show

Recent text-guided diffusion models provide powerful image generation capabilities. Currently, a massive effort is given to enable the modification of these images using text only as means to offer intuitive and versatile editing. To edit a real image using these state-of-the-art tools, one must first invert the image with a meaningful text prompt into the pretrained model's domain. In this paper, we introduce an accurate inversion technique and thus facilitate an intuitive text-based modification of the image. Our proposed inversion consists of two novel key components: (i) Pivotal inversion for diffusion models. While current methods aim at mapping random noise samples to a single input image, we use a single pivotal noise vector for each timestamp and optimize around it. We demonstrate that a direct inversion is inadequate on its own, but does provide a good anchor for our optimization. (ii) NULL-text optimization, where we only modify the unconditional textual embedding that is used for classifier-free guidance, rather than the input text embedding. This allows for keeping both the model weights and the conditional embedding intact and hence enables applying prompt-based editing while avoiding the cumbersome tuning of the model's weights. Our Null-text inversion, based on the publicly available Stable Diffusion model, is extensively evaluated on a variety of images and prompt editing, showing high-fidelity editing of real images.

1. Introduction

The paper addresses the difficulty of inverting real images for text-guided diffusion editing. It introduces an inversion scheme combining pivotal inversion with null-text optimization to achieve high-fidelity reconstruction while preserving editability.

  • 1. Introduction: Text-guided editing of real images requires inversion that reconstructs the input while preserving editing capabilities, but classifier-free guidance makes existing DDIM inversion inadequate.The guidance mechanism amplifies accumulated inversion errors, causing artifacts and reduced editability.
  • 1. Introduction: The method achieves near-perfect reconstruction while retaining the original model’s text-guided editing capabilities.
  • 1. Introduction: Null-text optimization updates only the unconditional embedding used in classifier-free guidance, leaving the conditional text embedding and model weights unchanged.The optimized embedding replaces the empty-text embedding.
  • 1. Introduction: Diffusion Pivotal Inversion uses an initial DDIM trajectory as an anchor and optimizes locally around it rather than mapping many noise vectors to one image.The initial inversion is inadequate by itself but provides a promising starting point.
  • 1. Introduction: The approach enables Prompt-to-Prompt text editing on real images without tuning model weights.The paper reports high-fidelity reconstruction alongside meaningful and intuitive editing abilities.

2. Related Work

Prior work established powerful text-to-image generation and several image-editing strategies, but real-image editing remained constrained by inversion, masking, global edits, or model fine-tuning. This work positions its approach as enabling intuitive Prompt-to-Prompt editing directly on real images without fine-tuning.

  • 2. Related Work: Large-scale diffusion models substantially advanced text-conditioned image synthesis, motivating efforts to adapt their semantic capabilities for image editing.
  • 2. Related Work: Existing inversion and textual-inversion methods can regenerate concepts but struggle to preserve unedited details when editing a specific real image.
  • 2. Related Work: Noise-based editing methods may lose input details, while mask-based methods require burdensome precise masks and can discard information used by inpainting.
  • 2. Related Work: Prompt-to-Prompt preserves spatial layout and geometry through cross-attention injection but, without inversion, is limited to synthesized images.
  • 2. Related Work: Competing approaches such as Imagic and UniTune require model fine-tuning, whereas the proposed method applies Prompt-to-Prompt editing to real images without fine-tuning.

3. Method

The method inverts a real image into the latent domain of a text-guided diffusion model, then supports text-only editing from the recovered representation. It combines a DDIM-derived pivot trajectory with optimized unconditional embeddings while retaining the model and conditional prompt.

  • 3. Method: The task is to reconstruct a real image from a source prompt while retaining intuitive text-based editing through an edited prompt.The setup follows Prompt-to-Prompt and can use an automatically generated source caption.
  • 3.1. Background and Preliminaries: The implementation uses Stable Diffusion’s latent image encoding and decoder, with deterministic DDIM sampling for reconstruction.
  • 3.1. Background and Preliminaries: Classifier-free guidance combines conditional and unconditional noise predictions using guidance scale w, with Stable Diffusion’s default set to w = 7.5.
  • 3.2. Pivotal Inversion: DDIM inversion with classifier-free guidance accumulates amplified errors, producing artifacts and potentially leaving the noise vector outside the Gaussian distribution needed for editability.
  • 3.2. Pivotal Inversion: The method uses a w = 1 DDIM inversion as a highly editable pivot, then optimizes around that trajectory with the standard guidance scale w > 1.Separate optimizations are performed for timestamps in the order t = T → t = 1.
  • 3.3. Null-text optimization: Null-text optimization updates the unconditional embedding in classifier-free guidance while keeping the model weights and conditional embedding intact.This yields high-quality reconstruction and supports multiple Prompt-to-Prompt edits after one inversion.

4. Ablation Study

The ablation study evaluates pivotal inversion and null-text optimization, showing that their combination improves reconstruction fidelity and supports effective editing.

  • Ablation Study: Per-timestamp null-text embeddings converge more effectively than one global embedding shared across all diffusion timestamps.The global embedding is less expressive and struggles to converge.
  • Ablation Study: The method remains robust to the input caption for reconstruction, although editable regions must appear in the source caption to obtain suitable semantic attention maps.Random captions can still produce optimal reconstruction with respect to the VQ auto-encoder, but editing quality depends on captioning the relevant content.
  • Ablation Study: Pivotal inversion substantially improves optimization compared with random noise, while preserving better editability than textual inversion with a pivot.Textual inversion with a pivot achieves comparable reconstruction but reduces Prompt-to-Prompt editability because its attention maps are less accurate.
  • Ablation Study: Optimizing null-text embeddings from random noise completely breaks the method and performs worse than the DDIM inversion baseline.This result supports using a single optimized trajectory rather than independently sampled noise vectors.

5. Results

The results show that null-text inversion enables high-fidelity, text-only editing of real images across textures, structured objects, and multiple editing operations. Comparisons indicate stronger detail preservation than competing text-only and mask-based approaches in the reported examples.

  • Real Image Editing: Prompt-to-Prompt editing, previously limited to synthesized images, is applied to real images through the proposed inversion technique.The method retains high reconstruction quality and editability without requiring a separate inversion for each demonstrated edit.
  • Real Image Editing: The method achieves realistic editing of textures and structured objects while preserving original image details and identity.Examples include changing clothing, hair, glasses, expression, backgrounds, lighting, and objects from a single inversion.
  • Qualitative Comparison: Compared with text-only baselines, VQGAN+CLIP produces less realistic results, Text2LIVE struggles with structured objects, and SDEdit exhibits artifacts and identity drift.The comparison covers 100 images including humans and animals.
  • Comparison to Mask-Based Methods: Mask-based methods preserve regions outside the mask but often lose structural details inside the edited region, whereas the proposed method preserves those details.The reported example notes that mask-based editing does not preserve basket size.
  • Quantitative Comparison: Quantitative evaluation uses a user study because ground truth is unavailable for text-based editing of real images.Participants rated fidelity to both the input image and the textual edit instruction across the compared methods.
  • SDEdit Extension: The method improves SDEdit results by preserving the baby’s identity at matched CLIP similarity.The SDEdit parameter controls a trade-off between reconstruction fidelity and text alignment, so different parameters were selected to equalize CLIP scores.

6. Limitations

The method has practical limitations in inference speed, Stable Diffusion artifacts and attention quality, and the scope of Prompt-to-Prompt structural edits.

  • Inverting one image takes approximately one minute on GPU, preventing real-time applications despite later edits taking only ten seconds each.The method supports unlimited editing operations after inversion, but the initial inversion remains time-consuming.
  • Stable Diffusion’s VQ auto-encoder can introduce artifacts, especially in images containing human faces.Optimizing the VQ decoder is outside the paper’s scope because this issue is specific to Stable Diffusion.
  • Stable Diffusion attention maps may associate words with incorrect regions, weakening text-based editing compared with Imagen.The paper identifies attention-map accuracy as a limitation of the underlying Stable Diffusion model.
  • Prompt-to-Prompt cannot handle complicated structural changes such as turning a seated dog into a standing dog.The inversion approach is presented as orthogonal to the particular model and editing technique.

7. Conclusions

The paper concludes that pivotal inversion followed by null-text optimization reconstructs real images accurately while preserving text-based editing. This avoids model-weight tuning and supports efficient repeated edits.

  • Pivotal inversion and null-text optimization bridge accurate reconstruction and editable latent representations for real images.DDIM inversion supplies noisy codes as a pivot, while null-text optimization compensates for classifier-free guidance errors.
  • Prompt-to-Prompt editing can be applied after the image-caption pair is accurately embedded in the model’s output domain.The paper describes editing as immediately available at inference time after inversion.
  • The approach reconstructs arbitrary images without computationally intensive model tuning, preserving the trained model’s prior.The method modifies the unconditional embedding rather than the model weights or conditional embedding.
  • The method enables Prompt-to-Prompt editing on real images while retaining meaningful and intuitive manipulation abilities.The paper presents this capability as a first for applying Prompt-to-Prompt to real images.

B. Ablation Study

The ablations show that guidance scale, pivotal inversion, and the choice of conditioning embedding affect reconstruction and editability. The method is robust to different captions but requires editable source content for targeted edits.

  • DDIM Inversion: Increasing DDIM inversion guidance scale worsens initial reconstruction and editability, motivating the choice w = 1.The evaluation uses log-likelihood and PSNR, with larger guidance values producing less editable latent vectors and poorer reconstruction.
  • Robustness to different input captions: The inversion remains robust across multiple input captions, but edited regions must appear in the source caption for semantic attention maps.For example, editing a shirt print requires a source caption mentioning a shirt with a drawing or similar content.
  • Inference time comparison: The method allows multiple editing operations after a single inversion, unlike approaches requiring repeated inversion or slower processing.Table 2 compares inversion and editing time across methods and identifies the proposed method as more efficient than the other listed baselines.
  • Null-text optimization without pivotal inversion: Null-text optimization without pivotal inversion produces low-quality reconstructions.The non-pivotal variant fails to provide the efficient anchor needed for the optimization.
  • Textual inversion with a pivot: Textual inversion around a pivot achieves comparable reconstruction but reduces editability relative to null-text optimization.Its attention maps are less accurate, causing desert edits to introduce artifacts over the goats.

C. Additional results

Additional results compare the method with SDEdit and Imagic on fidelity, editing, and runtime. The reported comparisons favor the proposed approach for preserving original details and supporting repeated edits.

  • Additional editing results: Additional editing examples for the proposed method are provided in Figure 10.
  • Inference time comparison: The method provides accurate reconstruction in approximately one minute while allowing multiple editing operations after one inversion.It is reported as more efficient than Text2Live, VQGAN+CLIP, and Imagic, whereas SDEdit is faster but fails to preserve unedited details.
  • Comparison to Imagic: Lower LPIPS indicates that the method preserves original image details better than Imagic in the quantitative comparison.Imagic also struggles to retain background content and is sensitive to the interpolation parameter α.

D. Implementation details

The implementation uses Stable Diffusion with deterministic DDIM sampling and optimizes null-text embeddings across diffusion timestamps. The appendix also specifies baseline implementations and contrasts timestamp-specific optimization with a slower global variant.

  • Implementation settings: Stable Diffusion experiments use a DDIM sampler with T = 50 diffusion steps and guidance scale w = 7.5.Stable Diffusion uses a pretrained CLIP network as its language model.
  • Optimization settings: The full inversion procedure uses N = 10 optimization iterations, a 0.01 learning rate, and early stopping at ϵ = 1e −5.An input image and caption take 40s −120s to invert on a single A100 GPU.
  • Null-text inversion: The method optimizes the null-text embedding across diffusion timestamps, producing a noise vector zT and optimized embedding ∅ from a source prompt and input image.The algorithm computes intermediate results z∗ T, . . . before optimizing across timestamps.
  • Global null-text inversion: Global null-text inversion converges much more slowly, requiring 7500 optimization steps, or about 30 minutes, to accurately reconstruct the input image.This variant optimizes one null-text embedding shared across all timestamps.
  • Diffusion background: DDPMs define a forward Markov process that gradually adds Gaussian noise from x0 to xT and a learned reverse process that generates data from noise.DDIM sampling makes the reverse process deterministic by setting σt = 0 at every step.

F. User-Study

The paper provides an illustration of the user study in Figure 18.

  • User study: Figure 18 illustrates the user study setup.

G. Image Attribution

The appendix provides source attributions for the real images and captions used in the figures, alongside additional robustness, ablation, comparison, and user-study visuals.

  • Image sources: The appendix attributes example images to Unsplash, Pixabay, and Flickr, including the blue-haired woman image used with its input caption.The listed input caption is “A woman with a blue hair.”
  • Robustness and ablations: Robustness and ablation figures examine caption wording, optimization iterations, and replacing null-text optimization with textual inversion using a pivot.The textual-inversion comparison reports less accurate attention maps and distorted goat heads during editing despite high-fidelity reconstruction.
  • Comparisons and study materials: Additional figures compare the method with Text2LIVE, VQGAN+CLIP, SDEdit, and Imagic, and include an illustration of the user study.The Imagic comparison includes Stable Diffusion interpolation values α = 0.6, 0.7, 0.8, 0.9 and Imagen values α = 0.93, 0.86, 1.08.
Loading 2211.09794v1…