Source-linked AI summary

Forget-Me-Not: Learning to Forget in Text-to-Image Diffusion Models

Eric Zhang, Kai Wang, Xingqian Xu, Zhangyang Wang, Humphrey Shi

arXiv:2303.17591v1cs.CVcs.AIcs.LG

TL;DR

Text-to-image models raise privacy, copyright, and safety concerns because they learn unauthorized identities, content, styles, and potentially harmful materials. Forget-Me-Not targets specified concepts through attention re-steering, while M-Score and ConceptBench evaluate forgetting and memorization; experiments report successful concept diminution and correction, with scope limitations for abstract concepts and possible manual tuning.

  • Problem

    Text-to-image models can learn and reproduce unauthorized personal identities, artistic content, copyrighted material, and harmful content, creating privacy, copyright, and safety concerns.

  • Method

    Forget-Me-Not finetunes attention-related components to minimize attention maps associated with target concepts, and introduces M-Score and ConceptBench for evaluation.

  • Results

    Experiments demonstrate successful diminution and correction of target concepts in Stable Diffusion, while the method also supports concept manipulation through lightweight patches.

  • Takeaways & Limitations

    The approach provides a lightweight basis for ad-hoc concept forgetting, correction, disentanglement, and distribution through model patches.

  • Takeaways & Limitations

    The method faces challenges with abstract concepts, and successful forgetting may require manual concept-specific hyperparameter tuning.

Abstract

from arXiv · show

The unlearning problem of deep learning models, once primarily an academic concern, has become a prevalent issue in the industry. The significant advances in text-to-image generation techniques have prompted global discussions on privacy, copyright, and safety, as numerous unauthorized personal IDs, content, artistic creations, and potentially harmful materials have been learned by these models and later utilized to generate and distribute uncontrolled content. To address this challenge, we propose \textbf{Forget-Me-Not}, an efficient and low-cost solution designed to safely remove specified IDs, objects, or styles from a well-configured text-to-image model in as little as 30 seconds, without impairing its ability to generate other content. Alongside our method, we introduce the \textbf{Memorization Score (M-Score)} and \textbf{ConceptBench} to measure the models' capacity to generate general concepts, grouped into three primary categories: ID, object, and style. Using M-Score and ConceptBench, we demonstrate that Forget-Me-Not can effectively eliminate targeted concepts while maintaining the model's performance on other concepts. Furthermore, Forget-Me-Not offers two practical extensions: a) removal of potentially harmful or NSFW content, and b) enhancement of model accuracy, inclusion and diversity through \textbf{concept correction and disentanglement}. It can also be adapted as a lightweight model patch for Stable Diffusion, allowing for concept manipulation and convenient distribution. To encourage future research in this critical area and promote the development of safe and inclusive generative models, we will open-source our code and ConceptBench at \href{https://github.com/SHI-Labs/Forget-Me-Not}{https://github.com/SHI-Labs/Forget-Me-Not}.

1. Introduction

Text-to-image models create growing privacy, copyright, safety, and fairness concerns because their large-scale training data are difficult to filter safely. Forget-Me-Not addresses this with low-cost concept forgetting, quantitative evaluation tools, and extensions for content removal and correction.

  • Large-scale text-to-image models raise security, fairness, regulation, copyright, and safety concerns as their applications expand.
  • Web-scraped public datasets often lack human-level quality assurance, while private training sources cannot be determined at scale.
  • Domain adaptation to cleaner datasets remains laborious and can severely impair out-of-domain image synthesis.
  • Forget-Me-Not provides plug-and-play concept forgetting in as few as 35 optimization steps, typically taking about 30 seconds.
  • M-score and ConceptBench quantitatively measure models’ capacity to synthesize target concepts and assess memorization and forgetting.
  • Extensive studies report the method as simple, low-cost, and effective, with applications including harmful or NSFW removal and biased concept correction and disentanglement.

2. Related Works

Prior work developed increasingly capable conditional generative models and efficient adaptation methods, but this diversity also introduced privacy and copyright risks. Forget-Me-Not differs by targeting concepts rather than only specific data points.

  • Conditional generative models use class labels, image instances, and other signals as conditioning for image generation.
  • DreamBooth adapts models by finetuning all weights, whereas Textual Inversion and LoRA add smaller sets of extra weights.
  • Stable Diffusion’s diversity can enable privacy leakage and copyright infringement through retrieval of samples faithful to training examples.
  • This work studies text-to-image forgetting and deletes concepts behind reference data rather than only requested data points.
  • Concept forgetting is defined as disentangling concept prompts from visual contents while retaining generative abilities as much as possible.

3. Method

Forget-Me-Not defines concept forgetting as disentangling concept prompts from visual contents, then uses attention resteering to reduce target-concept attention while preserving unrelated generations. The method is designed as a low-cost, plug-and-play procedure applicable across concepts and models, with optional inversion for unclear or out-of-vocabulary concepts.

  • Concept Forgetting: Concept forgetting disentangles designated concept prompts from visual contents so the model retains its generative abilities as much as possible.
  • Concept Forgetting: Forget-Me-Not targets four goals: removing designated concepts, preserving other concepts, covering diverse concepts, and supporting varied models and tasks.
  • Forget-Me-Not: Forget-Me-Not is presented as a low-cost, plug-and-play approach that can remove varied concepts while limiting changes to other outputs.Its attention-resteering methodology is described as fitting major text-to-image models and potentially extending to other conditional multimodal generators.
  • Baselines: Token Blacklisting can be reversed through token inversion and may affect shared concepts, whereas Naive Finetuning can corrupt unrelated ideas.These baselines illustrate the performance-or-integrity dilemma that attention resteering is designed to address.
  • Attention Resteering: Attention resteering locates target context embeddings, computes their attention maps, minimizes those maps, and backpropagates through selected cross-attention layers.In Stable Diffusion, the method applies attention resteering across all UNet cross-attention layers and decouples finetuning from the diffusion variational lower-bound objective.
  • Optional Concept Inversion: Optional textual inversion strengthens generality when target concepts are out of vocabulary, lack model vocabulary, or have unclear descriptions.

4. Experiments

The experiments evaluate Forget-Me-Not across benchmark design, qualitative concept forgetting and correction, NSFW removal, memorization measurement, and ablations. Results show targeted concepts can be reduced while preserving other content, though closely related concepts and some model capabilities may be affected.

  • ConceptBench: ConceptBench measures memorization and forgetting across identity, object, and style concepts, spanning discrete-to-abstract and easy-to-hard instances.Its identity coverage includes people, franchises, animals, and brands, while styles include Van Gogh, Picasso, doodle, pixel art, neon, and sketch.
  • Baselines: Naive token removal and unrelated-image finetuning can deteriorate or shift generation, with unrelated-image finetuning distorting other concepts through selected visual details.The paper also notes that relation-based concepts cannot be exhaustively tested for blacklisting or finetuning.
  • Qualitative Comparison: Forget-Me-Not forgets multiple target concepts while preserving content and visual quality on controls, although minor pose and style changes appear in closely related concepts.The qualitative study targets Elon Musk and Taylor Swift and evaluates man, woman, Bill Gates, and Emma Watson as related concepts.
  • Memorization Measurement: The Memorization Score measures concept-embedding changes between original and forgetting models using textual inversion of anchor images.For Elon Musk, the procedure compares the reference concept embedding with embeddings from original-model and forgetting-model inversions.
  • Concept Correction: Concept correction makes lesser semantics more prominent after diminishing a dominant concept, including alternative James Bond actors and competing Mulan or apple interpretations.The method also produces a new painting style after forgetting Picasso and Van Gogh, and is reported to provide more comprehensive forgetting than negative prompts.
  • NSFW Removal and Ablations: Forget-Me-Not removes the “naked” concept from NSFW generations without additional data or third-party detectors, while cross-attention finetuning avoids the greater sensitivity of full-UNet optimization.The NSFW results change or remove sensitive cues, whereas the ablation reports successful forgetting for both UNet and cross-attention settings.

5. Conclusion

The study introduces Forget-Me-Not for ad-hoc concept forgetting in text-to-image models, extending it to concept correction and disentanglement. Experiments show successful target-concept diminution and correction in Stable Diffusion, supported by new evaluation metrics.

  • Forget-Me-Not enables ad-hoc concept forgetting using only a few real or generated concept images.The approach is lightweight and can be distributed through model patches.
  • The method naturally extends to concept correction and disentanglement.
  • Experiments demonstrate successful diminution and correction of target concepts in Stable Diffusion.
  • ConceptBench and Memorization Score provide evaluation metrics for concept forgetting and manipulation.
  • The work provides a foundation for further research on concept forgetting and manipulation in text-to-image generation.The authors state that it may extend to other conditional multimodal generative models to improve accuracy, inclusion, and diversity.

6. Social Impact & Limitations

The work presents Forget-Me-Not as a cost-efficient way to remove and correct harmful or biased concepts in text-to-image models. Its lightweight patches are designed for convenient distribution, while the approach remains challenged by abstract concepts and may require manual tuning.

  • Social Impact: Forget-Me-Not offers an effective and cost-efficient method to remove and correct harmful and biased concepts.
  • Social Impact: Lightweight model patches can be conveniently distributed to text-to-image model users.
  • Social Impact: The research aims to support fairness and privacy protection in AI tools.
  • Limitations: The approach performs well on concrete concepts in ConceptBench but faces challenges identifying and forgetting abstract concepts.
  • Limitations: Successful forgetting may require manual interventions such as concept-specific hyperparameter tuning.
Loading 2303.17591v1…