Source-linked AI summary
I-CARE: Analysis of interference-related phenomena in a controllable, diverse and representative unlearning setting for text-to-image models
Leonardo Santiago Benitez Pereira, Marcos Escudero Viñolo, Luis Herranz Arribas
TL;DR
Generative unlearning can damage retained semantically related concepts, but this interference has lacked systematic characterization. I-CARE formalizes tasks, metrics, and reporting templates, and its feasibility demonstration reveals method-dependent interference patterns while supporting reusable analysis and open-source exploration.
Problem
Interference in generative unlearning has not been systematically analyzed, and existing evaluations do not precisely quantify degradation of retained concepts.
Method
I-CARE provides a reusable methodology with formal concepts, constrained study designs, standardized Result Templates, and open-source software with the Forgety interface.
Results
The feasibility demonstration found initial differences among SPARE, MUNBa, and UCE in the extent and predictability of interference across three tasks.
Takeaways & Limitations
I-CARE provides a starting point for standardized, reusable interference studies and supports future comparisons across models and unlearning methods.
Takeaways & Limitations
General conclusions about interference remain difficult because all-vs-all evaluations found non-obvious patterns, including cases where similar entities did not interfere most.
Abstract
from arXiv · showhide
Machine unlearning studies the removal of knowledge from an AI model, making the system forget a concept it previously learned. Despite rapid progress in generative machine unlearning, the unintended degradation of semantically related concepts that should have been retained (henceforth, interference) remains poorly characterized and inconsistently evaluated. This paper introduces I-CARE, a methodology that formalizes interference as a first-class object of study in generative unlearning. Rather than proposing a new benchmark or unlearning algorithm, I-CARE provides formal definitions for tasks, metrics, and templates for reporting results, enabling the systematic and reproducible study of interference across unlearning settings. While our methodology is designed to remain valid as models and unlearning algorithms evolve, decoupling long-term scientific insight from transient empirical results, we present a feasibility demonstration with state-of-the-art algorithms and frequently used datasets. The results demonstrate that I-CARE enables meaningful analysis of interference patterns across multiple unlearning settings, establishing the practical applicability of the framework. The software implementation of the methodology is provided in an open-source framework, together with a web-based graphical interface that enables exploration of the outcomes of this study without requiring direct interaction with the codebase or specialized data analysis tools.
Introduction
Machine unlearning removes learned knowledge but can unintentionally degrade retained, semantically related concepts. I-CARE addresses the lack of systematic evaluation with a reusable methodology, feasibility demonstration, and open-source tools.
- Motivation: Machine unlearning removes data or concepts from AI systems while aiming to retain knowledge learned from remaining data.Its importance includes data-control requirements, bias mitigation, and model interpretability.
- Motivation: Existing image-generation unlearning methods struggle to remove target concepts while preserving unrelated capabilities, and this interference has not been systematically quantified.Prior evaluations used few target classes and often assumed interference followed semantic similarity.
- Contributions: I-CARE proposes a controllable, diverse, and representative methodology for analyzing interference across unlearning settings.The work also demonstrates non-random interference patterns and provides an open-source framework plus the Forgety visual interface.
- Contributions: The conceptual scheme is designed for repeated use with new base models and unlearning methods as the field advances.This separates the methodology’s longer-term utility from results that may become outdated.
- Foundations: I-CARE adopts consistent notation across unlearning formulations involving datasets, prompts, or metadata.The notation covers forget, retain, and optional overwrite concepts.
Proposed methodology
I-CARE defines interference studies through a formal vocabulary, constrained task design, and standardized result templates rather than fixing particular entities, models, or methods.
- Proposed methodology: I-CARE formalizes interference as unintended degradation of retained, semantically related concepts without assuming a specific cause or consequence.It therefore uses “interference” instead of “concept adjacency” or “collateral forgetting.”
- Proposed methodology: The methodology specifies an ontology, admissible operations, instantiation constraints, and reporting interfaces while leaving entities, models, and methods open.Each unlearned model generates images for all task entities so unlearning effects can be analyzed finely.
- Result reporting: The framework defines standardized Result Templates for aggregating, plotting, comparing, and reporting interference outcomes reproducibly.Templates include matrices, metric comparisons, alignment analyses, attribute relationships, and minimum-cut analysis.
- Notation: I-CARE distinguishes entities, attributes, tasks, models, unlearning methods, latent embeddings, similarities, and interference metrics as formal study concepts.Examples include Stable Diffusion, UCE, CLIP embeddings, and CLIP-based quality or interference measures.
- Metrics: MetricQuality measures image-generation quality, while MetricInterferencePerEntityPair and MetricInterferencePerEntity quantify degradation between entities or for individual entities.Pairwise interference can compare an entity’s CLIP score before and after unlearning another entity.
3.3 Attribute choice
I-CARE treats attribute selection, balanced data, controlled unlearning quality, and standardized result templates as requirements for interpretable interference analysis.
- Attribute choice: Attributes should be class-level when possible, unambiguous, immutable per entity, and annotated through a documented procedure that discloses bias.Authoritative datasets are preferred over custom inference rules.
- Data preparation: Retain sets should contain all other entities, with similar forget and retain image distributions across entities to reduce newly introduced external biases.These controls reflect the direct influence of training images on unlearning.
- Data balancing: Datasets should be balanced across socially relevant attribute combinations, with preprocessing such as removing rare categories or binning numerical attributes documented.This process is called data balancing.
- Evaluation design: I-CARE quantifies interference by unlearning each entity and evaluating effects on every other entity using pairwise and per-entity metrics.For model comparisons, it recommends at least three base models and three or more unlearning methods.
- Evaluation design: Equalization configures unlearning methods to achieve comparable unlearning quality so comparisons target interference rather than differences in forgetting effectiveness.Non-equalized studies must account for unlearning quality during analysis.
- Result Templates: Result Templates standardize outputs such as distributions, matrices, correlations, attribute relationships, and graph-based minimum cuts.MinimumCutInterference interprets interference as directed weighted paths whose removal can block influence between entities.
Feasibility demonstration
The feasibility demonstration applies I-CARE across balanced tasks, multiple unlearning methods, and diverse interference analyses. It reveals structured interference patterns, while showing that similarity alone is generally insufficient to predict interference.
- Experimental setup: The demonstration used one Stable Diffusion 1.4 model, three unlearning methods, four image-generation seeds, four similarity functions, five pairwise-interference metrics, and thirty-nine entity-level metrics across three 100-entity tasks.Entities were intersectionally balanced across selected attributes, and all valid test-relation combinations were evaluated.
- Attribute relationships: Sheepdogs and cattledogs caused the most interference when UCE unlearned breeds, whereas Spitz and primitive types caused the least.This pattern was detected for EmitterWorstInterferedClipDiff across groups.
- Attribute relationships: UCE and SPARE exhibited more significant attribute–interference channels than MUNBa across the evaluated combinations.Attributes were ranked over 147 combinations of 49 entity-level metrics and three unlearning methods.
- Directional relationships: Interference in scenes concentrated within the emitter’s own sport group under SPARE, with sport scenes affecting sport scenes and non-sport scenes affecting non-sport scenes.The directional analysis used ΔClip and separately conditioned emitters on sport and non-sport scenes.
- Similarity relationships: No single similarity function reliably predicted any pairwise-interference metric across all tasks and unlearning methods.Prediction improved slightly for restricted settings; for ΔDino and Clip under SPARE in scenes, Pearson correlation was −0.278.
- Matrix and graph analyses: Similarity matrices revealed occupation-based blocks, but the study explicitly found that this organization in similarity space did not translate into interference space.The interference matrix mainly showed self-effects along the diagonal and no large-scale organization under the tested sorting attributes.
4.4 Effects of the equalization
Equalization did not place the unlearning methods at a common operating point: MUNBa incurred much greater collateral damage, while UCE and SPARE stayed near zero interference. Within SPARE on scenes, only a few sessions jointly achieved strong forgetting and little interference.
- MUNBa sessions showed large collateral damage, whereas UCE and SPARE remained close to zero on the retain-interference axis.The methods still varied substantially despite attempted hyperparameter equalization.
- Equalized sessions remained dispersed along both forget-quality and retain-interference axes, so the methods did not share a common operating point.
- SPARE sessions on scenes formed a diffuse cloud, with only a handful on the Pareto front for strong forgetting and little interference.
- Pareto-optimal SPARE sessions tended toward higher embeddingSpecificityRatio values, with medians of 5.87 versus 5.66 for other sessions, although distributions overlapped substantially.The metric was therefore at best suggestive in this comparison.
4.5 Effects of retain data
The retain-data and evaluation analyses examine how interference varies with retained concepts, prompts, similarity, and session-level conditions. They reveal heterogeneous interference patterns, including similarity-linked harm in some sessions, occasional constructive changes, and important interpretive and implementation caveats.
- Retain-set composition was varied in a focused football-field experiment using Standard, Sport, and No sport compositions.The experiment measured ∆Clip on 10 entities, with lower values indicating stronger interference.
- Single-prompt and varied-prompt evaluations largely overlapped for per-emitter ∆Clip, with no systematic pattern observed.
- For football field, the negative relationship between DINOv2 similarity and ∆Clip persisted under both prompt settings, with approximately equal correlations.No statistical comparison between prompt variations was performed because of the small sample size.
- Unlearning ice skating rink indoor harmed several related scenes, ranging from lower image quality to images becoming moon-like after replacement by the moon.The five most interfered entities included arena hockey, ski slope, wrestling ring indoor, velodrome outdoor, and boxing ring.
- Some least-interfered entities showed small Clip-quality improvements, and similarity alone did not consistently identify the most affected entity.In the selected session, waterfall cascade and oil refinery outdoor were more similar to the target than boxing ring, yet boxing ring suffered more interference.
- The analysis compares ∆Clip with similarity metrics for Michael Schumacher and visualizes selected sessions through alignment plots, rankings, and generated images.The figures highlight the most and least interfered or similar entities and compare original with unlearned outputs.
- For giant schnauzer dog, the four most harmed entities matched the four most similar entities, while breed-specific features disappeared although a large dog remained generated.
Forgety architecture and features
Forgety is a no-code web interface for exploring I-CARE results and running unlearning sessions, with visualizations and session management. Its implementation is modular, open to future extensions, and currently requires Slurm for unlearning execution.
- Forgety enables no-code experimentation with unlearning so broader audiences can explore its benefits and shortcomings.
- Users can compute result templates on demand, list entities and attributes, inspect InterferencePerEntity values, and review prior sessions.
- A new unlearning session requires configuring a Slurm cluster, with the workflow documented through interface captures and a sequence diagram.
- The frontend architecture orchestrates these capabilities through multiple services and is intended to expand into a broader I-CARE-compatible results platform.
- Future work targets process-aware similarity, more sensitive domain-specific metrics, simpler controlled domains, matched seeds, and easier declarative tooling.
6.1 Mapping existing benchmarks into I-CARE
The paper proposes mapping existing unlearning benchmarks into I-CARE so their concepts, metrics, and reporting structures become more comparable and reusable. It also emphasizes that the methodology is intended as a durable framework, while the feasibility results remain limited in generality.
- 6.1 Mapping existing benchmarks into I-CARE: I-CARE notation could make heterogeneous benchmarks mutually comparable and allow reusable software components and result templates to be applied without reimplementation.
- 6.1 Mapping existing benchmarks into I-CARE: UnlearnCanvas is proposed as an initial mapping candidate, with 60 styles and 20 objects represented as 80 entities in one task.
- 6.1 Mapping existing benchmarks into I-CARE: Its Unlearning Accuracy can be formalized as InterferencePerEntity, while in-domain and cross-domain retain accuracies filter receiver entities by domain.
- 6.1 Mapping existing benchmarks into I-CARE: HUB’s dynamic receiver selection conflicts with I-CARE’s fixed all-versus-all design and would require additional sessions to recover that structure.
- 6.2 Generalization beyond unlearning: The authors present Specification-Driven Science as a future philosophy centered on reusable conceptual schemes rather than solutions for individual problems, but do not claim its foundations or novelty are established here.
- Conclusions: The feasibility study validates core methodology aspects, but its specific interference findings cannot support broad conclusions or generalization to other datasets.
- Conclusions: Across three methods, SPARE caused more ∆Clip than UCE, while MUNBa performed worse by all analyzed criteria.
Appendix A
Appendix A documents the experimental infrastructure, standardized artifacts, and formal procedures used to analyze attribute-related interference. It covers data organization, statistical tests, and stored outputs for reproducible computation.
- Experiments used Stable Diffusion v1.4 and approximately 1200 GPU hours on NVIDIA V100 and A100 hardware.
- The implementation used Python 3.10 with PyTorch, Diffusers, Transformers, PEFT, Vision-Unlearning, Hugging Face Hub, Poetry, Apptainer, and Makefile workflows.
- Standardized artifacts include filtered metadata, pairwise similarity matrices, forget and retain splits, unlearned models, generated data, interference dictionaries, entity-level metrics, and result-template outputs.
- InterferencePerPair stores emitter-to-receiver metrics, InterferencePerPairInverse aggregates received interference by entity, and InterferencePerEntity stores entity-level metrics with metadata.
- For categorical attributes, entities are partitioned by attribute value and statistical tests evaluate whether entity-level interference distributions differ across groups.
- The categorical procedure iterates over attribute values, collects metric values for entities in each group, and returns a statistical test result.
C.1 Considerations about numerical attributes
The appendix extends interference analysis from aggregate attribute associations to directional and similarity-based relationships. It also defines analyses for testing whether similarity predicts interference and for measuring shifts in protected-attribute associations.
- C.1 Considerations about numerical attributes: Numerical attributes are analyzed through correlations between attribute values and entity-level interference metrics, using statistics such as Pearson and Spearman coefficients.
- C.2 Directional analysis: Aggregate entity-level metrics cannot identify which source entities cause interference toward particular receiver groups.
- C.2 Directional analysis: The directional extension uses pairwise interference metrics to retain emitter and receiver identities and condition received interference on a chosen source group.
- C.2 Directional analysis: Directional group comparisons can be visualized with flow-based diagrams that highlight how interference moves between source and receiver groups.
- RT MetricSimilarityAlignment: MetricSimilarityAlignment represents each ordered entity pair with similarity features and pairwise interference outcomes, then estimates a regression model with validation.
- RT MetricSimilarityAlignment: For a removed entity, similarities to remaining entities are computed and the fitted model predicts expected receiver interference.
- RT MostSimilarInterferedMatrix: MostSimilarInterferedMatrix counts cases where the most-similar receiver appears among the top-k most-interfered receivers, comparing method-by-task cells against a null baseline.
- Association analysis: The iEAT analysis constructs original and unlearned embedding-alignment matrices, with ∆B capturing association shifts between target and protected attributes.
Appendix G
The appendix represents interference as a directed weighted graph whose edge weights encode pairwise interference. It restricts analysis to strong local relations and uses minimum cuts to identify emitter-side partitions separating two entities.
- Each task entity is a graph vertex, and directed edges represent ordered interference relations between distinct entities.
- An edge from e_i to e_j is included when the selected pairwise interference metric exceeds the positive threshold λ.
- Local-neighbourhood filtering retains the strongest interference relations, reducing computation and mitigating measurement noise.
- A minimum e1–e2 cut partitions the vertices into disjoint sets containing e1 and e2 separately while minimizing the total weight of directed edges crossing between them.
- The framework defines flows, cuts, residual capacities, and augmenting paths for computing cut structure in the directed interference graph.
APPENDIX H. MAX-FLOW / MIN-CUT AND ISOLATION DUALITY
The appendix establishes the max-flow/min-cut basis for interference isolation and describes how the resulting graph analysis is applied to datasets and attributes. It also records dataset-selection constraints and balancing choices for people, breeds, and scenes.
- Max-flow / min-cut: Ford–Fulkerson repeatedly augments feasible flows until no s–t path remains in the residual graph.
- Max-flow / min-cut: The reachable vertices in the final residual graph define a cut whose crossing edges are saturated.
- Max-flow / min-cut: The max-flow/min-cut theorem equates maximum feasible flow with minimum cut capacity.
- Isolation duality: Isolation duality equates the minimum total weight of an edge set blocking s→t paths with minimum cut capacity and maximum flow.
- Dataset selection: Dataset construction balances selected attributes into grids supporting 100 entities, while finer-grained attributes could leave insufficient statistical power.
- Dataset selection: Breed selection is constrained by the lack of sufficiently large authoritative sources for balanced selection at the required scale.
I.3 Metrics
I.3 defines entity similarities and pairwise interference measures, then aggregates them into emitter-, receiver-, and difference-based metrics. It additionally introduces an embedding-specific ratio to distinguish target-specific visual change from background drift.
- Four entity similarities use CLIP text embeddings, DINOv2 image embeddings, UNet cross-attention activations, and annotated-attribute overlap.
- Five pairwise interference metrics compare same-seed images before and after unlearning, covering semantic alignment, image quality, representation change, pixel change, and structural change.
- The framework derives 49 entity-level interference metrics by grouping five pairwise metrics into 15 buckets and applying aggregation functions.
- The metrics distinguish emitter impact on other entities, receiver impact from other entities, and the difference between those values.
- For ΔClip metrics, NumberOfInterferedWorseThanZero uses zero as the boundary between degradation and improvement.
- EmbeddingSpecificityRatio compares the forgotten entity’s visual shift with mean shifts for other entities, where a ratio ≫1 indicates a specific unlearning effect.
I.4.1 Hyperparameters
The appendix documents method-specific hyperparameters, reference-based equalization, and illustrative interference analyses. These examples show destructive failures, attribute-linked interference, distributed effects, and weak metric or layer alignments.
- Method hyperparameters: MUNBa exposes forget and retain image sets, learning rate, and epoch count, using unlearning by subtracting LoRA adapters.
- Method hyperparameters: UCE uses forget and retain concepts plus erase, preserve, and λ scales under a recommended guide-concept configuration.
- Method hyperparameters: SPARE overwrites dog breeds with “cat” or people with “child,” intentionally increasing destructiveness to expose unlearning and interference effects.
- Equalization procedure: Hyperparameters are manually selected on five entities before full analyses across 100 entities and three tasks, with UCE serving as the reference point.
- Illustrative analyses: Interference can spread across the embedding space rather than remain near the forgotten entity, with some nearby entities showing improvement.
- Illustrative analyses: The strongest layerwise correlations with ΔClip remain weak, while attribute and group analyses associate interference with fame, occupations, and within-group flows.