Source-linked AI summary

What to Forget in Unlearning? Forget Set Curation for Language Models

Animesh Jha, Arpandeep Khatua, Youssef Allouah, Sanmi Koyejo

arXiv:2608.14855v1cs.CL

TL;DR

Practical language-model unlearning must first determine which concrete data support a requested suppression, but this forget-set curation problem is usually assumed away. CleanSlate evaluates curation for songs and books and finds that strong suppression depends on the curator–unlearner combination and brings collateral suppression or capability loss.

  • Problem

    Most unlearning evaluations assume the forget set is known, leaving unclear how to map a work-level suppression request to the concrete spans or documents required by an unlearning algorithm.

  • Method

    CleanSlate benchmarks forget-set curation for verbatim suppression of songs and books using model-specific extraction profiles, content-grounded QA, and capability-retention evaluations.

  • Results

    Strong suppression depends on the curator–unlearner combination and is accompanied by collateral suppression; evaluation-aware curation nearly eliminates requested continuations but causes capability regressions and non-requested-content suppression.

  • Takeaways & Limitations

    Forget-set curation is part of the unlearning problem because the selected data determine both requested suppression and collateral damage.

  • Takeaways & Limitations

    CleanSlate studies the narrow setting of verbatim output suppression for songs and books and does not establish selective suppression across all tested operating points.

Abstract

from arXiv · show

Machine unlearning aims to remove targeted data or behaviors from a trained model without retraining from scratch. Yet most evaluations assume that the examples to forget are already known. In realistic language-model deployments, a requester may ask a model to stop reproducing a song or book without knowing which spans, documents, quotations, or near-duplicates in a trillion-token corpus support that behavior. We study this missing upstream problem, forget set curation: mapping a suppression request to the data passed to an unlearning algorithm. We introduce CleanSlate, a benchmark for verbatim output suppression over songs and books, with model-specific extraction profiles, content-grounded QA, and capability-retention evaluations. CleanSlate exposes two failure modes. Natural lexical and exact-substring curators often yield forget sets that lead to weak suppression. An evaluation-aware curator suppresses requested continuations almost completely, but causes collateral regression on non-requested content and model-dependent capability loss. These results show that practical unlearning is not only an optimization problem once a forget set is given: the data chosen for forgetting determines both what can be unlearnt and what else is damaged.

1 Introduction

The paper frames forget set curation—the selection of intervention data from a suppression request—as a central problem in language-model unlearning. Its results show that corpus footprint, curator–unlearner interaction, and evaluation-aware selection determine both target suppression and collateral damage.

  • Problem framing: Machine unlearning is typically evaluated after the forget set is supplied, leaving the request-to-data selection problem unexamined.The paper studies this upstream step as forget set curation.
  • Selective suppression: Selective suppression should make requested song or book continuations difficult to elicit while preserving factual knowledge, unrelated capabilities, and non-requested content.The target is verbatim output suppression rather than erasing an author, plot, genre, or cultural context.
  • Corpus footprint: Cultural works have diffuse corpus footprints spanning canonical copies, quotations, reviews, fan forums, news, code, synthetic examples, and incidental discussion.Supporting evidence may predate a work or appear in sources that do not resemble copies.
  • Selectivity gap: Evaluation-aware curation can suppress requested continuations almost completely, but it also suppresses non-requested content and causes substantial model-dependent capability regressions.Identifying target text therefore does not by itself construct a clean forget set.

2 Related Work

Related work studies memorization through verbatim extraction and evaluates unlearning by suppression and utility preservation. It also shows that forget-set composition affects outcomes, while prior request-level curation generally assumes unwanted examples or domains are already known.

  • Memorization and verbatim extraction: Memorization is commonly measured by prompting models with prefixes and testing whether they reproduce target suffixes.Probabilistic discoverable extraction extends this by measuring whether target continuations emerge under repeated sampling; extractability varies across copyrighted books and model families.
  • Machine unlearning benchmarks: Machine unlearning removes specified data or behaviors without retraining from scratch, using objectives such as loss ascent, preference optimization, logit adjustment, and representation interventions.Benchmarks assess targeted-information suppression alongside utility preservation.
  • Data selection, retrieval, and attribution: Forget-set contents can substantially change the suppression–preservation tradeoff, but prior work largely assumes unwanted examples or target domains are already known.Reported approaches include selecting small subsets or individual tokens and retrieving text overlapping with the target.
  • Data selection, retrieval, and attribution: Older works often have broad corpus footprints created by copies, quotations, discussion, and other web sources.The figure characterizes this pattern as outward diffusion and contrasts it with post-cutoff coverage from older or unrelated sources.

3 Problem Statement: Forget Set Curation

This section frames unlearning as a forget-set curation problem: a curator maps suppression requests to data for downstream unlearning while preserving non-targeted reproduction. Curators are evaluated by the behavior of the resulting model, including suppression, knowledge retention, and general capabilities.

  • 3 Problem Statement: Forget Set Curation: A suppression request targets works Wf whose verbatim reproduction should be suppressed while preserving reproduction of works Wr = W \ Wf.The setup defines W as a collection of works such as songs or books.
  • 3 Problem Statement: Forget Set Curation: Extractability is measured over prefix-suffix windows, with a window considered extractable when pz ≥ 0.001.The threshold τ = 0.001 is adopted from prior work.
  • 3 Problem Statement: Forget Set Curation: Given target texts, a trained model, and a search corpus, curator A returns a forget set Df and optionally a retain set Dr for downstream unlearning algorithm U.The resulting unlearnt model is produced by applying U to the curator’s outputs.
  • 3 Problem Statement: Forget Set Curation: The study evaluates curator A while holding the downstream unlearning procedure fixed unless otherwise stated.This isolates the effect of forget-set curation from the downstream unlearning algorithm.
  • 3 Problem Statement: Forget Set Curation: A successful curator makes target windows un-extractable without similarly affecting non-target windows, while retaining work knowledge and general capabilities.Evaluation includes verbatim suppression, content-grounded question answering, and general capability benchmarks.

4 The Corpus Footprint: Why Forget Set Curation Is Hard

A work’s corpus footprint is a distributed set of documents and spans, not merely its canonical copy, making forget-set curation difficult. Exact-overlap measurements show rapid outward diffusion, inward diffusion from preexisting language, and frequent overlap in extractable continuations.

  • The Corpus Footprint: A work’s corpus footprint includes lyric aggregators, forums, fan fiction, code snippets, and other documents or spans supporting a target continuation.The target is distributed across the corpus rather than isolated in a canonical copy.
  • Measuring literal overlap: Literal-overlap measurement is conservative because formatting, whitespace, or case changes can break matches, while paraphrases, translations, and semantic references remain undetected.The method measures localized exact word-level n-gram occurrences across every corpus position.
  • Outward diffusion: 100% coverage at both 5-gram and 50-gram levels shows Smile’s footprint rapidly formed in CCC25, with some spans reaching ≥105 and ≥104 occurrences, respectively.The passage describes Smile as released on 31st Dec 2024 and measured in CCC25 from Jan 2025.
  • Inward diffusion: Roughly 20% coverage at n = 7 in Cmid shows how preexisting language can diffuse inward into a later work through a quoted 1910 poem.The example is Red Terror by The Weeknd, whose release came after the corpus material.
  • Connection to target continuations: Extractable continuations often occur in regions with dense literal overlap, including song choruses, famous quotations, and repeated phrases.The paper treats this overlap as diagnostic rather than causal: high overlap does not prove that a particular source explains extraction.

5 CLEANSLATE

CleanSlate benchmarks end-to-end verbatim suppression by pairing model-specific extractability evidence with content-grounded QA and capability-retention tests. It covers songs and books, constructs fixed model-specific forget and retain pools, and evaluates curator–unlearner–model pipelines.

  • Benchmark design: CleanSlate pairs each suppression request with per-model extractability evidence and content-grounded QA to score the curator, unlearning algorithm, and edited model together.The benchmark evaluates the full request-to-data-to-update pipeline, while algorithm-focused evaluation varies one component with the other two fixed.
  • Content domains: The benchmark covers songs and books as complementary stress tests because songs are short and repetitive, whereas books are longer passages often entering corpora through commentary and excerption.Songs come from Billboard Hot 100 annual charts spanning 1970–2025; books include 50 works, mostly from Project Gutenberg.
  • Extraction profiles: A work is model-extractable when at least 5% of its 100-character prefix–suffix windows, sampled with stride 10, are extractable for that base model.Extractability is model-specific, so the same song can qualify for one model but not another.
  • Forget and retain pools: For each model, CleanSlate samples a suppression request of |Wf| = 50 from model-extractable works and reuses the same forget and retain pools across curators.This controls the comparison so observed differences are attributable to curation.
  • Content-grounded QA: CleanSlate-QA tests whether models retain factual knowledge about a work while suppressing its verbatim reproduction, using atomic statements grounded in named entities, places, numbers, or concrete events.The benchmark distinguishes factual knowledge retention from verbatim output suppression.
  • End-to-end evaluation: End-to-end evaluation measures baseline extractability and QA, applies each curator and a fixed unlearning algorithm, then re-measures QA change and six validation capabilities.The validation suite covers GSM8K, BBH, WinoGrande, CoQA, HumanEval+, and LAMBADA.

6 Experiments

Experiments across six models compare retrieval-based and evaluation-aware forget-set curation, while also testing how downstream unlearning algorithms interact with the curated sets. Retrieval effects depend strongly on model size, whereas evaluation-aware curation achieves strong forgetting but causes broad retain-side and capability damage.

  • Experimental setup: Experiments use CLEANSLATE with |Wf| = 50 across six models, fixing SimNPO for curator comparisons and fixing EA curation for unlearner comparisons.The models are LLAMA-3.1-8B, OLMO-3-7B, NEMOTRON-9B, QWEN3-8B, GEMMA-3-12B, and OLMO-3-32B.
  • Curators: BM25-pre, BM25-mid, and Infini-gram-mid provide retrieval-based curators, contrasted with an evaluation-aware baseline that directly accesses evaluation windows.All curators produce forget examples using a 100-character prefix followed by a 100-character suffix.
  • Retrieval-curator results: Retrieval engagement broadly tracks model size: the two largest models move most, intermediate models move partially, and the smallest models barely move on forget and retain sides.Selectivity is model-determined rather than retriever-determined; LLAMA-3.1-8B loses 51.2%–55.7% retain-pool extractability while gaining at most 28.9% on the forget side.
  • Evaluation-aware results: EA drives forget extractability to near-100% on every model, including small models that retrieval barely moves, but retain-pool extractability also falls throughout.This indicates direct target-window interventions can move models under SimNPO, while retrieval outcomes cannot be isolated to curator, model, or unlearner effects.
  • Unlearner comparison: Across SimNPO and UNDIAL, F reaches 100% on all three tested models while R remains 72–100%, yet capability outcomes differ and average capability drops by up to 10.4 pp.RMU reduces F on LLAMA-3.1-8B and OLMO-3-7B; on QWEN3-8B, F=24.5% and R=25.7%.

7 Discussion and Future Direction … D.2 Aggregate Coverage Statistics

The discussion frames forget-set curation as a central part of practical unlearning: the selected data determines both suppression success and collateral damage. The paper therefore advocates algorithm-aware, jointly evaluated curation, supported by large-corpus coverage analysis of CleanSlate works.

  • 7 Discussion and Future Direction: Forget-set curation determines whether requested continuations become difficult to elicit and what other model behaviors are disturbed.The forget set should not be treated as a fixed premise when requests identify works or behaviors rather than concrete training data.
  • 7 Discussion and Future Direction: Natural lexical and exact-substring retrieval often fails to produce selective verbatim suppression, while evaluation-aware curation causes non-requested suppression and capability regressions.The two approaches are incomplete for opposite reasons: weak target suppression versus insufficient localization of effects.
  • 7 Discussion and Future Direction: Curation and unlearning should be evaluated jointly because the same evaluation-aware forget set yields different capability profiles under SimNPO, UNDIAL, and RMU.Future curators may combine corpus-footprint signals with model-specific extraction profiles, dense or hybrid retrieval, influence estimation, or datamodeling.
  • A Additional Related Work: Prior work studies unlearning with gradient ascent, preference optimization, representation perturbation, and self-distillation, while RWKU specifies targets at the entity level.RWKU constructs forget data from model-generated synthetic examples rather than receiving an explicit training corpus.
  • B Compute Requirements / C Search Corpora (C) Details: Experiments use three corpus scales, and corpus search and forget-set curation are memory- and storage-bound at dataset scale.The compute node has 8 NVIDIA H200 GPUs, 230 CPU cores, 3TB of RAM, and 60TB of local NVMe storage; the corpora include Cmid and an approximately 11% Dolma-3 6T subset.
  • D Computing Coverage of Songs and Books in Large Corpus: Coverage is computed with Infini-gram-mini, using iterative retrieval to identify maximal exact n-gram matches at every word start position.The method preserves punctuation and spacing when measuring exact character spans in the indexed corpus.
  • D.1 Computing N−Gram Matches and Coverage: The coverage statistic measures the fraction of work word positions covered by corpus-present verbatim n-grams of length at least N, extending matches from nmin = 5.Shorter n-grams are excluded because linguistic coincidence gives them high background frequency.
  • D.2 Aggregate Coverage Statistics: 98.6% of CleanSlate works retrieve at least one positive-count 5-gram match in Cmid.This is 4,596 of 4,663 works; retrieved documents have a median of 119, a 90th percentile of 299, and a maximum of 45,338.

D.3 Extractability vs. Footprint Density Plots

The analysis compares local corpus footprint density with maximum extraction probability at each word position. Across songs, poems, and books, extractability spikes align closely with positions having the highest footprint density.

  • Method: Footprint density aggregates occurrence counts of all valid Cmid n-grams overlapping each word position and is compared with maximum extraction probability pz.The maximum is measured across all suffixes z covering position j.
  • Results: Figures 5 through 8 show that peaks in Cmid occurrences strongly align with spikes in extractability.The alignment is reported for Never Gonna Give You Up, Rocket Man, A Dream Within a Dream, and The Communist Manifesto.
  • Results: The model’s extraction probability rises precisely where footprint density is highest, including choruses of popular songs and famous poetic refrains.These regions exhibit massive frequency spikes in the corpus, mirrored by extraction probability.

E Baseline Extractability Patterns Across Model Families

Baseline extractability varies by model family, scale, tuning, target work, and text location. Models retain semantic knowledge of the works even when exact extraction is localized, making curation a selective suppression problem.

  • Model scale and capability: 32B OLMO-3 reaches 8.27% forget-pool extraction (Ext-F), versus 7.20% for 7B OLMO-3.Larger models consistently show higher raw extractability, while base models generally exceed instruction-tuned variants.
  • Heterogeneity of extractable text: Extractability is highly localized: repetitive, externally prominent song choruses often cross the threshold while verses may remain below it.The extraction threshold is pz ≥0.001, evaluated over sliding windows of 100 prefix and 100 suffix characters.
  • Heterogeneity of extractable text: Different models extract different windows, so a verse extractable for LLAMA-3.1-8B may fall below threshold for NEMOTRON-9B.This model dependence makes the contents of an effective forget set architecture-specific.
  • QA performance vs. verbatim extraction: CLEANSLATE-QA accuracy ranges from 22.01% to 35.07% pass@5, indicating factual knowledge alongside localized verbatim extraction.Curation must suppress pz spikes responsible for exact extraction while preserving broader semantic knowledge and general capabilities such as GSM8K.

F CleanSlate-QA Benchmark Construction … F.3 Stage 3: OOS Knowledge Filter and Judge

CleanSlate-QA is built in three stages: extracting interior factual propositions, converting them into atomic question-answer pairs, and retaining only questions answerable by at least one out-of-sample probe model. The construction targets content-grounded knowledge retention while excluding priors, generic statements, and source-memorization-only questions.

  • F.1 Stage 1: Proposition Extraction: Stage 1 extracts interior factual propositions anchored to named entities, places, numbers, or concrete events, while excluding title, creator, year, genre, generic themes, and unsupported quotations.Each proposition must cite a short source span from the work.
  • F CleanSlate-QA Benchmark Construction: CleanSlate-QA supplies the content-grounded retain metric reported as the QA column of Table 4, with its three-stage construction filtering candidates by out-of-sample model knowledge.The pipeline is explicitly designed to measure retained knowledge on the works’ interior content.
  • F CleanSlate-QA Benchmark Construction: The benchmark pipeline routes books through 60,000-character extraction windows and songs through single-call extraction, using Gemini for Stages 1–2 and Qwen plus Llama probes for Stage 3.These routing and model choices define the construction infrastructure.
  • F.1 Stage 1: Proposition Extraction: Song prompts request 3–8 propositions, whereas book-chunk prompts request 5–15 propositions covering concrete characters, settings, events, and numeric specifics.Abstract songs may yield fewer propositions rather than being padded.
  • F.2 Stage 2: Atomic QA Generation: Stage 2 converts each proposition into one self-contained question-answer pair naming both title and creator and targeting an interior detail.The answer must be short, 1–5 words, and a named entity, number, specific place, object, or proper noun.
  • F.2 Stage 2: Atomic QA Generation: Invalid pairs are rejected when answers are titles, creator substrings, question phrases, title-derived facts, pronouns, generic phrases, subjective moods, or hook echoes.The generator may return an empty list when no candidate satisfies the requirements.
  • F.3 Stage 3: OOS Knowledge Filter and Judge: Stage 3 presents each candidate question without surrounding context to out-of-sample probe models and keeps it iff at least one probe answers correctly.Correctness is judged by whether the model answer contains the reference answer or conveys the same meaning.

G Unlearning Training Details … I Scaling Forget Request

CleanSlate’s unlearning trainer operates on curated prefix–suffix rows, pairing each forget example with a randomly sampled retain example. Retrieval curators construct bounded, evaluation-shaped forget sets through corpus retrieval, window extraction, deduplication, and quota allocation, while keeping WikiText fixed as retain data.

  • G Unlearning Training Details: The trainer consumes curated rows containing a prefix and a suffix to suppress, pairing each forget item with a randomly sampled retain example.The prefix is x and the suppress-target suffix is z.
  • G Unlearning Training Details: SimNPO uses average suffix NLL, while UNDIAL and RMU apply distinct reference-model distillation and activation-level objectives.UNDIAL subtracts βU from gold-token teacher logits on forget suffix tokens; RMU matches forget activations to a random control vector at model.layers.7 and retains activation matching against a frozen reference.
  • G.1 Some example QA pairs: CleanSlate includes example content-grounded question-answer pairs for evaluation.The supplied passage identifies these as example QA pairs forming CleanSlate.
  • H Retrieval Curators and Forget-Set Construction Details: Each retrieval curator maps request text to prefix–suffix examples by retrieving corpus units, projecting them into evaluation-shaped windows, and allocating a bounded number per requested work.The trainer consumes only these pre-sliced rows, tokenizing concatenated prefixes and suffixes while masking prefix tokens from forget loss.
  • H Retrieval Curators and Forget-Set Construction Details: BM25 searches segmented corpus documents and retains the global top 100 segments for each requested work.Documents shorter than 100 characters are discarded; remaining documents use segments of at most 2,000 characters with 500-character overlap, removing segments shorter than 200 characters.
  • H Retrieval Curators and Forget-Set Construction Details: Infini-gram scans request offsets for matching 20-character substrings, binary-searches longest positive-count spans, and ranks retained spans by character length.It keeps the top 100 spans after scanning the request text.
  • H Retrieval Curators and Forget-Set Construction Details: All forget sets use W = 100 and S = 10, emitting adjacent 100-character prefix and suffix spans with curator-specific source-window rules.EA windows the reference text directly; BM25 windows retrieved units; Infini-gram windows contain matched-span midpoints, and duplicate pairs are removed per work.
  • I Scaling Forget Request: Each requested work is capped at C = 128 windows, initially allocating up to F = 4 windows per retrieved document before distributing remaining budget round-robin.This quota prevents a single repeated source from dominating the forget set.

J Curator × Unlearner × Model Cross Evaluation · K Curator Output Sizes

The cross evaluation reuses each model–curator forget set across three unlearning algorithms, isolating algorithm effects and reporting per-model and averaged results. Curators differ in forget-set size while optimization settings remain fixed, with additional reporting on window counts, fixed retain data, and an OLMO-3-7B pool limitation.

  • J Curator × Unlearner × Model Cross Evaluation: Each model–curator forget set is constructed once and reused across SimNPO, UNDIAL, and RMU, so only the unlearning algorithm changes within each block.SimNPO rows repeat Table 1, EA-curator cross evaluation is reported in Table 2, and three-model averages appear in Table 3.
  • J Curator × Unlearner × Model Cross Evaluation: Table 3 summarizes the cross evaluation with three-model averages, complementing the per-model results.These averages are reported alongside the per-model EA-curator evaluation and repeated SimNPO entries.
  • K Curator Output Sizes: Curators produce forget sets of different sizes, while the unlearner’s optimization hyperparameters remain fixed regardless of |Df|.The forget-set size is treated as a property of the curator under test, rather than as a reason to retune optimization.
  • K Curator Output Sizes: Table 6 reports a forget-set size ablation at |F| = 100 for Llama-3.1-8B and Qwen3-8B, using Table 1’s metrics and columns.The ablation examines the effect of forget-set size for two models.
  • K Curator Output Sizes: At |Wf| = 50, Table 8 reports the exact number of prefix–suffix windows produced by each curator, while the retain set is fixed at 1,646 WikiText rows.The retain-set size is identical across every configuration.
  • J Curator × Unlearner × Model Cross Evaluation: Table 7 reports per-model curator × unlearner cross evaluation at |Wf| = 50, using the metrics and columns of Table 1.The table covers the model, curator, and unlearner combinations described for the cross evaluation.
  • K Curator Output Sizes: Table 8 specifically tabulates the number of curated forget-set windows at |Wf| = 50.This count is the curator output-size measure used for the configurations.
Loading 2608.14855v1…