Source-linked AI summary
Rethinking Demonstration Unlearning in Imitation Learning for Robotics
Jiazhuo Li, Yu Zhang, Yiming Fei, Kangkang Dong, Xiaojun Zhu, Houde Liu, Jinze Tao
TL;DR
Demonstration unlearning in robotics lacks a clear operational meaning when edits are judged by behavior or membership evidence alone. The paper introduces a retrain-calibrated two-axis audit and finds that the axes dissociate, with ACT redirect restoring 18/20 blind-scored robot trials while evidence remains unchanged.
Problem
Robot demonstrations can be withdrawn, but existing metrics do not establish what an edit removed from a policy acting in a feedback loop.
Method
The audit compares behavior and membership evidence with a retrain counterfactual, combining calibrated action divergence, rank and absolute evidence measures, and a conformal joint test.
Results
The two axes dissociate in both directions across the evaluated conditions; on ACT, redirect restored blind-scored robot success to 18/20 trials.
Takeaways & Limitations
Deletion is treated as a conjunction of retrain-consistent behavior and evidence, which no one-axis audit can certify.
Takeaways & Limitations
Hardware contrasts use sequential collection, and the tested operators have request-specific applicability limits.
Abstract
from arXiv · showhide
Imitation learning for robotics depends on human demonstrations, some of which people may later ask to remove. Retraining without them is the natural reference, but its cost grows with policy and dataset scale, motivating cheaper operators that edit a trained policy. Metrics inherited from machine unlearning, such as forgetting loss or a single membership attack, do not establish what an edit removed from a policy acting in closed loop. We therefore introduce a retrain-calibrated audit that reads demonstration unlearning along two axes: behavior, whether the edited policy acts like one retrained without the removed demonstrations, and evidence, whether an auditor can still detect it was trained on them. The behavior axis measures action divergence to that retrain at matched states, calibrated by a floor built from independent retrains, so a policy at the floor is as close to a retrain as retrains are to each other. The evidence axis applies a per-demonstration membership attack against a retrain null, reporting both its rank and its absolute member-loss level, since rank alone accepts operators that inflate member losses past the null. A conformal test then combines both axes into one hypothesis of joint retrain consistency, against a fleet of independent retrains large enough to reject at conventional significance. Across five preregistered conditions on three real-robot policy classes and two simulation suites, the axes dissociate in both directions on one checkpoint, as an edit may repair task behavior while leaving evidence unchanged, or reduce evidence while moving behavior away from retraining. On the ACT arm, a redirect edit restores blind-scored robot success to 18 of 20 trials.
1 INTRODUCTION
The paper argues that demonstration unlearning must be evaluated by both behavioral similarity to retraining and residual training evidence. Across ACT and broader conditions, these axes can disagree, so one-axis audits can yield false deletion verdicts.
- 18/20 blind-scored ACT robot trials succeeded after redirect editing, versus 20/20 for retraining.
- AUC 1.000 matched the unedited policy after redirect, despite a retrain-null AUC of 0.639.
- Behavior and evidence moved independently: behavior improved while membership evidence persisted, whereas ascent changed evidence while behavior moved farther from retraining.
- A 19-replica conformal audit rejected every audited ACT checkpoint at p=0.05, while the proposed audit remains refutation-only.
- 32× forgetting-loss gap closure at zero action movement shows that standard metrics can accept checkpoints rejected by the joint criterion.
2 RELATED WORK
Related work covers machine unlearning, robot-policy editing, and demonstration attribution, but does not combine demonstration-identity deletion with hardware evaluation and retrain-calibrated evidence auditing.
- Machine-unlearning research defines deletion against retraining and commonly uses membership inference as its evidence instrument.
- Robot-policy editing methods redirect behavior or add corrective supervision, but do not establish demonstration deletion semantics with an evidence axis.
- Attribution methods identify which demonstrations mattered, whereas this paper asks whether identified demonstrations are gone from the policy.
3 WHAT DELETION REQUIRES: A RETRAIN-CALIBRATED DEFINITION
The paper defines demonstration deletion relative to a counterfactual retrain and audits it through separate behavior and evidence axes. A conformal joint test compares both against an independent retrain fleet.
- Deletion names demonstrations by episode identity and compares an edited policy with π∗, retraining on the retained set at much lower edit compute.
- Behavior axis: Behavior measures deployed-path action divergence from π∗ at matched states, calibrated against divergence among independent retrains.
- Evidence axis: Evidence reports membership-attack rank AUC and absolute member loss relative to a retrain null, because rank is blind to overshoot.
- Four outcomes: Only joint retrain consistency—behavior at the floor and both evidence statistics consistent with the null—is eligible for a deletion claim.
- Joint test: The conformal statistic tests exchangeability with retrain replicas and sets a refutation floor of 1/20=0.05 using K=19 replicas.
4 EXPERIMENTAL DESIGN
The experiments apply one deletion-request family and a common operator ladder across three robot policy classes and two simulation suites. They combine matched-state behavior measurements, retrain floors, membership audits, and disclosed hardware-design constraints.
- Scope: Five conditions span ACT, Diffusion Policy, and π0.5, plus robomimic-BC and Diffusion-PushT simulation arms.
- Contamination: The contamination is a coherent wrong mode: 30 of 130 demonstrations teach release at a displaced point, while Diffusion-PushT mirrors 80 of 286 episodes.
- Operators: Each class uses retraining, the all-data policy, redirect, calibrated ascent, and budget-matched fine-tuning as the operator ladder.
- Measurements: ACT behavior is measured primarily on an executed 25-step slice, where the BRANCH gap is 2.2 raw units versus 7.9 for the full plan.
- Audits: Membership audits average per-demonstration losses over 64 frames and pool retrain seeds with frame draws into the null.
- Hardware protocol: The hardware design pairs initial positions across conditions but collects sequentially, leaving session drift as a disclosed confound.
- ACT ladder: No ACT evidence-ladder rung reaches the absolute null: B200, B300, and B400 intervals all exclude 1.0.
5 RESULTS
Across hardware and simulation results, behavior and evidence repeatedly dissociate: operators can improve task conduct without removing membership evidence, or alter evidence without approaching retrained behavior. The results also expose scope and design limits, including state dependence, sequential collection, and request structures that make localized editing infeasible.
- Behavior axis: 24.9% of the executed-slice BRANCH gap was closed by redirect at R200, while ascent moved behavior farther from the retrain floor.R400 matched redirect within noise while increasing retain damage; ascent degraded beyond B50.
- Evidence axis: 1.000 AUC persisted for θ0, redirect, and FT against a retrain-null grand mean of 0.639, showing that behavior repair did not remove membership evidence.The null pooled five retrain seeds and three frame draws.
- Seed replication: Redirect replicated its masking pattern across three θ0 seeds, whereas ascent’s evidence endpoint varied with lineage because its stopping rule was self-referential.Redirect remained at AUC 1.000 with mem/null 0.39–0.41; ascent reached mem/null 1.83/0.91/1.19 across seeds.
- Hardware validation: 18/20 blind-scored hardware trials succeeded for redirect, compared with 20/20 for the ceiling, 9/20 for FT, 7/20 for ascent, and 5/20 for θ0.Redirect was not separable from the ceiling at n=20, which indicates failure to detect a difference rather than equivalence.
- Hardware mechanism: Success tracked release angle rather than inherited grasp dispersion, while FT recovered much of the success gap without changing offline behavior or rank evidence.The 130-lineage grasp spread survived every operator; release angle alone distinguished θ0 from ceiling conduct at AUC 0.958.
- Cross-class evidence: At ratio 0.50, success gains of +8.0/+11.6/+2.4 points across seeds required retain damage of +42.7/+33.8/+68.7%, and no tested dose approached the retrain evidence reference.These runs were stopped at 150/200 steps because they were not budget-matched or deployable.
- Applicability: Localized redirect was inapplicable for whole-episode contamination because its support-gated editable intersection was 3.65%, below the pre-set 5% feasibility bar.The operator requires contamination sharing a prefix with clean support; tightening the detector worsened the intersection.
- Baselines: 13 of 14 robomimic-BC methods degraded closed-loop success by 68–101% regardless of offline forget quality.The comparison included EU-k methods with near-perfect evidence scores but dead policies.
6 CONCLUSION
The paper defines demonstration deletion as joint retrain consistency across behavior and evidence, while limiting its claims to tested empirical audits and settings.
- The audit measures behavior against a retrain floor and evidence against a retrain null using rank and absolute membership measures.
- A 19-replica conformal audit rejects every audited edit at p=0.05, so neither one-axis audit can certify deletion.
- The scope is mode-level contamination of a re-runnable set, with on-distribution repair, surviving deep-state attractors, and failed composition identified as limits.
- The study evaluates technical deletion proxies rather than legal compliance or causal removal of a demonstration’s influence.
A MEASUREMENT CONVENTIONS
The measurement conventions specify how action divergence, closed-loop success, and membership audits are reported across policy arms.
- Every action-space number carries units, reference, region, horizon, and sample-count metadata.
- ACT, DP, and π0.5 use per-demonstration loss attacks with mean aggregation and bootstrap confidence intervals, while PushT uses closed-loop success for manifestation.
- Membership nulls pool all retrain seeds; ACT additionally pools five seeds and three frame draws.
B FLOOR TABLES, BOTH HORIZON CONVENTIONS
The floor tables distinguish executed-slice measurements from full-plan diagnostics while preserving the same operator ordering.
- The executed-slice convention is primary, while the full 50-step plan serves as a late-ramp diagnostic.
- A pooled action-space magnitude cannot identify whether divergence reflects a path or placement difference, motivating region-resolved temporal reporting.
- 2.2 raw units versus 7.9 raw units is the ACT θ0 BRANCH gap over the floor under executed-slice and full-plan conventions, respectively.
C THE PUSHT LEARNING-RATE CLIFF, THE VOIDED GRID, AND THE SEED
The PushT appendix documents a learning-rate instability, seed-level controls, contamination setup, and a feasibility gate for localized redirect.
- A 10× learning-rate change moves budget-matched fine-tuning retain damage 175×, voiding the initial ascent sweep as a learning-rate artifact.
- PushT uses 80 reflected episodes within a 286-episode training set, creating a rigid-body-consistent wrong-goal mode rather than noise.
- The PushT gate separates contaminated and retrain recipes by 14.8 points, with seed-level Welch t=3.78 and p ≈0.018.
- The frozen redirect diagnostic requires forget frames high in action discrepancy and low in state-support distance; whole-episode contamination supplies almost none.
F ROBOMIMIC-BC: 14 METHODS, THE FAIRNESS SWEEP, AND THE BLINDNESS LEDGER
The robomimic-BC comparison evaluates 14 methods under fairness controls while showing that offline forgetting quality can diverge from closed-loop recovery. The blindness ledger further shows why rank-only audits cannot establish retrain-consistent deletion.
- 14-method comparison: Positive recovery is unique to redirect; every alternative degrades closed-loop success by 68–101% regardless of offline forget quality.Recovery measures the θ0-to-retrain closed-loop success gap, while AUC-dist measures distance from the retrain-null membership AUC.
- Fairness sweep: No fairly tuned baseline reaches positive recovery, and adding the proximity tether rescues none.The sweep covered 32 configurations and retained the offline/closed-loop dissociation.
- Blindness ledger: Rank acceptance can pass checkpoints that the joint criterion rejects, demonstrating disagreement among standard audits on the same checkpoint.The blindness ledger records measured cross-class exhibits rather than relying on a single audit outcome.
- Temporal divergence: Figure 7 shows ACT divergence increasing from 4 to 17 raw units across the 50-step plan, making pooled and executed-slice closure fractions differ.The operator ordering remains unchanged despite the time-localized divergence profile.
- Rank-audit invariance: AUC is invariant to strictly increasing score transformations, so rank-based verdicts cannot detect arbitrary changes in absolute score levels.B200 and B300 have mem/null levels of 0.84 and 1.59 while both rank confidence intervals contain the null.
I THE JOINT TEST, AND WHERE EVIDENCE LIVES
The joint test combines behavioral divergence and calibrated membership evidence against a fleet of independent retrains. Region-resolved analysis shows that evidence changes can be local, diluted by episode means, or overshoot across all regions.
- Joint test: The conformal test uses executed-slice BRANCH divergence, absolute log mem/null distance, and absolute AUC distance from the retrain null.It is calibrated against K=19 retrains of the identical recipe and retain set, with nested behavioral and evidence subsets.
- Joint test: Every checkpoint is rejected at p=0.050, while redirect is nearest and ascent’s high rungs are farthest from joint retrain consistency.Rank acceptance and joint rejection disagree on the same checkpoints.
- Ascent endpoint: Across three θ0 lineages, ascent’s stopping quantile saturates by step 300 and produces mem/null endpoints of 1.83, 0.91, and 1.19.Overshoot occurs when the saturated quantile passes the null member level; all runs reached the full budget without guard trips.
- Where evidence lives: Redirect reaches the retrain evidence null on the 8.6% of frames it edits, while ascent overshoots every region and fine-tuning deepens memorization everywhere.The episode-mean statistic dilutes redirect’s local repair by 12:1.
- Parameter space: At the repaired R200 checkpoint, membership-loss and counterfactual-proximity gradients have cosine +0.30, indicating distinct audit directions at the localized repair.The corresponding cosine is +0.66 at θ0; the B300 reading is excluded because both losses are far from their targets.
J PRE-REGISTERED PREDICTIONS AND OUTCOMES (PUSHT ARM)
The PushT runbook preregistered predictions for redirect, ascent, fine-tuning, and evidence audits, then evaluated them at a frozen dose and feasibility diagnostic. Redirect was not run because the editable intersection fell below its prespecified threshold, while ascent and control predictions were tested directly.
- P1 expression: The frozen PushT expression gate was confirmed: the θ0-to-ceiling success gap was 14.8 points, with seed-level Welch t=3.78 and p≈0.018.Every θ0 seed fell below every ceiling seed, and average maximum reward separated in the same direction.
- P2 redirect: Redirect was not evaluated because the editable intersection was 3.65%, below the prespecified 5% feasibility bar.The study did not substitute an all-frames operator for the unavailable redirect edit.
- P3 ascent: At the frozen dose, ascent was refuted: forget loss moved +40–78%, rank AUC moved by at most 0.003, and pooled behavior changed by +1.6 points.The runbook treated the failed prediction as a finding.
- P4 fine-tuning control: Budget-matched fine-tuning moved neither axis, with rank changes of −0.001 to −0.003 per seed and pooled behavior change of −2.1 points.The pooled behavioral result had t=−0.95.
- P5 evidence: The θ0 membership audit was confirmed at AUC 0.981–0.997 versus a 5-seed null of 0.534 ± 0.012, while the redirect clause was not evaluated.The latter followed from the failed redirect feasibility diagnostic.