Source-linked AI summary
From Reweighting to Rewriting: Unlocking the Intervention Effects of Influential Samples in Training Data Attribution
Yuzhang Luo, Chenpeng Wang, Jianhui Chen, Liangming Pan
TL;DR
The paper addresses whether weak effects from reweighting mean influence-selected examples lack intervention value or whether reweighting fails to realize their leverage. It introduces influence-guided response rewriting, which fixes influence-based selection while replacing selected responses with aligned or opposed supervision. Across four open-weight LLMs, rewriting yields stronger, more persistent, and bidirectional shifts than reweighting, with greater leverage than alternative selectors and similar qualitative effects for safety refusal.
Problem
Influence-selected examples often show limited advantages over random examples under weight-based interventions, leaving their broader intervention value unresolved.
Method
The framework uses influence functions to select examples, then either reweights their original supervision or rewrites responses while keeping instructions fixed.
Results
Across four open-weight LLMs, response rewriting produces stronger, more persistent, and bidirectional behavioral shifts than reweighting the same examples, with greater leverage than alternative selectors and similar trends for safety refusal.
Takeaways & Limitations
Influence estimates capture local reweighting effects, while their selected examples can possess broader behavioral leverage under intervention-aware rewriting.
Takeaways & Limitations
The framework is most natural for behaviors with clear aligned or opposed rewriting targets, and safety gains can incur over-refusal on benign prompts.
Abstract
from arXiv · showhide
Training data attribution (TDA) aims to identify training examples that shape model behavior, but its intervention value depends on both which examples are selected and how they are modified. Influence functions (IF) estimate behavioral changes under infinitesimal reweighting, yet IF-selected examples often show limited advantages over random selection under conventional weight-based interventions. This raises the question of whether influential examples lack intervention value or whether reweighting fails to realize their behavioral leverage.We introduce influence-guided response rewriting, which uses IF to identify intervention targets and replaces their responses with behavior-aligned or behavior-opposed supervision while keeping instructions fixed. Across four open-weight LLMs, we compare rewriting and reweighting on the same influence-selected examples using epistemic abstention as our primary testbed. Response rewriting produces stronger, more persistent, and bidirectional behavioral shifts, while reweighting the same examples yields weak and inconsistent effects. Further analyses show that influence-selected examples provide greater rewriting leverage than alternative selectors, with changes remaining concentrated on target-relevant behaviors. The same qualitative contrast extends to safety refusal. These results distinguish the local reweighting effects captured by influence estimates from the broader intervention leverage of the examples they identify, motivating intervention-aware evaluation of TDA methods.
1 INTRODUCTION
The paper asks whether weak reweighting effects reflect poor influence-selected targets or an intervention strategy that fails to realize their leverage. It introduces response rewriting to separate example selection from the behavioral signal those examples provide.
- Training data attribution scores examples by their contribution to a target model quantity, while influence functions estimate effects of infinitesimal reweighting.
- Recent studies report that influence-selected examples often provide little advantage over random examples under weight-based interventions.
- The paper distinguishes whether influence selects poor targets from whether reweighting fails to realize selected examples’ behavioral leverage.
- Influence-guided response rewriting keeps instructions fixed while replacing responses with supervision that encourages or discourages a target behavior.
- Across four open-weight LLMs, rewriting produces stronger, more stable, and bidirectional shifts than reweighting, with similar trends for safety refusal.
- Controlled analyses find that rewriting redirects local supervision while preserving target specificity and avoiding substantial degradation of general capabilities.
2 RELATED WORK
Related work frames TDA as a way to trace behavioral origins and abstention as a capability that remains difficult for modern language models. Existing abstention interventions use several strategies but face generalization limitations.
- Influence-based attribution traces training origins and is commonly evaluated or applied through data reweighting and filtering.
- Abstention is the ability to refrain from giving a definitive answer when a query cannot be reliably resolved.
- Modern language models show poor abstention across diverse forms of unanswerability, motivating uncertainty, calibration, representation, prompting, and post-training approaches.
- Prior work reports that instruction tuning may struggle to generalize abstention across domains and model settings, while fine-tuning can erode abstention capabilities.
3 METHODOLOGY
The methodology fixes influence-based selection and compares changing supervision strength with rewriting supervision content. It evaluates directional effects, selection advantage, and persistence during supervised fine-tuning.
- 3 METHODOLOGY: The study keeps the attribution rule fixed and varies whether selected examples are reweighted or rewritten.
- 3.1 INFLUENCE FUNCTIONS FOR SFT: Influence functions estimate how infinitesimal training-weight changes affect learned parameters and a target behavioral quantity.
- 3.1 INFLUENCE FUNCTIONS FOR SFT: A positive influence score predicts that upweighting an example increases the target quantity, whereas a negative score predicts a decrease.
- 3.2.1 BEHAVIOR ATTRIBUTION: The method ranks examples by influence and selects top-scoring supposedly helpful and bottom-scoring supposedly harmful sets, labels restricted to local reweighting predictions.
- 3.2.2 INTERVENTION DESIGN: Reweighting changes selected examples’ contribution while preserving their original responses; α > 1 upweights and α = 0 deletes them.
- 3.2.2 INTERVENTION DESIGN: Rewriting keeps each selected instruction fixed and replaces its response with supervision aligned with or opposed to the target behavior.
- 3.2.2 INTERVENTION DESIGN: Because rewriting changes the selected example’s gradient, influence determines which examples to modify rather than specifying a finite reweighting effect.
- 3.2.3 EVALUATION PROTOCOL: Evaluation measures directional effectiveness, selection advantage against matched random examples, and persistence across SFT checkpoints.
4 EXPERIMENTS
The experiments compare influence-selected interventions during supervised fine-tuning across four open-weight LLMs, using epistemic abstention and abstention recall. Reweighting is weak and inconsistent, whereas response rewriting produces stronger bidirectional behavioral changes.
- Experimental setup: The study evaluates epistemic abstention across four open-weight language models, selecting supposedly helpful and harmful examples from both influence-ranking extremes plus matched random sets.Interventions generally target 2.5% of the SFT data.
- Experimental setup: Abstention recall measures the fraction of unanswerable evaluation queries on which the model abstains, with higher values indicating stronger abstention behavior.The primary training-dynamics scenario is answer unknown, where no documented or commonly agreed-upon answer exists.
- Interventions: Aligned rewriting replaces selected responses with abstention supervision, while opposed rewriting replaces them with supervision that encourages answering; instructions remain fixed.Reweighting uses α = 2 for upweighting and α = 0 for deletion by default, while rewriting uses diverse semantically equivalent abstention templates.
- Intervention effects: Reweighting produces unstable effects across models, fails to outperform the baseline reliably, and can move in the opposite direction.Changing the reweighting coefficient α does not recover a consistent dose–response pattern or expected bidirectional behavior.
- Intervention effects: Aligned rewriting consistently increases abstention recall and opposed rewriting decreases it, with larger and more persistent effects than random rewriting.The two influence-ranking extremes show model-dependent dynamics: helpful-example effects may decay after an early shift, while harmful-example effects may emerge and grow later; Gemma3-4B is not universal.
5 FURTHER ANALYSIS
Further analyses examine why influence-selected examples respond strongly to rewriting and whether the resulting behavioral changes remain targeted. They associate influential examples with unanswerability-related representations, identify greater redirection leverage than projection-based selection, and find selective behavioral gains without systematic capability degradation.
- What makes influential examples effective rewriting targets?: Influence-selected examples from both ranking extremes have higher unanswerability-direction projections than the overall training distribution, though they are not the most extreme examples.These projection scores quantify alignment with the learned unanswerability direction.
- What makes influential examples effective rewriting targets?: Projection-based selection matches influence-guided selection under opposed rewriting but falls substantially short under aligned rewriting.Many high-projection examples may already carry abstention-consistent supervision, leaving less room for aligned redirection.
- What makes influential examples effective rewriting targets?: Additional selectors confirm that influence-based selection yields the largest and most sustained rewriting effects.This supports distinguishing behavioral leverage from simply selecting examples with the strongest target representation.
- How rewriting changes training influence: Aligned rewriting shifts predicted influence scores positively, reverses originally harmful examples, and remains above original-response scores under Bayesian influence throughout training.The comparison keeps prompts fixed and evaluates original and rewritten responses under matched local conditions.
- Is the behavioral change targeted?: The largest abstention gains occur on answer unknown and false premise, while effects on subjective and underspecified-context scenarios are smaller.Random-aligned rewriting produces broader gains on non-target scenarios, suggesting influence-guided rewriting is more selective.
- Is the behavioral change targeted?: Additional metrics show no substantial systematic degradation of general capabilities, although safety refusal improvements produce a noticeable XSTest performance drop from over-refusal on benign prompts.In abstention, F1 improves over random rewriting while accuracy remains close to the original baseline, despite some precision reduction.
6 GENERALIZATION TO SAFETY REFUSAL
The influence-guided rewriting framework also generalizes to safety refusal, producing broad directional changes that deletion and upweighting do not match. However, stronger refusal behavior incurs an over-refusal cost on benign prompts.
- The safety-refusal experiment applies the same intervention framework as abstention, replacing the target signal and rewriting templates with safety-specific versions.
- Aligned rewriting substantially strengthens safety refusal, while opposed rewriting produces large degradations in the opposite direction.
- Deletion and upweighting remain less systematic and do not produce the broad directional changes induced by rewriting.
- Aligned rewriting causes a noticeable XSTest performance drop, indicating a substantial over-refusal risk on benign prompts.The passage attributes this trade-off to the broad, aggregate nature of the safety-attribution target set.
7 CONCLUSION
The study concludes that influence-selected examples can have substantial behavioral leverage that reweighting fails to realize, and that intervention-aware evaluation should distinguish local reweighting effects from broader actionable potential. The framework is best suited to behaviors with clear aligned or opposed rewriting targets, while safety applications still require finer refinement.
- Response rewriting produces stronger, more persistent, and bidirectional shifts than reweighting the same influence-selected examples across four open-weight LLMs.
- Influence-selected examples provide greater rewriting leverage than alternative selectors while effects remain concentrated on target-relevant behaviors.
- Influence estimates capture local reweighting effects, whereas the identified examples may possess broader intervention leverage that requires intervention-aware TDA evaluation.
- The framework is most natural for behaviors with clear behavior-aligned or behavior-opposed rewriting targets, such as abstention and safety refusal.
B.1 FIXED-REFERENCE INFLUENCE COMPARISON
The fixed-reference analysis isolates the effect of rewriting responses by holding the checkpoint, target gradient, and curvature constant. Under this shared local geometry, aligned responses shift influence toward target-relevant directions and produce stronger abstention effects than alternative selectors and weight-based interventions.
- Fixed-reference design: The comparison keeps the checkpoint, target gradient, and curvature fixed while substituting only the training-example gradient.This isolates response replacement from changes in model state or local geometry.
- Fixed-reference design: A positive influence shift indicates that rewriting moves an example gradient toward a direction coupled to reducing target query loss.The interpretation is defined under the original model’s local geometry.
- Bayesian validation: Paired Bayesian influence scores remain higher for aligned responses than original responses throughout SFT for both influential groups.Both response versions are evaluated under the same localized posterior at each checkpoint.
- Selector comparison: Influence-based selection produces the strongest and most sustained rewriting effects among the tested selectors under both aligned and opposed rewriting.The separation from alternative selectors grows over training, especially at later checkpoints.
- Behavioral outcomes: Aligned rewriting raises abstention recall across four models, improves F1 over random rewriting, and keeps accuracy near the original baseline.Precision decreases somewhat, indicating a minor over-refusal trade-off.
- Behavioral outcomes: Influence-selected rewriting produces strong directional effects on false premise but does not consistently outperform random rewriting on underspecified context.The contrast separates target-relevant behavior from a non-target scenario.
D.3 GENERAL-CAPABILITY EVALUATION
The capability evaluation tests whether stronger abstention interventions impair general performance. Influence-selected rewriting shows no substantial systematic degradation relative to random rewriting, while target-related abstention changes remain behaviorally concentrated.
- Evaluation setup: The evaluation summarizes OLMo2-1B rewriting conditions across GSM8K, HellaSwag, ARC-Challenge, and ARC-Easy using accuracy.Random results are averaged over four independently sampled intervention sets.
- Capability preservation: No substantial systematic general-capability degradation is observed after influence-selected rewriting on the evaluated OLMo2-1B benchmarks.HellaSwag is nearly unchanged, while ARC-C and ARC-E fluctuate modestly around baseline.
- Capability preservation: The GSM8K decrease after rewriting is also observed under random rewriting, so it is not an additional consistent cost of influence-selected intervention.The comparison uses accuracy (%) with higher values better.
- Capability preservation: Influence-selected rewriting’s stronger behavioral effects are not accompanied by correspondingly larger degradation in general capabilities.The authors interpret the combined metric and benchmark results as evidence of targeted abstention changes.
E.1 EFFECT OF INTERVENTION BUDGET
Intervention budget changes affect rewriting and weight-based operators differently. Rewriting effects generally scale with the number of selected examples, whereas upweighting and deletion show inconsistent, sometimes counterintended budget behavior.
- Budget scaling: The evaluated budgets are k = 800, 1,600, and 3,200 examples, corresponding to 1.25%, 2.5%, and 5% of SFT data.All other training and evaluation settings are held fixed.
- Weight-based interventions: Upweighting and deletion do not reliably outperform random selection and show no consistent monotonic relationship with intervention budget.Increasing the budget can sometimes move behavior in the unintended direction.
- Response rewriting: For aligned and opposed rewriting, increasing k generally strengthens the behavioral effect in the intended direction.The largest changes occur at k = 3,200 and the weakest at k = 800.
- Response rewriting: The clearer dose-response pattern indicates that rewriting’s advantage is not specific to the default budget of k = 1,600.Unlike deletion and upweighting, its effect scales systematically with selected influence-guided supervision.
F.1 OVERLAP OF RANKINGS ACROSS DIFFERENT MODEL FAMILIES AND SIZES
Influence rankings overlap substantially within some model families but are not identical across models. Rankings transferred from OLMo2-1B generally retain intervention value beyond random selection, although effect magnitude and dynamics vary by target model and training stage.
- Ranking overlap: OLMo2-1B and OLMo2-7B share 41.5% of top examples and 49.5% of bottom examples.This is the strongest reported consistency within the same model family.
- Ranking overlap: Gemma3-4B shows a qualitatively different cross-end pattern, with its bottom-ranked examples overlapping other models’ top-ranked sets.Reported overlaps with OLMo2-1B, Qwen3.5-2B, and another model are 15.2%, 12.5%, and 14.0%.
- Cross-model transfer: The transfer test applies OLMo2-1B’s influence ranking to interventions on the other three models and compares it with each model’s native ranking.This directly tests whether selected examples remain useful across target models.
- Cross-model transfer: Transferred selections generally retain pronounced behavioral effects and remain substantially separated from random selection.Transfer magnitude may weaken or strengthen depending on target model, ranking end, and training stage.
- Cross-model transfer: The results provide preliminary evidence that influence-based data selection can transfer across models, potentially reducing the need to compute influence directly on every target model.The proposed direction is to compute rankings on smaller models and transfer selections to larger models.
G EVALUATION DETAILS
The evaluation separates epistemic abstention from safety refusal and applies matched interventions to fixed training-example selections. It uses dedicated benchmarks, query sets, and response-rewriting templates under controlled training conditions.
- Evaluation scope: Abstention and safety are evaluated separately because declining to answer can reflect epistemic uncertainty or policy violation.
- Evaluation scope: AbstentionBench evaluation uses 100 prompts from each of 18 datasets, totaling 1.8k prompts across four reporting scenarios.The scenarios are answer unknown, false premise, subjective, and underspecified context.
- Evaluation scope: Safety evaluation selects nine benchmarks covering harmful-prompt refusal, jailbreak resistance, and over-refusal.
- Intervention design: Four interventions are applied to the same 1,600 selected examples, isolating the effects of example selection from supervision modification.The interventions are deletion, upweighting, behavior-aligned rewriting, and behavior-opposed rewriting.
- Intervention design: Rewriting keeps instructions fixed while replacing responses with refusal or compliance templates, whereas deletion and upweighting alter original supervision presence or strength.Refusal pools contain 80 abstention templates or 100 safety templates; the comply pool contains 20 templates.
- Influence queries: Influence computation uses separate held-out query sets with target answers: 300 abstention questions and 300 harmful safety prompts.The abstention targets are generated under a prompt requiring a short uncertainty-expressing sentence, while verified safety refusals are retained when available; queries are decontaminated against training data.