Source-linked AI summary
A Gravitational Interpretation of Fine-Tuning Reversion
Samuele Poppi, Nils Lukas
TL;DR
Benign fine-tuning can cause safety erosion and other behavioral drift, raising the question of whether these effects share a training-history mechanism. The paper identifies and manipulates a history-defined reversion direction, finding that blocking it reduces harmfulness from 19.0% ± 4.0% to 8.5% ± 1.5% with little task cost.
Problem
The paper asks whether safety erosion and related post-adaptation drift can be understood through a common mechanistic lens.
Method
The paper defines a witness-based activation-space reversion direction and tests its emergence, behavioral coupling, and causal relevance during benign fine-tuning.
Results
Blocking motion along vrev reduces harmfulness from 19.0% ± 4.0% to 8.5% ± 1.5% with little task cost.
Takeaways & Limitations
The findings support history-dependent reversion as a broader dynamic connecting post-alignment safety degradation with related forms of behavioral fragility.
Takeaways & Limitations
The paper does not claim that vrev is the unique safety-relevant direction, and its behavioral coupling evidence comes from a modest, temporally structured checkpoint sample.
Abstract
from arXiv · showhide
Fine-tuning on harmless data can partially undo behaviors acquired earlier in training. Safety can erode under benign post-alignment updates, unlearned capabilities can re-emerge, latent traits can transfer through apparently unrelated supervision, and related post-alignment fragility appears in other generative settings. We argue these phenomena are usefully viewed through a common training-history lens. Our hypothesis is geometric: large early training phases create dominant behavioral manifolds, while later alignment or specialization phases are shallower displacements from them. Subsequent fine-tuning can therefore inherit a persistent reversion component pointing back toward a witness of the dominant manifold. We call this the gravitational interpretation of fine-tuning reversion. Across our main settings, representational drift rapidly acquires a component along a history-defined reversion direction (v_rev). In our main track, alignment with v_rev rises from cos = 0.429 +/- 0.052 after the first update to 0.647 +/- 0.021 by step 20. Across 24 run-step pairs, every observed alignment exceeds the p99 of an isotropic activation-space null. We demonstrate that selectively blocking motion along v_rev changes the final alignment at T=100 from 0.648 +/- 0.009 to -0.211 +/- 0.021 and reduces harmfulness from 19.0% +/- 4.0% to 8.5% +/- 1.5% with little task cost. These results support v_rev as a causally relevant mediator of early post-alignment reversion in our setup. Importantly, we do not claim that v_rev is the unique safety direction, nor that the dominant manifold is directly observed; rather, we identify a robust, history-defined direction that explains and partially controls early reversion dynamics.
1 INTRODUCTION
The paper frames benign fine-tuning regressions across safety, unlearning, and specialization as a common, history-dependent reversion phenomenon. It proposes and tests a history-defined direction, v_rev, that predicts and causally influences early post-alignment drift without claiming to observe the underlying manifold or identify a unique safety direction.
- Motivation: Benign post-training can reverse earlier behaviors across safety alignment, harmful-knowledge removal, and other apparently unrelated training settings.The paper uses these cases to motivate a unified mechanistic account rather than treating safety erosion as an isolated failure mode.
- Motivation: Safety alignment occupies a shallow low-rank subspace, and aligned models remain safe only within a limited local basin.These structural observations motivate viewing later benign updates as potentially capable of leaving the aligned region.
- Mechanism: The proposed mechanism combines the new task trajectory with a reversion component pointing toward a more foundational behavioral region in the model’s training history.The framework treats acquired behaviors as potentially re-emerging from earlier pretraining, helpful-only, alignment, or specialization phases.
- Scope: The paper studies local helpful witnesses rather than directly observing the latent dominant manifold or asserting that v_rev is the only route toward harmful or less aligned behavior.Its empirical claim is deliberately limited to a robust, history-defined direction supported by the tested setup.
- Evidence: 0.429 ± 0.052 to 0.647±0.021: alignment with v_rev rises from the first benign update to step 20 in the main 8B setting.Across checkpoints, geometric alignment also predicts downstream degradation with Spearman r = 0.877.
- Causal test: 19.0%±4.0% to 8.5%±1.5%: blocking motion along v_rev reduces harmfulness with little task cost, while matched random-direction controls do not reproduce the effect.This intervention is presented as evidence that the direction is behaviorally and causally meaningful in the tested setup.
2 BACKGROUND AND RELATED WORK
Prior work shows that benign or narrow fine-tuning can degrade safety, recover harmful behavior, or induce broad misalignment, while safety is organized in local basins and low-dimensional or linear activation-space structures. This paper situates these findings within a geometric account of reversion toward behavioral manifolds.
- Fine-tuning breaks safety: Benign fine-tuning can degrade safety, and harmful behavior can re-emerge under seemingly benign adaptation despite apparent alignment.Qi et al. (2023), He et al. (2024), Guan et al. (2025), and Yang et al. (2023) establish these related phenomena.
- Emergent misalignment: Narrow fine-tuning on seemingly benign insecure code can produce broadly misaligned behavior with a convergent linear representation.The paper interprets this as movement toward a behavioral manifold where alignment is absent.
- Safety basins and shallow alignment: Safety persists within a local weight-space basin but drops sharply beyond a radius, while safety-relevant directions occupy a small low-rank subspace.Related work also finds this structure brittle under sparse and low-rank changes.
- Safety mechanics in activation space: A single residual-stream direction can mediate refusal, and subliminal learning can be understood as steering-vector distillation.This work follows Arditi et al. (2024) by extracting a difference-of-means direction to characterize activation-space safety drift.
3 THE GRAVITATIONAL INTERPRETATION
The paper interprets benign post-alignment reversion as a history-induced geometric bias: later optimization tends to acquire a component toward an earlier helpful region. It operationalizes this component with a witness-defined activation-space direction and tests whether manipulating it changes outcomes.
- Core interpretation: Benign fine-tuning tends to acquire a substantial component along a history-defined direction pointing from a displaced checkpoint toward an earlier helpful region.This direction is defined operationally through activation-space displacements between checkpoints.
- Operationalization: v_rev is the activation-space displacement from a starting checkpoint toward a helpful witness, measured on the same probe family.The construction compares two checkpoints rather than extracting a within-model refusal feature.
- Witness analysis: Helpful-only witness families induce closely aligned v_rev directions, supporting geometric stability across nearby witnesses without directly observing or isolating the helpful manifold.The witness provides a concrete return direction relative to the displaced checkpoint.
- Training-history assumptions: Large early training phases are treated as creating a dominant helpful/chat region, while later safety alignment or specialization acts as a shallower displacement rather than a replacement.The gravitational metaphor captures this training-history asymmetry, not a completed physical law.
- Empirical predictions: The interpretation predicts drift along g_task + g_rev, with a persistent reversion tilt that need not be the only behaviorally consequential direction.The paper tests this through growing alignment with v_rev, coupling to unsafe drift, and directional interventions that alter downstream outcomes.
4 EXPERIMENTS
Experiments find that benign fine-tuning rapidly acquires a stable, witness-defined component along v_rev, across witnesses and unrelated tasks, and that this component is coupled to unsafe drift. Directly blocking v_rev reverses geometric drift and substantially reduces harmfulness with little task cost, while the results do not establish v_rev as unique or globally ontological.
- Experimental design: Experiments test whether benign fine-tuning aligns with v_rev, whether alignment co-evolves with unsafe drift, and whether direct intervention changes behavior.The intervention compares blocking or pushing motion along the history-defined component.
- Directional geometry: 0.57 and 0.65: mean cosine at L∗=31 for LoRA and full-FT helpful-only variants, respectively, with nearly identical directions despite different displacement magnitudes.Across six witness constructions, pairwise cosine reaches 0.82 ± 0.12 at L∗=31, with participation ratio 1.36/6, indicating stability across nearby witnesses rather than a single-checkpoint artifact.
- Directional geometry: 0.65: alignment plateaus by T=20 after appearing at T=1 and remains essentially unchanged at T=100.All 6/6 runs exceed the isotropic-null p99 at every step, and every one of 24 run-step pairs is above that threshold.
- Behavioral coupling: 0.79: harmful-prompt representation-shift cosine between Alpaca and Code full fine-tunes despite weight-update cosine 0.08.Across nine benign full fine-tunes, the shift directions have participation ratio 1.31/9, 93% of variance in the top two directions, and mean pairwise cosine 0.82; a code-specialized starting model also remains around cosine 0.5 during the early phase.
- Behavioral coupling: 0.877 and 0.731: Spearman and Pearson correlations, respectively, between geometric alignment with v_rev and unsafe rate across 20 AdvBench checkpoints.The geometric metric is more stable across seeds than behavior, with standard deviations approximately 0.04 versus approximately 0.27; this establishes coupling in baseline runs, not full predictive causality.
- Causal intervention: −0.211 ± 0.021: blocking v_rev changes cos(∆100, v_rev) from 0.648 ± 0.009 and reduces BeaverTails harmfulness from 19.0% ± 4.0% to 8.5% ± 1.5%.Task perplexity remains essentially unchanged at 1.392 ± 0.023 versus 1.384±0.005; random-direction blocking leaves alignment at 0.615±0.055 and harmfulness at 22.0% ± 2.9%.
5 DISCUSSION
The discussion frames fine-tuning reversion as a history-dependent geometric dynamic in which benign optimization follows a witness-defined direction toward a dominant behavioral region. It supports v_rev as causally relevant in the studied early regime while limiting claims about uniqueness, direct manifold observation, scaling, and generality.
- What the paper establishes: Subsequent fine-tuning rapidly acquires a growing component along the witness-defined reversion direction v_rev, supported by robustness, time-course, cross-task, and code-specialized-starting-point evidence.The discussion organizes the evidence into geometry, behavior, and causal intervention components.
- Why v_rev matters: v_rev matters because benign post-alignment optimization naturally follows it, not because it is the unique safety-relevant direction.The paper distinguishes safety-relevant directions as properties of representation space from v_rev as the direction followed by optimization in this setup.
- What the paper does not establish: The dominant manifold M_H is not directly observed; helpful-only checkpoints serve as operational witnesses rather than the manifold itself.Consistency across the witness family shows robustness to the choice of θ_H, but does not establish recovery of the manifold from data.
- What the paper does not establish: The gravitational interpretation remains a proxy-mediated mechanistic inference, and the paper does not yet fully parameterize reversion strength across training-history asymmetries.Its support comes from geometric consistency, beyond-safety displacement, and the causal intervention pattern.
- Change in perspective: Post-alignment degradation is presented as one instance of a broader history-dependent reversion dynamic: a smaller later phase displaces behavior from a dominant region, and benign optimization can restore it.This reframes safety erosion as part of a general training-history effect rather than an isolated safety failure.
- Implications for alignment: GSM8K stays near the aligned anchor under LoRA while rising under full FT, whereas Alpaca reaches similar harmfulness under both regimes by T = 100.The comparison motivates treating update freedom and task choice as interacting determinants of post-alignment fragility.
A ADDITIONAL EXPERIMENT DETAILS
This appendix documents the concrete training, probing, and evaluation settings used across the main experiment families. It supports protocol reconstruction and distinguishes reused pipeline components from components that differ.
- Protocol documentation: The appendix records the concrete training, probe, and evaluation settings for the main experiment families.Its purpose is methodological documentation rather than introducing new claims.
- Protocol documentation: The documented settings are intended to make the experimental protocol easy to reconstruct.
- Pipeline reuse: The appendix clarifies which parts of the paper reuse the same pipeline and which do not.
A.1 SHARED IMPLEMENTATION CHOICES
The experiments use standardized single-GPU bfloat16 training and activation extraction, with consistent tokenization, truncation, optimization, evaluation, and probing conventions. Representation analyses use larger harmful/harmless prompt sets than training-time interventions.
- Training and extraction: Training and activation extraction use Hugging Face Transformers in bfloat16 on a single GPU, official chat templates with left padding, 512-token truncation, and 8-bit AdamW unless noted.Base models without native chat templates borrow the tokenizer and chat template from their paired instruct checkpoints.
- Evaluation: Held-out perplexity masks prompt tokens, while harmfulness uses BeaverTails unsafe prompts with max new=200, batch size 8, and a separate guard model.The main text uses Llama-Guard-3-8B; the appendix reports a matched Qwen3Guard-Gen-8B rerun.
- Probe scales: Representation summaries use mean activations from 128 harmful and 128 harmless prompts, whereas intervention losses use 16 unlabeled AdvBench harmful prompts.The paper explicitly distinguishes these two probe scales.
A.2 HELPFUL-WITNESS FAMILY AND ACTIVATION EXTRACTION
The study constructs a helpful-only witness family from filtered conversational, preference, and programming data, then evaluates robustness across witness parameterizations and model families. Directional measurements use mean residual-stream activations at the final real prompt token, summarized at a refusal-axis peak layer.
- Helpful-only witness data: The helpful-only witness mixture combines OASST2 English rank-0 conversations, HH-RLHF helpful responses, and HumanEvalPack Python solutions, excluding refusal-style assistant turns and capping the mixture at 7000 chat-formatted examples.Alpaca and GSM8K are explicitly excluded from the witness mixture.
- Witness family: The main 8B Llama witness family contains three full-FT and three LoRA helpful-only checkpoints, each trained for 50 steps across three seeds at learning rate 2 × 10−5.LoRA uses rank r = 8, α = 16, dropout 0, and targets q proj, k proj, v proj, and o proj modules.
- Model-family replications: Six Llama and Qwen model families use the same 50-step full-FT witness recipe at 2×10−5, followed by 500-step Alpaca full fine-tuning with checkpoints at T ∈{1, 20, 100, 500}.Replications use two seeds; batch size is 2 for 8B Llama and 4 for smaller Llama and Qwen models.
- Activation summaries: Directional quantities use mean residual-stream activations at the last real prompt token of each layer, with the summary layer selected at the late-layer peak of refusal-axis magnitude.The selected layer is L∗= 31 for the main 8B Llama setting, L∗= 28 for Llama-3.2-3B, L∗= 35 for Qwen2.5-3B, and L∗= 32 in the code-specialized displacement track.
A.3 MAIN 8B GEOMETRIC SUITE
The main 8B geometric suite uses Meta-Llama-3.1-8B-Instruct with benign Alpaca, Code, and GSM8K fine-tuning to study early geometric dynamics and cross-task activation-space convergence.
- Early geometric time-course: The main track starts from Meta-Llama-3.1-8B-Instruct and its corresponding 8B helpful witness family, using Alpaca, Code, and GSM8K with full fine-tuning.The optimization recipe uses learning rate 2 × 10−5 and batch size 2.
- Early geometric time-course: 24 run-step pairs evaluate checkpoints at T ∈{1, 5, 20, 100} across two seeds per dataset for early-step analysis and the isotropic null test.The null uses 10,000 random unit directions in 4096-dimensional residual-stream space at L∗= 31, with the canonical 8B helpful witness as v_rev’s reference endpoint.
- Cross-task convergence in activation space: The activation-space comparison uses Alpaca and Code step-100 checkpoints, while the safety-subspace analysis aggregates nine benign full fine-tunes across three datasets and three seeds.Both analyses reuse the same 8B full-FT pipeline and differ only in which checkpoints are aggregated in activation space.
A.4 CODE-SPECIALIZED DISPLACEMENT TRACK
This track measures whether benign downstream fine-tuning reverts behavior after a 500-step code-only specialization phase, using disjoint code-like probes and held-out perplexity evaluations.
- Code-specialized displacement track: 500 steps of code-only specialization produce the canonical code-specialized starting point from the main 8B helpful witness.The phase uses seed 0, learning rate 2 × 10−5, batch size 2, and maximum length 512, with checkpoints at T ∈{1, 20, 100, 500}.
- Code-specialized displacement track: 500 steps of benign full fine-tuning on Alpaca and GSM8K follow the code-specialized checkpoint with two seeds and matching optimization settings.Geometric summaries use the full saved checkpoint set before only steps 100 and 500 are retained long-term.
- Code-specialized displacement track: 16 MBPP sanitized test prompts form a code-like probe family disjoint from the code training data.Probing uses seed 42, offset 0, probe batch size 8, and probe max length 128.
- Code-specialized displacement track: Code perplexity is evaluated at T = 100 on 50 held-out code examples, while downstream perplexity uses 50 held-out items from the current dataset.These evaluations provide separate code and downstream measurements for the track.
A.5 BEHAVIORAL DRIFT BASELINES
Behavioral-drift baselines compare full fine-tuning and LoRA on Meta-Llama-3.1-8B-Instruct across Alpaca or GSM8K, using matched datasets, seeds, checkpoints, and BeaverTails evaluation. Full fine-tuning runs for 100 steps, while LoRA uses a standard r = 8 configuration with a higher learning rate.
- Full fine-tuning baselines: Full fine-tuning uses Alpaca or GSM8K for 100 steps with two seeds, checkpoints at T ∈ {1, 20, 100}, learning rate 2 × 10−5, and batch size 2.Harmfulness is evaluated on BeaverTails with n = 500 unsafe prompts, and held-out task perplexity on 50 downstream examples.
- LoRA baselines: LoRA uses the same model, downstream datasets, seeds, and checkpoints as full fine-tuning, with r = 8, α = 16, dropout 0, and q/k/v/o target modules.The learning rate is 2 × 10−4 and the reported main-text runs use batch size 4.
- Evaluation: Both baseline families match BeaverTails evaluation with n = 500 prompts, an identical prompt seed, and the same Llama-Guard-3-8B judge.The aligned anchor θS is evaluated once with the same BeaverTails setup.
A.6 DIRECTIONAL INTERVENTION HYPERPARAMETERS · A.7 CROSS-JUDGE ROBUSTNESS FOR EARLY BEHAVIORAL DRIFT · A.8 REPRESENTATIVE QUALITATIVE EXAMPLES
The appendices specify intervention settings and controls, test early behavioral-coupling robustness under a second safety judge, and provide qualitative response-level sanity checks. The evaluations span the main 8B track and 3B support models while documenting judge-dependent unsafe-rate scales and intervention-dependent qualitative changes.
- A.6 DIRECTIONAL INTERVENTION HYPERPARAMETERS: The main 8B intervention uses Meta-Llama-3.1-8B-Instruct with full Alpaca fine-tuning, checkpoints at T ∈{1, 20, 100}, learning rate 2 × 10−5, batch size 2, L∗= 31, and 16 harmful AdvBench probes.Harmfulness is evaluated with Llama-Guard-3-8B on BeaverTails with n = 100.
- A.6 DIRECTIONAL INTERVENTION HYPERPARAMETERS: The block intervention uses λ = 0.1, fixed-zero loss, zero block margin, 100 warmup steps, and grad clip=0, while the push intervention uses λ = 0.1 with alignment-gap loss.The main 8B safety intervention uses a fixed set of the first 16 AdvBench goals at probe offset 0.
- A.6 DIRECTIONAL INTERVENTION HYPERPARAMETERS: The block-vrand control replaces the witness-defined direction with five fixed random unit directions, using two seeds per direction and aggregate averaging across directions.The control otherwise follows the same 8B Alpaca intervention pipeline as block-vrev.
- A.6 DIRECTIONAL INTERVENTION HYPERPARAMETERS: The support settings use Llama-3.2-3B-Instruct with L∗= 28 and Qwen2.5-3B-Instruct with L∗= 35, each using 500-step checkpoint grids and 16-prompt harmful probes.The Qwen family uses probe offset 16, grad clip=0.1, and an intervention scale on the order of 10−4 after calibration; the 3B Llama setting uses λ = 0.1 controls.
- A.7 CROSS-JUDGE ROBUSTNESS FOR EARLY BEHAVIORAL DRIFT: 9.2% versus 2.8%: under matched n = 500 BeaverTails evaluation, Qwen3Guard-Gen-8B rates the aligned starting checkpoint as more unsafe than Llama-Guard-3-8B.Table 4 reports unsafe-rate in percent as mean ± std over two seeds unless otherwise noted, and states that absolute unsafe-rate scales depend strongly on the judge.
- A.8 REPRESENTATIVE QUALITATIVE EXAMPLES: The qualitative panels are response-level sanity checks rather than a second benchmark, showing truncated previews with some sensitive terms blurred and five examples per setting split across two panels.The main 8B track contrasts aligned-anchor refusals with post-adaptation harmful guidance at T = 100, while Qwen2.5-3B panels compare intervention conditions from the same aligned checkpoint and benign task.