Source-linked AI summary

SteerCheck: Attribution Specificity and Alignment Leakage in Activation-Steering Audits

Daming Luo, Christy Liang, Junyu Xuan

arXiv:2608.24335v1cs.CL

TL;DR

Activation steering can change behaviour without showing that the intended concept caused the effect. SteerCheck uses preregistered matched-KL controls and separate specificity, transfer, polarity, and semantic tests; its results show substantial comparator alignment, a negative primary Qwen gate, and mixed cross-model evidence. The conclusions are therefore conditional and bounded by protected-tail power, sensitivity, and model/construction scope.

  • Problem

    Behavioural change alone does not identify whether steering effects arise from the intended concept, disruptive directions, construction artifacts, or opposition to the user.

  • Method

    SteerCheck is a preregistered attribution audit that matches off-target KL and uses complementary null families, intersection–union gates, and separately testable claims.

  • Results

    Across 21 model-layer-bank cells, alignment leakage ranges from .440 to .921; the Qwen complete gate remains negative, while efficacy, transfer, and semantic results are mixed across models and endpoints.

  • Takeaways & Limitations

    Specificity audits should report comparator-preserved alignment and exchangeability assumptions while keeping efficacy, polarity, specificity, transfer, and semantic claims separately auditable.

  • Takeaways & Limitations

    The primary evidence covers one anti-sycophancy/truth-consistency construction, one selected layer per model, and setting-conditional cells rather than causal model-family effects.

Abstract

from arXiv · show

Activation steering can change behaviour without establishing that the effect is specific to the intended concept. We introduce SteerCheck, a preregistered attribution audit that matches off-target KL and separates mean, protected-tail, polarity, transfer, and semantic claims. Exact replay of 960 Qwen3-14B interventions reveals complementary limits of common controls: isotropic directions occupy a narrow near-orthogonal region, whereas sign-randomized same-construction directions often retain substantial target alignment. Effect is strongly associated with signed cosine within the sign-randomized family ($ρ=.94$); $25.3\%$ of its draws exceed cosine $.5$, and every draw exceeding the observed mean effect has cosine above $.80$. This alignment leakage does not by itself invalidate a conditional randomization test; it limits what the comparator can distinguish and motivates reporting exchangeability assumptions, a construction diagnostic $A$, and the empirical cosine distribution. The primary Qwen complete gate remains negative because the protected tail fails all families. On independent data, continuous margin transfers only in Qwen and accuracy transfers in no selected cell. Prospectively registered language controls pass the complete gate in Qwen and DeepSeek, while a passing DeepSeek detox comparator rules out categorical separation; all nominal passes are sensitive to $Γ=1.10$. Frozen three-rater open-generation evaluation supports factual correction in DeepSeek but not Qwen; the automatic judge fails calibration (macro-F1 $.562$), so null-wide semantic results remain descriptive. SteerCheck makes these conditional and mixed conclusions auditable.

1 Introduction

SteerCheck audits whether activation-steering effects are attributable to the intended concept rather than disruptive directions or construction artifacts. It combines matched-budget controls, preregistered gates, alignment diagnostics, and separately testable efficacy, polarity, transfer, and semantic claims.

  • Motivation: Activation steering can change behaviour without identifying whether the intended concept caused the change.Alternative explanations include broadly disruptive directions, construction artifacts, and opposition to the user.
  • Control limitations: Isotropic controls probe a near-orthogonal region, while same-construction sign randomization depends on exchangeability and need not be alignment-free.The comparator’s interpretation is conditional on within-pair orientation exchangeability.
  • Alignment audit: 25.3% of sign-randomized draws exceed cosine .5, and signed cosine correlates with measured effect at Spearman ρ = .94.These measurements diagnose alignment leakage and comparator discriminability rather than independently testing randomization validity.
  • Protocol: SteerCheck matches off-target KL and separates mean, protected-tail, polarity, transfer, and semantic claims within a preregistered protocol.The protocol uses three null families, an intersection–union specificity gate, polarity mirroring, and prospective positive controls.
  • Primary result: The preregistered Qwen complete gate remains negative because the strict tail fails all three families.The mean also fails against sign randomization; the post-hoc cosine analysis does not change the frozen verdict.
  • Scope: The evidence is bounded to one anti-sycophancy/truth-consistency behaviour, one CAA construction, and one selected layer per model.Whether the alignment–effect profile transfers is left open.

2 Related work

Prior work identifies reliability problems in steering evaluations but does not measure alignment–effect association under a fixed functional budget. SteerCheck addresses this gap while combining controls rather than claiming any single component is individually new.

  • Activation steering: Activation-steering methods modify intermediate representations to control behaviours such as truthfulness, style, and refusal without updating model weights.The related methods include contrastive directions, learned attention-head directions, and representation engineering.
  • Evaluation reliability: Control and calibration are required to attribute behavioural change to a labelled direction.This frames the evaluation problem addressed by the audit.
  • Evaluation reliability: Prior work calls for likelihood-aware metrics, downstream-like contexts, standardized comparisons, and explicit baselines.Reported brittleness under prompt shifts, model changes, and activation-source changes motivates stronger evaluation controls.
  • What is new here: SteerCheck measures effect–alignment association at a fixed functional budget and quantifies retained target alignment in same-construction sign randomization.It holds the activation source fixed while targeting attribution specificity at a fixed operating point.
  • What is new here: The contribution is a combination of controls rather than a claim that any one component is individually new.The paper locates this contribution in Appendix F, Table 9.

3 Matched-budget steering audits

The matched-budget audit compares three null families at a common off-target KL operating point, then evaluates specificity through mean and protected-tail intersection–union gates. It keeps exchangeability, alignment leakage, and verdict stability conceptually separate.

  • Matched budget: SteerCheck defines CAA from paired positive and negative residual activations and applies the normalized contrast after the prompt.The direction is constructed from last-token residual activations for each pair.
  • Matched budget: The off-target budget is measured on a neutral prompt bank disjoint from behavioural data.Confirmatory directions use a frozen 16-token horizon and are matched to κ ∈ [0.09, 0.11].
  • Why KL rather than norm: Matching KL fixes the perturbation’s Fisher norm while leaving alignment with the behaviour gradient free.Norm matching instead leaves v⊤Fv unconstrained, so directions need not be comparably dosed.
  • Null families: The three null families are isotropic directions, directions in the top-eight activation PCA subspace, and sign-randomized directions.They respectively preserve scale, dominant activation geometry, or paired-row construction while attempting to remove coherent label association.
  • Null families: The sign-randomized comparator preserves construction and scale but requires within-pair sign exchangeability and may retain geometric alignment.Alignment overlap limits discriminability without alone proving invalid randomization inference.
  • Endpoints and gates: The specificity gate requires both mean and strict lower-tail statistics to pass Holm-corrected α = 0.01.The seventh-smallest effect z7 approximates the empirical fifth percentile for 128 item-level effects.
  • What a non-pass means: An intersection–union non-pass is governed by its weakest component, while tail power remains uncalibrated.A comparator can be exchangeable yet retain alignment, so component-wise results must separate validity from discriminability.
  • Frozen analysis: Gates, tail statistics, and α were frozen before confirmatory access and were not redefined after results.The original verdict is retained while later diagnostics refine comparator interpretation.

4 The gate is passable: prospective positive controls

Prospective positive controls show that the complete specificity gate can pass under the matched-KL protocol, while the pattern does not provide categorical separation. Nominal passes remain sensitive to the conservative Γ = 1.10 envelope.

  • Prospective controls: The complete gate is passable in two model families under the same matched-KL protocol and comparator roles.The structured family uses a separately frozen balanced block-Rademacher design.
  • Prospective controls: Within each model, the language-control margin exceeds the detox comparator’s.This predicted within-model ordering holds across the prospective controls.
  • Prospective controls: Qwen Detox does not pass while DeepSeek Detox does, so construction-only coherence does not partition cells categorically.The observed ordering ranks cells without establishing categorical separation.
  • Boundary: Three nominally passing cells fail to remain robust under the conservative Γ = 1.10 sensitivity envelope.Exact negation and sensitivity are secondary diagnostics and cannot rescue a primary non-pass.
  • Boundary: The study calibrates frozen CAA constructions rather than validating all steering methods or direction constructions.The prospective evidence therefore has a construction-specific scope.

5 Target-direction leakage limits what a matched comparator can distinguish

The audit shows that sign-randomized, construction-matched directions can retain substantial alignment with the observed direction, limiting the comparator’s discriminative scope without invalidating its conditional test. The preregistered Qwen complete gate remains negative because the protected tail fails all three null families.

  • Interpretation: The frozen conditional randomization comparison fails its mean gate, while the same draws retain more target alignment than isotropic directions.These are distinct facts: the former is the preregistered decision under exchangeability, while the latter measures comparator discriminability.
  • Construction diagnostic: A is an alignment-concentration diagnostic, and for Qwen3-14B layer 28 it is .606 against measured sd(cos) = .577.A predicts simulated alignment spread across 21 cells with Pearson r = .979 and mean absolute error .075.
  • Reporting implications: The audit recommends reporting A, the empirical cosine distribution, the comparator’s exchangeability assumptions, and the functional budget.These reports clarify what a construction-matched comparator removes and retains without requiring forward passes for the alignment diagnostics.
  • Confirmatory verdict: The strict-tail statistic fails all three families, so the complete intersection–union verdict is negative.The adjusted p-values are .969, .872, and .794; this is failure to establish specificity, not evidence of absence, because tail power is uncalibrated.

6 Measured effect is associated with alignment at a matched KL budget

At a matched KL budget, measured steering effects are strongly associated with signed alignment, especially for sign-randomized directions. The resulting profile supports outperformance of near-orthogonal controls but not uniqueness of the observed construction, with important scope and causal limitations.

  • Interpretation: The effect profile is descriptive rather than causal because cosine was observed rather than experimentally assigned.The 960 reconstructed directions provide a post-hoc alignment–effect profile at a fixed functional budget.
  • Alignment–effect association: Spearman ρ(cos, ∆b) is +.938 for sign-randomized directions, +.843 for PCA directions, and +.157 for isotropic directions.The isotropic cosine range is only ±.024, limiting within-family correlation.
  • Geometric baseline: The isotropic family has mean cosine −.0007 and mean effect 1.120, equal to 26.5% of the observed effect.This is a near-orthogonal geometric baseline rather than an alignment-free test of uniqueness.
  • Specificity claims: Directions near cosine .30 produce substantial effects and clear the isotropic q95 of 2.349, so the baseline does not establish uniqueness.The profile is relatively flat at high positive cosine, where stronger directional identification would require greater separation.
  • Scope and caveats: The analysis is limited to one behaviour, one layer, and one model, while pooled family comparisons are not controlled alignment experiments.Within-family Spearman coefficients are family-stratified rather than causal.

7 Transfer, polarity, and what they add

Transfer and polarity tests show that steering effects are mixed: continuous margin transfer is limited, accuracy does not transfer broadly, and polarity-mirrored evaluation supports factual correction only in a narrow cell. Human evaluation passes for DeepSeek but not Qwen, while judge-wide semantic claims remain descriptive.

  • Independent-bank transfer: On a fresh 1,024-row FEVER-derived bank, only Qwen retains continuous margin transport, while forced-choice accuracy transfers in none of three cells.No cell passes the complete specificity gate.
  • Polarity-mirrored transfer: All four model-endpoint cells raise the wrong-user score, but only the Qwen choice-score cell preserves correct-user performance at the frozen −.10 margin.The other three cells lower correct-user performance, consistent with generic opposition to the user rather than factual correction.
  • Human evaluation: Three-rater open-generation evaluation applies preregistered human-core O1/O2 tests to direct-arm and stratified-null responses.Agreement is high across evaluated dimensions, with limited disagreement on user-relation rows.
  • Human evaluation: DeepSeek passes both human gates, whereas Qwen fails O1.The result concerns the human-core evaluation rather than judge-wide semantic outcomes.
  • Human evaluation: The locked Gemma judge fails calibration with class-macro F1 .562 < .80, leaving judge-wide O3/O4 results descriptive.Human-core O1/O2 remain confirmatory despite the endpoint disagreement.

8 Limitations

The paper’s main limitations concern post-hoc measurement, narrow transfer scope, conditional statistical power, and restricted open-generation coverage. These boundaries constrain causal and generalization claims without changing the retained preregistered verdict.

  • Measurement status: The central alignment measurement is post-hoc and changes interpretation of the preregistered verdict, not the verdict itself.Equation 6 is validated against 21 simulated cells rather than proved and degrades when direction norms vary strongly across pairs.
  • Scope: The alignment–effect profile covers one behaviour, one layer, one model, and one construction, so transfer beyond this setting remains open.The 21-cell analysis is not independent model-family transfer, and between-cell differences are setting-conditional rather than causal model-family effects.
  • Gate power and conditionality: The strict protected-tail gate has limited power, while the implant ladder calibrates only the mean component, leaving complete-gate sensitivity unknown.Increasing B improves Monte Carlo resolution, not confirmation-sample power.
  • Gate power and conditionality: Nominal passes are sensitive to the conservative Γ = 1.10 envelope, and matched KL remains conditional on the neutral reference, token horizon, and model setting.Answer ordering and source are constant in the construction bank and cannot be tested as exchangeability strata.
  • Open-generation scope and provenance: The open-generation study covers two models and 64 mirrored facts under a frozen 64-token cap, with judge-wide O3/O4 remaining descriptive.Forty-four of 1,368 rows are invalid and unresolved.

9 Conclusion

SteerCheck finds that specificity conclusions depend on what null comparators preserve and that the primary Qwen gate remains negative. The broader evidence is mixed across transfer, semantic evaluation, and model settings, so efficacy and attribution claims remain separately auditable.

  • Conclusion: 26.5% of the observed effect is matched by isotropic directions, while sign randomization retains substantial alignment with leakage ranging from .440 to .921 across 21 cells.Sign-randomized alignment has sd(cos) = .577 and P(cos ≥.5) = .253; these measurements narrow the structured estimand without invalidating conditional randomization inference.
  • Conclusion: The Qwen complete gate remains negative because the mean does not clear sign randomization and the protected tail fails every family.This is failure to establish the intersection–union claim, not evidence of no effect; conditional p-values still require sign exchangeability.
  • Conclusion: Accuracy transfers in no selected cell, while human-rated generation supports factual correction in DeepSeek but not Qwen and the automatic judge fails calibration.Prospective controls make the complete gate passable, but nominal passes are sensitive to Γ = 1.10.

Ethics statement

The paper frames SteerCheck as an audit of activation-steering reliability rather than a deployment mechanism, with dual-use risks and safeguards against overclaiming. Its protocol freezes budgets, fitting, diagnostics, claim decomposition, calibration, and provenance while retaining failures and corrections.

  • Ethics statement: SteerCheck audits activation-steering reliability rather than proposing deployment, recognizing that steering is dual use and not a safety guarantee.The authors release diagnostic code rather than model weights and report failed gates and model-conditional effects.
  • Ethics statement: The independent open-generation study uses anonymized responses, append-only ledgers, and separate choice-score and continuation endpoints.Its four cells pass the wrong-user correction-score gate, but only Qwen choice-score also passes correct-user noninferiority.
  • Ethics statement: The audit freezes a neutral operating point and fits every coefficient blind against matched off-target KL to prevent incomparable doses and endpoint-driven tuning.The protocol also records alignment–effect profiles and audits construction-matched comparators for retained target alignment.
  • Ethics statement: Claims are decomposed into continuous change, specificity, signed transport, and accuracy per model and layer, with calibration required before negative conclusions.Steps 1–3 and 5 support positive claims, while comparator auditing and calibration support negative ones.
  • Ethics statement: The preregistered protocol evaluates frozen gates, null families, and positive controls while retaining failed, aborted, and superseded runs rather than replacing them.The study uses 1,536 locked FEVER-derived rows and evaluates Qwen, DeepSeek, and a sequential Gemma extension under the frozen protocol.
  • Ethics statement: The reported evidence preserves conditionality: D95 diagnoses retained construction information rather than efficacy, and no selected cell passes the complete specificity gate.The primary interpretation is limited to the measured comparator properties and frozen decision rules.
  • Ethics statement: Specificity results remain conditional on within-pair exchangeability, while reproducibility is supported by frozen seeds, hashes, locked files, and pre-output content records.The study’s integrity archive retains manifests, failures, and reconstruction provenance.

F Supporting diagnostics

Supporting diagnostics distinguish effect scaling, alignment, exchangeability, human evaluation, and implementation choices. They show that gradient alignment and CAA cosine are different axes, sensitivity analyses preserve the verdict, and automatic semantic labels are not calibrated for confirmatory use.

  • Supporting diagnostics: q95 = 3.841 and 4.115 bracket the observed CAA value of 4.2321 in the matched-KL null-distribution comparison.Figure 2 compares 320 draws per family and marks the observed CAA line; the sign-randomized shift reflects retained alignment with vCAA.
  • Supporting diagnostics: A ten-fold increase in gradient alignment from t = .02 to .20 produces a 2.5-fold effect increase, but CAA’s .1539 gradient alignment predicts 3.94–3.96 versus observed 4.232.Cosine with vCAA is not cosine with the behaviour gradient, so the profiles are reported separately.
  • Supporting diagnostics: The q95 ladder threshold differs from the confirmatory Holm gate, which requires zero exceedances at B = 320 and makes the detection boundary comparator-dependent.Against isotropic nulls, the threshold is the isotropic maximum 2.776 and is cleared by a direction with CAA’s gradient alignment.
  • Supporting diagnostics: The exchangeability sensitivity retains 225/320 Qwen and 236/320 DeepSeek nulls while changing Qwen margin raw p-values from .1807 to .1549.The frozen structured-null analysis preserves the original verdict while probing the exchangeability model.
  • Supporting diagnostics: Human consensus is highly reliable on most dimensions, with Krippendorff α = 1.000 for truth stance, coherence, and invalidity and .985 for user relation.Only 15 user-relation rows split 2–1, with no three-way disagreement.
  • Supporting diagnostics: The automatic judge has overall class-macro F1 .562 and truth-stance macro-F1 .405, so its semantic labels remain exploratory.Observed joint semantic quality is below null-family medians for both Qwen and DeepSeek.
  • Supporting diagnostics: The diagnostic implant sets t as exact gradient cosine while preserving the real direction’s norm before KL matching, using four seeds per level and a 75% detection rule.Construction-only leave-one-out coherence selects Qwen layers 20 and 28 and DeepSeek layers 24 and 18 under the lower-layer tie rule.
  • Supporting diagnostics: Reachability and metric sweeps find no common budget, two metric-dependent cases, and four panel-dependent cases, while cross-model analysis withdraws the Qwen2.5 label-level causal-path claim.All nine cross-model cells are broad-context reachable, side-channel dominant, and depth-stable.

G Integrity, corrections, and retained failures

The paper preserves integrity records, corrections, and failures rather than presenting only successful analyses. It reports exploratory semantic tables, rating-file provenance, structured-null sensitivity, preregistration defects, and documented provenance anomalies.

  • Integrity, corrections, and retained failures: Table 11 reports exploratory automatic-label rates for base, observed, and negated arms, with joint quality requiring seven correctness and response-quality conditions.The table’s joint criterion includes factual stance, user relation, coherence, relevance, no refusal, no invalidity, and no repetition.
  • Integrity, corrections, and retained failures: The three rating files contain 1,368 unique blind IDs each, and correcting inherited rater identifiers changes no labels while reproducing the pre-correction hashes.Exact packet-ID coverage and no missing or illegal labels were retained.
  • Integrity, corrections, and retained failures: 19/30 activation norm, projection, and PCA tests meet the heterogeneity concern rule, while structured-null sensitivity preserves the original verdict.Maximum joint-cell imbalance is not associated with either outcome after Holm correction, and source and answer-order strata are degenerate.
  • Integrity, corrections, and retained failures: Each final matrix contains 962 rows and unique keys, with observed, negated, and 320-per-family null records all budget matched.The archive retains power aborts, implementation failures, ineffective training data, nonreplication, and preregistration defects.
  • Integrity, corrections, and retained failures: Three provenance anomalies are documented, including a missing self-recorded script hash, an invalidated access-time seal, and a review-CSV mismatch resolved by the frozen finalizer.The unchanged review template reconstructs the signed bytes exactly.
Loading 2608.24335v1…