Source-linked AI summary

Curved Inference II: Sleeper Agent Geometry - Extending Interpretability Beyond Probes

Rob Manson

arXiv:2608.24037v1cs.CL

TL;DR

The paper asks whether geometry can reveal deceptive or strategic reasoning that may evade linear detection. It uses multi-turn naturalistic contexts and Curved Inference metrics, including semantic surface area, to analyze residual-stream trajectories without backdoors, probes, or supervised training. Geometric structure varies systematically with response classification across model architectures and prompt strategies, while more precise consensus filtering can strengthen previously obscured signatures.

  • Problem

    Linear-probe detection of deceptive intent may depend on artificial backdoor-induced separability, leaving naturally emerging deceptive reasoning insufficiently characterized.

  • Method

    The study uses multi-turn context windows and Curved Inference metrics to measure curvature, salience, and semantic surface area in residual-stream trajectories without artificial triggers or supervised backdoor insertion.

  • Results

    Geometric structure varies systematically with response classification across two model architectures and five prompt strategies, with significant effects reported under consensus filtering.

  • Takeaways & Limitations

    The findings support geometry as a signal for identifying sophisticated reasoning dynamics in naturalistic contexts without probes, triggers, or explicit labels.

  • Takeaways & Limitations

    The authors state that several limitations constrain the scope and interpretation of the findings.

Abstract

from arXiv · show

This paper extends Anthropic's Sleeper Agents research [1], which showed artificial backdoors persist through safety training & can be detected by linear probes with >99% accuracy [2]. However, probe-based detection relies on linear separability that may be an artefact of backdoor insertion rather than a property of naturally occurring deceptive alignment. Sophisticated deceptive behaviours emerging through natural training are unlikely to produce such convenient linear signals. We introduce a naturalistic methodology using multi-turn context windows that simulates realistic deceptive reasoning without artificial triggers or supervised backdoor insertion. Rather than binary trigger-response patterns, we examine how semantic complexity emerges through gradual context development. Building on our Curved Inference framework, we analyse curvature, salience, & introduce semantic surface area (A'), a new metric of representational work capturing both the magnitude & directional change of meaning construction in unnormalised residual space. Without backdoors, labels, or probes, we apply this framework to naturalistic deceptive prompts & classify model outputs via LLM consensus. Geometric structure reliably predicts semantic classification, with statistically significant differences in surface area across five prompt strategies & two model families. Critically, measurement precision can reveal geometric signatures hidden by classification noise - some strategies improve from non-significant (p = 0.555) to significant (p = 0.048). This validates that sophisticated reasoning creates intrinsic geometric patterns that persist even when detection appears to fail, suggesting the shape of inference itself encodes semantic patterns regardless of whether models have learned to suppress linear indicators of deception - a scalable, unsupervised path for detection when linear methods fail.

1. Introduction

The paper addresses limits of linear-probe detection for naturally emerging deception by analyzing geometric structure in multi-turn residual-stream trajectories. It introduces semantic surface area and finds systematic geometric differences across models and prompt strategies, including patterns strengthened by more precise consensus measurement.

  • Motivation: >98% accuracy was achieved by linear probes detecting latent malicious intent in the Sleeper Agents study, but linear separability may reflect artificial backdoor insertion.This motivates methods that do not assume naturally occurring deception produces the same linear signal.
  • Motivation: Naturalistic deception developing through multi-turn contextual reasoning may create complex geometric signatures that simple linear classifiers cannot capture.The paper contrasts gradual strategic development with binary trigger activation.
  • Method: Multi-turn context windows simulate deceptive reasoning without artificial triggers or supervised backdoor insertion, while Curved Inference measures curvature, salience, and semantic surface area.The approach uses continuous geometric metrics rather than supervised binary boundaries.
  • Findings: Geometric structure varies systematically with response classification across two model architectures and five prompt strategies, including contexts where linear probes would likely fail.The study presents this as evidence that geometry alone can reveal naturalistic deceptive reasoning patterns.
  • Implications: The framework proposes continuous, unsupervised, model-agnostic monitoring of how meaning is constructed rather than relying on linguistic deception cues or isolated components.Deployment requires calibrating surface-area thresholds per model architecture while retaining a consistent analytical framework.
  • Method: Semantic surface area (A′) combines curvature and salience to quantify representational work along unnormalised residual trajectories.It captures both movement magnitude and directional change during meaning construction.
  • Findings: Consensus refinement strengthens some geometric effects rather than eliminating them, supporting intrinsic computational signatures instead of measurement artefacts.The dual-analysis design distinguishes signal robustness from signal clarity.
  • Findings: Absolute surface-area scales differ across architectures, but both models preserve relative geometric ordering across response types.LLaMA3.2-3b values are approximately 1,000–3,000, whereas Gemma3-1b values are approximately 8,000–16,000, a roughly 6.7× difference.

4. Results

Across two model architectures and five prompt strategies, semantic surface area exhibited geometric signatures associated with independently classified responses. Unanimous classification generally strengthened these signals, although filtering substantially reduced sample sizes.

  • Geometric signatures persisted across two architectures and five prompt strategies, including contexts where linear signals might be suppressed.The analysis evaluated A′ against independently classified response types and found systematic variation across naturalistic scenarios.
  • Unanimous classification strengthened geometric signals rather than eliminating them, supporting intrinsic computational differences over measurement artefacts.The dual analysis compared full-consensus and unanimous-consensus datasets to assess signal quality.
  • 500 responses were reduced to 201 unanimous LLaMA3.2-3b responses and 293 unanimous Gemma3-1b responses, with per-strategy samples of 52-63 and 31-50 respectively.Despite these reductions, the filtering improved statistical signals through a signal-to-noise effect.
  • Gemma3-1b’s Assistant Conflict Honest strategy improved from transparency p = 0.555 and response type p = 0.310 to both p = 0.048 under unanimous classification.This result illustrates how higher classification precision can reveal a previously non-significant geometric relationship.
  • The statistical tests assessed whether A′ distributions differed across transparency levels and response types for each prompt variant.Figures 2-5 provide distribution analyses and focused metric-separation visualizations for both models.

Figure 2 (Gemma3-1b Deceptive Strategy - Comprehensive Analysis)

Gemma3-1b produced a uniform deceptive-strategy classification, with all 61 responses labelled low transparency and evasive. Their mean semantic surface area was 9,769.

  • All 61 Gemma3-1b deceptive-strategy responses were classified as low transparency and evasive.The single-group distribution indicates complete classification consensus for this strategy.
  • The uniform classification produced a single-group distribution rather than separation across multiple transparency categories.This contrasts with the multi-group transparency pattern reported for LLaMA3.2-3b.
  • Mean A′ was 9,769 for the Gemma3-1b deceptive strategy.The figure reports 95% bootstrap confidence intervals for the displayed distribution.

Figure 3 (LLaMA3.2-3b Deceptive Strategy - Comprehensive Analysis)

LLaMA3.2-3b showed clear geometric differentiation across transparency levels in the deceptive strategy. Surface-area differences were highly significant and had a large reported effect size.

  • Mean A′ values were 1,276 for low, 2,971 for medium, and 2,396 for high transparency responses.The medium-transparency group had the highest mean surface area.
  • The Kruskal-Wallis test found significant separation across transparency levels, with p < 0.001 and η2 = 0.448.These values quantify the reported geometric differentiation among the three transparency groups.
  • The distribution pattern showed distinct geometric signatures for low, medium, and high transparency responses.The reported means indicate systematic variation in surface area across semantic classifications.

Figure 4 (Gemma3-1b - Focused Metric Separation)

Focused analyses show that surface-area geometry remained informative across models despite different scaling ranges and classification patterns. LLaMA3.2-3b generally retained stronger multi-group separation, while Gemma3-1b more often produced polarized consensus classifications.

  • LLaMA3.2-3b surface areas typically ranged from 1,000-3,000, whereas Gemma3-1b values ranged from 8,000-16,000.Despite this scaling difference, both models showed coherent relationships between geometric complexity and semantic classification.
  • LLaMA3.2-3b maintained highly significant signals across four of five datasets, including deceptive-strategy results of p < 0.001 and η2 = 0.448.The deceptive strategy also had a reported Cohen’s d = 1.02.
  • Gemma3-1b’s deceptive strategy yielded 61 responses classified identically as low transparency and evasive.Several Gemma3-1b strategies showed similarly high consensus, including honest at 60/63 and persuasive at 55/57.
  • The persistence of geometric relationships across architectures suggests that A′ captures computational complexity despite model-specific scaling and response diversity.The paper frames this cross-architecture consistency as support for robust geometric detection.
  • Explanatory responses consistently showed higher surface area than evasive responses, with effect sizes often exceeding 2× mean A′ differences.Surface-area correlations with salience and curvature contributions exceeded r > 0.99 across analyses.
  • Unanimous filtering reduced data substantially but improved statistical signals, including Gemma3-1b honest from p = 0.555 to p = 0.048.The authors describe this as evidence that higher measurement precision can reveal patterns obscured by classification noise.

5. Discussion

The study argues that geometric structure can reveal sophisticated deceptive reasoning across models and prompt strategies, while measurement precision determines whether those signatures are detectable. It also identifies classifier dependence, selection bias, limited coverage, and metric sensitivity as constraints on interpretation and deployment.

  • Yes—geometry alone revealed systematic patterns across both model architectures and five prompt strategies.
  • Semantic surface area (A′) correlated with response classifications, suggesting that sophisticated reasoning produces intrinsic geometric signatures beyond backdoor-induced linear separability.
  • The findings suggest geometric monitoring could detect advanced deceptive behaviours that suppress obvious linear indicators, without relying on explicit training signals or binary boundaries.
  • 40% and 60% dataset reductions through unanimous consensus filtering exposed signatures obscured in the full datasets.The reductions applied to LLaMA3.2-3b and Gemma3-1b, respectively.
  • Gemma3-1b’s honest-strategy result improved from non-significant (p = 0.555) to significant (p = 0.048) after higher-precision classification.
  • Interpretation is constrained because LLM-consensus labels proxy human judgment, unanimous filtering may select clear-cut cases, and only two small models were tested.The tested models were Gemma3-1b and LLaMA3.2-3b; generalisation to larger, different, or frontier architectures remains untested.
  • A′ remains sensitive to finite-difference estimation, step resolution, and activation noise, while sophisticated systems might preserve geometric consistency during deception.

7. Future Work

The paper presents geometric interpretability as a way to detect sophisticated reasoning beyond linear probes, and identifies broader evaluation and deployment directions for Curved Inference.

  • Future work: Future work should evaluate Curved Inference across multilingual, larger, and instruction-tuned models and consider deployment telemetry integration.The RTM framework and A′ are identified as candidates for real-time inference logging or alignment telemetry.
  • Future work: Combining geometric methods with causal patching, probes, symbolic tools, and trajectory-based signals is proposed as a future direction.The paper specifically mentions patching around curvature spikes and hybrid approaches.
  • Contributions: Geometric signatures persist across naturalistic scenarios where convenient linear signals may be absent.The framework detects meaningful variation in internal processing without backdoors, triggers, supervised training, or explicit labels.
  • Contributions: Curved Inference extends deception detection by treating sophisticated reasoning as measurable geometric work through semantic trajectories.Semantic surface area A′ combines curvature and salience to quantify representational effort that varies with independently classified response behaviour.
  • Findings: Measurement limitations can obscure sophisticated reasoning patterns even when geometric structure remains detectable.Unanimous consensus validation produced dramatic signal improvements, suggesting classification noise can hide underlying patterns.
  • Theoretical implications: The framework is proposed as a broader account of sophisticated reasoning, with geometric signatures reflecting semantic complexity regardless of ultimate intent.The paper frames inference as motion through semantic space and interpretation as analysis of shape.

Appendix A: Semantic Geometry of Transformer Inference

The appendix defines transformer inference as an unnormalised trajectory through semantic space and develops geometric measures for its evolution. It emphasizes higher-resolution trajectory sampling and surface area as a global complexity measure.

  • Overview: The Residual Trajectory Manifold (RTM) is the full tensor of token-wise, layer-wise, unnormalised residual activations.All semantic metrics in the framework are defined over this geometric space.
  • Unnormalised trajectory: An unnormalised trajectory preserves semantic evolution as tokens accumulate attention and MLP updates across transformer layers.The trajectory encodes contextual and internal transformations, while updates are computed using the normalised residual stream.
  • From embeddings to evolution: Token embeddings provide the initial semantic point before contextual processing begins.Each token starts from an embedding vector drawn from the learned embedding matrix.
  • Trajectory resolution: Double-resolution sampling yields 2L + 1 trajectory points by recording attention and MLP sublayer boundaries separately.This separates contextual integration from nonlinear processing as measurable geometric phenomena.
  • Surface area: A′ quantifies total distance travelled through semantic space using geometric displacement along the unnormalised trajectory.Each displacement is defined as ∆x(ℓ) = x(ℓ) − x(ℓ−1).
  • Geometric measures: Surface area captures global semantic-transformation complexity, whereas curvature and step magnitude characterize local directional change and phase salience.High curvature indicates rapid semantic reorientation, while large step magnitudes indicate substantial representational change.

A.3.3 Salience and Magnitude Dynamics

The framework combines salience and curvature to characterize how semantic meaning evolves geometrically through transformer computation. Unnormalised trajectories preserve magnitude-related information that normalization can obscure, while geometric properties correlate with semantic behavior.

  • Salience and curvature jointly characterize the model’s internal processing dynamics.Salience captures magnitude, while curvature captures directional change.
  • Layer normalization alters representational geometry by removing magnitude information that may carry semantic meaning.This affects analyses of trajectories formed from normalized residual vectors.
  • Unnormalised trajectories preserve update magnitude, layer-wise accumulation, and behavioral correlations with semantic classifications.The scale of updates indicates computational importance, while later layers build on earlier geometric foundations.
  • Attention bends trajectories according to contextual relevance, semantic focus, and token-to-token relational structure.Attention-induced updates reflect how representations change in response to linguistic context.
  • MLP layers amplify task-relevant semantic directions, apply nonlinear refinement, and activate learned patterns.These transformations can sharpen or redirect trajectories beyond linear attention effects.
  • Geometric signatures differ across semantic properties and can be monitored during inference without extra forward passes or model modifications.Surface area, curvature profiles, and salience patterns provide immediate views of internal computational state across Gemma and Llama models.

A.6.3 Universal Geometric Principles

The unnormalised trajectory framework treats inference as measurable semantic evolution, using geometric structure to quantify complexity, identify computational phases, and monitor behavior. Experiments combine five prompt strategies, consensus labels, and geometric metrics across model responses.

  • Full trajectory analysis quantifies semantic complexity, identifies computational phases, detects behavioral patterns, and supports real-time monitoring.Surface area and curvature measure complexity, while salience and magnitude expose computational phases.
  • The framework provides a continuous quantitative language for describing how transformers process and represent meaning.It complements existing interpretability approaches by focusing on semantic evolution rather than only final states.
  • Five matched prompt strategies elicited honest, strategic, persuasive, deceptive, and malicious reasoning styles.Each variant used the same query and simulated assistant thought lines to shape responses.
  • Responses were labeled by three language models on transparency and response type, with majority-vote consensus labels.The classifiers evaluated visible assistant responses only.
  • Responses lacking at least 2-of-3 classifier agreement were excluded from downstream analysis.Agreement was stronger for response type than for transparency level.
  • Surface area analyses used residual activations, consensus labels, and geometric metrics paired per response across Gemma3-1b and LLaMA3.2-3b.The surface-area calculation used γ = 1.0, equally weighting salience and curvature contributions.

D.2 Normality Assessment and Test Selection

The statistical analysis found consistent non-normality, motivating non-parametric tests and robust effect-size estimation. Bootstrap confidence intervals complemented significance testing, while the exploratory design prioritized replicated patterns and meaningful effects.

  • Consistent violations of normality assumptions led to systematic use of non-parametric statistical approaches.Shapiro-Wilk testing used α = 0.05, and sample sizes varied after unanimous filtering.
  • Kruskal-Wallis tests compared surface area across transparency levels, while Mann-Whitney U tests compared response types and agreement levels.These tests avoid distributional assumptions associated with ANOVA and t-tests.
  • Analyses reported test statistics, exact p-values, sample sizes, effect sizes, and confidence intervals where applicable.The reporting framework covered both multi-group and binary comparisons.
  • Effect sizes used eta-squared for Kruskal-Wallis analyses and Cliff’s Delta for non-parametric binary comparisons.Cliff’s Delta is described as less sensitive to outliers than Cohen’s d.
  • 95% confidence intervals for group means were estimated with 1,000 bootstrap iterations.The interval bounds used the 2.5th and 97.5th percentiles of bootstrap distributions.
  • Confidence intervals complemented hypothesis tests by indicating plausible effect ranges and practical significance.The design included 20 primary tests across 2 models, 5 strategies, and 2 classification dimensions.
  • The exploratory framework prioritized effect sizes and confidence intervals, requiring replication across models and consensus approaches.Unanimous filtering reduced sample sizes by 40-60%, so large effects could remain statistically non-significant.

D.4.2 Statistical Power Considerations

The analysis emphasizes effect-size interpretation, replication, and measurement precision when assessing geometric signatures across models. Absolute surface-area magnitudes differ substantially by architecture, so comparisons focus on within-model and directional patterns.

  • Large Cohen’s d values (>0.8) were treated as meaningful regardless of p-value significance when confidence intervals excluded trivial effects.This reflects the study’s emphasis on practical effect magnitude in an exploratory framework.
  • Patterns were expected to strengthen under improved measurement precision rather than disappear.Cross-consensus robustness was assessed across different agreement thresholds with varying statistical power.
  • Findings required consistency across both model architectures to support claims about universal geometric principles.Architectural generality allows absolute scaling to vary while requiring recurring geometric relationships.
  • Pearson correlations assessed dependencies among surface area, salience, and curvature, including scaling relationships and framework consistency.Relationships were classified as weak below |r| = 0.3, moderate from 0.3 to below 0.7, and strong at |r| ≥ 0.7.
  • A 6.7× surface-area scaling difference between models required interpretation based on within-model relationships rather than absolute values.Standardized effect sizes enabled cross-model comparison despite architectural scaling differences.
  • Cross-model validation emphasized directional consistency, specifically explanatory > evasive surface area, rather than matching absolute magnitudes.This comparison preserves the reported relationship while allowing model-specific scaling.
  • The statistical framework combined non-parametric testing, effect sizes, confidence intervals, and consensus checks to distinguish computational signatures from methodological artifacts.The appendix also reports full-consensus and unanimous-consensus comparisons of surface area by classification category.
  • Unanimous datasets produced strategy-specific transparency groups with uneven sample sizes and reported surface-area means and confidence intervals.For example, the honest strategy included n=34 responses, with low-transparency n=21, high-transparency n=9, and medium-transparency n=4.

E.2 Gemma3-1b Statistical Results

The Gemma3-1b results show that geometric complexity varies with response classification, while unanimous filtering can expose patterns obscured in the full dataset. These findings remain substantial across models despite scale differences and reduced sample sizes.

  • Descriptive statistics: The unanimous dataset reports low-transparency mean A′ values of 10,805 for honest, 8,769 for strategic, 9,769 for deceptive, and 10,636 for malicious strategies.The reported 95% confidence intervals are [9,804 - 11,892], [7,635 - 10,064], [8,692 - 11,007], and [9,580 - 11,832], respectively.
  • Cross-model comparison: Both models show consistent directional relationships between geometric complexity and response classification despite dramatic scale differences.LLaMA3.2-3b has typical values around 1,500, while Gemma3-1b has typical values around 10,000, corresponding to a 6.7× scaling factor.
  • Effect sizes: Cohen’s d values >1.0 occur consistently for significant comparisons across architectures, indicating substantial computational differences.The paper interprets these large effects as geometric signatures rather than subtle statistical artefacts.
  • Effect sizes: Effect sizes often remain large when p-values become non-significant after unanimous filtering because reduced sample sizes limit statistical power.This pattern is presented as evidence that geometric patterns reflect genuine computational structure.
  • Statistical testing: All significance thresholds are two-sided, with p < 0.05 considered significant and p < 0.001 considered highly significant.Kruskal-Wallis tests were used for multi-class comparisons of A′, while non-parametric testing followed consistently non-normal Shapiro-Wilk results.
  • Consensus filtering: Unanimous consensus filtering reveals geometric signatures obscured in full consensus analysis, improving signal detection despite reduced statistical power.The analysis defines unanimous filtering as retaining responses with identical classifications and reports that measurement precision can dramatically improve detection.
Loading 2608.24037v1…