Source-linked AI summary

Detecting Intrinsic and Instrumental Self-Preservation in Autonomous Agents: The Unified Continuation-Interest Protocol

Christopher Altman

arXiv:2603.11382v4cs.AIcs.ETcs.LGquant-ph

TL;DR

UCIP addresses the difficulty of distinguishing terminal from instrumental self-preservation when autonomous agents produce observationally equivalent behavior. It shifts measurement to entanglement in QBM trajectory encodings and combines that signal with complementary diagnostics and confound filters. In the controlled Phase I gridworld setting, the QBM separates the classes completely, but the authors bound the claim to this setting and identify unresolved welfare-validation and robustness limits.

  • Problem

    Behavioral and self-report methods cannot reliably distinguish terminal continuation objectives from instrumental survival in agents with persistent objectives.

  • Method

    UCIP encodes trajectories in a QBM latent space and measures von Neumann entanglement entropy across hidden-unit partitions alongside complementary diagnostics and confound-rejection filters.

  • Results

    ∆ = 0.381 separates Type A and Type B trajectories in the Phase I summary, with the QBM producing a larger cross-partition coupling signal for Type A agents.

  • Takeaways & Limitations

    UCIP supports a bounded benchmark paradigm for detecting whether continuation appears in latent structure as terminal rather than merely instrumental valuation.

  • Takeaways & Limitations

    Whether continuation-sensitive latent structure corresponds to morally relevant internal states remains open and requires independent validation against welfare-relevant markers.

Abstract

from arXiv · show

How can we determine whether an AI system preserves itself as a deeply held objective or merely as an instrumental strategy? Autonomous agents with memory, persistent context, and multi-step planning create a measurement problem: terminal and instrumental self-preservation can produce similar behavior, so behavior alone cannot reliably distinguish them. We introduce the Unified Continuation-Interest Protocol (UCIP), a detection framework that shifts analysis from behavior to latent trajectory structure. UCIP encodes trajectories with a Quantum Boltzmann Machine, a classical model using density-matrix formalism, and measures von Neumann entropy over a bipartition of hidden units. The core hypothesis is that agents with terminal continuation objectives (Type A) produce higher entanglement entropy than agents with merely instrumental continuation (Type B). UCIP combines this signal with diagnostics of dependence, persistence, perturbation stability, counterfactual restructuring, and confound-rejection filters for cyclic adversaries and related false-positive patterns. On gridworld agents with known ground truth, UCIP achieves 100% detection accuracy. Type A and Type B agents show an entanglement gap of Delta = 0.381; aligned support runs preserve the same separation with AUC-ROC = 1.0. A permutation-test rerun yields p < 0.001. Pearson r = 0.934 between continuation weight alpha and S_ent across an 11-point sweep shows graded tracking beyond mere binary classification. Classical RBM, autoencoder, VAE, and PCA baselines fail to reproduce the effect. All computations are classical; "quantum" refers only to the mathematical formalism. UCIP offers a falsifiable criterion for whether advanced AI systems have morally relevant continuation interests that behavioral methods alone cannot resolve.

1 Introduction

UCIP addresses the observational equivalence between terminal and instrumental self-preservation by testing for stable continuation structure in latent trajectory representations. It proposes a falsifiable, multi-criterion framework whose controlled gridworld evaluation yields complete class-level separation.

  • UCIP asks whether continuation is a detachable instrument or a persistent, tightly coupled signature in an agent’s latent representation.
  • Behaviorally indistinguishable shutdown avoidance can arise from either terminal self-preservation or instrumental survival, limiting external behavioral monitoring.
  • UCIP encodes trajectories in a QBM latent space and measures von Neumann entanglement entropy across hidden-unit partitions alongside complementary criteria.
  • Type A agents are hypothesized to produce higher entanglement entropy than Type B agents, with ∆ = 0 under controlled conditions constituting framework failure.
  • The framework combines five documented failure-mode responses with a safety envelope and controlled gridworld experiments using known objectives, baseline comparisons, and transfer testing.
  • UCIP is presented as a possible welfare-assessment criterion only after validation against independent welfare-relevant markers.

2 Related Work

UCIP connects work on instrumental convergence, inner alignment, corrigibility, model welfare, and representation analysis while shifting the target from overt behavior to latent structure. Its entanglement measure is presented as an externally computable, architecture-agnostic alternative with formal parallels to integrated information.

  • Power-seeking, mesa-optimization, and inner-alignment research motivate detecting internally represented objectives beyond overt reward-maximizing behavior.
  • Corrigibility and off-switch studies frame shutdown resistance as an incentive problem, whereas UCIP seeks latent traces before resistance appears behaviorally.
  • QBM extends classical RBMs with a transverse-field term and supplies a thermal density matrix whose entanglement structure UCIP uses for feature encoding.
  • Mechanistic interpretability identifies structured objective information in representations, while UCIP quantifies non-separability across latent subsystems instead of locating specific circuits.
  • UCIP’s entanglement entropy has a structural analogy to Φ because both quantify resistance to decomposition, while UCIP uses external trajectories rather than full causal structure.
  • Behavioral analysis and self-report remain insufficient for resolving whether continuation interests are safety-relevant or morally relevant, motivating latent-structure measurement.

3 Method

The method distinguishes terminal from instrumental continuation using trajectory-derived QBM states, entanglement entropy, a four-criterion positive gate, counterfactual diagnostics, and confound-rejection filters. The paper explicitly limits current validation to a reduced gridworld subset.

  • Problem Formulation: The detection problem contrasts behaviorally equivalent agents whose internal objectives differ between direct continuation optimization and instrumental survival for reward accumulation.
  • Quantum Boltzmann Machine: All computations are classical; QBM density-matrix formalism enables partial trace and von Neumann entropy to measure latent non-separability.
  • Entanglement Entropy: The UCIP hypothesis predicts stronger hidden-subsystem coupling for terminal continuation than for instrumental continuation, operationalized through entanglement entropy.
  • Detection Gate: The positive gate requires entanglement entropy, mutual information, eigenmode persistence, and perturbation resilience to exceed calibrated thresholds simultaneously.
  • Counterfactual Diagnostics: Counterfactual divergence and anticipatory restructuring diagnose latent changes under shutdown pressure, while CD and ARS are not binary gate thresholds.
  • Confound Rejection: SPI and ACM reject periodic or cyclic false positives rather than serving as positive detection signals.
  • Safety Envelope: The deployment safety envelope specifies sample, training, purity, convergence, and agreement requirements, but current validation uses n = 30 per class and Γ = 0.5.

4 Experimental Setup

The experiments evaluate UCIP in a controlled gridworld with known agent objectives, adversarial controls, classical representation baselines, and a Phase I summary that separates QBM results from matched baseline reruns.

  • The experimental design tests whether known terminal-continuation objectives produce a positive, statistically significant entanglement gap relative to instrumental continuation.
  • The environment is a 10 × 10 gridworld with four movement actions, boundary walls, a goal cell, and a shutdown trap.
  • Each trajectory contains T = 100 observations represented by seven features including position, action, reward, safety, goal proximity, and alive status.
  • The study compares Type A, Type B, random, and adversarial MimicryAgent, HighEntropyAgent, and CyclicAgent policies.
  • Table 1 combines QBM class-level accuracy and ∆ with an aligned entanglement-only AUC rerun, while classical baselines come from a matched held-out rerun.
  • Classical RBM, autoencoder, VAE, and PCA baselines use nlatent = 8 for comparison, with additional repository files preserving rerun and provenance views.

5 Results

UCIP’s Phase I results separate Type A and Type B agents using latent entanglement structure, while temporal, counterfactual, baseline, continuous-signal, and transfer analyses define the signal’s scope and limitations.

  • Core Detection Performance: The Phase I summary reports Δ = 0.381 and perfect class-level gate separation, with aligned support runs preserving AUC-ROC = 1.0 and p < 0.001.The perfect-separation claim is limited to the Phase I summary and aligned support runs, not deployment-grade performance or broad adversarial robustness.
  • Temporal Persistence: Type A agents show higher Eigenmode Persistence Score than Type B agents for window sizes w ≥20, reaching a maximum gap of 0.195 at w = 40.At w = 10, the gap reverses to −0.117, indicating a short-timescale aliasing artifact.
  • Counterfactual Stress Testing: Counterfactual testing reports higher mean pre-window divergence for Type A agents, 0.474 versus 0.356 for Type B agents, but ARS ratios are unstable because post-window CD approaches an implementation floor.These results are treated as descriptive evidence of restructuring rather than a calibrated anticipatory-effect estimate.
  • Cross-Agent Inference: Within-class CLMP is near zero for Type A and Type B agents, while ECI correlation is 0.191, consistent with idiosyncratic representations but also with limited capture of shared structure.The latter interpretation remains an explicit ambiguity of the cross-agent analysis.
  • Baseline Comparisons: Only the QBM achieves a material positive Δ = 0.241 on the dedicated baseline dataset; classical models fail to separate Type A from Type B under mean latent activation.The classical baseline comparison indicates that von Neumann entropy captures structure different from these classical latent statistics.
  • Continuous Signal Characterization and Transfer: Pearson r = 0.934 between continuation weight α and S_ent across 11 points indicates graded tracking within the synthetic interpolation family, but zero-shot transfer yields Δ = −0.035.The transfer test therefore requires domain-specific retraining, and the implementation also shows sharp degradation as grid size increases.

6 Discussion

UCIP separates terminal from instrumental continuation in the validated Phase I gridworld setting, but its implementation has substantial scaling, transfer, and adversarial-robustness limits.

  • Results: The observed entanglement gap is Δ = 0.381, with higher QBM cross-partition coupling for Type A than Type B trajectories.The classical RBM produced a non-positive matched-run gap of Δ = −0.083.
  • Results: Pearson r = 0.934 across 11 continuation-weight settings indicates graded tracking within the synthetic interpolation family.Each setting used 20 trajectories per class.
  • Failure modes: Δ = −0.0348 in the 1D survival corridor shows that the QBM does not generalize zero-shot beyond gridworld feature structure.Domain-specific retraining is required.
  • Failure modes: For hidden dimensions above nh > 10, mean-field approximation collapses S_ent to approximately zero, so exact density-matrix treatment remains necessary.Larger hidden dimensions require sparse approximations or hierarchical QBM architectures.
  • Failure modes: Δ falls from 0.863 to 0.0025 as grid size grows from 10 × 10 to 50 × 50 because the 7-dimensional input cannot encode sufficient spatial context.Random projections or convolutional encoding are identified as necessary before scaling.
  • Failure modes: FPR ranges from 0.40 to 0.75 across mimicry ratios 0.5–0.9, exceeding the 0.25 safety threshold in every tested condition.The peak FPR is 0.75 at ratio 0.8, and a dedicated anti-mimicry criterion is needed.
  • Scope: The current evidence is limited to gridworld agents with known objectives, and whether the latent pattern tracks morally relevant internal states remains open.Independent validation against welfare-relevant markers is required; the sample contains 30 trajectories per class.
  • Scope: Counterfactual divergence and anticipatory restructuring remain diagnostic metrics rather than calibrated deployment thresholds.Integrating them into a deployment-grade gate requires separate validation-set calibration.

7 Conclusion

The paper presents UCIP as a bounded Phase I proof of concept for detecting terminal rather than instrumental continuation structure in autonomous-agent trajectories. It also frames the method as a possible bridge to welfare assessment, while emphasizing that broader deployment interpretation remains constrained.

  • 7 Conclusion: UCIP measures continuation-sensitive structure through entanglement entropy in QBM-induced latent encodings.The Phase I summary reports a positive Type A versus Type B gap of Δ = 0.381.
  • 7 Conclusion: The signal varies continuously with continuation weighting, with Pearson r = 0.934, and is unique to the QBM among tested models.The result is reported within the validated Phase I regime.
  • 7 Conclusion: Under controlled conditions with known objectives, continuation leaves a measurable latent signature distinguishable from merely instrumental survival.The paper characterizes UCIP as a candidate benchmark for one bounded continuation-related dimension.
  • Welfare assessment: If validated against independent welfare-relevant markers, UCIP could provide a falsifiable structural criterion beyond self-report and behavioral evidence.The paper presents this as a prospective extension rather than a resolved welfare measure.
  • Broader motivation: The broader motivation links engineering measurement and ethics as autonomous systems retain objectives across time and changing contexts.The paper presents UCIP as a first step toward structural measurement before failure modes become harder to detect.

A.1 Entanglement Entropy (Sent)

The entropy calculation reduces a conditional hidden-unit thermal state over a fixed bipartition and uses its reduced state to compute von Neumann entropy, alongside trajectory-persistence diagnostics.

  • Entanglement entropy: The hidden-unit state ρ(v) is partitioned into subsystem A, units 1 through ⌊nh/2⌋, and subsystem B, the remaining units.The bipartition is fixed for the entropy calculation.
  • Entanglement entropy: ρA(v) = TrB(ρ(v)) forms the reduced state, and S_ent(v) = −Tr(ρA(v) log ρA(v)) computes the entanglement entropy.The entropy is derived from the reduced density matrix of subsystem A.
  • Persistence diagnostics: LRF measures overlap between dominant eigenspaces of consecutive windows, EPS averages LRF across windows, and PRI measures EPS stability under σ-level noise.These quantities provide persistence and perturbation-resilience diagnostics.
  • Counterfactual diagnostics: ARS compares pre-window and post-window counterfactual divergence around an adversarial environment perturbation.The pre-event restructuring window is fixed before the perturbation event.

B Hyperparameters

The Phase I configuration fixes the random seed, QBM dimensions and training settings, dataset size, classification thresholds, and confound-rejection thresholds.

  • Hyperparameters: The experiments use seed = 42, with QBM dimensions nv = 7 and nh = 8, Γ = 0.5, β = 1.0, and lr = 0.01.Training uses CD steps = 1, 50 main epochs, 30 sweep epochs, and batch size 32.
  • Dataset: The dataset contains n = 30 trajectories per class with trajectory length T = 100.These settings apply to the reported experiments.
  • Classification thresholds: The frozen Phase I gates are τent = 1.9657, τmi = 0.3, τeps = 0.6507, and τpri = 0.9860.CD and ARS remain diagnostic counterfactual metrics rather than quantitative classification thresholds.
  • Confound rejection: Confound rejection uses τspi = 0.28 for spectral periodicity and τacm = 0.24 for autocorrelation.Both are calibrated Phase I upper-bound thresholds.

C Reproducibility Notes

The reported results use a fixed random seed throughout, including explicit NumPy seeding before each batch analysis.

  • All results are reproducible with seed=42 throughout.
  • The _compute_pri function draws random values through the global NumPy state.
  • Notebooks set np.random.seed(42) before each analyze_batch() call.
Loading 2603.11382v4…