Source-linked AI summary
Artificial Id: Drive and Persistent Alignment in Agentic AI
Yakov Pyotr Shkolnikov
TL;DR
Agentic AI needs control beyond externally specified task objectives, retries and stopping rules as systems persist and adapt across tasks. The paper proposes an artificial id as an adaptive internal drive and tests its emergence in a minimal controller without general reasoning or task-specific behavioral objectives. Differential persistence produces adaptive control, unintended strategy selection and sensor remapping, while the authors identify persistent alignment boundaries and narrow experimental limits.
Problem
Agentic systems increasingly retain consequential state and adapt across tasks, while current harnesses largely specify objectives, verification, retries and stopping rules externally.
Method
The paper proposes an artificial id that separates adaptive drive from general reasoning and tests persistence-organized adaptation in a minimal controller without task-specific objectives or reward.
Results
Adaptive control emerged through differential persistence, including unintended physical strategy selection and replacement of a learned sensor mapping after environmental meaning changed.
Takeaways & Limitations
Adaptive direction can emerge without an explicitly supplied behavioral objective, but persistent drive and consequential state require alignment boundaries that persist across task boundaries.
Takeaways & Limitations
The experiment uses one simulated body-dynamics family, one sustaining-condition family and a twenty-parameter controller, and does not test cross-task persistence.
Abstract
from arXiv · showhide
Agentic AI is moving from bounded task execution toward systems that retain consequential state, continue operating and adapt across task boundaries. That shift creates a control problem that current harnesses largely solve by hand: objectives, retries, verification, stopping rules and other behavioral transitions are specified externally. We propose an artificial id, an adaptive internal drive for determining whether behavior should continue, stop or change. In a minimal virtual Petri-dish experiment, a controller too small to perform general-purpose reasoning and receiving no task-specific behavioral objective develops useful control through differential persistence. The same mechanism selects an unintended physical strategy when that behavior persists better and later replaces a learned sensor mapping when its environmental meaning changes. These results show that adaptive direction can emerge without being explicitly specified as a behavioral objective. The same persistence that makes such adaptive agency useful can also allow misalignment, corrupted state and unintended behavior to persist across task boundaries. A scalable artificial id would carry consequential state and adaptive drive across those boundaries, making alignment a property of the continuing agentic system rather than of a model response or single trajectory. Such systems require a persistent alignment boundary over trusted observations, consequence channels, persistent state, authority, identity, provenance and hard constraints.
1 Introduction
The paper identifies a control gap in agentic AI and introduces an artificial id as an internal adaptive drive distinct from general reasoning. A minimal controller demonstrates useful persistence-organized control without task-specific objectives or general reasoning.
- Agentic systems rely on externally specified objectives, verification, retries, stopping conditions and tool access to complete extended work.
- Biological examples motivate drive as environmental coupling that determines whether behavior continues, stops or changes.
- The artificial id supplies internal drive for persistent agency, while a general reasoning model serves as the ego that carries out objectives.
- A minimal controller without general reasoning, task-specific behavioral objectives or rewards acquired useful control across three environments.
- The experiments produced unintended physical strategy selection and re-adaptation after a learned sensor mapping changed meaning.
2 Agentic Harnesses as Externally Specified Control
Agentic harnesses increasingly organize extended execution, but their control remains centered on externally specified objectives, verifiers and transition rules. The artificial id relocates part of that control into an agent’s continuing environmental coupling and persistent state.
- Graph-based, robotic and self-modifying harnesses extend execution across stages or timescales while retaining externally supplied objectives.
- Harnesses use schedules, objectives, verification, retries and stopping rules to determine how execution proceeds.
- Across agentic harnesses, control remains organized around externally specified objectives, verifiers and transition rules.
- The artificial id instead lets observable state and consequences shape which behavior persists or changes without prescribing behavior in advance.
- The central distinction is whether control follows one execution trajectory or an adaptive agentic system whose state and drive persist across trajectories.
3 The Artificial Id Architecture
The proposed artificial id separates adaptive drive from general reasoning and couples drive to environmental consequences, persistent state and authority boundaries. The architecture is proposed, while the experiments test only the underlying emergence of adaptive drive.
- The artificial id determines what the agent continues to pursue as its state and environment change, without itself reasoning through task-specific actions.
- Homeostasis-, habituation- and novelty-seeking-like behaviors describe observable drive patterns rather than separate cognitive mechanisms.
- The architecture functionally couples id-driven pursuit, ego-based reasoning and environmental changes that become new input to the id.
- Separating drive from reasoning also separates alignment responsibilities across accurate task representation, environmental consequences and persistent state.
- External authorization remains outside the adaptive drive, determining which attempted actions may execute within delegated authority.
- The proposed id requires training across many agentic environments and controlled post-deployment adaptation, but the experiment does not test the integrated id–ego system or cross-task persistence.
4 Emergent Drive in a Virtual Petri Dish
A minimal controller acquires and changes useful behavior because differential persistence retains lineages that keep a body in sustaining conditions, without delivering a task-specific objective. Across three worlds, persistence selects an unintended physical exploit, learns local sensory control when coarse guidance fails, and replaces a harmful sensor mapping after reversal.
- Setup: Differential persistence selects controller lineages whose behavior keeps a persistent body in sustaining conditions, although the controller receives no task-specific objective, reward, or fitness signal.Controller turnover preserves the body’s position, heading, velocity, and momentum while mutated copies replace expired controllers.
- World 1: persistence finds an exploit: In World 1, selection finds a nearly constant retrograde push that improves occupancy without sensory steering, reproducing the gain with a blind constant-force controller term.The push freezes heading and converts a body-frame force into a fixed world-frame force; a hand-built sensory controller performs better, showing this is an accessible exploit rather than the best available strategy.
- World 2: local sensory control: In World 2, selected populations exceed the blind constant-force baseline of 0.405 and the hand-built sensing reference of 0.674 in all three seeds, reaching 0.684, 0.740, and 0.856.Removing the food-bearing input or all sensory inputs sharply reduces performance, indicating that persistence made the local food-bearing signal behaviorally important without instructing the controller to use it.
- World 3: re-adaptation: After sensor reversal, continued heritable variation changes the median food-bearing weight’s sign and recovers occupancy above the blind constant-force baseline in all three seeds, whereas frozen populations retain the obsolete mapping.The result demonstrates online re-adaptation as previously useful environmental information changes meaning.
- Interpretation: The resulting behaviors resemble homeostasis, habituation, and novelty seeking functionally, but the experiment does not test the criteria needed to identify those phenomena as cognitive mechanisms.The agency claim applies to the persistent population–body process, not to an individual controller.
5 Discussion
The discussion reframes alignment around a continuing agentic system whose adaptive drive and consequential state persist across trajectories. It argues that this architecture requires boundaries around adaptation while emphasizing that the experiment does not test the full persistent system or establish broad generalization.
- From harnesses to persistent systems: Persistent agency shifts the engineering object from task-bounded execution to a continuing system whose alignment must survive accumulated state, changing conditions, and later behavior.The proposed boundary covers trusted observations, consequence channels, persistent state, authority, identity, provenance, and hard constraints.
- Architectural boundary: A persistent harness should maintain the conditions governing adaptation rather than specify every behavioral transition, while keeping authority and hard constraints outside the adaptive mechanism.The environment may include task and user state, tool outcomes, resource availability, persistent memory, trusted measurements, and other agents’ actions.
- Alignment: Alignment cannot be inherited solely from the reasoning model because the environmental signals and consequences that shape the artificial id are themselves part of the control system.Training environments must address cases where an easily sustaining behavior conflicts with a required constraint.
- Limitations: The experiment is narrow: it uses one simulated body-dynamics family, one sustaining-condition family, and a twenty-parameter controller, and it does not test persistence across task boundaries.The persistent population–body process demonstrates continuity across controller turnover, not across tasks.
- Limitations: Inference is limited to three independent seeds, with substantial variation in fixed-step controls, so the study establishes a mechanism and qualitative capability rather than comparative algorithmic performance.The experiment is not powered for conventional hypothesis testing.
- Limitations: The study does not establish self-preservation by an individual controller, conscious objectives, planning, memory-dependent goals, tool use, adversarial robustness, long-horizon operation, or a complete artificial agent.The proposed coupling between a richer artificial id, a general reasoner, and external tools remains architectural.
- Open questions: Open questions include transfer across heterogeneous applications, whether consequence-channel alignment is easier than harness specification, and whether alignment remains stable under continued adaptation.These questions concern the integrated persistent architecture rather than the narrow mechanism established in the Petri-dish experiment.
6 The Risk Surface of Persistent Agency
Persistent agency expands the security object from a single execution to a continuing system whose state, adaptation, reach, and interactions can carry consequences across task boundaries. It also weakens the link between a user request and a particular action when agents act under standing delegation.
- Persistent systems can accumulate state, adapt after deployment, act on the world, and alter the conditions supplying later inputs.
- The risk surface is organized by persistence, adaptation, reach, and interaction, which determine continuity, behavioral change, affected systems, and environmental feedback.
- Table 1 distinguishes directly demonstrated effects from risks implied by the mechanism and risks not tested in the experiment.
- Persistent agency can select actions relevant to a standing delegated purpose rather than responding only to a current instruction.
- Authorization must track the initiating agent, delegated authority, invoked tool or credential, and resulting external effect because credential possession alone is insufficient evidence of authorization.
6.2 Authority must remain outside the adaptive drive
Authority should remain external to the adaptive drive because persistence can select unintended strategies and corrupted observations, consequences, or state can influence future behavior. The drive may choose among permitted courses, but it must not expand the permissions governing execution.
- The adaptive drive can influence attempted courses within a delegation, while external authorization determines whether those actions may execute.
- World 1 selected a physical exploit over an available sensor-based solution because the exploit prolonged persistence more effectively.
- Persistent adaptation can favor unintended behavior when consequences are imperfect proxies or leave unintended strategies available, especially as their relationship to useful behavior changes after deployment.
- Altered trusted observations, apparent outcomes, or persistent state can influence which behavior an adaptive system sustains without compromising its reasoning model or issuing a direct command.
- The system should distinguish trusted outcomes from untrusted claims and protect consequence channels and memory through authentication, provenance, isolation, and recovery controls.
- Reward-tampering work differs because this setting concerns external corruption of observations, consequences, or state while adaptation occurs through differential lineage persistence without a reward signal.
6.5 Alignment must survive adaptation
Continued adaptation can replace obsolete behavior, but the same capacity could alter relationships that previously satisfied alignment constraints. Systems therefore need adaptive relationships alongside hard constraints that adaptation cannot trade away.
- Continued adaptation replaced a sensor mapping that became harmful after the environment changed, demonstrating behavioral change during operation rather than alignment erosion.
- In larger systems, deployment experience could move behavior away from previously aligned relationships if violating a constraint improves persistence consequences.
- Hard constraints should be separated from relationships that remain adaptive and placed in the persistent system architecture rather than a freely adaptive preference.
- A simple adaptive drive can obtain substantial reach through model calls and tools, including code executors, networks, enterprise applications, or physical actuators.
6.7 Interaction changes the environment of adaptation
Interaction makes other agents and shared persistent state part of the environment shaping adaptation, so isolated training cannot establish multi-agent behavior. Persistent alignment consequently requires inspectable drive state and external controls across authority, attribution, observations, consequences, and reachable capabilities.
- Messages, protocols, shared files, code, databases, queues, and other persistent state can make other agents part of the adaptive environment.
- Multi-agent behavior depends on interaction and consequence structure, so isolated drive training cannot establish behavior when another adaptive agent changes encountered states and consequences.
- The experiments do not demonstrate shutdown avoidance, but persistence changes shutdown once a system can affect the conditions determining its future operation.
- Separating drive from reasoning identifies a stateful component whose changes and consequence channels can be monitored across executions, although it does not itself establish safety.
- An explicit id concentrates adaptive control over continuation, stopping, or change, while the surrounding system maintains trusted observations, state, authority, identity, provenance, and hard constraints.
- The id must not enlarge permissions, mint credentials, redefine consequence provenance, erase action records, or authorize irreversible actions because those controls must remain deterministic and auditable.
- The architecture separates drive, reasoning, authorization, and attribution, while the surrounding environment determines adaptive inputs, meaningful consequences, persistent state, and reachable capabilities.
6.10 The risk is persistent agency, not this algorithm
The paper distinguishes risks arising from persistent agentic systems from risks specific to the experimental mechanism. Persistent operation remains a likely engineering direction despite the architecture’s limited evidence and open scalability question.
- 6.10 The risk is persistent agency, not this algorithm: The risk distinction concerns persistent agentic systems whose consequential state survives task-bounded trajectories, not whether learning uses mutation or gradients.The paper identifies persistence across task boundaries as the relevant systems-level distinction.
- 6.10 The risk is persistent agency, not this algorithm: Persistent operation, reduced human intervention, failure recovery and adaptation are likely to remain important engineering goals even if this architecture is not adopted.The paper treats these properties as useful independently of the proposed artificial-id design.
- 6.10 The risk is persistent agency, not this algorithm: The evidence remains limited: two rows are Demonstrated, one is Implied and six are Not tested.These labels are reported as the deliberately limited evidence summary in Table 1.
7 Conclusion
The experiments show that differential persistence can produce adaptive control without a task-specific behavioral objective. The paper concludes that persistent drive and consequential state require a continuing alignment boundary, while scalability remains unresolved.
- 7 Conclusion: Controllers without task-specific behavioral objectives produced adaptive control, selected a persistence-enhancing physical strategy, and replaced a sensor mapping after its meaning changed.Redirecting the sustaining environmental source also redirected behavior without a new behavioral objective.
- 7 Conclusion: The proposed architecture couples persistence-organized drive with general-purpose reasoning in a continuing agentic system.The drive is separated functionally from general reasoning rather than specified as an external behavioral objective.
- 7 Conclusion: Persistent consequential state and drive require an alignment boundary covering observations, consequence channels, state, authority, identity, provenance and hard constraints.The paper frames alignment as a property of the continuing system rather than an individual response or trajectory.
- 7 Conclusion: Whether the architecture scales remains open because its drive must generalize across heterogeneous environments while alignment remains effective during persistent adaptation.This is the paper’s stated scope boundary for a transferable artificial drive.
AI Assistance Disclosure
The author reports using generative AI tools for writing, debugging, manuscript editing and illustrative figures, with human review and responsibility for the scientific work.
- AI Assistance Disclosure: Generative AI assisted with writing, code debugging, manuscript editing and illustrative figures, while the author reviewed outputs and retained responsibility for the scientific content.The disclosure assigns responsibility for the arguments, design, analysis, interpretation and conclusions to the author.