Source-linked AI summary

Software Engineering in the Agent Era From Trustworthy Change to Human Agent Software Organizations

Zhongjie Wang, Mingyi Liu

arXiv:2609.04630v1cs.SE

TL;DR

Software agents make execution elastic, but semantic commitment, verification, integration, and residual-risk acceptance remain bounded by human and organizational capacity. The paper constructs Trustworthy Change, Responsibility Topology, and the Human–Agent Cell to govern this mismatch, while limiting its claims to a testable framework whose empirical validity remains open.

  • Problem

    Agent-scaled execution can expand faster than the human cognition, organizational authority, and economic capacity needed to frame, verify, integrate, and accept software changes.

  • Method

    The paper defines Trustworthy Change as the lifecycle engineering object, Responsibility Topology as the distribution of residual-risk acceptance authority, and the Human–Agent Cell as an execution boundary.

  • Results

    The framework makes residual-risk acceptance authority an explicit classification axis and derives consequences for change state, shared engineering facts, verification, and flow control.

  • Takeaways & Limitations

    Agentic execution should be governed around complete changes and explicit acceptance authority rather than treated as equivalent to engineering completion.

  • Takeaways & Limitations

    The paper’s empirical commitments cover only single-center and multi-independent-center topologies, not more complex governance regimes.

Abstract

from arXiv · show

Software agents make digital execution elastic: repository analysis, code generation, testing, migration, tool use, and operations can be replicated and parallelized without proportional human headcount. Problem framing, semantic commitment, verification, integration, attention, and residual-risk acceptance remain bounded by human cognition, organizational authority, and economic capacity. How should scalable execution be governed so organizations can accept and sustain its changes? Our testable framework has two constructs and one execution abstraction. Trustworthy Change (TC) is the engineering object moving from intent through delegated execution, verification, integration, acceptance, and operation. Responsibility Topology classifies organizations by the distribution of independent residual-risk acceptance authority. A single-center topology has one final baseline responsibility anchor; a multi-anchor topology requires joint acceptance across independently governed domains. The Human-Agent Cell (HAC) produces candidates, proposals, and evidence; execution grants no acceptance authority. As execution and authority scale differently, distributed HACs create context-coherence and invalidation pressures, while multi-anchor governance adds joint acceptance and explicit responsibility closure. Responsibility, accountability, change management, specification, verification, and human oversight predate this work; our claim is only that agent-scaled execution changes how they fit together. We make that authority an explicit classification axis and derive consequences for change state, shared engineering facts, verification, and flow control. Progressive Specification and bounded-capacity analysis remain hypotheses to test, not laws. We contribute theory construction and operationalization; empirical validity remains open to controlled, longitudinal, and field studies.

1 Introduction: The Scale of Software Execution Is Changing

Software agents make digital execution replicable, concurrent, and more elastic, shifting scarcity toward semantic judgment, verification, integration, and responsibility. The paper responds by framing Agent Software Engineering around a common engineering core and explicit responsibility topology.

  • Agents can analyze repositories, generate code, test, migrate APIs, configure environments, and support deployment or runtime diagnosis.
  • Digital execution can scale through concurrent agents, isolated workspaces, and token, compute, and tool budgets faster than human staffing mechanisms.
  • Execution Capacity ≫ Judgment, Verification, Integration, and Responsibility.
  • As implementation capacity grows, bottlenecks move toward semantic decisions, verification, integration, resource management, and formal risk acceptance.
  • The paper shifts the control object from human process activity to Trusted Change Governance across intent, delegated execution, verification, integration, acceptance, and operation.
  • Agent SE combines a common engineering core with Responsibility Topology, which classifies how residual-risk acceptance authority is organized.

2 From Human Process Control to Trusted Change Governance

Agentic development expands execution capacity while leaving system constraints and organizational coordination necessary. The paper therefore makes the complete software change—not isolated tasks, commits, or releases—the control object of Trusted Change Governance.

  • Traditional software engineering constrains development activity, augments construction capability, and organizes collaboration across people and time.
  • Agentic execution reduces the marginal cost of search, generation, modification, testing, and local diagnosis through model calls, agent instances, and compute resources.
  • State, authorization, data consistency, interface compatibility, performance, and security still depend on system behavior rather than implementation origin.
  • Execution can complete while the software change’s engineering conditions remain open, producing locally correct candidates that compose into system-level errors.
  • A complete agentic change may span agents, tasks, commits, specifications, data contracts, configuration, instructions, and runtime policy.
  • Trustworthy Change connects intent, delegation, verification, integration, and operation into one lifecycle-level control object.

3 Execution Units and Responsibility Topology

The paper separates software execution from residual-risk acceptance. Human–Agent Cells execute within defined boundaries and produce candidates, proposals, and evidence, while Responsibility Topology classifies how acceptance authority is distributed.

  • Human–Agent Cell: A Human–Agent Cell combines one human execution subject with agents, local context, tools, permissions, and a resource budget.Changing agents or token budgets usually does not create a new HAC; changing the human subject or formal engineering boundary can trigger a handoff, termination, or redefinition.
  • Human–Agent Cell: A HAC consumes authoritative-state and task inputs, then produces candidate changes, local evidence, and proposals rather than organizational fact.Local memory remains execution state, so state-changing outputs are constrained to candidates, proposals, and evidence.
  • Responsibility Anchor: A Responsibility Anchor is a person or institutional role with formal authority to accept or reject residual risk for a baseline or responsibility domain.Agents may execute within delegated boundaries, but residual-risk acceptance remains assigned to an identifiable human or institutional legal subject.
  • Responsibility Topology: Responsibility Topology classifies the distribution of independent residual-risk acceptance rights, not generic organizational roles, reviewer counts, or agent counts.The paper distinguishes execution and acceptance roles and narrows the term to authority over software baselines or affected responsibility domains.
  • Responsibility Topology: Single-center topology has one final anchor, whereas multi-anchor topology has at least two independent anchors whose domains require joint closure.A multi-anchor organization can still have changes requiring only one effective anchor; topology describes the wider governance structure.
  • Execution and Responsibility Distribution: Execution distribution and responsibility distribution are independent: distributed execution creates context divergence and invalidation pressures, while distributed authority adds joint acceptance and responsibility closure.The paper therefore treats delegation and risk acceptance as different structures: the Delegation Graph is not the Responsibility Graph.

4 Trustworthy Change: A Unified Object for Software Change

Trustworthy Change (TC) is a lifecycle-spanning, cross-artifact identity for governing consequential software changes from intent through operation and residual-risk acceptance. Its Candidate–Eligible–Accepted progression separates engineering qualification from formal responsibility closure while preserving traceability, evidence, resources, and authority.

  • TC as a unified object: Trustworthy Change spans a software change’s intent, delegated execution, verification, integration, acceptance, and operation.TC addresses gaps among commits, pull requests, change requests, and releases by organizing the complete lifecycle of one consequential change.
  • TC as a unified object: TC provides one traceable identity across production, verification, configuration, measurement, economic accounting, and responsibility.The framework does not require one physical record; the relevant information must remain traceable to the same underlying change.
  • State progression: A Human-Agent Cell first produces a Candidate, which consumes token, compute, tool, and human-attention resources regardless of eventual acceptance.A Candidate is digital execution’s artifact awaiting engineering evaluation.
  • State progression: Eligibility is engineering qualification, whereas acceptance formally closes residual risk; neither implementation completion nor organizational willingness substitutes for the other.Only Eligible changes may enter formal acceptance, and only Accepted TC enters the authoritative baseline.
  • State progression: Accepted changes remain relative to their version, evidence set, and time condition because new runtime facts can invalidate assumptions and trigger a new TC.Eligibility may also require revalidation when an authoritative fact affecting the change changes.
  • Responsibility checkpoints: Responsibility follows TC across Intent, Delegation, Verification, Integration, and Operation rather than appearing as a separate final stage.These checkpoints ask who controls goals, authorizes action, judges evidence, admits the change to the baseline, and accepts runtime residual risk.
  • Relation to change management: Agent-scaled execution adds delegated permission, resource envelopes, provenance, evidence independence, runtime drift, and Responsibility Topology to traditional change management.TC is therefore an extensible abstraction for reorganizing change management under agentic execution conditions; the running order-state example is illustrative rather than empirical validation.

5 Common Engineering Methods for Agent Software Engineering

The paper defines common engineering methods for agent software engineering around progressively specified intent, bounded execution, evidence-based verification, and explicit human control. These methods preserve human authority while making agent actions auditable, reversible, and constrained by resources and risk.

  • The common method layer applies the same specification, task, verification, integration, and recovery methods across distributed execution settings.The paper treats topology-specific differences as differences in closure conditions, not separate engineering methods.
  • Intent and Progressive Specification: Progressive specification balances under-specification, which increases interpretation and rework, against over-specification, which increases elicitation and context costs.The paper frames a Minimum Sufficient Specification as a task- and risk-dependent boundary, not a universal minimum.
  • Intent and Progressive Specification: Semantic commitment remains with authorized human responsibility subjects, while agents support ambiguity detection, clarification, drafting, candidate tests, and consistency checking.Facts, evidence, assumptions, preferences, and decisions receive different epistemic and governance status.
  • Intent and Progressive Specification: The shapes and minimum-sufficient region in Figure 6 are conceptual hypotheses, not empirical results of this paper.The research agenda will test whether additional specification eventually yields diminishing reductions in ambiguity, rework, or verification uncertainty relative to lifecycle cost.
  • Architecture and Change Locality: Agent-native systems should be agent-operable, human-understandable, and human-takeoverable while representing change locality across artifact, context, and responsibility spans.Poor change locality is expected to correlate with higher context, integration, and verification cost, but this remains an empirical claim.
  • Construction and Verification: Execution should produce comparable, verifiable, and revocable candidate changes within controlled environments, with explicit task-resource limits and stop conditions.Security controls should bound action consequences even when interpretation is compromised, and candidate exploration must respect verification capacity.

6 Single-Center Software Engineering

Single-center software engineering combines distributed execution capacity with one final responsibility anchor. Its scale is bounded not only by execution throughput but by verification, attention, integration, continuity, and the anchor’s ability to sustain accepted changes.

  • Topology and Responsibility: The structural property of single-center software engineering is distributed execution capacity with concentrated final responsibility.Peripheral specialists and services may provide execution, consultation, or evidence, while formal residual-risk acceptance remains at the final anchor.
  • Capacity and Verification: Adding execution units can increase candidate arrival rates without increasing verification and acceptance service rates, creating partially understood changes and compressed review.The bottleneck remains structural because one responsibility center must understand the baseline coherently before accepting residual risk.
  • Capacity and Verification: When generation exceeds verification, a verification queue forms and can pressure the responsibility center toward summaries, rapid approvals, and weaker independent judgment.The verification queue is presented as an observable backpressure signal in agentic software production.
  • Evidence Independence: Single-center organizations can increase evidence independence through deterministic analysis, property-based testing, mutation testing, hidden tests, heterogeneous models, user validation, and external review.Shared models, specifications, and contexts can otherwise cause multiple agents to miss the same defect.
  • Attention and Change Locality: Change locality limits scalability when a small feature forces the final anchor to reconstruct many modules, protocols, data models, and agent traces.A scalable single-center architecture therefore requires bounded semantic scope, review scope, failure radius, and long-term responsibility obligations.
  • Complete Cost: Reducing creation cost does not proportionally reduce lifecycle cost, because verification, integration, operation, and long-term responsibility remain part of the complete cost.Resource optimization should focus beyond the price of an individual model call.
  • Topology Transition: A responsibility topology changes when governance creates multiple mutually independent responsibility anchors, such as through regulation, separation of duties, or independent safety or security acceptance.This transition is distinct from merely adding capacity within an unchanged topology.
  • Continuity and Scale: The effective scale of single-center software engineering is the number of accepted changes one center can continuously understand, verify, integrate, finance, and sustain.High automation alone creates neither takeoverability nor resilience to anchor absence; both require externalized knowledge, state, authority, and emergency procedures.

7 Multi-Anchor Team Software Engineering

Multi-anchor governance separates distributed execution from distributed authority: distributed HACs create context-coherence pressures, while independent anchors require shared authoritative facts, invalidation handling, and joint responsibility closure.

  • Authoritative State and Context Coherence: Multi-anchor governance makes authority, versioning, and invalidation semantics necessary when cross-anchor closure depends on shared engineering facts.The organization need not share all local memory, but it does need jointly recognized formal engineering facts.
  • Authoritative State and Context Coherence: Shared memory is not shared authoritative state: proposals become binding only after relevant anchors accept and commit them in versioned form.Typical authoritative content includes specifications, contracts, invariants, ownership, accepted changes, risk state, and release state.
  • Authoritative State and Context Coherence: Context Coherence requires formal actions to identify the authoritative facts and versions they depend on and respond appropriately when those facts change.Exploratory plans may remain local or eventually consistent, while public interfaces, accepted invariants, ownership, and release state may require stronger coordination.
  • Context Invalidation and Closure: Authoritative changes invalidate dependent work selectively: tasks may Continue, become MarkStale, Refresh, Replan, Reverify, or Stop according to impact.For a v3-to-v4 contract change, payment race analysis must Reverify, while invalidated historical backfill work can Stop.
  • End-to-End Closure: Execution Closure is distinct from Responsibility Closure because delegation and execution do not grant acceptance authority.The Delegation Graph describes assigned work, whereas the Responsibility Graph identifies subjects with formal acceptance authority.
  • End-to-End Closure: The running order-state example requires separate Order and Payment anchor acceptance before the v4 contract enters shared authoritative state.The example illustrates the framework rather than empirically validating it.
  • Augmented Conway Hypothesis: The Augmented Conway Effect hypothesis proposes that Agent Delegation and Context Dependency may add explanatory power for software co-change and integration beyond Human Communication.The proposed relation is explicitly testable rather than established by the framework.

8 A Reference Tool Architecture for Agent Software Engineering

The reference architecture projects the framework’s theoretical objects onto existing software-engineering infrastructure without prescribing a fixed product decomposition.

  • Three Layers and Their Theoretical Roles: Figure 10 shows where theoretical objects can be represented, not a mandatory product decomposition.The architecture is a reference mapping rather than a component-by-component design.
  • Three Layers and Their Theoretical Roles: The architecture has HAC, Team, and Organization layers with distinct execution, coordination, governance, and cost-control roles.Concrete platforms may merge or split these capabilities; the framework does not require one service per theoretical object.
  • Three Layers and Their Theoretical Roles: The HAC layer supports local execution, permission boundaries, and candidate evidence, while the Team layer supports shared facts, invalidation, acceptance, and integration.The Organization layer supports identity, policy, compliance, model and tool governance, and cost control.
  • Industrial Primitives: Existing agent-development platforms already provide primitives such as instruction files, skills, hooks, tools, sandboxes, subagents, and specification workflows.A source-code study of eleven production coding-agent harnesses found recurring subsystems for execution, tools, context, safety, orchestration, and extensions.
  • Implementation Plausibility: The mapping establishes implementation plausibility, but not empirical validation of utility, cost, organizational impact, or developer experience.Prototypes can be built by extending existing engineering infrastructure.

9 Measurement, Engineering Economics, and Project Control

The measurement and control framework treats Accepted Trustworthy Change as the accounting boundary while modeling execution, verification, integration, attention, authority, and economic capacity as distinct constraints.

  • Measurement Principles: Measurement should derive observables from theoretical constructs rather than begin with convenient activity counters.Responsibility Topology can be operationalized through the number and independence of required acceptance anchors, then related to latency and waiting outcomes.
  • Accounting Boundary: Accepted-TC accounting distinguishes changes entering the formal baseline from work that remains Candidate, Rejected, Abandoned, or Eligible-but-not-Accepted.It is intended for within-project lifecycle analysis, not as a cross-project scalar of software value or universal productivity score.
  • Architecture and Verification Cost: Context staleness and revalidation load should be measured alongside artifact, context, and responsibility span because broader changes can increase integration, review, and acceptance costs.One proposed measure is ContextStalenessRate = #{active tasks affected by stale authoritative context}.
  • Verification Economics: Verification investment should scale with risk and impact because independent evidence sources are more expensive than local self-check.Candidate sources include external experts, different-model review, cross-cell verification, and formal verification.
  • Effort Estimation: Agentic effort is hypothesized to depend jointly on ambiguity, context, change locality, tool steps, risk, verification requirements, and Responsibility Topology.No particular functional form, monotonicity, or unit scale is assumed; the relation remains an empirical estimation problem.
  • Bounded Capacity: Agent execution concurrency remains bounded by downstream verification, integration, architecture, decomposability, and economic capacity.The capacity expressions are schematic rather than calibrated performance laws.
  • Responsibility Topology: A new responsibility anchor requires a legitimate, stable acceptance right that is not normally unilaterally overridable by the current anchor.Mechanical verification bottlenecks can instead be addressed by adding agents, reviewers, or delegated HACs while preserving one responsibility center.
  • Trustworthy Change Flow: Coding acceleration does not equal complete Trustworthy Change: acceptance and production observation remain part of the flow.The order-status example also requires semantic clarification, invariant verification, legacy-client testing, analytics changes, and anchor acceptance.

10 Research and Industrial Landscape as of September 2, 2026

The landscape positions the framework among agent-native engineering, governance, specification, evaluation, organizational, and project-control research while emphasizing that its central questions remain empirical.

  • Adjacent Research: Prior work already extends software engineering beyond code toward intent, evidence, authority, verification, responsibility, and human oversight.The paper’s novelty is not the general claim that Agent SE should address these concerns.
  • Specification and Harness Engineering: Specification and harness research supports treating structured specifications, context management, safety controls, orchestration, and extension surfaces as engineering objects.The paper presents its methods layer as synthetic rather than claiming these infrastructure primitives are new.
  • Evaluation and Security: Existing evaluation and security work motivates attention to process quality, execution security, claims, evidence, and evaluation validity beyond final test success.The cited landscape includes lucky passes, benchmark signal problems, confabulation, task-alignment controls, and tool-poisoning risks.
  • Organizations and Project Control: Research on human-agent organizations, responsibility boundaries, federated governance, and agentic project management provides adjacent coordinates for the framework.These lines address roles, commitments, decision rights, cross-organizational interaction, concurrency, task selection, and bounded capacity.
  • Evidence Boundaries: The implementation substrate for testing the framework already exists, but product capabilities and conceptual proximity do not establish causal effects or empirical validity.The landscape is used to locate constructs and open questions rather than convert proximity into evidence.
  • Open Empirical Questions: Open questions include whether Responsibility Topology adds explanatory power beyond team size, whether Minimum Sufficient Specification is identifiable, and whether Context Invalidation reduces stale-context failures.The paper also leaves open whether Accepted-TC-centered accounting improves measurement of agentic software work.

11 Empirical Commitments and a Falsifiable Research Agenda

The paper operationalizes its constructs as measurable variables, directional hypotheses, and falsification conditions rather than treating them as established laws. It proposes controlled studies of specification, concurrency, effort, verification, context coherence, topology, communication, attention, and takeover.

  • Empirical commitments: The framework maps theoretical constructs to observables, predicted relationships, and results that could weaken, modify, or reject each claim.Empirical testing is left to controlled, field, and longitudinal studies using goal–question–metric logic.
  • Specification: Progressive Specification is tested by comparing specification levels from natural-language requests through executable contracts or formal properties.The proposed levels add examples, counterexamples, non-goals, state representations, and executable guarantees.
  • Concurrency and economics: Concurrency experiments vary parallel-agent counts while measuring candidate production, accepted Trustworthy Change, verification queues, human attention, rework, and token cost.A rise followed by saturation or decline would identify an effective concurrency region conditioned on task decomposability.
  • Concurrency and economics: Agentic effort studies test whether ambiguity, context, locality, tools, risk, verification, and topology jointly explain token use, attention, lead time, and rework.No fixed functional form is assumed, and ambiguity proxies require calibration across model and harness conditions.
  • Verification and context: Verification experiments compare generator, model-based, deterministic, cross-cell, and expert reviewers using defect detection, omissions, and escaped defects.Evidence independence is separated into model, context, criterion, and responsibility dimensions to test whether diversity reduces common-mode failure.
  • Topology and human oversight: Topology and organization studies measure decision latency, accepted-change throughput, defects, responsibility ambiguity, recovery, context invalidation, and human cognitive fragmentation.They compare single-anchor and multi-anchor systems while controlling relevant domain, size, team, and risk factors.

12 Discussion and Conclusion: What Software Engineering Governs in the Agent Era

Agentic execution expands software-production capacity, but acceptance, responsibility, and organizational continuity remain governing concerns. The paper therefore frames agent-era engineering around Trustworthy Change, Human–Agent Cells, and Responsibility Topology while presenting its distinctions as empirically provisional.

  • Execution and organization: Agents can substantially increase the software execution that one responsibility center mobilizes, potentially enabling small human cores to develop and operate more products.This expansion concerns execution capacity, not independent acceptance authority.
  • The changing value of teams: A team’s marginal value increasingly lies in independent knowledge, context, evidence, and acceptance rather than additional execution labor.Shared models, contexts, and assumptions mean that more HACs do not automatically increase assurance.
  • The changing value of teams: Professional difference, independent rejection rights, and responsibility continuity remain design goals for team software engineering under high execution capacity.Multi-person structures remain necessary where independent professional responsibility, separation of duties, regulation, or continuous operations require them.
  • Risk-adaptive autonomy: Autonomy should be constrained by impact, reversibility, detectability, evidence, permissions, recovery, and responsibility rather than measured by autonomy alone.Requiring human approval for every important action can also create verification bottlenecks and rubber-stamping.
  • Scope and governance: Responsibility Topology classifies authority independently of human or agent headcount, while multi-anchor governance requires explicit closure across independent responsibility domains.The paper covers only single-center and multi-independent-center forms and does not validate more complex topologies.
  • Conclusion: The paper defines agent-era software engineering as organizing humans and scalable digital execution so software changes are formed, verified, integrated, accepted, operated, and sustained.Its central shift is from code production toward trustworthy software change and inspectable residual-risk acceptance.
Loading 2609.04630v1…