Source-linked AI summary
Model-Based Agentic Software Engineering
James C. Davis, Kelechi Kalu, Huiyun Peng, Parth V. Patil
TL;DR
Coding agents expand implementation capacity without making intent, structure, or acceptance evidence explicit, leaving consequential properties to repeated reconstruction. MAGE proposes a governed engineering environment that externalizes purposeful representations and gives obligations operational authority, grounding the theory in one longitudinal case and six industrial accounts. The evidence supports recurring structures and reports a reduction in unmodeled implementation elements from 56% to 7.89%, while not supporting population estimates or causal claims about effectiveness.
Problem
Coding agents increase implementation capacity without automatically making consequential engineering properties explicit, leaving agents and humans to reconstruct them from dispersed project artifacts.
Method
MAGE combines purposeful modeling with alignment mechanisms—constraints, sensors, validators, and gates—and develops the theory from DocAble plus comparative reconstructions of six industrial systems.
Results
Unmodeled implementation elements decreased from 56% to 7.89% across a nine-stage modeling sequence, while industrial accounts showed recurring externalization, bounded action, separated evaluation, and retained human authority.
Takeaways & Limitations
MAGE frames trustworthy agentic engineering as a governed environment in which explicit knowledge, bounded action, independent evaluation, and human authority compose around autonomous implementation.
Takeaways & Limitations
The industrial accounts support recurrence and variation within stated scope conditions, but not population estimates or causal claims about MAGE’s effectiveness.
Abstract
from arXiv · showhide
Coding agents increase implementation capacity without automatically making project intent, system structure, or acceptance evidence explicit. As implementation becomes abundant relative to engineering judgment, the scarce work shifts toward choosing useful abstractions, producing evidence, and determining which obligations govern acceptance. Existing workflows address parts of this gap through larger prompts, repository retrieval, or perchange review, but still require agents and engineers to reconstruct consequential properties. As an alternative, we present Model-Based Agentic Software Engineering (MAGE). MAGE is a framework and a theory for building trustworthy autonomy from commodity intelligence. MAGE addresses a representation problem and an authority problem: it externalizes the smallest purposeful representation needed to answer an engineering question, then gives settled obligations proportionate authority through constraints, sensors, validators, and gates. It keeps uncertain intent open and turns recurring reconstruction and judgment into durable engineering structure that later work can inherit. We developed MAGE from a longitudinal case and refined it through six independently reported industrial accounts. Across these sources, MAGE explains how externalized knowledge, bounded action, independent evaluation, and retained human authority can compose into a governed engineering environment, and proposes tests of when that environment turns commodity intelligence into durable engineering progress.
1 Introduction
Coding agents increase implementation capacity, but trustworthy development still depends on making consequential engineering knowledge explicit and giving obligations enforceable authority. MAGE addresses this gap through modeling, alignment, and evidence from a longitudinal case plus six industrial accounts.
- Coding agents can plan, implement, test, and optimize software, shifting bottlenecks toward human specification, understanding, and validation.
- Existing approaches improve information access through retrieval, memory, tools, and harnesses, but agents and humans may still reconstruct system properties from dispersed artifacts.
- MAGE makes consequential knowledge and intent explicit through modeling and gives engineering obligations authority through constraints, sensors, validators, and gates.
- Failures can be converted into durable engineering structure that reduces repeated reconstruction and human judgment required for later governance.
- The theory was developed from the longitudinal DocAble case and refined through comparative reconstructions of six independently reported industrial systems.
- MAGE contributes a theory of governed engineering environments, empirical grounding, and a falsifiable agenda for testing when agentic capacity becomes durable engineering progress.
2 Background
MAGE combines engineering traditions for explicit, enforceable system knowledge with AI mechanisms for retaining state, reasoning structurally, and acting through tools. Its novelty lies in composing these capabilities within a governed engineering environment.
- Governance: Governance determines what work may proceed, who may authorize it, what evidence changes must produce, and which conditions admit them.
- Engineering Traditions: Software and systems engineering externalize structure, behavior, interfaces, specifications, boundaries, hazards, and assurance relationships into stable objects for reasoning and verification.
- AI Traditions: AI contributes knowledge representation, memory, planning, structured reasoning, and tools that connect intermediate state to repositories, tests, permissions, and feedback.
- Governed Engineering Environment: The Governed Engineering Environment surrounds autonomous implementation with explicit knowledge and enforceable obligations, making composition—not any single model or control—the claimed novelty.
3 Motivation
The paper argues that evaluating coding-agent effects requires studying the governed engineering environment through which capabilities are enacted, not adoption alone. It motivates exploratory theory development to identify that environment’s constructs and mechanisms.
- Agent capability alone cannot establish assurance at velocity because implementation capacity and governing engineering structures are distinct.
- Adoption indicators can capture when a tool enters a project without specifying the software-engineering method or environment through which it operates.
- The missing empirical object is the governed engineering environment of representations, evidence, constraints, procedures, and authority surrounding agentic work.
- Characterizing this environment enables distinction between access to a general-purpose capability and a reproducible software-engineering intervention.
4 Methodology
The study develops MAGE through an exploratory, interpretive design combining a longitudinal DocAble case with comparative reconstruction of six independent industrial accounts. It uses these sources to identify recurring mechanisms, alternative realizations, scope conditions, and quantitative evidence of engineering hardening.
- Research design: MAGE was developed through a longitudinal DocAble case followed by comparative reconstructions of six independently reported industrial systems.The case identified mechanisms and temporal ordering; the comparison examined whether related structures recurred under different organizational and technical pressures.
- Research design: The exploratory, interpretive design used DocAble as the empirical starting point and industrial accounts to examine variation across organizational and technical conditions.The analysis treated DocAble as an empirical seed rather than a complete derivation of the theory.
- DocAble case: DocAble was developed over approximately 20 weeks with coding agents as the primary implementation workforce, reaching approximately 540,000 lines of production code and 1.6 million lines of supporting infrastructure.The project used 3–4 Claude Max 20x accounts during most weeks, and the present study combined the earlier case analysis with subsequent repository evolution.
- Comparative analysis: The comparative sample comprised first-party accounts from Cloudflare, Spotify, Shopify, Docker, Siemens, and Zenseact, analyzed with a common frame linking pressures, representations, action boundaries, evaluation, authority, inheritance, and scope.Table 1 distinguishes observed source claims from the MAGE interpretation; the Siemens source describes prototypes rather than production deployment.
- Comparative analysis: Across the industrial accounts, recurring structures included externalized knowledge, bounded action, separated generation and evaluation, and retained human authority when decisions were not adequately mechanized.The comparison supports recurrence and variation, not causal effectiveness, because several accounts lack model correspondence, longitudinal adaptation, or outcome measures.
- Empirical grounding: DocAble’s support apparatus grew from 0.85× production source after the prototype to approximately 3× in mature snapshots, peaking at 3.68× during hardening.Project-specific lint files grew from 0 to 747 and gate scripts from 0 to 102; derived checks caught six instances of a prior model–code drift class, with no observed recurrence across 56 subsequent feature implementations, while unmodeled elements fell from 56% to 7.89%.
5 MAGE Theory
MAGE explains durable agentic engineering through a governed environment that makes consequential properties answerable and selected obligations authoritative. Modeling reduces reconstruction, Alignment supplies operational force, and governance adaptation converts recurring judgment into reusable engineering capital.
- Core mechanisms: MAGE distinguishes Modeling, which externalizes consequential knowledge, from Alignment, which gives selected obligations authority through controls.Together they determine which properties are representable and operationally consequential.
- Core mechanism: The governed engineering environment mediates whether agentic capacity becomes durable throughput or churn.It includes representations, evidence, constraints, procedures, and authority through which work is understood, constrained, and admitted.
- Semantic gap: The semantic gap separates where obligations arise from where the environment has enough meaning, evidence, and authority to determine satisfaction.Modeling reduces this gap by exposing missing properties, while Alignment gives obligations operational force.
- Engineering capital: Engineering capital forms when recurring judgment becomes reusable structure whose future value exceeds construction and maintenance costs.Its return diminishes when inherited structure creates conflicts or costs more to maintain than it saves.
- Propositions: Representation leverage grows with reasoning burden, but its benefit fails when maintenance and interpretation costs exceed reconstruction savings.The representation must remain task-relevant and faithful.
- Propositions: Authority expands the consequential surface, while Modeling and Alignment contribute complementary rather than interchangeable effects.The semantic gap determines how much authority must be deferred or reconstructed instead of enforced in place.
6 Application of MAGE
MAGE applies by starting from an engineering question and externalizing only the model needed to answer it, while assigning proportionate authority to settled obligations. The same cycle adapts to project state, implementation stock, intent, scale, and assurance needs.
- Engineering questions → governed work: MAGE begins with an engineering question and creates the smallest model needed to make that question answerable.Authority is assigned proportionately to settled, consequential obligations.
- Starting matrix: The starting matrix is diagnostic rather than sequential: existing implementation and settled intent determine the initial representation and authority available.Projects and subsystems can move between cells as implementation accumulates or intent stabilizes.
- MAGE cycle: The recurring cycle models the question, aligns obligations with evidence and authority, performs work, and converts worthwhile recurring judgment into reusable structure.The cycle does not require a comprehensive system model; views can be connected selectively around the engineering question.
- Greenfield projects: Greenfield work uses provisional models during exploration, then models settled contracts and consequential boundaries before realization.Written hypotheses do not create authority until meaning is settled and adequate evaluation evidence exists.
- Brownfield projects: Brownfield work first separates actual system behavior from intended behavior, then reconciles models with the realized system before enforcement.Under uncertain intent, recovery is bounded to the structure needed for the immediate question.
- Why not just prompt?: Prompts, context, and skills guide work but do not establish correspondence, independent acceptance evidence, or binding obligations by themselves.Constraints, validators, gates, or accountable admission decisions provide operational force.
7 Research Agenda
The research agenda turns MAGE’s claims into falsifiable studies of representation leverage, integrity, and determinization. Preliminary DocAble evidence reports lower agent resource use when agents receive explicit model-use instructions.
- Hypotheses: H1 predicts that task-relevant models reduce reconstruction cost or increase task scope at matched quality, with larger benefits under greater reasoning burden.The proposed environment-by-agent experiment compares reasoning engines across environments with different representation support.
- Hypotheses: H2 predicts that stale, irrelevant, or poorly chosen models reduce or reverse representation benefits and predict downstream churn or defect escape.The proposed intervention compares synchronized, stale, irrelevant, and absent representations.
- Hypotheses: H3 predicts that explicit models and corresponding predicates make selected system-level obligations repeatedly evaluable without semantic reconstruction.The proposed study compares evaluation with and without the model.
- Preliminary evidence: Agents receiving explicit model-use instructions required fewer tokens, turns, and elapsed time than agents receiving no such instructions.The two-week in situ DocAble experiment randomly assigned agents, and the effect was more pronounced for Sonnet.
- Evaluation: The agenda evaluates reconstruction cost, task quality, task scope, correspondence, obligation coverage, churn, defect escape, and model maintenance cost.Each hypothesis is explicitly weakened by results that fail to show its predicted representation effects.
8 Discussion
MAGE argues that abundant implementation capacity makes the surrounding engineered environment increasingly consequential and experimentally separable from agent capability. Its broader pattern includes heterogeneous representations and enforcement mechanisms, with formal methods as one instantiation.
- Discussion: As implementation capacity becomes abundant, the engineered environment becomes increasingly consequential to realized performance.Commodity agents make implementation capacity more separable from that environment than human engineering historically allowed.
- Discussion: Weak representations, missing constraints, inadequate evidence, and costly human checkpoints increasingly register as bottlenecks when implementation becomes cheap.Environments that externalize knowledge and give obligations effective authority can convert more capacity into durable progress.
- Positioning: MAGE broadens the assurance pattern beyond formal methods by accommodating heterogeneous representations and enforcement mechanisms.Interactive proof assistants instantiate Modeling–Alignment at the formal end of the spectrum.
9 Conclusion
MAGE presents a theory of trustworthy agentic software engineering centered on a governed engineering environment. It combines explicit, task-relevant representations with mechanisms that give selected obligations authority, while framing its claims as hypotheses for systematic study.
- MAGE distinguishes Modeling, which makes consequential knowledge and intent explicit, from Alignment, which gives selected obligations authority through constraints, sensors, validators, and gates.
- The theory was developed from the longitudinal DocAble case and refined through comparative reconstructions of six independently reported industrial systems.
- MAGE hypothesizes that engineered representations can reduce agents’ reconstruction requirements and support more durable engineering progress.
- The framework hypothesizes that alignment mechanisms can translate selected engineering obligations into enforceable action boundaries.
- Testing MAGE requires systematic study of how representations, controls, and human decision boundaries shape agent behavior across tasks, projects, and environments.