Source-linked AI summary
Seven Sources of Physical AI Capability Formation
Gang Chen
TL;DR
The paper addresses the gap between observing what Physical AI systems can do and identifying how those capabilities were formed. It develops and tests a seven-source framework through reconstructive induction and challenge sampling, finding scope-qualified theoretical saturation without claiming logical completeness.
Problem
Existing ways of classifying Physical AI do not directly identify the formative sources that give rise to particular capabilities.
Method
The paper uses reconstructive induction, primary-study tracing, coding rules, and three rounds of maximum-difference and negative-case sampling to classify capability-formation sources.
Results
Under the specified scope and cutoff, 49 evidence records were explainable by seven non-exclusive sources, with no irreducible eighth source and no substantive new boundary rule in R3.
Takeaways & Limitations
The framework separates similarity in observed capability from similarity in formation, supporting analysis of dependencies, transfer, replication, governance evidence, and geoeconomic foundations.
Takeaways & Limitations
The evidence base is constrained by language, indexing, access, search cutoff, predominantly English-language literature, lack of independent dual coders, and representative rather than prevalence-oriented records.
Abstract
from arXiv · showhide
Capabilities relevant to Physical AI can arise from materially different formation histories, yet existing taxonomies organized by morphology, architecture, learning algorithm, task, or domain do not directly answer what gives rise to a capability. We define a capability-formation source as a factor materially contributing to capability formation, distinct from components or construction steps. We identify seven non-exclusive sources: Recorded-Experience (RE), Predictive-Modeling (PM), Evaluative-Interaction (EI), Surrogate-Environment (SE), Mechanism-Grounded (MG), Embodied-Coupling (EC), and Evolution-Driven (ED) Formation. Using reconstructive induction with theoretical saturation, we traced a research matrix to primary studies, deduplicated the literature, set coding rules, and conducted three rounds of maximum-difference and negative-case sampling. Challenges included curriculum and self-supervised learning, active inference, open-ended and developmental learning, planning and search, neuro-symbolic architectures, digital twins, generative physical world models, and morphology-control co-design. Within the scope and criteria fixed as of September 4, 2026, all 49 evidence records were explainable by the seven sources individually or in combination. No R1-R3 challenge produced an irreducible eighth source, and R3 required no new core definition or substantive boundary rule. We therefore claim theoretical saturation within the stated scope, not logical completeness or exhaustive future coverage. The framework distinguishes similarity in observed capability from similarity in how it was formed, supporting analysis of explanation, transfer, replication, dependencies, governance evidence, and geoeconomic foundations.
1. Introduction
The paper asks where capabilities relevant to Physical AI come from, treating capability formation as distinct from observed behavior, system components, or construction pipelines. It proposes seven non-exclusive sources and investigates whether challenge sampling reveals any irreducible additional source.
- Capability formation is analyzed relationally: factors contribute to acquiring, shaping, improving, transferring, or extending specific Physical AI capabilities.
- Observed similarity does not imply similar formative bases, which may change dependencies and conditions for transfer, replication, review, substitution, and reproduction.
- The framework matters scientifically, operationally, industrially, and geopolitically because it broadens analysis beyond model weights to formative assets and dependencies.
- The study asks which identifiable sources materially form Physical AI capabilities and whether progressively added boundary cases reveal irreducible sources beyond the existing categories.
- The paper does not redefine Physical AI, treat the sources as necessary components, or prescribe a mandatory construction pipeline.
- It proposes seven non-exclusive sources, establishes boundary pairs, and reports theoretical saturation within the specified literature scope.
2. Object of Analysis and Boundaries
The unit of analysis is a target capability together with the formative practice that materially contributes to it at a particular stage. Candidate sources must have an independently identifiable formative role, while labels for systems, algorithms, schedules, or deployment patterns do not automatically qualify.
- 2.1 Object of analysis: The unit of analysis is a specific target capability paired with an identifiable formative practice at a particular formation stage.
- 2.1 Object of analysis: A system may draw on different sources across formative stages, and multiple coding denotes combined formative sources rather than overlapping categories or seven system modules.
- 2.3 Exclusion rule: A candidate source must materially alter capability acquisition, shaping, improvement, transfer, or extension when removed or replaced.
- 2.3 Exclusion rule: A candidate factor also requires an independently identifiable material role at the adopted analytical level.
- 2.3 Exclusion rule: New algorithm names, architectures, schedules, deployment patterns, or organizational mechanisms are not new sources unless their formative contribution cannot be explained by existing sources.
3. Method
The study uses reconstructive induction: it traces a prior research matrix to primary literature, deduplicates and codes evidence, and applies fixed stopping rules through challenge sampling. Terminology calibration preserves seven source categories while defining boundaries for coding.
- 3. Method: The method is reconstructive induction rather than a retrospectively recast de novo grounded-theory study.
- 3. Method: The evidence base prioritizes primary studies, peer-reviewed versions, and authoritative reviews, with frontier preprints used only for boundary stress testing.
- 3. Method: Included records must identify a factor that formed or changed a capability, with support from the paper body, methods, or experimental setup.
- 3. Method: The coding excludes unsupported claims about intelligence or embodiment, deployment-only planning and filtering, hardware inventories, irrelevant digital tasks, marketing material, and duplicates.
- 3.4 Terminology calibration and coding rules: The seven labels are Recorded-Experience, Predictive-Modeling, Evaluative-Interaction, Surrogate-Environment, Mechanism-Grounded, Embodied-Coupling, and Evolution-Driven Formation.
- 3. Method: The stopping rule required three consecutive maximum-difference and negative-case rounds with no irreducible source, no substantive late definition revision, and explanation of uncoded records.
4. Seven Sources of Capability Formation
The paper defines seven non-exclusive sources that materially contribute to Physical AI capability formation, distinguishing formative histories from components, algorithms, or construction steps. Each source is illustrated through criteria, boundaries, and representative physical-system evidence.
- 4. Seven Sources of Capability Formation: The seven sources are analytical categories that may combine within one formation history but are not system modules or mandatory construction steps.Figure 1 relates sources analytically to a target capability; implementation maturity and system integration are separate questions.
- 4. Seven Sources of Capability Formation: The framework distinguishes source-specific evidence from complete-system evidence and treats source maturity, physical realization, and system integration as separate questions.Representative lineages span demonstrations, world models, interaction loops, and simulators, while individual systems may rely on several sources.
- 4.1 Recorded-Experience Formation (RE): Recorded-Experience Formation uses fixed demonstrations, trajectories, corpora, or other pre-existing experience to form capabilities.RT-2 provides real-robot evidence using recorded robot trajectories and internet vision-language data, while recorded experience need not be the sole source.
- 4.2 Predictive-Modeling Formation (PM): Predictive-Modeling Formation uses learned predictive structures, such as world or dynamics models, to participate materially in training, policy improvement, or planning.DayDreamer demonstrates participation in physical capability formation across four real-robot tasks; the predictive model remains one source within the complete system.
- 4.3 Evaluative-Interaction Formation (EI): Evaluative-Interaction Formation updates capabilities through action or candidate behavior, consequence, evaluation, and capability-update cycles.QT-Opt formed a closed-loop visual grasping policy from more than 580,000 real robot grasp attempts.
- 4.4 Surrogate-Environment Formation (SE): Surrogate-Environment Formation trains capabilities in an action-responsive simulated, modeled, or generative world that supplies formative interaction.The defining distinction is an interactive experience-generating substrate rather than static synthetic data or an internal predictor alone.
4.5 Mechanism-Grounded Formation (MG)
Mechanism-Grounded Formation enters capability formation through explicit physical laws, structures, constraints, or mechanisms, whereas Embodied-Coupling Formation depends on body-environment dynamics shaping what and how capabilities can be formed. Evolution-Driven Formation operates through variation, selection, and retention across generations or evolutionary cycles.
- 4.5 Mechanism-Grounded Formation (MG): Mechanism-Grounded Formation uses explicit physical mechanisms as priors, structures, losses, feasible sets, policy constraints, or direct computational definitions.The mechanism must materially alter the learnable space, objective, policy formation, or capability boundary.
- 4.5 Mechanism-Grounded Formation (MG): PINNs, Hamiltonian neural networks, and training-time safe reinforcement learning exemplify mechanisms entering capability formation rather than merely appearing in explanations or runtime filters.Physical experiments support MG in soft-robot motion prediction and Franka Emika Panda trajectory tracking, primarily for local capabilities and control loops.
- 4.6 Embodied-Coupling Formation (EC): Embodied-Coupling Formation assigns a material formative role to morphology, materials, sensorimotor loops, contact dynamics, or body-environment coupling.Having a robot body, URDF, or physical deployment alone is insufficient; the body must shape what, how, or at what cost capability is formed.
- 4.6 Embodied-Coupling Formation (EC): Passive dynamic walking and soft-body morphological computation establish bodily dynamics as behavioral or computational resources, but do not alone establish capability formation.Morphology-control co-design provides more direct evidence when bodily changes alter the path or cost of policy learning.
- 4.7 Evolution-Driven Formation (ED): Evolution-Driven Formation forms capabilities through variation, selection, and retention or inheritance across generations, populations, or evolutionary cycles.Within-lifetime learning, ordinary reinforcement learning, self-improvement, and iterative optimization without retention or selection are excluded.
- 4.7 Evolution-Driven Formation (ED): Real-world evolutionary experiments jointly searched morphology and control in mechanically reconfigurable quadruped robots using physical evaluation.Current representative evidence is concentrated in experimental evolutionary-robotics systems, many of which otherwise rely on physics simulators.
4.8 Definition-Boundary-Lineage Matrix
The matrix organizes each source by definition, positive criterion, negative boundary, and lineage anchors. Its distinctions separate recorded experience, predictive models, evaluative loops, surrogate worlds, mechanisms, embodiment, and evolutionary selection.
- Predictive-Modeling Formation (PM): Predictive-Modeling Formation requires a learnable predictive structure to materially participate in capability formation, unlike static classification or deployment-only model invocation.World models and latent dynamics are representative anchors.
- Recorded-Experience Formation (RE): Recorded-Experience Formation uses pre-existing recorded experience, while generic data-driven claims are insufficient and online feedback requires an Evaluative-Interaction check.Its lineage anchors include learning from demonstration, ALOHA, and RT-2.
- Evaluative-Interaction Formation (EI): Evaluative-Interaction Formation requires consequences or an evaluative loop to change capability; interaction without evaluative updating is excluded.Reward, preference, error, and success signals are listed as representative evaluative signals.
- Surrogate-Environment Formation (SE): Surrogate-Environment Formation uses an action-responsive surrogate world for formative interaction, distinguishing simulators and interactive twins from static synthetic samples or internal prediction alone.The boundary is the environment’s role as an interaction-generating substrate.
- Mechanism-Grounded Formation (MG): Mechanism-Grounded Formation requires explicit mechanisms to enter capability formation, with PINNs, Hamiltonians, and barrier functions as lineage anchors.Generic regularization and runtime-only constraints do not qualify.
- Embodied-Coupling Formation (EC): Embodied-Coupling Formation requires body, material, or coupling dynamics to shape formation; merely having a body, URDF, or deployment is insufficient.Morphology-control co-design and passive dynamics anchor the category.
- Evolution-Driven Formation (ED): Evolution-Driven Formation requires cross-generational or population-level variation, selection, and retention, excluding ordinary reinforcement learning, self-improvement, and iterative optimization.Evolutionary cycles, archives, evolutionary robotics, quality-diversity methods, and POET provide lineage anchors.
5. Results: Distinction, Combination, and Theoretical Saturation
The framework distinguishes adjacent capability-formation sources by fixed criteria, while allowing multiple sources to contribute at different stages without collapsing their analytical distinctions. Across R0–R3 reconstruction and challenge sampling, no irreducible eighth source emerged within the stated scope, and R3 confirmed stability.
- Critical adjacent boundaries: RE assigns formation to experience fixed before updating, whereas EI requires action consequences or evaluation to immediately change subsequent capability.Both sources may occur at different stages of one project.
- Critical adjacent boundaries: PM denotes predictive structure, whereas SE denotes an action-responsive surrogate world that hosts training.This prevents all predictive rollouts from being labeled SE.
- Critical adjacent boundaries: EI concerns within-lifetime evaluative updating, whereas ED concerns cross-generational or population-level variation, selection, and retention.Evolutionary fitness used only for selection is not double-coded.
- Combination without collapse: Systems can combine sources without erasing distinctions: PlaNet and Dreamer combine PM, SE, and EI, while Shadow Hand, ANYmal, and POET combine other source sets.The framework treats each contribution as independently meaningful for transfer, replication, and substitution requirements.
- Candidate concepts and boundary tests: Curriculum learning reorders RE, EI, and SE, while developmental learning remains within-lifetime RE, EI, or EC unless cross-generational selection is present.Other candidate concepts are reduced according to the formative process they contribute, rather than their technology label.
- Negative-case tests for an eighth source: Social or collective formation and open-ended task generation remain candidates for scrutiny but were reducible to existing sources and did not warrant an eighth source.The former was reduced to RE/EI/EC/ED and the latter to PM/SE/EI/ED.
- R0–R3 saturation trajectory: 49 cumulative evidence records produced seven irreducible sources, with no new source in R1–R3; R2 fixed substantive boundaries and R3 confirmed stability.The reconstruction included 28 core records at R0 and seven additional maximum-difference or negative-case records in each later round.
- R0–R3 saturation trajectory: The reported saturation is scope-qualified: it concerns additional irreducible source categories under the specified literature, coding rules, and September 4, 2026 cutoff.It does not claim exhaustion of subtypes, combinations, effect sizes, industry distributions, or future technologies.
6. Research Significance and Application Value
The framework shifts analysis from technology labels or final demonstrations to the formative dependencies behind capabilities. It supports traceable questions about assets, transfer, replication, validation, governance, and geoeconomic foundations.
- Research significance: The framework asks which formative sources a capability depends on and whether those sources can be accessed, transferred, reproduced, or substituted.It is not a score of performance, maturity, valuation, or national competitiveness.
- Investment and industry: For investors, the framework turns a vague capability moat into questions about irreplaceable experience, predictive models, simulation infrastructure, mechanism knowledge, physical platforms, and evolutionary experiments.These sources are treated as potentially critical formative assets.
- Investment and industry: For AI companies, RE dependence emphasizes experience coverage and data rights, PM or SE dependence emphasizes world-model quality and surrogate environments, and MG dependence emphasizes mechanism knowledge.EC and ED dependencies can make physical embodiments, evaluation cycles, population scale, retention, and search infrastructure part of capability formation.
- Deployment and governance: Industrial evaluation should examine formation conditions: proprietary experience can create data-transfer costs, SE requires validation against real conditions, and MG requires mechanism audits.The final demonstration alone does not reveal these dependencies.
- Deployment and governance: Regulatory assessment can separate formative evidence from runtime safety controls, since runtime mechanism constraints do not establish MG formation and SE training may leave uncovered deployment conditions.The framework therefore supports distinct examination of formation, deployment constraints, and reproduction evidence.
- Geoeconomic foundations: Geoeconomic analysis uses the framework to identify foundations such as data rights, compute, simulation infrastructure, mechanism knowledge, embodiments, supply chains, and evolutionary-search capacity.It is better suited to identifying capability foundations than ranking countries.
7. Limitations, Falsification, and Reopening Procedure
The study limits its claims through search, coding, evidence-quality, and interpretive boundaries, while defining a falsifiable procedure for reopening the framework if genuinely irreducible sources appear.
- Limitations: The evidence base is not an enumerable internet census and is constrained by language, indexing, access, search cutoff, and predominantly English-language technical literature.The study also used some 2025–2026 preprints only for boundary stress testing.
- Limitations: The study did not use independent dual coders and therefore reports no Cohen’s κ or other inter-coder reliability coefficient.Future validation should independently recode a substantial subset of records.
- Limitations: Evidence-record counts cannot indicate category importance, prevalence, or causal effect size because records are organized around representative methods.This limits quantitative interpretation of the framework’s evidence distribution.
- Falsification criterion: R4 should reopen the framework only if a formative contribution remains irreducible to the seven frozen sources and boundary rules, individually or in combination.A new name, algorithm, architecture, implementation, or technology school alone is insufficient.
- Reopening procedure: Reopening should freeze the current version and cutoff, independently code new records without prior judgments, and add a category only for a genuinely new irreducible source.Insufficient evidence or an existing boundary issue should update the audit record without adding a category.
8. Conclusion
The paper reframes Physical AI capability analysis around where capabilities come from rather than what systems are made of. It retains seven non-exclusive sources and reports scope-qualified saturation while preserving the framework’s falsifiability.
- Conclusion: The paper retains seven non-exclusive sources—RE, PM, EI, SE, MG, EC, and ED—using reconstructive induction, 49 evidence records, 45 public sources, and three challenge rounds.The sources are Recorded-Experience, Predictive-Modeling, Evaluative-Interaction, Surrogate-Environment, Mechanism-Grounded, Embodied-Coupling, and Evolution-Driven Formation.
- Conclusion: The framework distinguishes fixed experience from online evaluation, predictive structure from surrogate worlds, within-lifetime adaptation from cross-generational evolution, and embodiment from body-mediated formation.A system may combine sources, but combination does not erase these distinctions.
- Conclusion: Under the specified scope and September 4, 2026 cutoff, three challenge rounds produced no irreducible eighth source, and R3 added neither a core definition nor a substantive boundary rule.The result is scope-qualified theoretical saturation, not logical completeness or exhaustion of future manifestations.
- Conclusion: The resulting framework is auditable, comparable, and falsifiable, and can be overturned or extended by new evidence.Its claim concerns the emergence of additional irreducible capability-formation categories, not every internal subtype, combination, or effect size.
Supplementary Audit File
The accompanying auditable saturation-induction research archive preserves the complete evidence set and audit materials, while the main paper reports only the definitions, boundaries, representative lineages, and saturation results needed for its research question.
- The archive preserves 49 evidence records, 45 public sources, the R0-R3 saturation log, exclusion rules, candidate counterexamples, and prospective R4 triggers.
- The main paper reports only the definitions, boundaries, representative lineages, and saturation results necessary to support the research question.