Source-linked AI summary

Training, learning and inference: unified dynamics of neural systems

Mian Wang

arXiv:2608.20965v1cs.LG

TL;DR

The paper asks how complete generation facts and relations can be represented and analyzed. It develops a unified account of neural-system dynamics, reports irreducible generation-fact coordinates, and confirms the framework beyond nanoGPT in ResNet/CIFAR and diffusion/CIFAR experiments.

  • Problem

    The paper addresses the challenge of stating complete facts about concrete generation, including how results arose and the generation relation involved.

  • Method

    The paper represents generation histories through atomic facts and database computations, then analyzes training, learning, and inference as successive relations within unified neural-system dynamics.

  • Results

    The paper proves that five coordinates are individually irreducible and confirms the unified relations in ResNet/CIFAR and diffusion/CIFAR experiments beyond nanoGPT.

  • Takeaways & Limitations

    Training, learning, and inference are presented as successive relations within a unified dynamics of neural systems rather than as separate processes.

  • Takeaways & Limitations

    The findings do not argue against reinforcement learning; sustained asymmetric feedback may require monitoring functional support because reinforcement can be double-edged.

Abstract

from arXiv · show

We define an atomic generation fact f=(u,tau,omega,z;rho), recording the origin, realized transformation, concrete occurrence, generated result and relation role. Compiled into a Generation-Fact Graph (GFG), these facts provide an AI-native, compilable scientific fact substrate preserving generation histories. We establish a GFG-based recursive scientific process in which analysis, intervention, replay and validation form facts for later cycles. Using nanoGPT, we establish unified training-learning dynamics. Training is the evolution of a parameter-optimizer system with state and memory: each actual training action enters the receiving state and produces a finite-amplitude nonlinear functional response conditioned by that state and target-specific update geometry. Learning is the persistent reorganization of distributed functional support by these responses; capability formation, maintenance, decline or recovery becomes observable when target-specific states are evaluated against their readout boundaries. Three primary coordinates - target-boundary state, target-specific update geometry and parameter-Adam receiving state - yield a second-order predictor operating before post-update outputs are read. On held-out runs, it achieved 91.43% accuracy and 91.49% macro-averaged recall across four transitions. We further establish inference as a frozen projection of training-learning dynamics. Component gating and rollback show causal recruitment and non-additive combination of query-conditioned support formed during training, deriving organizational conditions realized by Attention. Controlled feedback indicates possible double-edged reinforcement effects. ResNet/CIFAR-100 and diffusion/CIFAR-10 experiments confirm receiving-state-conditioned responses, persistent support reorganization and frozen inference projection beyond nanoGPT.

1 Wuhan Polytechnic University, Wuhan 430023, Hubei, China

The paper introduces complete generation facts and compiles them into Generation-Fact Graphs, then uses this substrate to study unified training, learning, and inference dynamics. Across nanoGPT and additional vision experiments, training responses depend on receiving state and update geometry, while learning reorganizes distributed support that is projected during inference.

  • Training is the evolution of a parameter–optimizer system with state and memory, and learning is the persistent reorganization of distributed functional support.
  • 91.43% target-boundary accuracy and 91.49% macro-averaged recall were achieved on held-out confirmation runs across four target transitions.
  • Inference is described as a frozen projection that query-conditionally recruits and non-additively combines distributed support formed during training.
  • ResNet/CIFAR and diffusion/CIFAR experiments confirmed receiving-state-conditioned responses, persistent support reorganization, and frozen support projection beyond nanoGPT.
  • Generation facts record sources, realized transformations, concrete occurrences, outcomes, and relation roles, preserving complete generation histories.
  • Interconnected GFGs connect training executions, capability evaluations, interventions, support probes, response measurements, and frozen inferences within one traceable history.

2. Model of generation facts

The model defines an executable capture-and-binding procedure for atomic generation facts and validates them before constructing an AI-native Generation-Fact Graph. The resulting graph preserves occurrences, multiplicities, typed relations, and cross-stage provenance.

  • The executable model combines capture protocols, execution records, a generation binder, validated snapshots, and a GFG construction layer.
  • An atomic generation fact binds source information, realized transformation, concrete occurrence, outcome, and relation role.
  • The capture protocol specifies runtime records and associates them with a declared scope before native computation produces ordinary results.
  • Repeated facts remain independently significant because complete generation state is represented as a multiset rather than collapsed by set semantics.
  • Validation checks bindings, evidence, generator authorization, successful generation, capture coverage, and fixed protocol, implementation, and execution identities.
  • The constructed GFG retains atomic facts while organizing concrete occurrences and typed relations into an AI-traversable structure validated against source records.

3. The GFG-based recursive scientific process

The GFG-based process recursively turns questions, analyses, interventions, replays, and validations into accumulating scientific facts. Experiments then establish training as stateful, nonlinear response and learning as persistent reorganization of distributed support.

  • The GFG-based recursive scientific process: Scientific questions query and recompile existing GFGs, while interventions and replays produce executions whose validated results enter subsequent graphs.
  • Stateful capability dynamics: States with similar loss, accuracy, training step, margins, or normalization can have different capability futures, and capability can decline and recover after formation.
  • Stateful capability dynamics: Pausing parameter or optimizer evolution delayed capability formation by 1800, 800, and 1100 steps, whereas changing gradient clipping did not reproduce those delays.
  • Receiving-state-conditioned response: The same update produced different functional responses in different receiving states, showing that update effects depend jointly on the update and receiving state.
  • Nonlinear response: Complete response curves showed saturation, acceleration, turnback, and sign reversal, while fixed local approximations failed to recover complete-update endpoints reliably.
  • Primary conditioning factors: Three primary conditioning factors were identified: target-boundary state, target-specific update geometry, and parameter–Adam receiving state.
  • Distributed functional support: Capabilities were jointly sustained by distributed components whose necessity, substitution, redundancy, synergy, and failure tolerance varied with target and training state.
  • Distributed functional support: Across 72 realized-update sections, support reallocation correlated 0.765 with absolute capability change, while mean reallocation was 0.2080 when capability changed versus 0.0340 otherwise.

5. Direct prediction of target-level learning outcomes

The paper predicts target-level learning outcomes before post-update responses are observed using three state and update coordinates. The predictor distinguishes four target-boundary transitions with high held-out accuracy and recall.

  • The predictor uses target-boundary state, target-specific update geometry, and the parameter-optimizer receiving state to forecast post-update outcomes.The receiving state includes the parameter state and the Adam state involved in forming the actual update.
  • It distinguishes targets that remain correct, become incorrect, remain incorrect, or recover after an update.
  • 91.43% target-boundary accuracy and 91.49% macro-averaged recall were achieved across the four transitions on the frozen confirmation split.The predictor correctly classified 4652 of 5088 cases; balanced accuracy was 92.17%.
  • Across 12 eligible training runs, the same formula predicted 14069 of 15264 target-update outcomes, achieving 92.17% accuracy.
  • The algorithm was derived from the three primary coordinates and its predictive result is presented as validation of the training-learning theory.

6. From learning formation to inference

The paper treats inference as a frozen projection of training-formed functional support. Query-conditioned recruitment, joint component combination, rollback effects, scale-dependent support, Attention, and feedback are examined as parts of this relation.

  • 6. From learning formation to inference: The unresolved question is how a capability formed during training participates in a new inference occurrence.
  • 6. From learning formation to inference: Frozen inference causally recruits and combines functional support formed during training.
  • 6.1 Inference is a frozen projection of training–learning dynamics: Different targets use components differently, with each query recruiting a target-specific distribution of support rather than one fixed allocation.Across 13 runs, component-gating patterns changed from pre-formation to formation, and 23 target groups produced 23 different patterns.
  • 6.1 Inference is a frozen projection of training–learning dynamics: All 138 pairwise gating comparisons departed from independent addition, indicating that inference depends on joint participation of multiple components.
  • 6.2 The scale effect follows from the expansion of projectable functional support: Training scale affects inference through learning formation: more usable functional support yields richer query-conditioned projections and stronger combinations.Replacing a formed component with its earlier version weakened inference while architecture and query remained unchanged.
  • 6.3 Why Attention succeeds: Attention supports this projection by forming different active support configurations for different queries and combining them without persistently modifying learned parameters.
  • 6.4 Reinforcement as a potentially double-edged feedback process: Concentrated positive feedback increased support for the reinforced capability but reduced margins of unreinforced capabilities, with recovery after feedback rebalancing.Sufficiently concentrated feedback eventually produced observable capability loss; recovery occurred across all twelve seeds.

7. Cross-system validation

Experiments in a ResNet classifier and a diffusion model tested whether the identified training-learning and inference relations extend beyond nanoGPT. Both systems recovered the principal state-conditioned, persistent, and frozen-projection dynamics.

  • The study repeated frozen training-learning and inference protocols in a ResNet classifier and a diffusion model using CIFAR data.The systems were tested with three formal seeds per system.
  • Receiving-state-conditioned functional responses and persistent support reorganization were recovered across the two structurally different systems.The training-learning validation included 504 independently checked diffusion-response records.
  • Inference validation recovered exact state preservation, query-conditioned support recruitment, non-additive combination, and dependence on training-formed component versions in all six formal seeds.
  • The cross-system results exclude a nanoGPT-specific explanation and support the cross-system scope of the proposed dynamics.

8. Conclusion

The paper presents training, learning, and inference as successive relations within unified neural-system dynamics, supported by a recursive GFG-based scientific process. Training reorganizes distributed functional support, while inference projects and combines that learned support without further reorganization.

  • Training evolves a parameter–optimizer system with state and memory, while each action produces a finite-amplitude nonlinear response conditioned by receiving state and target-specific update geometry.
  • Learning persistently reorganizes distributed functional support, making capability formation, maintenance, decline, and recovery observable at target-specific readout boundaries.
  • Inference is a frozen projection that query-conditionally recruits, projects, and non-additively combines support formed during learning while leaving the learned state unchanged.
  • ResNet/CIFAR and diffusion/CIFAR experiments confirmed receiving-state-conditioned responses, persistent support reorganization, and frozen inference projection beyond nanoGPT.
  • The GFG-based recursive process preserves generation histories in compilable fact graphs and turns analyses, interventions, replays, and validations into facts for continued recompilation.

Methods

The methods capture executions as validated generation facts, compile them into GFGs, and use matched states, interventions, response measurements, gating, and replay to test unified dynamics. A pre-output predictor and frozen-checkpoint experiments connect internal responses and support allocation to capability transitions and inference behavior.

  • Generation-fact capture: Runtime records were bound into atomic facts, validated within scope, and compiled into GFGs that supported subsequent queries, replays, interventions, and analyses.
  • Training histories: Thirteen independently executed nanoGPT histories preserved inputs, computations, gradients, Adam states, parameter updates, versions, and capability evaluations.
  • State matching and intervention: Matched states with similar summaries but divergent futures were retained as counterexamples, and interventions paused parameter–optimizer evolution or altered clipping thresholds.
  • Response measurement: Functional response was measured by comparing full-update and skip branches, testing whether an actual training action depended on receiving state as well as the update itself.
  • Capability transitions: A target-level ledger connected internal functional change and support reallocation to four observable transitions: remaining correct, correct-to-incorrect, remaining incorrect, and incorrect-to-correct.
  • Pre-output prediction: Predictions were made after the actual parameter update but before post-update logits, margins, or predictions were computed, using mechanically derived transition rules.
  • Inference evaluation: Frozen inference checkpoints tested causal support recruitment, non-additive component combination, dependence on training-formed versions, and learned-state preservation.

The foundational role of the GFG-based recursive scientific process

The GFG-based recursive scientific process is presented as a foundational methodological contribution, not merely an additional experimental stage. Successive experiments use generation facts, causal replay, and recompilation to study long-horizon temporal credit assignment.

  • The process was used to address long-horizon temporal credit assignment through four successive experiments.
  • Matched causal forks showed that correct consequence binding and credit re-formed a reversed policy in all twelve formal seeds, whereas disrupting either relation prevented the same formation.
  • Formation-path retrieval reduced a 64-action history to nine candidates while retaining all six functional actions and three formation ancestors without causal effect.
  • Training with discovered credit achieved 96.48% terminal success on held-out evaluation, matching the hidden oracle result reported for that experiment.
  • Recompiling the formation structure preserved exact credit within 10^-12 tolerance, reduced native replay transitions by 90.64%, and produced a 2.38-fold end-to-end speedup.
  • With stochastic inputs introduced across the formation chain, matched replay separated action-contingent credit from environmental variation while preserving zero credit for noncausal ancestors.

AI-assisted scientific execution

AI systems assisted analysis and experimental execution within a process that retained human authority over scientific questions, boundaries, hypotheses, and decisions. Reported machine-learning results were required to satisfy predefined criteria and independent evidence checks.

  • Programs and execution environments, rather than generated by the AI systems, were used for the reported computational work.
  • The first author defined the scientific questions, research boundaries, decision criteria, theoretical hypotheses, and conceptual transitions, retaining authority over scientific decisions.
  • Under that direction, AI systems used domain knowledge and accumulated GFG evidence to perform analysis, implement and monitor experiments, conduct interventions and replays, search for counterexamples, and execute validation.
  • No conclusion was accepted without predefined criteria, native execution, generation-fact evidence, and independent verification.

Data availability

The study’s generation-fact evidence, frozen experimental records, validation results, and supporting data are publicly available, alongside protocols, implementations, executable experiments, validators, and reproduction entry points. The paper reports no specific external grant, no competing interests, and identifies Mian Wang for correspondence and material requests.

  • Data and evidence: Generation-fact evidence, frozen experimental records, validation results, and supporting data are available in the accompanying public evidence archive.The archive is linked to the frozen experimental repository and identifies evidence associated with each reported experiment.
  • Reproducibility: Frozen protocols, GFG implementations, executable experiments, independent validators, and reproduction entry points are publicly available.The materials are provided through the study’s public code repository and a tagged release.
  • Declarations: The research received no specific grant from public, commercial, or not-for-profit funding agencies.
  • Declarations: The author declares no competing interests.
  • Contact: Correspondence and material requests should be addressed to Mian Wang.
Loading 2608.20965v1…