Source-linked AI summary

Reproducible macroscopic dynamics in a closed-loop human-AI learning system

Minlin Wu, Xu Fang, Yicheng Zhang, Chenyu Zhou, Zhiyi Liu

arXiv:2608.30946v1cs.LGnlin.AO

TL;DR

The paper develops a procedure for identifying reproducible macroscopic dynamics in closed-loop systems using prespecified semantic observables, held-out cohorts, construction-aware nulls, and matched representations. It finds reproducible basin-like flow, state-heterogeneous metastable-like kinetics, and a transferable leading-order effective-field structure within the studied platform and event-time setting.

  • Problem

    The paper addresses how to identify reproducible macroscopic dynamics in closed-loop systems where agent actions reshape future observations.

  • Method

    The study prespecifies semantic observables, reconstructs user-disjoint cohorts, tests held-out dynamics against construction-aware nulls, compares matched gauges, and evaluates mesostate persistence and predictive representations.

  • Results

    The state shows reproducible contractive basin geometry and state-heterogeneous metastable-like kinetics, with training-validation drift agreement of r = 0.930 across 939 common cells.

  • Takeaways & Limitations

    Within one adaptive platform and conditional event-time closure, the identification procedure provides a falsifiable template for studying closed-loop systems.

  • Takeaways & Limitations

    The evidence is scoped to one adaptive platform and conditional event-time closure, and population-level reproducibility does not imply learner-level homogeneity.

Abstract

from arXiv · show

Closed-loop human-AI systems generate high-dimensional behavioural trajectories whose collective dynamics remain obscure. Using 297,915 learners' adaptive-tutoring histories, we define semantic order variables before model fitting and test them in user-disjoint cohorts. The state exhibits reproducible basin-like flow and operationally defined, state-heterogeneous metastable-like kinetics. A construction-matched null distinguishes normalised-memory relaxation from a reproducible excess field. A four-term conditional mechanism recovers population drift (r = 0.946; learner-bootstrap 95% CI, 0.935-0.955). Predictive event-level self-supervised learning recovers the state and learned-plane flow; null-referenced corrections retain directional, partial-amplitude excess-field structure without full calibration. Shuffled-order training reverses learned-plane flow on ordered trajectories; support-alignment randomisation selectively reduces inward transport. Both axes remain linearly accessible without state supervision. Without cross-model fitting, the models share leading population drift (r = 0.866; learner-bootstrap 95% CI, 0.857-0.875) and persistence ordering; residual directions remain model-specific. These results identify an externally anchored leading-order effective field linking empirical dynamics, an interpretable mechanism and neural computation.

Results

Prespecified semantic coordinates revealed reproducible basin-like flow, state-heterogeneous persistence, and an excess field beyond construction-implied relaxation. A compact mechanism and Event-SSL independently recovered important aspects of these dynamics across held-out users, while their shared field was strongest after coarse-graining.

  • Semantic coordinates: Response order M and exposure alignment Ψ were prespecified as distinct bounded coordinates from event semantics before fitting.M captures signed response residuals; Ψ captures whether exposure addresses unresolved demand.
  • Reproducible empirical effective dynamics: 77.6% of interior user-balanced occupancy had negative divergence, with weighted mean divergence −0.215 and a dominant low-Ψ quasi-potential minimum.Half the user-balanced mass occupied Ψ ∈[−0.875, −0.525].
  • Reproducible empirical effective dynamics: The construction-matched null rejected full-field relaxation and excess negative-divergence occupancy, localising replicated excess to directional structure rather than uniform core slowing.The full-field departure had pMC = 0.0099, while excess negative divergence had adjusted q = 0.0297; shell inward flow replicated only in confirmation.
  • Operational mesostate kinetics: All six mesostates were diagonal-dominant across 3,328,409 validation transitions, with mean self-transition probability 0.750; strict user-equal weighting altered only S0.S0 self-transition fell from 0.3623 to 0.0860, identifying a mesostate-specific weighting boundary.
  • Operational mesostate kinetics: Positive 10-step survival excess occurred in S0, S4 and S5, while statewise RMST exceeded geometric references in all six mesostates by factors of 1.127–6.473.The evidence supports operationally defined, state-heterogeneous metastable-like kinetics rather than a uniform positive shift.
  • Reproducibility: Training and validation agreed closely, with occupancy JS divergence 5.24 × 10−4, mean local drift cosine 0.966, and drift-speed correlation r = 0.930.Transition row-wise TV was 0.006, while recurring contractive regions were more reproducible than exact thresholded contours.
  • Mechanistic closure: The four-term mechanism recovered confirmation drift-vector correlation r = 0.9457 and preserved persistence ordering, although empirical residence tails retained structure beyond its one-step map.Across 954 cells, the bootstrap interval was 0.9351–0.9546; mean row-wise transition TV was 0.1021 and self-transition profiles correlated at r = 0.9870.
  • Predictive representation: Event-SSL recovered both coordinates and coarse dynamics across independent users, retaining directional and partial-amplitude excess-field structure without full calibration.Learned-plane drift correlation was r = 0.6877, and transition matrices matched dominant outgoing edges in all six mesostates.

Discussion

The semantic state supports reproducible basin-like flow and state-heterogeneous metastable-like kinetics, while controls and cross-model tests identify a leading organising backbone rather than a complete dynamical description. The evidence supports a transferable identification procedure for closed-loop systems, but remains scoped to one adaptive platform and conditional event-time closure.

  • Discussion: The state X = (M, Ψ) supports reproducible basin-like flow and operationally defined, state-heterogeneous metastable-like kinetics.The kinetics interpretation is conditional on the stated Kaplan–Meier censoring assumption.
  • Discussion: Shuffled-order training reversed learned-plane flow, while support-alignment randomisation reduced inward transport without changing state location and coarse transitions.The inward-transport reduction was 66.9% for seed 42 and occurred in every seed.
  • Discussion: The state is a leading organising backbone rather than an exhaustive dynamical description: residence tails, density mass, off-diagonal routes and local residual direction remain partly model-specific.Within Event-SSL, the two-coordinate bottleneck retained 90.8% of descriptive macrostructure, while residual activity retained 80.2% of drift structure.
  • Discussion: Held-out replication, construction-aware excess, low-order closure and approximately linear accessibility support the empirical status of the prespecified coordinates.Field geometry remained stable across memory and activity variants, with r = 0.962–0.981.
  • Discussion: The transferable contribution is an identification procedure rather than the numerical form of M and Ψ, demonstrated within one adaptive platform and conditional event-time closure.The procedure combines semantic observables, held-out tests against construction-aware expectations, matched gauges, low-order closure and predictive-representation tests.

Methods

The study defines a prespecified two-dimensional state from response evidence and exposure alignment, then evaluates its fields, kinetics, null departure, and mechanism under frozen cohort-wise specifications.

  • The analysis uses a common submitted-bundle event-time panel and prespecified (M, Ψ) phase plane, with confirmation held out until specifications were final.
  • Event panel: Submitted bundles define event time, with pre-states at bundle entry or submission and interval updates spanning responses and subsequent support activity.Missing steps are not interpolated, and final pre-states contribute to occupancy but not drift.
  • Null testing: The construction-matched null preserves current states, denominator increments, weights, support, and training-defined regions while permuting innovation pairs within opportunity-matched strata.Validation and confirmation each use 100 permutations analyzed separately.
  • Kinetics: Kinetic summaries use a fixed K = 6 training partition and statewise Kaplan–Meier estimates, whose interpretation assumes non-informative right censoring.Residence episodes are censored at user boundaries, sequence gaps, and unobserved mesostates.
  • Mechanism inference: The selected mechanism was calibrated on pooled data and then fixed before confirmation, with calibration parameters (η, r, γR, γA) = (20, 0.3865, 1, 0.9537).Confirmation used no further search, recalibration, or spatial redefinition.

Data availability

The EdNet-KT4 interaction logs and associated metadata are publicly available from the official EdNet repository under a non-commercial Creative Commons licence.

  • EdNet-KT4 logs and associated EdNet Contents metadata are publicly available through the official EdNet repository.The release is licensed under Creative Commons Attribution–NonCommercial 4.0 International for non-commercial scholarly research.

Ethics statement

This study is a secondary analysis of publicly released interaction records and reports only aggregate, non-identifying results.

  • No participants were recruited or contacted, no new human data were collected, and researchers did not re-identify users or link records to external personal data.

Supplementary Information

Supplementary analyses preserve the frozen empirical construction while testing null departure, alignment specificity, and matching quality across held-out cohorts.

  • The accounting null preserves current states, increments, weights, support, and frozen regions while permuting response–exposure innovation pairs without refitting downstream definitions.
  • Validation and confirmation each use 100 permutations with separate analyses.
  • After null subtraction, excess fields agree across 917 cells, with vector r = 0.7164, speed r = 0.7362, and occupancy-weighted local cosine 0.9604.The result supports reproducible excess structure beyond normalized-memory accounting.
  • The alignment-specificity control replaces Ψ with activity–idle balance Φ while retaining the same memory, denominator, mappings, weights, thresholds, and null strata.Its result indicates greater null-referenced prominence, not causal necessity or coordinate optimality.
  • Supplementary Table 1 reports prespecified validation tests, held-out confirmation replication, matching-tier composition, opportunity composition, and reconstruction checks.The full-field distance is tested separately from the three basin metrics, using Benjamini–Hochberg-adjusted values for basin tests.

b. Opportunity-matching audit

Opportunity-matching controls support persistent, reproducible dynamical regions, while benchmark analyses indicate sensitivity to threshold-defined boundaries.

  • 1.023 median vector agreement-to-benchmark and 0.976 median speed agreement-to-benchmark quantify post hoc whole-user attenuation benchmarks.Across 29 valid partitions, the 2.5–97.5% ranges were 0.963–1.067 and 0.939–1.017.
  • 0.785 selection-frequency overlap indicates persistent detection across learner reweightings, despite a median thresholded-region Jaccard of 0.456.Median centre separation was 0.110, and the prespecified core was unchanged.
  • The opportunity-matching audit distinguishes stable dynamical detection from threshold-sensitive region boundaries.Permutation-based and exact point estimates differed by at most 0.0041.
  • All four memory or activity variants retained contractive, inward and centrally slowed validation flow, with the slow-activity field correlating 0.9622 with the primary field.The slow-activity mapping displaced the training core most strongly.

Supplementary Note 2. Mesostate kinetics and construction-aware robustness

The fixed six-state partition provides an operational coarse-graining for heterogeneous residence and persistence, with construction-aware tests targeting state-specific excess kinetics.

  • All six validation mesostates had heterogeneous persistence strengths and residence profiles under the fixed K = 6 partition.The continuous field itself was estimated without clustering.
  • Reliable-horizon restricted mean residence exceeded the state-matched geometric reference in all six mesostates.The supplied passage also reports positive survival excess at the prespecified 10-step horizon, but the sentence is truncated.
  • The recursive surrogate preserved opportunity matching, denominators, segments and observation structure while propagating coherent state sequences.The all-state endpoint did not support a uniform 10-step lift, whereas the family-wise maxT test identified S0, S4 and S5.
  • Supplementary controls altered memory or activity mappings, re-estimated training fields and evaluated variants on validation without confirmation data.The controls also include observation-gap and alignment-specificity analyses.

b. Observation-gap restrictions

Observation-gap and partition controls frame mesostate kinetics as population-level, censor-aware estimates with broad reliable-horizon persistence but fixed-horizon excess concentrated in selected states.

  • 0.9963 occupancy mass and 0.8735 Jaccard are reported for the validation alignment-based state, compared with 0.9984 and 0.7851 for the activity–idle comparator.The corresponding confirmation values are 0.9961 and 0.8848 versus 0.9982 and 0.7851.
  • The fixed K = 6 table reports user-balanced occupancy and censor-aware Kaplan–Meier residence, with RMST lift defined at each state’s reliable horizon.Statewise lift magnitudes are not directly comparable.
  • RMST lift exceeded one in every mesostate across learner-cluster bootstrap results and remained above one for K = 4–8.Positive 10-step excess persisted in the same three mesostates.
  • The stated metastable-like interpretation is limited to reproducible population-level, state-heterogeneous basin and residence kinetics.It does not claim spectral metastability, autonomous Markov closure or causal platform effects.

Supplementary Note 3. Mechanism-family selection and post-selection kinetics

Mechanism selection favors a stable four-term offset dual-channel family, while post-selection transition and residence diagnostics preserve persistence ordering but reveal differences in long-horizon scale.

  • Four-term parsimony remained stable under equal weighting, with no family of three or fewer coefficients qualifying under the declared hierarchy and criteria.Selection remained bounded by the objective, hierarchy and search domain.
  • At 10 steps, observed D10 was family-wise significant for S0, S4 and S5, with pFWER = 0.0099 for each listed state.Learner-cluster RMST lift confidence intervals were above one for these states.
  • The construction- and denominator-inertia-matched surrogate retains the declared opportunity-matching scheme and observation structure for robustness testing.The supplied table reports a descriptive mean self-transition probability of 0.7496 versus a null median of 0.6728.
  • The confirmation evaluation used frozen specifications without candidate-family reconsideration, parameter updates, recalibration, region redefinition or partition refitting.Shared scales and accounting quantities remained part of the frozen update specification.
  • 0.9870 correlation shows confirmation self-transition probabilities preserved the empirical persistence ordering.Five of six states retained the same dominant destination; the sole dominant-edge deviation was S0.
  • Transition-implied residence preserved persistence ordering but differed in absolute tail scale, exposing long-horizon structure beyond one-step closure.Transition and residence diagnostics were excluded from mechanism selection.

Supplementary Note 4. Event-SSL seed robustness and hidden-state organisation

Event-SSL results were robust across seeds, while hidden-state analyses separated compact coordinate structure from higher-dimensional refinements and task information.

  • All six seeds preserved positive learned-plane field correlation, while shuffled-to-ordered transfer reversed its sign and support-alignment randomisation reduced inward transport.
  • The two-coordinate bottleneck retained nearly all coordinate and descriptive transition information, while closure and drift retention were approximately four-fifths.
  • The selected dual-channel core achieved the lowest reported mean score, 0.2715, among the listed mechanism families.
  • Seed-42 local cosine and coarse transition topology exceeded within-user permutation floors, but global field correlation and statewise persistence did not.
  • Alternative state-only closures improved matched-origin drift and transitions but transferred incompletely to the primary learned-plane gauge.

Supplementary Note 5. Null-referenced downstream recovery and cross-model spatial structure

Null-referenced analyses showed that the mechanism was more fully calibrated, whereas Event-SSL recovered directional and partial-amplitude excess-field structure. Independently frozen models agreed most strongly on leading population dynamics, with residual structure remaining model-specific.

  • The mechanism reduced occupancy-weighted field distance from 0.0907 for the construction-null expectation to 0.0217, achieving null-relative skill 0.9430.
  • Event-SSL retained positive Ψ-specific skill but negative overall and M-specific skill, localising the calibration gap to response order.
  • Without cross-model fitting, mechanism and Event-SSL fields shared drift-vector correlation r = 0.8661 and drift-speed correlation r = 0.8552.
  • Cross-model agreement strengthened after coarse-graining, while density mass, off-diagonal routes and local residual direction remained model-specific.

Supplementary Note 6. Learner-level uncertainty, estimands and field-grid robustness

Sensitivity analyses examined learner-composition assumptions, representation geometry, and diagnostic state alignment. The reported hidden representations remained substantially higher-dimensional than the two-coordinate bottleneck.

  • The supplementary diagnostics treated representation-space clusters as distinct from the empirical K = 6 states.
  • Participation ratio was 11.23 and effective rank was 18.84 for the hidden representation.
  • TwoNN intrinsic dimension decreased from 8.913 in training to 7.471 in confirmation.
  • Bottleneck state alignment exceeded full-hidden and residual-hidden alignment in confirmation, with NMI / ARI of 0.4030 / 0.2378.

b. Composite-score definitions

Composite and field diagnostics quantify empirical replication, mechanism recovery, Event-SSL recovery, and cross-model agreement under multiple floors and sensitivity definitions. The main limitation is a weighting boundary affecting response-order and one empirical state's persistence estimate.

  • Composite-score definitions: Strict user-equal weighting reduced S0 self-transition probability from 0.3623 under interval counts to 0.0860.
  • Composite-score definitions: Field-grid variation changed local resolution but did not reverse the reported field relationships.
  • Composite-score definitions: Event-SSL coordinate correlations were 0.8837 for M and 0.9905 for Ψ, while one-step RMSEs were 0.1100 and 0.0353.
  • Composite-score definitions: Shuffled-to-ordered transfer produced learned-plane drift correlation −0.4191, and support-alignment randomisation changed inward fraction by 0.2047.
  • Composite-score definitions: Mechanism drift-vector recovery reached r = 0.9457, while Event-SSL learned-plane drift-vector recovery reached r = 0.6877.
  • Composite-score definitions: Mechanism versus Event-SSL anchor drift-vector correlation was r = 0.8661, with occupancy-weighted local cosine 0.7851.
Loading 2608.30946v1…