Source-linked AI summary
World Model Science: Self-Organized Criticality, Weak Chaos, and Metastable Belief Dynamics in Long-Horizon LLM Agents
Xinyuan Song, Zekun Cai
TL;DR
Long-horizon agents must preserve task state across extended interactions, but terminal reward and step validity do not reveal trajectory-level state dynamics. The paper aligns agent-implied and benchmark-grounded states and evaluates stress, avalanches, temporal dependence, local–global mismatch, and structural propagation across diverse tasks. It finds finite, substrate-conditioned world-model dynamics while rejecting universal criticality claims.
Problem
Long-horizon evaluation lacks trajectory-level evidence about how agent state errors accumulate, persist, cluster, and propagate beyond terminal reward and stepwise validity.
Method
The framework aligns agent-implied states with benchmark-grounded states and applies stress, avalanche, temporal-dependence, finite-size, invariance, and null-model diagnostics across agent trajectories.
Results
The experiments identify stress-sensitive collapse, local–global divergence, temporally correlated errors, dependency-conditioned propagation, and horizon-bounded avalanche growth across controlled and natural tasks.
Takeaways & Limitations
Trajectory-level world-state diagnostics add information that terminal reward and stepwise validity do not record, supporting finite, substrate-conditioned world-model analysis.
Takeaways & Limitations
The results do not establish physical SOC, a universal power law or critical point, or a single intervention policy across substrates, and state extraction is substrate-specific.
Abstract
from arXiv · showhide
Long-horizon LLM agents must maintain task state across extended sequences of observations, actions, tool calls, and intermediate beliefs. We study these trajectories through three dynamical views: self-organized criticality, weak chaos, and metastable belief dynamics. Our framework aligns agent-implied states with benchmark-grounded states and measures stress accumulation, error avalanches, temporal dependence, local--global mismatch, bounded divergence, belief-basin transitions, and finite-size scaling under explicit null models. Across 22 experiments spanning controlled puzzles, tool use, embodied tasks, multi-hop retrieval, general-assistant reasoning, and Game of Life, we find that locally valid actions can persist after global state fidelity fails, stress can trigger abrupt collapse, error sequences exhibit long memory, dependency depth changes the propagation regime, and larger horizons support larger avalanches. At the same time, divergence remains bounded, belief states show metastable rather than fully chaotic behavior, and stronger claims of universal power laws, critical points, or shared intervention optima are not supported. These results suggest a science of agent world models based on trajectory-level dynamical diagnostics rather than terminal reward alone.
1 Introduction
The paper operationalizes finite world-model SOC by measuring benchmark-grounded agent trajectories rather than relying on terminal outcomes. It studies whether stress, correlated errors, local–global divergence, and structural propagation produce collapse signatures across long-horizon tasks.
- Core framing: Long-horizon agents accumulate assumptions, contradictions, retrieval conflicts, and tool debt even when individual actions remain locally valid.This motivates separating executable action validity from semantic task-state fidelity.
- Core framing: The framework measures agent world models as benchmark-grounded dynamical traces, asking whether stress rises before collapse and whether local validity can hide global divergence.Measured dimensions include progress, belief, constraints, uncertainty, risk, memory, and plan.
- Diagnostics: Matched null models test whether state errors are temporally dependent, structurally propagated, and shaped by finite horizons beyond independent per-step errors.The estimands include temporal dependence, structural propagation, and finite-size behavior.
- Scope: The study spans controlled puzzles, tool use, embodied navigation, retrieval, assistant reasoning, and Game of Life to examine stress-sensitive collapse and propagation regimes.The benchmark suite includes StatefulPuzzle-SOC, τ-bench, GAIA, ALFWorld, HotpotQA-RAG, and Game of Life.
- Contributions: The contributions define agent-implied state, world-state fidelity, the local–global gap, collapse episodes, and diagnostics for temporal dependence, structural propagation, and finite-size behavior.These contributions turn the SOC analogy into measurable trajectory-level quantities.
2 Preliminaries: Agent Trajectories as Finite-State Measurements
The paper represents each agent interaction as a finite logged trajectory mapped into a benchmark-grounded state vector. Fixed fidelity, validity, stress, and capability-matched measurements then define comparable collapse quantities.
- State representation: A substrate-specific extractor maps logged observations, actions, and environment responses into progress, belief, constraints, uncertainty, risk, memory, and plan state.A gold extractor obtains corresponding state from simulator, database, supporting-fact, or audit-log sources.
- State representation: Local action validity measures syntax, immediate admissibility, and one-step consistency, while world-state fidelity compares measured state with benchmark-grounded state.Their difference forms the local–global gap.
- Stress and avalanches: Frozen stress combines unresolved uncertainty, contradictions, retrieval conflicts, unverified assumptions, and tool-error debt using weights fixed before outcome analysis.These components define the stress score used for threshold and avalanche analysis.
- Stress and avalanches: For a threshold, avalanche sets yield event size, weighted size, and duration over contiguous stressed or erroneous trajectory segments.The duration is bounded by the first and last included event times.
- Measurement assumptions: The protocol fixes measurable maps and weights before observing outcomes, assigns zero to unavailable benchmark coordinates, and conditions claims on independently valid trajectories.Empty capability support is recorded as a capacity-boundary result rather than treated as dynamical evidence.
3 Theory and Diagnostic Claims
The theory defines finite world-model SOC through correlated errors, stress response, finite-size scaling, and structural conditioning under explicit measurement assumptions. It also establishes that local validity cannot generally identify global state fidelity.
- Scope of theory: The framework treats the remaining diagnostics as confirmatory tests after defining finite world-model SOC and recording the identification limitation.The measurement claims therefore depend on the stated assumptions and finite observability setup.
- Finite world-model SOC: Finite world-model SOC requires temporal dependence beyond an independent-error null, stress-sensitive perturbation response, horizon-dependent cascade scale, and graph-conditioned propagation.Tail universality, a unique power-law exponent, and a shared intervention optimum are not required.
- Identification limitation: Local validity is insufficient to identify global state fidelity in finite partially observable processes, so syntactically admissible actions may coexist with diverged policy, evidence, or belief state.This limitation is demonstrated for τ-bench Airline and GAIA.
- Structural conditioning: Surface-preserving prompt transformations are compared with matched task-structure perturbations to test whether macro collapse statistics depend more on task graphs than presentation surfaces.Horizon and task graph are treated as structural variables affecting the same statistics.
4 Related Work
Related work evaluates agent steps, intermediate states, criticality-inspired statistics, and long-horizon reliability. This paper connects these strands by analyzing benchmark-aligned state-error sequences over time and task graphs.
- Process-level evaluation: Existing agent evaluations identify step failures, while this paper tests whether trajectory errors are independent, persistent, and structurally propagated.The distinction shifts analysis from isolated breakdowns to sequence-level dependence.
- Agent state and belief: State-tracking and belief-evaluation studies make intermediate drift observable; this work aligns agent-implied state with benchmark gold state before analyzing errors.The resulting sequence is evaluated across time and task graphs.
- Criticality-inspired diagnostics: SOC and neural-criticality work supplies event-size, finite-size, temporal-dependence, and invariance estimands, which this paper transfers without importing physical criticality conclusions.The transfer is methodological rather than a claim that agents are physical SOC systems.
- Reliability under long horizons: Long-horizon studies show that error accumulation matters, while this paper asks whether errors are independent, persistent, clustered, and conditioned by depth, topology, and horizon.These conditions define the paper’s more specific statistical question.
5 Experimental Protocol
The protocol evaluates long-horizon agents with fixed multi-substrate experiments and benchmark-grounded trajectory diagnostics rather than terminal performance alone. It converts traces into world-state, stress, avalanche, and statistical measurements, using corrected tests and estimand-specific null models.
- Experimental design: Twenty-two fixed experiments use one model interface, temperature zero, and a fixed seed across controlled puzzles, tool use, assistant reasoning, embodied navigation, retrieval, and Game of Life.StatefulPuzzle-SOC provides stress, horizon, and dependency-depth interventions; other substrates test policy-constrained API use, evidence gathering, embodied state, multi-hop retrieval, and local-rule dynamics.
- Measurement pipeline: The framework converts agent traces into benchmark-grounded world-state estimates, stress components, avalanche events, and diagnostic summaries while separating local action validity from global state fidelity.This separation is intended to expose silent state collapse before terminal reward error.
- Null models and testing: Statistical tests compare each signature with an appropriate null, including independent Bernoulli errors, Markov persistence, task-difficulty predictors, shuffled spectra, and surface-perturbation controls.Confirmatory p-values use Benjamini-Hochberg false-discovery control at q = 0.05.
- Evidence accounting: Table 1 organizes the paper-level evidence into stress, local-global, memory, geometry, and scaling blocks, while separately recording partial, boundary, capacity-limited, and alternative-explanation results.The accounting is semantic rather than hypothesis-code based.
6 Results: Collapse Is Measurable and Not Independent Noise
Across controlled and benchmark settings, collapse is associated with externally injected stress, clustered errors, and local-global state mismatches rather than independent step noise alone. These diagnostics remain stable under task-preserving surface changes, while temporal dependence and dependency depth reveal structured, bounded propagation.
- 6.1 Controlled Stress Produces Sharp Collapse: AUROC 0.979 links injected stress to collapse in StatefulPuzzle-SOC, where the first nonzero stress level shifts behavior from a zero-stress floor to near-deterministic collapse.The intervention uses horizon 64 and dependency depth one, with stress manipulated before outcome analysis.
- Collapse Is Measurable and Not Independent Noise: Retail avalanches are too clustered for matched independent Bernoulli errors, although truncated power laws, lognormal tails, and Markov persistence each explain part of the finite-horizon pattern.The result supports temporal clustering with bounded tails under finite policy constraints rather than a universal heavy-tail claim.
- 6.2 Local Actions Can Stay Valid After the World State Has Diverged: ΔLG = 0.857 in GAIA Level-1 shows that intermediate conclusions can remain locally valid after the evidence state has collapsed, while information-only Airline tasks have gaps near zero.Thus step admissibility and latent task-state correctness can decouple across task classes.
- Macro stability under surface perturbation: Task-preserving paraphrase, renaming, distractors, order changes, and style changes keep avalanche summaries aligned with identity conditions, with KS distances below the same-distribution null threshold.HotpotQA and Retail show the same qualitative stability.
- Temporal dependence: Long-memory signatures persist across StatefulPuzzle horizons, whereas full-context HotpotQA nearly decorrelates errors and narrow top-2 retrieval produces flicker-like persistence.The retrieval comparison provides a mechanism check for temporal dependence.
- 6.5 Dependency Depth Creates a Bounded Divergence Transition: From dependency depth two onward, power-law fits are preferred in most paired trajectories instead of the exponential fits preferred at depth one, but finite-state divergence saturates.The result indicates a depth-dependent change in bounded propagation, not a positive Lyapunov exponent or unbounded chaos.
7 Results: Collapse Has Geometry, Scale, and Capability Limits
Collapse dynamics depend on topology, horizon, retrieval regime, and capability boundaries rather than following one universal pattern. Stronger interpretations receive partial support: natural stress is modestly predictive, macro regimes are continuous, and intervention optima are substrate-specific.
- 7.1 Geometry: Fractal dimensions are stable within fixed HotpotQA evidence topology but vary across ALFWorld task graphs, supporting topology-conditioned geometry.HotpotQA Df = 0.835–0.904 with spread 0.069; ALFWorld values are 0.639 for long-chain, 1.042 for container, and 1.50 for multi-room tasks.
- 7.2 Memory: Error sequences remain long-memory in StatefulPuzzle, whereas narrow top-2 retrieval pushes HotpotQA errors toward flicker-like persistence.The retrieval width changes the temporal regime, showing that information channels condition error dependence.
- 7.2 Propagation: Dependency depth changes the propagation regime: the AIC sign flips at depth two while divergence remains bounded rather than chaotic.The fraction of pairs favoring power-law fits rises above chance, but finite-state saturation rules out unbounded divergence.
- 7.2 Scale: Maximum avalanche cutoff grows monotonically from 7 to 490 as horizon increases, demonstrating bounded finite-size scaling.All adjacent-horizon tests remain statistically significant after false-discovery correction, while the finite task still constrains the process.
- 7.3.1 Natural Stress: Retail stress is only modestly above chance after removing circular labels, adding little predictive value beyond task difficulty.Figure 12 separates independent reward-error events from concurrent fidelity drops and reports only a small incremental gain.
- 7.3 Boundary Tests: Macro statistics remain stable across prompt variants, but clustering is continuous rather than sharply discrete, and intervention optima differ by substrate.Retail has KS distances 0.03–0.17; ALFWorld favors exploration and memory, whereas Retail favors no intervention or high verification.
- Capability Limits: Missing diagnostic signatures are interpretable only when the model generates valid trajectories on the measured substrate.For unsupported regimes, absence of a signature indicates a capability limit rather than measured dynamics.
8 Discussion
The discussion argues that trajectory-level world-state diagnostics reveal process dynamics that terminal reward and stepwise validity omit. It recommends measuring local-global mismatch, temporal dependence, and structure-conditioned propagation while maintaining explicit scope limits.
- Evaluation Practice: Agent evaluation should log implied task state and local-global mismatch instead of relying only on terminal reward and executable actions.The mismatch distinguishes locally executable actions from fidelity to the actual task state.
- Evaluation Practice: Error sequences should be tested against Bernoulli and persistence nulls because equal mean error rates can conceal different burst and memory structures.The framework separates independent noise from correlated collapse using event sizes, spectra, and explicit null models.
- Evaluation Practice: Propagation diagnostics should be repeated across horizon, dependency depth, and task topology because the same model can occupy different regimes under different structures.The paper treats graph-conditioned propagation and finite-size cutoffs as distinct from ordinary long-task error accumulation.
- Scope: The claims are limited to finite, substrate-conditioned diagnostics: they do not establish physical SOC, a universal power law, a universal critical point, or one intervention policy.State extraction is also substrate-specific, so the measurements are not fully substrate-independent.
9 Conclusion
Across twenty-two experiments, the paper finds consistent finite world-model signatures alongside mixed support for stronger universality claims. The conclusion favors trajectory-level, substrate-conditioned diagnostics over terminal outcomes alone.
- Conclusion: Across twenty-two experiments, the most consistent signatures are stress-sensitive collapse, local-global divergence, correlated errors, dependency-conditioned propagation, and horizon-bounded avalanche growth.These findings span controlled and natural tasks.
- Conclusion: Stronger interpretations remain limited: natural stress is modestly predictive, macro regimes are continuous, and intervention optima do not transfer across substrates.Persistence models explain only part of avalanche structure, so the results support finite rather than universal criticality.
- Measurement Scope: The measurement protocol fixes seven state components before computing fidelity, stress, or avalanche statistics for each substrate.This protocol makes the extracted world-state representation part of the empirical measurement design.
B Proofs and Formal Verifications
The formal analyses show that local action validity cannot identify global world-state fidelity, while dependency depth and null-model comparisons distinguish correlated bounded dynamics from independent errors. Additional diagnostics constrain interpretation through finite-horizon, finite-sample, and multiple-testing safeguards.
- B Proofs and Formal Verifications: Local validity cannot identify world-state fidelity when hidden constraints differ despite identical observations, actions, and immediate responses.The construction yields identical local-validity sequences but divergent fidelity sequences after a hidden constraint is violated.
- B Proofs and Formal Verifications: The independent first-error null gives no-error probability (1 − ε0)^H, with Cind = 0.133 at H = 64 implying ε0 ≈ 0.0022.
- B Proofs and Formal Verifications: Dependency depth changes propagation regime: mean ∆AIC switches sign after depth one, while power-law-fit preference rises from 0.133 to at least 0.667.Because divergence is bounded, fitted exponents describe shape, while saturation timing is reported separately.
- C Complete Evidence Accounting and Null Models: Matched Bernoulli errors reproduce mean event rates but underestimate clustering and large bursts, whereas Markov persistence explains substantially more Retail burstiness.The comparison separates independent noise, correlated persistence, and heavier bounded tails.
- C Complete Evidence Accounting and Null Models: Retail avalanche model selection favors bounded heavy-tail families, including lognormal and truncated power-law alternatives, over a pure power law.Finite horizons, tool policies, and domain constraints produce a bounded correlated-collapse regime.
- C Complete Evidence Accounting and Null Models: Retail early stress is above chance but modest, task difficulty nearly matches it, and adding stress yields only a small cross-validated gain.Controlled stress sensitivity and Retail precursor information therefore represent different estimands.
- C Complete Evidence Accounting and Null Models: Spectral exponents are the primary long-memory diagnostic across horizons because short-horizon DFA can be inflated by step-like series.DFA is treated as reliable only at longer horizons.
- C Complete Evidence Accounting and Null Models: Confirmatory tests use Benjamini-Hochberg correction, and headline finite-size and long-horizon spectral tests remain significant after correction.
D Additional Boundary and Direction Results
Boundary analyses show shallow rather than rigid belief-basin trapping, modest Retail stress precursor information, continuous prompt-regime structure, and substrate-dependent intervention preferences. The evidence accounting and replication guidance emphasize conditional, benchmark-specific interpretation.
- D. Boundary Conditions: World-state components are benchmark-specific, normalized by horizon or maxima, and set to zero when unavailable rather than imputed.
- D. Boundary Conditions: The dependency-transition evidence combines AIC sign, paired-trajectory power-law preference, and bounded-system saturation timing.
- D. Boundary Conditions: Complete evidence accounting distinguishes supporting, partial, boundary, capacity-limited, and alternative-explanation rows, limiting stronger SOC interpretations.
- D.1 Shallow-Basin Direction: ALFWorld agents escape wrong belief basins quickly but under-exploit success-associated basins, indicating over-escape and underexploitation rather than rigid lock-in.Visibility manipulations do not explain the effect, and Retail verification timing provides only a modest intervention signal.
- D.2 Retail Stress Precursor: Under an independent reward-error label, Retail early stress remains above chance but modest, while task difficulty nearly matches its predictive value.
- D.3 Prompt-Regime Clustering: Retail prompt variants yield six interpretable clusters, but the silhouette-optimal partition is coarser, supporting continuous dynamical regimes with fine interpretable partitions.
- D.4 Intervention Direction: Retail favors no intervention or high verification, whereas ALFWorld favors exploration-heavy or memory-heavy regimes, making intervention balance substrate-dependent.Interventions should be selected after measuring the substrate’s error dynamics.
- D. Boundary Conditions: Replications should combine controlled stress collapse, dependency-depth transitions, local-global gap extraction, long-memory estimators, and independent and correlated nulls.