Source-linked AI summary
PLDR-LLMs Reason At Self-Organized Criticality
Burc Gokden
TL;DR
The paper addresses how reasoning emerges in PLDR-LLMs and whether it can be characterized without curated benchmark evaluation. It studies PLDR-LLMs through self-organized criticality, measuring deductive-output steady states and relating them to reasoning. The authors report that critical pretraining produces reasoning, metastable deductive-output states, and an order parameter that approaches zero for models with higher reasoning and comprehension benchmark scores.
Problem
The paper investigates how reasoning emerges in PLDR-LLMs and seeks a characterization beyond traditional loss optimization and curated benchmark evaluation.
Method
The paper studies small PLDR-LLMs from a self-organized-criticality perspective and defines an order parameter from normalized RMSE between deductive outputs across stochastic and cached inference runs.
Results
PLDR-LLMs pretrained at criticality achieve reasoning, attain metastable deductive-output steady states, and show order parameters closer to zero alongside higher reasoning and comprehension benchmark scores.
Takeaways & Limitations
Reasoning in PLDR-LLMs can be quantified from global deductive-output statistics at steady state without curated benchmark datasets.
Takeaways & Limitations
Near-critical training ranges depend on the tokenizer, model hyperparameters, training framework, training setup, and pretraining dataset, while deductive tensors can differ from their physical analogies.
Abstract
from arXiv · showhide
We show that PLDR-LLMs pretrained at self-organized criticality exhibit reasoning at inference time. The characteristics of PLDR-LLM deductive outputs at criticality is similar to second-order phase transitions. At criticality, the correlation length diverges, and the deductive outputs attain a metastable steady state. The steady state behaviour suggests that deductive outputs learn representations equivalent to scaling functions, universality classes and renormalization groups from the training dataset, leading to generalization and reasoning capabilities in the process. We can then define an order parameter from the global statistics of the model's deductive output parameters at inference. The reasoning capabilities of a PLDR-LLM is better when its order parameter is close to zero at criticality. This observation is supported by the benchmark scores of the models trained at near-criticality and sub-criticality. Our results provide a self-contained explanation on how reasoning manifests in large language models, and the ability to reason can be quantified solely from global model parameter values of the deductive outputs at steady state, without any need for evaluation of curated benchmark datasets through inductive output for reasoning and comprehension.
1 Introduction
The paper explains PLDR-LLM reasoning through self-organized criticality, where critical training conditions produce global metastable steady states in deductive outputs. It proposes an intrinsic order parameter for quantifying reasoning and generalization without curated benchmark datasets.
- 1 Introduction: PLDR-LLMs exhibit reasoning at specific warm-up step counts and maximum learning rates, while other settings produce random token sequences at inference.The associated loss curves appear underfit-like when reasoning is achieved and overfit-like when it is not.
- 1 Introduction: The study uses small PLDR-LLMs as an experimental vehicle for investigating intelligence emergence through self-organized criticality and second-order phase transitions.It also relates the model dynamics to observations associated with the human brain and complex systems.
- 1 Introduction: At criticality, overlapping long-range interactions establish a global metastable steady state across the model's deductive outputs.The paper interprets warm-up rate and maximum learning rate as control parameters for extrinsic driving and intrinsic dissipation.
- 1 Introduction: A global order parameter measures how well a PLDR-LLM can reason without curated benchmark datasets and remains robust against stochastic sampling.An order parameter close to zero indicates high reasoning and generalization capabilities.
- 1 Introduction: The paper offers explanations for the dependence between LLM size and token amount and for performance improvements associated with rotary positional embeddings and GLUs.These explanations are presented as consequences of the self-organized-criticality perspective.
2 Background and Related Work
PLDR-LLMs use nonlinear power law graph attention with structured deductive outputs that form global model representations. The paper situates this architecture within self-organized criticality and uses it to study reasoning and generalization.
- 2 Background and Related Work: PLGA replaces predefined attention adjacency structures with input-driven, learnable adjacency matrix parameters.Its motivation draws on graph interpretations of attention and analogies to quantum mechanics and general relativity.
- 2 Background and Related Work: PLGA constructs deductive outputs from query-derived density matrices, residual-network transformations, and learned power law tensors.These transformations produce tensors associated with density, metric, potential, and energy-curvature representations.
- 2 Background and Related Work: The deductive outputs {A, ALM, AP, GLM} provide global model representations, while projected query and key vectors extract input-relevant attention for value-vector prediction.Equations 1-6 describe this attention-generation process.
- 2 Background and Related Work: The paper treats self-organized criticality as a framework for systems that reach critical states with power-law behavior and notes that power laws also occur in natural language.PLDR-LLMs provide controllable parameters and access to intrinsic characteristics through deductive outputs.
- 2 Background and Related Work: The study investigates small PLDR-LLMs trained near and below criticality and compares an analytical order parameter with curated benchmark scores.The benchmarks cover commonsense reasoning, question answering, and language understanding.
3 Approach
The approach trains PLDR-LLMs across warm-up and maximum-learning-rate conditions spanning near-critical and below-critical regimes. It measures steady-state deductive-output variation and compares the resulting order parameter with zero-shot benchmark performance.
- 3 Approach: PLDR-LLMs with 5 decoder layers, 14 heads, and 64 embedding dimensions per head were pretrained over ∼8B RefinedWeb tokens.The experiments varied maximum learning rate and warm-up step count, with a SwiGLU:LU ratio of 170:64.
- 3 Approach: A second model with the same architecture was pretrained over ∼41B RefinedWeb tokens under a maximum learning rate selected for near-critical training with 2000 warm-up steps.The training sample interval was [0, 80M].
- 3 Approach: Learning-rate and warm-up pairs were chosen across above- and below-criticality regions, producing underfit-like and overfit-like loss and accuracy curves.During warm-up, driving and dissipating forces interact and determine whether critical learning continues.
- 3 Approach: The order parameter is defined as normalized RMSE by mean magnitude between deductive outputs collected from two uncached runs and one cached run.The study used 100 IMDB test samples, nucleus sampling with top-p 0.8, temperature 1.0, and continuations up to 256 tokens.
- 3 Approach: Zero-shot evaluation covered ARC, Hellaswag, WinoGrande, TruthfulQA, OpenBookQA, PIQA, and SIQA using normalized accuracy measures.These benchmarks assess commonsense reasoning, question answering, and language understanding.
- 3 Approach: Near-criticality ranges depend on tokenizer, model hyperparameters, training framework, training setup, and pretraining dataset.The implementation update to value-vector initialization may also slightly change the warm-up and maximum-learning-rate range for criticality.
4 Results
Near-critical PLDR-LLMs follow a shared metastable training trajectory, while sub-critical models diverge and show less stable deductive-output behavior. Near-criticality is associated with lower output perturbations, an order parameter closer to zero, and higher benchmark scores.
- Training Loss and Accuracy Characteristics: Maximum learning rates from 1.2 × 10−3 to 8 × 10−4 produced near-critical models with almost identical loss curves and slight final-loss offsets.Their behavior suggests approximation of the same critical steady-state condition.
- Training Loss and Accuracy Characteristics: PLDRv51-SOC-110M-5 followed the same near-critical trajectory despite five times more training data and a more pronounced loss-and-accuracy offset.
- Distribution of Deductive Output Values: More training data made PLDRv51-SOC-110M-5's normalized RMSE zero for AP and GLM and orders of magnitude smaller for A and ALM than in other models.The authors associate this with greater invariance of deductive outputs to inference inputs.
- Distribution of Deductive Output Values: Near-critical and sub-critical models differ in deductive-output distributions: sub-critical A values span a wider range and lose repetitive, uniform last-layer heatmap structure.Global distributions are sparse for A, ALM, and GLM, while AP is centered around ∼1.
- Comparison of Order Parameter and Benchmark Scores: Models nearer criticality had order parameters closer to zero and higher average benchmark scores, whereas sub-critical models had larger order parameters and lower scores.PLDRv51-SOC-110M-5 had the smallest order parameter values and highest average benchmark scores among the pretrained PLDR-LLMs.
5 Ablation Studies
The ablation studies examine warm-up and maximum-learning-rate pairs associated with dragon king events, which appear as sharp loss peaks and simultaneous accuracy drops. These events occur in both near-critical and sub-critical conditions and are associated with reduced benchmark performance.
- The ablations vary warm-up step counts and maximum learning rates to identify conditions that produce dragon king events.Dragon kings often arise early in pretraining while the learning rate remains high.
- Dragon kings appear as sharp loss peaks accompanied by significant accuracy drops.The study describes these effects as extreme events rather than simply ordinary optimization noise.
- Dragon king behavior occurs in both near-critical and sub-critical models.
- Large normalized RMSE by mean magnitude values are associated with reduced benchmark scores.
- Dragon kings are described as predictable and avoidable events in PLDR-LLMs and other critical systems.The paper links them to imbalance between driving impulse and dissipation force and to deviations from power-law behavior at criticality.
6 Discussion
The discussion interprets PLDR-LLM reasoning through self-organized criticality, emphasizing metastable deductive-output dynamics and their possible implications for generalization and scientific modeling. It also highlights architectural and training considerations associated with maintaining criticality.
- A metastable, input-independent deductive-output state is proposed as a way to explain why scaling model size can improve reasoning.The paper argues that larger models can capture more details of higher-dimensional symmetries in pretraining data.
- Slower parameter changes are favored because self-organized criticality requires slowly driven updates to maintain the steady state.
- The discussion proposes studying whether criticality-based representations learned from high-resource NLP data can support prediction in low-resource physical domains.Earthquake dynamics is given as a possible controlled test domain.
- Deductive outputs remain negligibly perturbed under stochastic nucleus sampling, unlike the probabilistic inductive output selection process described for SDPA-LLMs.The paper suggests that this invariance reflects representations related to scaling functions, universality classes, and renormalization groups.
- The PLDR-LLM framework is presented as a controlled computational vehicle for investigating possible criticality-related dynamics of the brain.The paper notes that further experiments are needed to support the brain-criticality hypothesis.
7 Conclusion
The conclusion reports that PLDR-LLMs reason when pretrained at criticality, with deductive outputs reaching a metastable steady state. It introduces an order parameter whose proximity to zero aligns with stronger reasoning and comprehension benchmark scores.
- PLDR-LLMs exhibit behavior similar to second-order phase transitions during training and achieve reasoning when pretrained at criticality.
- At inference, deductive outputs attain a metastable steady state that is linked to representations equivalent to scaling functions, universality classes, and renormalization groups.
- The proposed order parameter is computed from normalized RMSE by mean magnitude of global deductive-output values across stochastic-sampling runs.
- An order parameter closer to zero is observed in models with higher reasoning and comprehension benchmark scores.
- The paper presents deductive outputs as sufficient for quantifying reasoning without curated benchmark datasets, while benchmarks remain useful for evaluating inductive output.
A Mean and Std Dev for Pretrained PLDR-LLM Deductive Outputs
This appendix section reports mean and standard-deviation statistics for deductive-output quantities across pretrained models, runs 1, 2, and cached evaluations.
- Mean values of A and ALM are reported across runs 1, 2, and Cached.
- Mean values of AP and GLM are reported across runs 1, 2, and Cached.
- Standard deviation values of A and ALM are reported across runs 1, 2, and Cached.
- Standard deviation values of AP and GLM are reported across runs 1, 2, and Cached.
B RMSE and Normalized RMSE for Pretrained PLDR-LLM Deductive Outputs
Tables 12 and 13 report error metrics for deductive outputs from all pretrained models across runs 1, 2, and Cached.
- Table 12 reports RMSE values for deductive outputs of all pretrained models across runs 1, 2, and Cached.
- Table 13 reports normalized RMSE, scaled by the mean magnitude of deductive outputs, across runs 1, 2, and Cached.
C Global Density Distributions and Heatmaps for Deductive Outputs of PLDR-LLMs
Figures 4–7 present global probability-density distributions for several deductive outputs across models, while Figure 8 presents averaged decoder-layer attention heatmaps.
- Global Density Distributions: Figures 4 and 5 show probability-density distributions for A and ALM across models, using 100 buckets for main plots and insets.
- Global Density Distributions: Figures 6 and 7 show AP and GLM probability-density distributions across models, generally binned in 100 buckets.
- Global Density Distributions: The AP and GLM distributions are plotted up to ±5σ to make their main distribution characteristics easier to see.
- Global Density Distributions: The inset for ABL-SOC-110M-1 in Figure 7 uses 50 buckets rather than the 100-bucket convention.
- Heatmaps: Figure 8 shows heatmaps from the last decoder layer and a single head, averaged over all samples for every model.
D Benchmark Datasets
The benchmark suite covers grade-school reasoning, commonsense, truthfulness, scientific question answering, physical and social intelligence, and sentiment analysis.
- Reasoning and Commonsense Benchmarks: ARC contains multiple-choice grade-school questions for grades 3–9, divided into easy and challenge sets.The challenge set contains questions answered incorrectly by both a retrieval-based and a word-co-occurrence algorithm.
- Reasoning and Commonsense Benchmarks: HellaSwag is a commonsense natural-language-inference dataset designed with adversarial filtering to challenge models while remaining easy for humans.
- Reasoning and Commonsense Benchmarks: WinoGrande evaluates commonsense reasoning through pronoun-resolution problems intended to defeat statistical models relying on selectional preferences or word associations.
- Truthfulness and Scientific QA: TruthfulQA measures truthfulness across 38 categories, including health, law, finance, and politics.Its questions are selected from cases involving human false beliefs or misconceptions.
- Truthfulness and Scientific QA: OpenBookQA contains about 6000 science questions requiring the combination of provided facts with additional common knowledge.
- Physical and Social Intelligence: PIQA tests physical commonsense about concepts traditionally learned through real-world experience.
- Physical and Social Intelligence: SIQA probes emotional and social intelligence in everyday situations using 38000 multiple-choice questions.
- Sentiment Analysis: IMDB Review contains 50000 highly polarized positive and negative movie reviews for sentiment analysis.Each movie has no more than 30 reviews, with ratings grouped as negative at ≤4 and positive at ≥7 out of 10.