Source-linked AI summary

Dimensionless Controls of Plasticity Under Alternating Tasks: From Evolutionary Biology to Continual Learning

Owen Skriloff

arXiv:2608.23889v1math.OCcs.LG

TL;DR

Plasticity under alternating environments raises the question of which biological controls persist when translated to gradient-based learning. Using alternating Boolean tasks, the paper identifies two dynamical controls that govern plasticity and derives a disagreement-based heuristic for tuning reach.

  • Problem

    The paper asks which controls of biological plasticity survive when alternating-environment dynamics are translated to deep learning.

  • Method

    The authors train a neural network alternately on Boolean label sets and recast biological factors as the task disagreement r and reach ηT.

  • Results

    Plasticity is governed by r and ηT, while neutral-set size is negligible; the optimal reach follows ηT* ∝ r^-1.18.

  • Takeaways & Limitations

    The results support a dynamical rather than geometric analogy and suggest setting the optimal reach from task disagreement alone.

  • Takeaways & Limitations

    The minimal Boolean-task setting leaves finer sweeps of reach interactions with r to sharpen the estimated scaling exponent.

Abstract

from arXiv · show

Plasticity under changing environments is central to both evolutionary biology and continual learning. Motivated by recent work on genotype--phenotype maps, we study a minimal deep-learning analogue where a network is trained alternately on two Boolean label sets, and ask which biological controls of plasticity survive the translation to gradient descent. Reinterpreting four proposed biological factors as quantities of training dynamics, we find the system reduces to two dimensionless controls: the task disagreement $r$, the fraction of disagreeing labels, and the reach $ηT$, the product of learning rate and switching period. We derive two bounds on plasticity: $r$ alone fixes an extremal geometric floor on the utopia distance, while $r$ and $ηT$ jointly bound forgetting. Across 9,720 trajectories, an ANOVA confirms that $r$, $η$, and $T$ dominate, while the effect of neutral-set size (emphasized in the biological setting) is negligible. The optimal reach itself follows an approximate inverse power law $ηT^{*}\propto r^{-1.18}$, yielding a heuristic that sets the optimal reach $ηT^*$ from the task disagreement alone. The analogy that survives is therefore dynamical rather than geometric, and our setting enables a view of plasticity through the lens of other driven systems in physics and engineering.

1. Introduction

This paper translates biological plasticity under alternating environments into a minimal deep-learning setting, identifying dimensionless controls, theoretical bounds, and empirical dynamical drivers. The results support a solvable mechanics of learning connected to periodically forced systems in physics and engineering.

  • Motivation: The study translates biological plasticity research into a deep network trained alternately on two label sets, building on four proposed biological controls: neutral set size, Hamming distance, switching period, and mutation rate.The broader motivation is evidence of shared structure between continual learning and evolutionary biology.
  • Toward a mechanics of learning: Interpreting learning rate as inverse resistance and switching period as relaxation time makes ηT a dimensionless reach that measures relaxation within one phase, linking the model to periodically forced systems.The analogy is drawn to driven systems in physics, electrical engineering, and chemical engineering.
  • Core contributions: Alternating-training plasticity reduces four biological factors to two dimensionless controls: task disagreement r and reach ηT, with optimal reach ηT* ∝ r^-1.18 ≈ 1/r.The remaining factors are subdominant, so ηT* can be set from r alone.
  • Core contributions: The task disagreement r fixes a tight geometric floor on utopia distance, dU ≥ 2(1 − 2^-r), while r and ηT jointly bound forgetting, Δa1 ≲ e^r ηT M^-1.These are stated as Theorem 2 and Theorem 1, respectively.
  • Empirical validation: Across 9,720 trajectories, ANOVA finds neutral-set size negligible, whereas task disagreement, learning rate, and switching period dominate.This contrasts the static landscape quantity emphasized in biology with dynamic training quantities.

2. Related Work

Related work connects neural networks, evolutionary biology, and continual learning through shared physical, geometric, and adaptive concepts. This study builds on prior work linking plasticity under alternating environments to four biological control factors.

  • Connections across fields: Evolutionary biology and learning have been connected through genetic algorithms, shared simplicity bias, and interpretations of the central dogma as a generative model.Neural networks also emerged from biomimicry and statistical physics, including artificial-neuron and spin-glass models.
  • Continual learning and plasticity: Plasticity loss in continual learning is well documented, with research largely targeting catastrophic forgetting under task change (Kirkpatrick et al., 2017; Yu et al., 2020).Recent work also relates forgetting to loss-landscape geometry (Mirzadeh et al., 2020).
  • Landscape geometry: Fitness landscapes with genes as parameters share geometric parallels with neural-network loss landscapes defined by weights and biases, including connected minima, flatness, and generalization.These parallels recall large connected neutral sets in genotype–phenotype maps that produce the same optimal phenotype.
  • Biological controls of plasticity: García-Galindo and Ahnert (2025) found that plasticity emerges under alternating environments and identified neutral-set size, Hamming distance, switching period, and mutation rate as four controlling factors.The present study tests whether this biological analogy transfers to continual learning.

3. Notation and Key Quantities

This section defines performance proxies and two plasticity metrics for alternating training, then identifies task disagreement r and reach ηT as the key dimensionless controls. It also introduces Type II ANOVA to compare the influence of training and task factors.

  • Plasticity: Plasticity is defined as sustained learning in continual learning, whereas the evolutionary usage denotes simultaneous performance across environments or rapid adaptation to change.
  • Accuracy proxies: Accuracy proxies aj(t) = e^-ℓj(t) map task losses to bounded performance values in (0, 1], placing alternating-training trajectories in the unit square.
  • Utopia distance: The utopia distance dU averages Euclidean distance from (a1, a2) to (1, 1) over the final two switching periods, with lower dU indicating higher plasticity.It penalizes trajectories that excel at one task while failing to sustain performance on the other.
  • Forgetting amplitude: Forgetting amplitude Δa1 measures the peak-to-peak loss in task-1 performance after switching to task 2, paralleling catastrophic forgetting and an evolutionary robustness measure.
  • Dimensionless controls: The two dimensionless controls are task disagreement r = dH/N, the differing-label fraction, and reach ηT, the parameter-space travel during one training phase.Here η is the learning rate and T is the switching period.
  • Analysis of variance: Type II ANOVA uses ω2 to quantify each factor’s explained outcome variance, enabling comparisons among r, η, T, neutral-set size, and complexity.

4. Experimental Setup

The experiments approximate 7-bit Boolean functions with a neural network and systematically vary neutral-set size, complexity, Hamming distance, learning rate, and switching period. Alternating SGD across these conditions produces 9,720 trajectories for measuring utopia distance and forgetting.

  • Function-pair design: The study selects 27 function pairs spanning low, medium, and high neutral-set size, complexity, and Hamming distance in a 3 × 3 × 3 grid.Pairs are sampled within three complexity bins to decouple complexity from neutral-set size; each pair has total neutral-set size ν = P(f1) + P(f2).
  • Training protocol: Each experiment pretrains on one function, alternates SGD between two label sets every T steps, and continues until reaching steady state.The network approximates 7-bit Boolean functions, with architecture and training details provided in Appendix A, Table 2.
  • Measurements: Sweeping six learning rates and six switching periods with ten Monte Carlo trials per setting yields 9,720 trajectories, each evaluated using dU and Δa1.Neutral-set sizes are estimated by Monte Carlo sampling from a Gaussian prior p(θ).

5. Theory

The theory bounds forgetting through the task disagreement r and reach ηT, while a separate theorem gives a tight lower bound on the utopia distance depending only on r. The forgetting bound is primarily a scaling result because one assumption is close but not universal and its forcing estimate is loose.

  • Forgetting bound: For trajectories on a period-2T steady-state orbit with monotonic y2-phase decay and smooth loss/network dynamics, Theorem 1 upper-bounds forgetting as Δa1 ≤ a1(0)(e^(rηTM+O(η^2T))−1).For sufficiently small η, the bound becomes approximately a1(0)(e^(rηTM)−1).
  • Forgetting bound: Theorem 1 is mainly useful for scaling: monotonic decrease held in 97.5% of 56,748 steady-state y2 phases, while the estimate M ≤ LC is loose.The bound’s empirical scaling was checked using the uniform proxy M = 100.
  • Utopia-distance bound: Theorem 2 derives a tight lower bound on the utopia distance dU that depends only on task disagreement r, with both theorem bounds demonstrated across 9,720 trajectories.Figure 2 uses the uniform estimate M = 100 for Theorem 1.

6. Results

Results show that task disagreement r sets the geometric floor for utopia distance, while reach ηT determines where trajectories sit relative to that floor. Across experiments, ηT* decreases with r approximately as r^-1.18, whereas neutral-set size and complexity have negligible practical effects.

  • ANOVA results: The Type II ANOVA confirms that η and T dominate Δa1, whereas r dominates dU, matching the dependencies predicted by Theorems 1 and 2.The bounds capture the dominant dependence of each plasticity outcome on the controls.
  • Control factors: Neutral-set size ν and complexity K̃ have statistically significant but negligible effect sizes, unlike their stronger influence in biological systems.This indicates a structural difference between evolution on genotype–phenotype fitness landscapes and learning on neural-network loss landscapes.
  • Optimal reach: ηT* ∝ r^-1.18 across 37 task pairs, with 95% CI [−1.34, −1.01] and R2 = 0.86.As r increases, the minimum dU rises while the optimal reach decreases and the curves become more convex.
  • Robustness: The reach-governed picture persists under minibatch SGD, momentum, Adam, held-out continuous-input tasks, and 16× variation in parameter count.For held-out tasks, the mean train–test gap in dU is below 0.01, while varying width and depth leaves dU essentially unchanged.

7. Discussion

The discussion identifies reach ηT and task disagreement r as transferable dynamical controls linking deep learning with physics, evolutionary biology, and continual learning. It also explains why neutral-set geometry diverges across substrates and proposes theory-guided hyperparameter selection for alternating tasks.

  • Plasticity across substrates: Across evolutionary and deep-learning substrates, Hamming distance, learning rate, and switching period control plasticity, whereas neutral-set size is negligible in learning.A proposed explanation is that populations sample neutral-set volume through many trajectories, while one SGD run follows a locally forced point; ensembles or strong noise may restore its role.
  • Reach as a physical control parameter: Reach ηT is interpreted as distance traveled during one task phase: η acts as inverse curvature-based resistance, while T is the relaxation time.Under deterministic square-wave forcing, linearized response scales like tanh(x), and Theorem 1 bounds this response, leaving room for tighter linear-regime bounds.
  • Reach as a physical control parameter: The optimal reach follows ηT*∝r^-1.18, suggesting a near-inverse universal scaling that estimates switching periods from task disagreement and curvature-based learning-rate estimates.This heuristic avoids probing the loss landscape or architecture and, with the two theorems, reduces plasticity outcomes to dimensionless controls.
  • Connections to continual learning: The results connect curvature-weighted displacement and finite per-phase learning to continual-learning mechanisms including forgetting analyses, EWC, and CLEAR.Because optimal reach under dU is nonzero, alternating disagreeing tasks require finite learning in each phase; the heuristic may estimate CLEAR hyperparameters theoretically.
  • Toward a mechanics of learning: Together, the findings support four mechanics-of-learning strands: solvable plasticity, governing laws, two dimensionless controls, and substrate links across biology, chemistry, and learning.The discussion frames these links as evidence for substrate universality while retaining geometric differences between genes and network weights.

8. Conclusion and Future Work · Appendix A. Computational Experiment Details · Appendix B. Proofs for Theoretical Results

The study identifies task disagreement r and reach ηT as the core controls of plasticity under alternating Boolean tasks, with theoretical bounds and empirical dominance over static landscape quantities. Future work targets disagreement-matched reach selection, gradient-flow non-commutativity, and extensions to generalization and broader task settings.

  • 8. Conclusion and Future Work: Plasticity under alternating Boolean tasks is governed by two dimensionless controls: task disagreement r and reach ηT.Theoretical results show that r sets a tight geometric floor on utopia distance, while r and ηT jointly bound forgetting.
  • 8. Conclusion and Future Work: An ANOVA found that r and ηT dominate empirical plasticity, whereas static landscape quantities emphasized in the biological setting do not.
  • 8. Conclusion and Future Work: The authors conjecture that setting ηT ≈1/r provides near-optimal conditions for alternating-task continual learning.The proposed regime uses phases short enough to avoid overfitting the current task but long enough to sustain learning; finer sweeps could sharpen the exponent.
  • 8. Conclusion and Future Work: Future analysis should relate reach, forgetting amplitude, and the Lie bracket [∇ℓ1, ∇ℓ2] governing non-commuting gradient flows.Because the bracket vanishes for commuting flows, this direction recovers the r →0 limit of Theorem 1 and could extend the framework to general losses and continuous outputs.
  • 8. Conclusion and Future Work: The framework may extend to generalization, with ∆a1 and dU serving as potential measures of robustness to distribution shift and noisy labels.This motivation connects forgetting to performance loss under task distribution shift and to out-of-distribution forgetting under small intra-class shifts.
  • Appendix A. Computational Experiment Details: Appendix A documents the architecture and training hyperparameters used in the computational experiments.

B.1. Proof of Theorem 1 … Appendix C. Choice of Accuracy Proxy

The appendices prove the dynamical and geometric bounds, extend the forgetting bound to EWC, and test whether conclusions depend on the accuracy proxy. The results establish explicit conditions for the bounds and show that key empirical trends persist across proxies.

  • B.1. Proof of Theorem 1: ς ≤ rηTM + O(η^2T) bounds the loss change during the second task phase, completing Theorem 1 under η < 2/L and a uniform Jacobian bound.The proof uses monotonic loss and accuracy behavior, gradient disagreement on differing labels, Cauchy–Schwarz, and periodicity.
  • B.2. Proof of Theorem 2: For disagreeing labels, the term-wise BCE losses satisfy ℓ1,i + ℓ2,i ≥ ln 4, yielding the total-loss constraint used in Theorem 2.Agreement points contribute a nonnegative sum, while disagreement points produce the logarithmic lower bound.
  • B.2. Proof of Theorem 2: Theorem 2’s utopia-distance floor follows from the pointwise constraint a1a2 ≤ c^2, with the minimum attained symmetrically at a1 = a2 = √c^2.The optimization reduces to the hyperbola a1a2 = c^2; the alternative stationary case is ruled out for 0 < r < 1.
  • B.3. EWC Corollary: Under diagonal Fisher information, a near-optimal quadratic regime, sufficiently small η, and Theorem 1’s assumptions, EWC modifies the forgetting bound through λ.The BCE result is recovered at λ = 0, while rM ≥ E(1 + λ) ensures the resulting exponent is nonnegative and the bound is meaningful.
  • B.3. EWC Corollary: The EWC penalty lowers forgetting because the Fisher-weighted displacement term is approximately aligned with the current BCE gradient and contributes a nonnegative inner product.This alignment follows from the approximate equality between the BCE Hessian and diagonal Fisher information near the optimum.
  • Appendix C. Choice of Accuracy Proxy: Across 7,212 runs, four accuracy proxies rank-correlate strongly, while the optimal-reach power-law trend replicates for exponential and reciprocal proxies with R^2 ≈ 0.97–0.99.Theorem 2’s floor is asserted only for a = e^-ℓ, so proxy robustness does not extend that geometric claim.

Appendix D. Theorem 1 Monotonicity Audit · Appendix E. Robustness Experiments · E.1. Held-out generalization on continuous inputs

Theorem 1’s monotonicity assumption is strongly supported across the Boolean reach grid, while continuous-input experiments show that plasticity geometry generalizes beyond training truth tables with a small train–test gap.

  • Appendix D. Theorem 1 Monotonicity Audit: 97.5% of 56,748 steady-state y2 phases were strictly decreasing, while 98.8% satisfied the weaker endpoint condition required by Theorem 1’s bound.The median normalized positive variation was 0.
  • Appendix E. Robustness Experiments: Continuous robustness experiments replaced Boolean truth-table pairs with linear-separator tasks on isotropic Gaussian inputs in R10.Two unit weight vectors were placed at angle θ = πr, making expected label disagreement exactly r.
  • Appendix D. Theorem 1 Monotonicity Audit: The normalized positive-variation distribution was sharply concentrated at 0, with a heterogeneous right tail.
  • E.1. Held-out generalization on continuous inputs: Over 454 continuous-input runs, the train–test gap in dU was small: mean |∆| = 0.007, median 0.004, and maximum 0.064.This supports the conclusion that the plasticity geometry was not a training-set memorization artifact.
  • E.1. Held-out generalization on continuous inputs: The continuous-input evaluation used independent train, validation, and test sets containing 4096, 2048, and 4096 examples, respectively.
  • E.1. Held-out generalization on continuous inputs: The held-out grid was small, and some small-model pretraining runs did not fully converge.

E.2. Optimizer and architecture

The reach-governed picture persists across minibatch SGD, momentum, and Adam, while varying architecture size leaves steady-state dU essentially unchanged. Reach curves, including optimum location and depth, are preserved across a 16× parameter range.

  • Optimizer and architecture: Across minibatch SGD, momentum, Adam, and architectures spanning 3,641 to 59,041 parameters, steady-state dU remains essentially unchanged.This supports the network-size independence assumed by the architecture-agnostic bounds, though those bounds do not prove it.
  • Optimizer and architecture: The reach-governed picture persists under minibatch SGD, momentum, and Adam.Table 5 holds the reach sweep fixed while varying optimizer and architecture across 9 task pairs and 5 seeds.
  • Optimizer and architecture: Figure 5 shows dU reach curves with preserved optimum location and depth across a 16× parameter range.The comparison covers three architectures.
Loading 2608.23889v1…