Source-linked AI summary

Functional compatibility as a determinant of persistent neural learning

Hossein Javidnia

arXiv:2608.22462v2cs.LGcs.AI

TL;DR

Whether new learning can persist without damaging existing behaviour remains unclear. The paper experimentally varies functional compatibility from identical neural states while matching learning opportunity and retention requirements, finding that persistent learning increases with compatibility across tasks, architectures, modalities, and seeds.

  • Problem

    Continual learning lacks a clear account of what determines whether newly acquired information persists while existing behaviour is preserved.

  • Method

    The study intervenes on functional compatibility from identical neural states while matching learning opportunity and retention requirements, alongside AFM-based measurement and multi-layered stress tests.

  • Results

    Persistent learning increases with functional compatibility across independent directions, convolutional and transformer architectures, vision and text, and a ten-seed replication.

  • Takeaways & Limitations

    Functional compatibility controls the local persistent-learning frontier, while retention constraints determine how much compatible opportunity is retained and nonlinear geometry limits its extension.

  • Takeaways & Limitations

    The experiments evaluate continual adaptation over a shared learned representation rather than unrestricted end-to-end continual representation learning, and do not establish universal empirical dominance.

Abstract

from arXiv · show

Neural networks can acquire new capabilities while damaging existing ones, but what determines whether new learning persists remains unclear. We identify functional compatibility, the extent to which incoming learning can coexist with behaviour that must be preserved, as an experimentally manipulable causal determinant of persistence. From identical neural states, we vary compatibility while matching unrestricted learning opportunity and imposing a common retention requirement. Persistent learning increases with compatibility across independent directions, convolutional and transformer architectures, vision and text, and a ten-seed replication. Learning rules and retention constraints determine how much compatible opportunity is retained, whereas nonlinear geometry limits the matched intervention at larger update norms. Functional compatibility therefore reframes stability-plasticity from preventing forgetting to determining which new learning can coexist with existing function and persist.

1 Introduction

The introduction reframes continual-learning stability around functional compatibility: whether incoming learning can coexist with protected behaviour. It predicts that changing compatibility from an identical neural state will causally change persistent learning under the same retention requirement.

  • Motivation: Continual learning is framed as a competition between acquiring new information and preserving existing knowledge, motivating methods that reduce interference through parameter, data, or update constraints.These approaches include regularization, replay, constrained gradients, and function-space protection.
  • Functional compatibility: Functional compatibility measures how much incoming learning can coexist with behaviour that must be preserved.The paper introduces this property as the target of its analysis.
  • Functional compatibility: Compatibility is assessed at the pre-update neural state from the overlap between current active learning geometry and directions compatible with protected behaviour.The measure is explicitly distinguished from a post-hoc forgetting score.
  • Method: Adaptive Functional Metaplasticity provides a constructive framework for quantifying and exploiting functional compatibility.The framework compares learning paths involving compatible projected updates with a same-state no-protection endpoint.
  • Causal prediction: The central prediction is that deliberately changing compatibility from the same neural state will change how much new learning persists under the same retention requirement.This tests whether compatibility is causal rather than merely a retrospective correlate of interference.
  • Novelty: The study presents itself as the first controlled causal investigation to intervene on a quantified functional-compatibility variable while matching neural state and separating the intervention from retention requirements.The introduction distinguishes this design from prior work on gradient alignment, protected subspaces, and interference-related update geometry.

2 AFM framework and theoretical foundations … 2.3 Theoretical guarantees

AFM separates persistent learning from finite deployed protection and population behavior, then uses compatible projected updates, certified residual completion, and task-free mechanisms to preserve declared function. Its guarantees establish normalized persistent assimilation and exact finite counterfactual completion while making capacity, observability, compatibility, and finite-rank leakage explicit limits.

  • 2.1 Problem formulation and design principle: AFM distinguishes persistent parameter adaptation, finite exact output protection, and population behavior, which requires additional assumptions or certificates beyond finite evidence.This separation prevents finite interpolation from being mistaken for learning stored in the persistent base model.
  • 2.1.1 Same-state no-protection comparator: The same-state no-protection comparator measures how much learning the unrestricted accepted transaction could contribute, rather than comparing protection with a nominal gradient or arbitrary schedule.The comparator matches parameters, state, minibatch, active coordinates, optimization rules, buffers, and declared randomness.
  • 2.1.2 Persistent compatible assimilation: AFM allocates a requested assimilation fraction ηt along a compatible projected comparator path, while κt measures the active gradient energy compatible with the protected geometry.A persistent compatible step generally cannot reproduce the unrestricted current endpoint; finite endpoint completion is therefore handled separately in function space.
  • 2.1.3 Finite endpoint completion: Finite residual completion restores requested logits or protected outputs only on certified finite support, while incompatible requests, exhausted capacity, failed certificates, or conflicting addresses cause atomic rejection.Away from finite support, the deployed predictor follows the protected persistent base endpoint.
  • 2.2 Adaptive Functional Metaplasticity: AFM combines sensitivity sketches, multiscale metaplastic traces, task-free routing and consolidation, evidence-based reopening, and function-preserving structural renewal within bounded resources.Dormant zero-gated modules can be reset without changing deployed predictions or active protected behaviors, then activated only after certified nonzero motion.
  • 2.2.5 Protected update: A protected round routes observations, updates private candidate and memory states, computes unrestricted and compatible references, accepts only certified persistent motion, and then applies finite completion and separately controlled evidence updates.Failed certificates never authorize installing the unrestricted base endpoint or reducing the declared persistent-assimilation requirement.
  • 2.3 Theoretical guarantees: Under the theorem’s projector, smoothness, and retention-charge conditions, normalized persistent assimilation is guaranteed, and ηt = 1 makes the complete projected comparator feasible.The guarantee is not generally an ηt fraction of unrestricted loss decrease because κt is unavoidable in the selected protected geometry.
  • 2.3.1 Retention, stationarity, and local optimality: The integrated guarantees provide exact equality on declared finite protected evidence and a sharp first-order leakage floor of σr+1(J) for any (d−r)-dimensional plastic subspace, while retaining explicit broader-behavior mismatch charges.The framework is obstruction-aware: identical observable laws prevent exact semantic routing, finite-rank protection cannot remove all leakage under nonzero movement, and nonlinear stationarity differs from the separate convex dynamic-regret specialization.

2.4 AFM framework synthesis · 2.5 Full AFM specification and finite endpoint construction

AFM combines persistent compatible adaptation with an exact finite deployed endpoint, using explicit certificates, residual correction, and rollback conditions. Its specification fixes the measurable routing, retention, certification, shielding, and resource rules governing protected updates.

  • 2.4 AFM framework synthesis: AFM treats continual learning as a certified frontier between persistent compatible adaptation and exact finite deployed completion.Protected updates reference the genuine same-state no-protection endpoint, while compatible motion is retained under an explicit retention charge and finite residual correction.
  • 2.4 AFM framework synthesis: Across CORe50, CLEAR-10, and CLAD-C, the evaluated AFM frontier improves on the strongest tested classical non-oracle family and remains competitive with six modern challengers.Frozen five-seed protocols and a non-vacuous execution audit support the reported comparison and endpoint transaction.
  • 2.5.1 Admissible specification: The full AFM specification fixes every object affecting protected decisions before the protected horizon, including routing, trust, certification, shielding, renewal, and resource policies.The representation is frozen after any optional unprotected initialization prefix, and components must be measurable with respect to learner history.
  • 2.5.2 Losses and retained behaviors: At each round, AFM processes arbitrary data through differentiable losses and fixed within-batch routing, updating consolidation and refinement statistics at their declared resolutions.The learner observes a data item or finite ordered minibatch, with every member routed in fixed order.
  • 2.5.2 Losses and retained behaviors: Before any nonzero update, AFM requires finite trust and curvature certificates; absent a valid certificate, it sets Rcap_t = 0 and makes no update.Finite computation graphs can yield certificates through interval automatic differentiation and recursive chain, product, and composition rules.
  • 2.5.3 Task-free routing and semantic mismatch: Task-free routing protects router-assigned behaviors, while explicit mismatch terms quantify disagreement with external evaluator semantics and vanish when the semantics agree.The mismatch terms are not assumed small, preventing shield replacement from being hidden inside a parameter-only routing bound.
  • 2.5.4 Streamed whole-behavior sensitivity memory: AFM stores complete frozen commit-block behavior through streamed Jacobian memory and charges the measured deployed activation-transfer gap explicitly.The gap is zero only when the validated behavior is already deployed; otherwise ordinary protected updates must reduce it below a predeclared threshold.
  • 2.5.5 Guarded compact-cardinal shielding of certified candidates: The compact-cardinal transaction preserves active finite deployed behaviors exactly, restores current-minibatch comparator logits, and completes functionally consistent finite transfers without a compatibility-rate assumption.With ξt = 1, the selected finite transfer objective is zero after one accepted service, while failures restore predictive components and emit typed obstructions.

2.6 Task-free routing, consolidation, and reopening

The task-free controller routes observations using frozen, separated signatures, consolidates records with anytime-valid risk certificates, and reopens or splits them under separately controlled evidence. These guarantees remain conditional on observable information and preserve existing retention protections when permitted splits occur.

  • Task-free routing: Under separated observable contexts, convex-EMA centroids and nearest-threshold routing assign every subsequent signature to its unique correct context slot without task labels.The proof keeps each centroid within its context’s signature ball, yielding within-context distance at most h_Z and cross-context distance greater than h_Z.
  • Frozen-signature calibration: Calibration must replay the declared initialization prefix through the final frozen representation before fixing signature statistics and routing thresholds.If a requested threshold exceeds the operational ceiling, clipping is allowed only while withholding the positive routing conclusion; retention and reopening results remain valid for the actual routed sequence.
  • Finite signature selection: A predeclared finite signature family can provide error-controlled task-free route identification, with a positive expert incurring only the fixed evidence overhead log(1/w_m⋆).If no observable expert has positive information, AFM makes no finite route-identification claim.
  • Operational consolidation: Records are committed only when their anytime-valid evidence certifies average conditional validation risk at most τ, without requiring stationarity or iid data.The transfer target becomes immutable after certification, while a separately controlled staleness process governs later cancellation of the candidate.
  • Reopening and retention: Outcome and observable evidence independently control splitting and reopening errors, while permitted splits leave source records’ existing retention guarantees unchanged until explicit release.Finite route-splitting delay requires positive observable signature drift; semantic-only change can justify reopening from outcomes but cannot justify a new route when the signature law is stable.

2.7 Whole-behavior memory and spectral protection

This section establishes deterministic and population certificates for bounded sensitivity sketches, including adaptive commitment times. It also identifies the exact spectral rank-plasticity frontier and the unavoidable first-order forgetting price.

  • Streamed-memory certification: Theorem 2.22 gives a deterministic streamed-memory certificate for every round and every unit vector v ∈ range(Πt).The result quantifies leakage represented by bounded sensitivity sketches and characterizes the exact first-order rank-plasticity frontier.
  • Population certification: Corollary 2.23 provides a constructive population certificate under iid commit inputs and operator-norm bounds, including data-dependent commitment times.The population term is constructible under declared sampling conditions rather than assumed to vanish; simultaneous bounds remain valid at adaptive commitment times.
  • Spectral protection: For compact J with singular values σ1 ≥ · · · ≥ σd ≥ 0, the exact rank-r optimum is the span of right singular vectors associated with σr+1, . . . , σd.The plastic subspace has dimension d − r.
  • Spectral protection: No rank-r linear protection rule can guarantee first-order leakage below σr+1(J) in every unit plastic direction.Thus, the spectral tail in the AFM certificate is not merely a proof artifact.

2.8 Retention and persistent adaptation

This section establishes exact finite protected retention, decomposes activation and release effects, and derives conditions under which compatible safe-base movement can persist. It also separates shield-emulated deployed progress from the base stationarity analysis.

  • One-step and active-interval retention: Accepted protected updates preserve declared finite protected evidence exactly, while broader deployed behavior is controlled by a joint retention bound.The finite protected map has zero deployed drift on each declared protected-evidence block, with only the stated numerical envelope applying to the stronger identity.
  • Activation and release decomposition: Activation transfer, exact active-interval retention, and post-release drift are reported as separate terms rather than combined as protected forgetting.The final post-release term is intentionally uncontrolled by a bounded registry, while delayed activation and capacity release remain separately identified.
  • Projected stationarity: The projected-stationarity identity holds pathwise for smooth non-convex losses, while every accepted deployed update separately achieves current-batch ratio one.Shield-emulated deployed progress is maintained as a separate exact ledger and is not substituted into the base stationarity proof.
  • Persistent compatible movement: Under finite cumulative safe-base behavior charge, exact finite deployed retention and ratio-one current endpoint emulation coexist with indefinitely available compatible persistent movement.A broader joint behavior additionally requires summable structural and routing charges for finite total drift.

2.9 Multi-timescale metaplasticity and allocation

The section establishes that logarithmic EMA trace banks approximate scale-free relevance with logarithmic memory, while allocation balances retention leakage against blocked plastic gradient energy. It further shows that accurate prediction enables near-oracle rank allocation and that exponential-weights control achieves high-probability regret guarantees.

  • Trace-bank approximation: K = O(log H) EMA traces approximate a power-law relevance profile over ages 1, . . . , H with constant multiplicative distortion.This follows from Theorem 2.32’s logarithmic trace-bank approximation result.
  • Trace-bank approximation: Every single exponential incurs polynomial worst-case distortion when approximating the power-law profile f(τ) = τ −α.Theorem 2.33 formalizes this limitation for odd horizons H ≥3.
  • Adaptive allocation: The allocation frontier jointly penalizes retention leakage and plastic gradient energy blocked by protection, so larger rank cannot dominate simply by freezing more coordinates.The cost explicitly couples preservation with continued adaptation.
  • Adaptive allocation: Accurate prediction of realized retention weights yields near-oracle functional leakage, and exact prediction makes rank-r allocation oracle optimal for that round.The result compares the policy’s selected projector with the oracle top-r projector.
  • Adaptive control: The timescale/rank controller is coupled simultaneously to retention and compatible adaptation through frontier costs containing spectral residual and blocked gradient fraction.This coupling prevents timescale and rank selection from being analyzed as separate objectives.

Give atom a the explicit realized signal

A PSD record can be decomposed into finitely many certified atoms with individual traces and policy weights, whose weighted sums define predicted and realized covariances. This atomization preserves controller guarantees, but strengthens memory certification only when atom-specific sensitivity and error certificates are supplied.

  • Atom construction: Each finite PSD atom carries its own K traces and policy weight, and nonnegative weighted sums form predicted and realized covariances.The construction replaces record-level PSD contributions with atom-level weighted contributions.
  • Memory certification: Atom-specific memory certification remains valid when each atom provides a certified PSD sensitivity/error decomposition whose weighted sum bounds the original derivative, sketch-error, and anchor-transport terms.The required bound covers the same conceptual whole-behavior derivative energy, Frequent-Directions error, and anchor-transport terms used in Theorem 2.22.
  • Memory certification: Algebraic atomization alone does not strengthen the derivative certificate; without atom-specific certificates, the original record-level certificate remains in force.Atoms may still be used by the allocation controller without independently weighted memory guarantees.
  • State representation: Persistent state scales with the fixed number of retained atoms rather than elapsed lifetime, including atoms defined as sketch eigen-directions, parameter groups, or rank-one synaptic sensitivities.These are examples of certified atoms supported by the construction.
  • Controller guarantees: Controller proofs transfer verbatim after relabeling records as atoms because they use only nonnegative weighted PSD sums, trace residuals, operator-norm perturbations, and bounded expert losses.The preserved arguments include trace, spectral-cost, allocation-transfer, and Hedge results.

2.10 Function-preserving structural renewal

Function-preserving structural renewal resets dormant modules without changing predictions, loss, or protected behaviors, then certifies activation through the ordinary compatible update. Renewal success is governed by realized conditional richness, with zero richness making success impossible.

  • Exact reset and certified activation: A zero functional gate removes a module from the realized predictor and every active protected behavior map, defining functional dormancy.The gate is applied to both the module contribution and each protected behavior map.
  • Exact reset and certified activation: Resetting any internal parameter under a zero gate changes the predictor, current loss, and protected behaviors by exactly zero.After reset, the certificates and full gradient are recomputed before proposing the update.
  • Exact reset and certified activation: Accepted positive-step renewal produces nonzero gate motion, making activation functionally nonvacuous and covered by the same behavioral and loss certificates.If no accepted motion occurs, rolling back restores the exact pretrial state.
  • Richness-controlled renewal: Renewal success is controlled by actual conditional richness: once cumulative richness reaches log(1/δe), failure probability is at most δe.The guarantee uses fresh conditional randomness and does not require a known lower bound on richness.
  • Richness-controlled renewal: If every richness value is zero, no sampling algorithm using that renewal family can succeed.Under an optional isotropic projector model, h > 0 and ω > 0 imply explicit positive richness.
  • Resource use: Failed or unactivated trials reuse zero-gated slots, while exhausted dormant capacity triggers a structural-capacity obstruction rather than additional parameter allocation.Successful activations retain their slots until a separately certified move restores exact zero-gated dormancy.

2.11 Integrated frontier theorem and convex specialization

The section establishes a convex dynamic-regret specialization and an integrated adaptive frontier theorem covering retention, adaptation, routing, renewal, timescales, evidence, and resource guarantees. Its guarantees are pathwise when certificates accept, probabilistic at probability at least 1 −δ, and explicit about compatibility and oracle obstructions.

  • Integrated frontier theorem: Theorem 2.42 integrates pathwise, statistical, resource, and obstruction-aware guarantees across routing, retention, non-convex adaptation, renewal, evidence, and timescale control.The theorem applies conditionally on accepted certificates and combines the listed frontier components into one maximal adaptive functional metaplasticity statement.
  • Convex specialization: Under Assumption 2.40, convex differentiable losses, bounded gradients, nonempty convex retention sets, and a feasible initial point support the convex branch.The compact convex domain K has diameter D, each loss gradient is bounded by G, and projected iterates remain retention feasible.
  • Convex specialization: Theorem 2.41 gives dynamic regret with an unavoidable compatibility price, whose first term no retention-feasible learner can remove.The convex branch provides a global dynamic-regret statement when an exact projection oracle is available.
  • Integrated frontier theorem: With probability at least 1 −δ, all probabilistic conclusions hold simultaneously, while pathwise conclusions hold on every realized trajectory.The failure budget is allocated across consolidation, reopening, route splitting, calibration, population certificates, randomization, refresh, append events, and renewal episodes.
  • Integrated frontier theorem: Accepted protected rounds preserve the safe base, satisfy projected stationarity, retain a compatibility-weighted current-loss fraction, and reproduce one hundred percent of the genuine no-protection current endpoint.The theorem also states that exact finite deployed protection and lifelong compatible base plasticity can coexist on tails covered by Corollary 2.30.
  • Integrated frontier theorem: Structural renewal preserves the realized predictor and active protected behaviors, while exhausted capacity, zero richness, uncertified feasibility, and oracle failure are reported as explicit obstructions.The theorem also provides bounded structural resources, including W = O(d) + O(Lmaxd) + O(Cmaxℓcandd) + O((Jmax + Cmax)d) + O(CmaxdZ).

2.12 Obstructions and maximality

The section establishes that each obstruction term in the universal theorem is necessary: deleting any associated term falsifies a corresponding guarantee. It therefore proves componentwise maximality of the theorem’s structure and quantifiers, but not sharpness of numerical constants or minimax optimality on restricted subclasses.

  • Routing and consolidation: Routing unidentifiability makes a nonzero routing-mismatch term or outcome-only release unavoidable when semantic contexts are observationally indistinguishable.No task-free router can identify the correct context below error probability 1/2 under an equal prior.
  • Routing and consolidation: Finite commit histories cannot provide distribution-free future-risk guarantees without recurrence, exchangeability, drift, or another predictive assumption.Two data streams can be identical through commitment yet yield valid or arbitrarily poor future predictor risk.
  • Geometric and compatibility obstructions: Curvature, anchor drift, and compatibility impose unavoidable costs on functional preservation and adaptation under nonzero movement or directly incompatible objectives.The curvature obstruction creates change H ∥d∥2 /2 along d ∈null(J), while compatibility lower bounds learner regret by Γt on round t.
  • Information and memory limits: Zero observational information, zero compatible richness, and bounded persistent state make reliable reopening, renewal success, and exact recall impossible in their respective settings.The results also include the finite-delay reopening obstruction at I = 0 and the inability of B persistent bits to recall every sequence of T > B independent unbiased bits.
  • Maximality: Componentwise maximality follows because each deleted obstruction term is refuted by a construction in the universal stream class, without establishing sharp constants or restricted-class minimax optimality.The corollary covers routing, consolidation, leakage, curvature, anchor drift, compatibility, reopening, richness, memory, and infinite-horizon testing.

3 Controlled causal compatibility experiment

The controlled within-state experiment shows that functional compatibility causally increases retention-constrained persistent learning, with an approximately unit slope for the projected compatible proposal across vision and text systems. Retention budgets and learning rules determine how much compatible geometry is used, while the causal probe is deliberately local and does not establish the same slope for large updates.

  • Causal intervention: The six interventions share the same pre-update state, protected evidence, and current inputs while varying only the current learning-signal direction.This within-state construction is the central causal control.
  • Causal intervention: 0.27% to 0.28%: average mismatch across interventions, never exceeding 0.40% in any system, indicating essentially matched finite unrestricted progress.This matching supports interpreting persistent-progress differences as consequences of the controlled directions rather than unequal unrestricted opportunity.
  • Compatibility effect: 0.993: the projected compatible proposal’s cross-system mean slope over 15 system-seed slopes, with a 95% t interval of [0.980, 1.006].All five seed-level slopes were positive in each of the three systems, and every system-level 95% interval lay above zero.
  • Retention interaction: For every positive β, the matched compatibility slope is positive in all three systems, while the text transformer requires a more permissive retention budget before approaching unit conversion.The zero-budget condition is a sanity check with zero persistent progress and zero compatibility slope; the two CIFAR-10 systems approach unit slope by β = 0.05 to 0.10.
  • Natural-state validation: Natural validation shows that compatibility alone is insufficient: the text system has approximately 0.95 natural compatibility, mean native AFM path fraction 0.107, and median persistent ratio approximately 0.119.The joint interpretation is that persistent progress depends on available compatibility multiplied by the accepted retention-constrained path.
  • Scope and limitations: The causal probe is local because the weakest unrestricted direction sets a small common positive learning target, so its normalized slope need not hold at arbitrary finite step size.Natural-state validation is complementary, operating at ordinary stream-generated decreases that are much larger, particularly for vision.

4 Generality, natural-state behaviour, and finite-scale boundary

Across architectures, modalities, directions, and retention budgets, greater functional compatibility generally increased persistent learning, with matched slopes near one when retention budgets were nontrivial. Natural-state compatibility also emerged spontaneously, but finite nonlinear geometry and retention limits constrained how much compatible learning could be retained.

  • Generality: Matched compatibility slopes were near one across four systems at β = 0.5, including 1.015 for the convolutional network, 1.001 for the baseline vision transformer, 0.999 for the stronger vision transformer, and 0.960 for the text transformer.The equal-system pooled slope was 0.994, and all four AFM/projection slopes were strongly positive and near one.
  • Natural-state behaviour: Across 750 natural states, mean κ was 0.228 in the convolutional network, 0.422 in the vision transformer, and 0.950 in the text transformer, yet the text model accepted a mean persistent path fraction of only 0.107.Native AFM satisfied the common recorded retention criterion at all 750 states, while the text model’s median persistent-progress ratio was 0.119.
  • Retention budget: At β = 0.01, CNN matched slopes rose from 0.10272 at s = 0.05 to 0.96935 at s = 0.9, whereas at β = 0.05 they remained near one across the same range.The corresponding β = 0.05 slopes were 1.01171, 1.01032, 1.00359, and 0.99390, showing that retention budget limits attainable persistent progress.
  • Generality: At β = 0.5, 95% seed-level confidence intervals for matched slopes were [1.03046, 1.06024] in the CNN, [1.02516, 1.03442] in the ViT, and [0.97849, 0.99454] in the text transformer.The mean fraction of states with within-κ direction SD below between-κ variation was 1.0 for all three systems.
  • Finite-scale boundary: The fixed-norm bridge was feasible for 231/500 ViT states at 1% of natural update norm, but for neither architecture at 10%, 50%, or 100%.Among feasible ViT states, mean within-state ∆0 coefficient of variation was 0.22586 and maximum was 0.58485, reflecting finite nonlinear variation across compatibility directions.

5 Experimental methods and reproducibility

The methods combine three linked experimental layers with controlled compatibility interventions, matched unrestricted learning opportunity, and reproducibility checks across architectures, modalities, seeds, and validation settings. AFM implements protected functional updates through comparator-based projection, certified assimilation, finite output correction, and function-preserving structural renewal.

  • Experimental design: Three linked layers comprise chronological AFM evaluation, controlled causal compatibility intervention, and stress tests of direction, seed, architecture, scale, and natural validation.The causal layer uses fully unfrozen networks, while follow-up tests assess generality and intervention properties.
  • Experimental design: The controlled suite uses CIFAR-10 convolutional and vision-transformer models plus a WikiText-2 character-level transformer, with 50 causal states across five seeds per system.A saved parent checkpoint preserves model, optimizer, stream, replay, and random-number-generator state for reproducible interventions.
  • Compatibility intervention: Compatibility is manipulated in function space by mixing generalized-eigenproblem residual modes, targeting six levels from 0 to 1 and re-estimating realized κ for analysis.The requested-zero condition is a boundary condition rather than exact zero in vision: realized means are 0.010681 for the convolutional network and 0.041266 for the vision transformer.
  • Compatibility intervention: A one-dimensional bisection matches the common target decrease across compatibility conditions without rotating directions, making ρ interpretable as persistent learning divided by approximately equal unrestricted opportunity.Mean within-state relative spread of ∆0 is 0.283%, 0.282%, and 0.268% in the convolutional network, vision transformer, and text transformer, respectively.
  • AFM implementation: AFM compares protected updates with same-state unrestricted learning, projects updates onto compatible directions, applies normalized persistent assimilation and endpoint checks, and rejects failed transactions atomically.Protected behaviour is represented with bounded Jacobian sensitivity sketches, while finite residual correction restores selected protected outputs and comparator logits.
  • Reproducibility and validation: Function-preserving structural renewal resets dormant zero-gated modules without changing deployed predictions or protected behaviours, then reactivates modules only through accepted protected updates.The natural validation samples 50 pre-update states per system and seed without compatibility-based selection and discards measurement branches after ordinary parent updates.

6 Chronological AFM evaluation and benchmark evidence

AFM outperforms the strongest eligible classical non-oracle family across all three frozen protocols and remains competitive across six recent challengers under a common predictive-base interface. The benchmark evidence also shows explicit preservation–acquisition trade-offs, substantial execution overhead, and limits on what the controlled comparison establishes.

  • Primary comparisons: AFM exceeds the strongest eligible non-oracle family on every paired seed across all three protocols, with paired confidence intervals above zero.Five-seed mean hypervolumes are 0.248 on CORe50, 0.381 on CLEAR-10, and 0.513 on CLAD-C; paired mean advantages are approximately 0.0090, 0.0068, and 0.0110, respectively.
  • Primary comparisons: No protection is the strongest classical non-oracle family on CORe50 and CLEAR-10, whereas A-GEM is strongest on CLAD-C; task-aware oracle EWC is ineligible.The comparison is framed against the strongest eligible non-oracle family rather than raw hypervolume across datasets, because raw hypervolume is protocol-specific.
  • Operating frontier: AFM’s frontier shifts from stronger preservation at η = 0.10 toward greater acquisition at η = 1.00, reducing forgetting on CLEAR-10 while retaining similar plasticity.On CLAD-C, the conservative coordinate substantially reduces forgetting and improves final-test accuracy, while full assimilation approaches unrestricted acquisition with more peak forgetting.
  • Benchmark comparisons: AFM has the highest mean primary score on CLEAR-10 and CLAD-C, while FGH exceeds it on CORe50 by about 2.8% at a more acquisitive, lower-retention point.In the descriptive aggregate comparison, AFM is about 7.0% higher than CCL-DC in primary score, with about 4.7% greater plasticity and 3.5% higher online accuracy; against six challengers, its primary score is 19.9% higher but mean retention is lower.
  • Execution checks: Protection geometry, finite counterfactual completion, and endpoint verification comprise 94.82% on CORe50, 91.70% on CLEAR-10, and 95.78% on CLAD-C of component-wise median execution cost.Ordinary learning accounts for only 0.36%, 0.44%, and 0.24%, respectively; accepted finite restorations keep maximum endpoint error below 10−6.
  • Scope and limitations: The empirical claim is restricted to continual adaptation over a frozen common representation prefix and does not treat architecture-changing methods as controlled comparisons.SinglePrompt, Online-LoRA, SERENA, EG-CNN, S6MOD, and Dual-Arch are excluded because their architectural or representation choices fall outside the shared predictive-base interface.

7 Discussion

The discussion frames functional compatibility as controlling the local persistent-learning frontier, while retention constraints determine how much compatible opportunity is used and nonlinear geometry limits its extension. It distinguishes finite protection from population retention and empirical neural executions from stronger convex-theoretic guarantees.

  • Conceptual implications: AFM separates requested assimilation η_t from compatibility fraction κ_t by comparing protected learning with the actual same-state endpoint accepted without protection.This comparison supplies a meaningful denominator for persistent plasticity; a large η_t does not imply incompatible gradients can be stored.
  • Scope and limitations: Exact restoration of protected outputs and finite evidence does not establish that unrestricted learning was stored or that unseen populations retain their behavior.The CLAD-C Pedestrian result illustrates how protected observations can be preserved while classifier interference overtakes the unseen Pedestrian population.
  • Scope and limitations: Task-free routing requires observable separation, and no learner can identify a semantic distinction that produces the same observation law.AFM can control false route refinement with separately allocated sequential evidence but withholds positive routing claims when observable separation is insufficient.
  • Scope and limitations: The main practical limitation is computational overhead concentrated in protection geometry, finite counterfactual completion, endpoint verification, and repeated protected-projector and shield construction.These costs exceed ordinary learning or construction of the same-state comparator in the profiling and CLAD-C analyses.
  • Scope and limitations: Compatibility does not independently determine persistent learning: retention allowance, method, finite curvature, and nonlinear geometry constrain the causal intervention and its extension.The fixed-norm construction does not establish a four-level causal continuum at ordinary full update magnitude.
  • Conceptual implications: Functional compatibility controls the local persistent-learning frontier, whereas retention constraints determine how much of that frontier learning rules can use.The discussion presents compatibility as an experimentally controllable property of incoming learning rather than merely a retrospective description of forgetting.

Funding

The author received no specific funding for this work.

  • No specific funding was received for this work.

Data availability

The study uses publicly available datasets from the original cited sources and generates no new primary dataset. Numerical summaries are provided in the paper, while checkpoints and row-level derived outputs are available from the corresponding author on reasonable request.

  • All datasets were obtained from publicly available original sources, and no new primary dataset was generated.
  • The public benchmark datasets are not redistributed with this submission.
  • Numerical summaries underlying the reported figures and tables appear throughout the paper, while saved checkpoints and row-level derived outputs are available from the corresponding author on reasonable request.
Loading 2608.22462v2…