Source-linked AI summary
Fractional Optimizers Meet Fractal Activation Functions: An Empirical Study of Multi-Scale Optimization in Neural Network
Sebastian Raubitzek, Georg Goldenits, Sebastian Schrittwieser, Philip König, Kevin Mallinger
TL;DR
Fractal activations and fractional optimizers have been studied independently, leaving their interaction insufficiently evaluated. This paper tests their pairings across perturbed benchmark surfaces and neural-network tasks, finding that benefits are selective rather than universal.
Problem
Fractal activations and fractional optimizers have largely been studied independently, leaving their interaction insufficiently evaluated.
Method
The study systematically evaluates fractional optimizer families with conventional and fractal activations across controlled multi-scale surfaces and feed-forward neural networks.
Results
Fractional methods are most reliable on harder cases, while perturbations reduce Himmelblau reliability, preserve optimizer-family rankings, and favor adaptive fractional Adam.
Takeaways & Limitations
Benefits arise from specific pairings, especially subcritical fractal activations with local-scaling or adaptive fractional optimizers, rather than from indiscriminate gradient memory.
Takeaways & Limitations
Conclusions are limited to feed-forward networks on ten tabular datasets and two perturbed surfaces with fixed budgets, initializations, and representative hyperparameters.
Abstract
from arXiv · showhide
Fractional optimization methods and fractal activation functions are two independent directions for improving neural network training. Fractional optimizers extend first-order optimization through fractional derivatives and memory effects, whereas fractal activations introduce multi-scale nonlinear representations based on self-similar Weierstrass- and Blancmange-type functions. Here, we investigate their interaction within a unified experimental framework. We evaluate fractional optimizer families on Ackley and Himmelblau benchmark surfaces, in standard form and with additive Weierstrass-type perturbations, and then in feed-forward neural networks with conventional and fractal activations on ten classification datasets. The comparison includes standard methods, regularization-style optimizers, explicit and adaptive memory-based fractional optimizers, and other representative literature methods. Overall, fractional optimization and fractal activations show useful but selective pairings. Regularization-style fractional scaling performs well with selected fractal activations in network training, while Grünwald--Letnikov memory is most relevant on perturbed surfaces. Adaptive memory improves plain memory substitution in several cases, supporting controlled fractional memory as a promising direction rather than a universal replacement.
1. Introduction
This study addresses the previously separate use of fractal activations and fractional optimizers by evaluating their interaction across perturbed benchmark surfaces and neural-network classification tasks. It introduces adaptive memory-based fractional optimizers within a unified mathematical and empirical framework.
- Objectives: The study evaluates fractional optimizers on Ackley and Himmelblau surfaces, with and without additive Weierstrass-type perturbations, and in networks using fractal and conventional activations across ten classification datasets.These experiments test optimizer behavior on controlled landscapes and in feed-forward neural-network training.
- Contributions: It introduces adaptive memory-based fractional optimizers that fix the fractional order and adapt a bounded trust coefficient controlling the memory contribution.At one boundary of the coefficient, the method exactly reduces to the underlying classical optimizer.
- Contributions: The paper unifies fractional derivatives of Weierstrass-type functions with discrete Grünwald–Letnikov constructions, presenting fractal activations and fractional optimizers as realizations of one mathematical mechanism.The connection extends from continuous fractional-derivative theory to discrete optimization and automatic differentiation.
- Experimental scope: A systematic comparison covers 21 optimizers spanning classical, Herrera-type fractional, memory-based fractional, adaptive memory-based fractional, and published related-work methods.The comparison is conducted first on fractally perturbed benchmark surfaces with known minima and then in neural-network classification experiments.
- Motivation: Fractal activations add multi-scale structure to representations, while fractional optimizers introduce memory into update directions.The paper frames both approaches as addressing the same multi-scale problem from different sides.
2. Related Work
The related work connects classical fractal functions, fractional derivatives, fractality in neural networks, and fractional optimization. The present study bridges these largely separate areas by evaluating fractional optimizers on controlled fractal landscapes and with fractal activations.
- Research scope: The section organizes prior research into four lines: classical fractal functions, fractional derivatives, fractality in neural networks, and fractional optimization for neural-network training.Each line is related to the present study’s mathematical or experimental framework.
- Classical fractal functions: Weierstrass and Takagi (Blancmange) functions provide canonical fractal constructions, while Mandelbrot framed such objects as models of irregular natural structure.The Weierstrass function is continuous and nowhere differentiable under established parameter conditions.
- Fractional derivatives: Fractional calculus supplies non-integer-order derivatives and integrals, including Grünwald–Letnikov, Riemann, Liouville, and Caputo formulations for rough-function analysis.The study uses this theory to relate its perturbed objects and optimization operators.
- Fractality in neural networks: Fractality in neural networks has been studied in training dynamics and in architectures and features, including fractal behavior emerging without explicitly built-in fractal structure.Sohl-Dickstein identified a fractal boundary between trainable and divergent hyper-parameter configurations.
- Fractional optimizers: Prior fractional-optimization work includes Caputo-based fractional backpropagation and later optimizer formulations incorporating bounded, hysteresis-controlled memory contributions.The reviewed approach can reduce exactly to the underlying classical optimizer at one coefficient boundary.
- Positioning of the present study: The study addresses the literature gap by evaluating fractional optimizer families on controlled Weierstrass-type landscapes and combining fractal activations with 21 optimizers under identical classification conditions.The reviewed communities have largely developed fractal neural structures and fractional optimizers separately.
3. Intuition Behind Fractal Functions and Fractional Derivatives
This section provides intuition for fractals, multi-scale fractal functions, fractal activations, and noninteger derivatives. It motivates pairing fractal activations with memory-based fractional optimizers because higher-scale derivatives can become strongly oscillatory while memory incorporates information from previous steps.
- Geometric fractals: Fractals arise from repeating simple rules across scales, producing increasingly fine detail and often self-similar structure.The section introduces geometric examples before transferring the idea to functions and neural-network activations.
- Fractal functions: Weierstrass-type functions add smaller, faster oscillations at successive scales, approaching continuous graphs that can become increasingly rough and nowhere differentiable.Parameter a controls amplitude decay, while b controls frequency growth.
- Fractal activations: Modified Weierstrass–Tanh activations retain a smooth tanh backbone while adding localized, alternating oscillatory scales whose depth increases with M.The factor am reduces higher-scale amplitudes, bm increases frequencies, and e−µe|x| localizes the correction around the origin.
- Fractional derivatives: Fractional derivatives measure change at noninteger orders by combining information over an interval or history rather than using only local slope information.The article uses the discrete Grünwald–Letnikov viewpoint because optimization proceeds iteratively.
- Motivation for the pairing: When ab > 1, higher-frequency derivative contributions grow with m, making fractal or near-fractal functions difficult for local first-order methods.Memory-based optimizers can include information from previous steps instead of reacting only to the current local gradient.
4. Weierstrass-Type Functions and Fractional Derivatives
This section links Weierstrass-type fractal structure and fractional differentiation through shared multiscale organization. It establishes conditional design principles for derivative definitions, numerical approximations, fractal activations, and optimizer memory.
- Fractional regularity: Fractional differentiation measures multiscale roughness nonlocally, with differentiability expected below the roughness exponent and failure above it.For roughness exponent γH, orders ν < γH are admissible, ν = γH is critical, and ν > γH is expected to fail.
- Choice of derivative: Weyl–Marchaud derivatives are appropriate because they use increments directly, unlike Caputo derivatives, which require an ordinary derivative absent from Weierstrass-type functions.The increment-based form is also compatible with periodicity.
- Fractional regularity: Weyl–Marchaud differentiation preserves Weierstrass multiscale structure while reducing the roughness exponent from γH to γH −ν.When 0 < ν < γH < βH ≤1, the derivative remains a fractal object of the same family and retains positive reduced roughness.
- Numerical implementation: Discrete fractional derivatives depend on the sampling grid, which may omit fractal scales and thereby alter the computed derivative.Short-memory truncation introduces a controllable horizon: larger K better approximates the full fractional model but increases computation and numerical sensitivity.
- Fractal activation and optimizer design: The modified Weierstrass–Tanh activation combines stable tanh behavior with a truncated geometric-frequency, geometric-amplitude oscillatory ladder.A representative configuration uses a = 0.5, b = 1.5, µe = 0.75, and truncation length N = 100.
- Fractal activation and optimizer design: Pairing is conditional: hereditary Grünwald–Letnikov updates preserve potentially useful oscillatory history, whereas non-hereditary scaling only rescales the current step.The choice depends on whether fine-scale oscillations carry directional information or mainly introduce noise; ordinary backpropagation still computes the derivative of the implemented activation.
5. Fractal Activation Functions
The study uses four truncated fractal activations built from stable backbones and self-similar oscillatory ladders. Their underlying ladders deliberately span subcritical, critical, and supercritical regularity regimes to provide controlled multi-scale gradient structure for optimizer evaluation.
- Construction and implementation: The ladder construction varies amplitude decay and frequency growth, with differentiability governed by the product ab and roughness represented by γH.The generic series uses amplitude decay rate a and geometric frequency base b, with a = b^-γH.
- Regularity spectrum: The selected activations cover one subcritical ladder, two critical ladders, and one supercritical ladder with parent roughness exponent γH = 1/2.This selection provides a graded range of regularity rather than a single fractal regime.
- Regularity spectrum: The Takagi-type activation uses a = 1/2 and b = 2, placing its underlying ladder at the critical boundary ab = 1, while added gates and offsets alter the full parent’s regularity.Its finite implementation remains practically usable despite the classical infinite Takagi function being continuous and nowhere differentiable.
- Regularity spectrum: The subcritical oscillatory ladder has ab = 0.75 < 1 and γH ≈1.71, so its differentiated series converges uniformly to a bounded, rapidly varying derivative.The complete activation also includes an envelope that is not differentiable at x = 0.
- Construction and implementation: All four activations combine a tanh-type bend or linear drift with a geometric oscillatory ladder truncated at N terms for automatic differentiation.The finite truncation avoids directly differentiating ideal infinite nowhere-differentiable constructions.
6. Fractional Optimizers
This section defines fractional optimizers as gradient-based methods that replace, rescale, or extend first-order derivatives with non-integer-order operators. It compares 21 optimizers spanning classical, local fractional-scaling, memory-based, adaptive-memory, and related-work families, with controlled history contributions in the adaptive framework.
- Definition: Fractional optimizers replace, rescale, or extend the ordinary first-order gradient using a non-integer derivative order ν.Their motivation is the nonlocal nature of fractional derivatives.
- Optimizer families: The comparison includes 21 optimizers across five groups: classical baselines, Herrera-type fractional, memory-based fractional, adaptive memory-based, and related-work methods.The classical baselines are SGD, RMSprop, Adam, and Adadelta; memory-based groups use Grünwald–Letnikov gradients.
- Fractional constructions: Two routes are used: Caputo-inspired local scaling changes only current-gradient magnitude, whereas Grünwald–Letnikov memory forms updates from recent gradients.The memory route uses a finite-history weighted sum with an algebraically decaying kernel.
- Adaptive memory: Adaptive memory keeps ν fixed and adjusts a bounded mixing coefficient λt to regulate the contribution of the finite-history fractional gradient.The coefficient is initialized at λ0 = 0.10 and clipped to [0, 0.30], limiting the memory contribution to 30%.
- Control mechanisms: Descent safeguards, norm matching, and λt control how strongly history is trusted while permitting exact recovery of the base optimizer.λt = 0 reproduces the base optimizer, while λt = 1 reproduces a purely fractional-memory optimizer.
7. Surface Optimization Experiments
Surface experiments show that optimizer effectiveness depends on landscape geometry and perturbation structure: Himmelblau is tractable, whereas Ackley remains difficult. Fractional and memory-based SGD variants are strongest on harder or perturbed cases, but added reliability can require greater runtime and may conceal instability in mean losses.
- Landscape difficulty: Himmelblau was largely tractable, while Ackley resisted optimization: no Ackley method exceeded a success rate of 0.15, and best mean final losses remained 4.6–5.8.The standard Himmelblau surface was completely solved by FCSGD_GL, with mean final loss zero and success rate 1.00.
- Himmelblau results: On standard Himmelblau, FCSGD_GL, MemoryFSGD, AdaptiveMemoryFSGD, and SGD all reached a success rate of 1.00.AOFGD_SGD reached 0.88 and AOFGD_Adam 0.83.
- Perturbed surfaces: On additive Himmelblau, AOFGD_Adam led with a success rate of 0.73 and mean final loss of 1.98, followed by FCSGD_GL and MemoryFSGD at 0.65 each.The perturbation reduced reliability for all methods but preserved the overall optimizer-family ordering.
- Ackley results: On additive Ackley, FCSGD_GL achieved the lowest mean final loss (4.59), while FCSGD_GL and memory-based methods shared the lowest mean best loss (1.21).On standard Ackley, FSGD led with a success rate of 0.15, ahead of SGD and AOFGD_SGD at 0.08 each.
- Cost and limitations: FCSGD_GL required roughly 1.9–2.8 times SGD’s standard-surface step time, while some MemoryFAdadelta runs diverged and inflated mean losses to the order of 10^16 and above.Runtime, success distributions, target distance, and mean-versus-median losses exposed costs and instability beyond average performance.
8. Neural Network Experiments with Fractal Activations and Fractional Optimizers
This section evaluates fractal activations and fractional optimizers in feed-forward neural networks across ten OpenML classification datasets. The experiments compare selected activation and optimizer families under dataset-specific architectures and repeated evaluation protocols.
- Experimental benchmark: The benchmark covers ten OpenML datasets spanning binary and multi-class classification tasks.The datasets form a compact, varied testbed for optimizer–activation combinations.
- Experimental benchmark: Binary datasets use one hidden layer, whereas multi-class datasets use two hidden layers with widths (n, 2n).Dataset-specific widths, batch sizes, and training epochs follow the earlier fractal activation study.
- Activation comparison: The activation comparison includes ReLU, tanh, and four selected fractal activations chosen for strong prior performance.The full fractal activation catalogue was excluded to keep computational costs manageable.
- Optimizer comparison: Twenty-one optimizers are organized into standard baselines, Herrera-style fractional variants, memory-based methods, adaptive-memory methods, and related-work fractional optimizers.The named standard baselines are SGD, Adam, RMSprop, and Adadelta; the related-work group includes AdaGL, FCSGD_GL, FCAdam_GL, AOFGD_SGD, and AOFGD_Adam.
- Evaluation protocol: All configurations are repeated over 40 random seeds, with accuracy, macro-F1, precision, recall, timing, completed epochs, and best validation loss recorded.Results are ranked by each optimizer’s best mean test-accuracy configuration, while activation figures aggregate all evaluated optimizer configurations and derivative orders.
1 AdaptiveMemory FAdadelta
AdaptiveMemoryFAdadelta achieved the highest accuracy in one setting with modified Weierstrass–tanh, but its advantage was small and came with higher training time. Across the reported settings, performance depended strongly on optimizer–activation pairings, with modified Weierstrass–tanh especially effective in aggregate but not uniformly dominant.
- AdaptiveMemory FAdadelta: 0.91250 ± 0.01740 was the highest mean accuracy, achieved by AdaptiveMemoryFAdadelta with modified Weierstrass–tanh at ν = 1.00.Standard Adadelta and MemoryFSGD followed at 0.91204 with standard deviation 0.01789, only 0.046 percentage points lower.
- AdaptiveMemory FAdadelta: All five optimizer groups appeared among the top six, while the complete top-15 range was only 0.108 percentage points.The strongest methods were therefore closely grouped despite representation from Standard, GL-Memory, Herrera, Adaptive-Memory, and Related-Work families.
- AdaptiveMemory FAdadelta: ReLU had the highest mean across the complete optimizer grid, whereas modified Weierstrass–tanh produced the best individual configuration.The modified Weierstrass–tanh activation therefore depended on suitable optimizer pairing rather than providing a uniform advantage.
- AdaptiveMemory FAdadelta: 0.6844 was the aggregated mean accuracy of Weierstrass–Mandelbrot x + sin, lowering the overall fractal-activation mean to 0.85095 versus 0.90934 for standard activations.The other three fractal activations remained within approximately 0.58 percentage points of ReLU.
- AdaptiveMemory FAdadelta: 0.77197 ± 0.02215 was the highest mean accuracy in the other reported setting, achieved by FSGD with ReLU at ν = 1.25.AdaptiveMemoryFAdam with modified Weierstrass–tanh ranked second at 0.76981 ± 0.02222, with a 0.216-percentage-point gap and overlapping confidence intervals.
- AdaptiveMemory FAdadelta: Modified Weierstrass–tanh reached the highest aggregated mean accuracy at approximately 0.755, exceeding tanh at approximately 0.732 by approximately 2.3 percentage points.However, standard activations averaged 0.72942 versus 0.69828 for fractal activations because the latter combined one strong activation with three weaker variants.
1 AdaptiveMemory FAdadelta
AdaptiveMemoryFAdadelta achieved the highest mean accuracy with decaying cosine, but its advantage over the faster FAdadelta–tanh configuration was small relative to run-to-run variation. Results indicate a selective interaction between decaying cosine and adaptive optimizers rather than a uniform fractal-activation benefit.
- Best configuration: 0.59538 ± 0.06739 was the highest mean accuracy, achieved by AdaptiveMemoryFAdadelta with decaying cosine and ν = 1.00.FAdadelta with tanh and ν = 1.50 ranked second at 0.58846 ± 0.07388, while AdaptiveMemoryFAdam with decaying cosine ranked third at 0.58654 ± 0.05051.
- Uncertainty and cost: The 0.692-percentage-point gap between the top two configurations was small relative to run-to-run standard deviations, with overlapping 95% confidence intervals.The winning configuration achieved mean macro-F1 of 0.49421, whereas FAdadelta reached 0.49777 and required 1.49 s versus 19.22 s for the winner.
- Activation interaction: Decaying cosine appeared in nine of the first 11 positions, including the best Standard, Adaptive-Memory, and Related-Work configurations.Tanh was used by the second- and fourth-ranked Herrera-type methods, while modified Weierstrass–tanh appeared only from rank 12 onward.
- Optimizer families: The strongest AdaptiveMemory, Herrera, and Standard configurations reached 0.59538, 0.58846, and 0.58385, respectively, while the best explicit GL-Memory method reached only 0.50885 ± 0.06741.The four GL-Memory methods occupied ranks 16, 18, 19, and 20, whereas AOFGD_Adam was the strongest Related-Work optimizer at 0.56692.
- Aggregate activation results: Tanh led aggregate activation accuracy at approximately 0.457, ahead of decaying cosine at approximately 0.443 and modulated Blancmange at approximately 0.431.Across families, standard activations reached 0.43471 accuracy versus 0.41951 for fractal activations, while fractal activations led macro-F1, 0.26190 to 0.24427.
9. Discussion
The study finds that fractional optimization and fractal activations can improve training selectively, with performance determined by specific optimizer–activation pairings rather than either method family alone. Fractional memory helps perturbed benchmark surfaces, whereas local scaling and safeguarded adaptive memory are more effective and robust in network training.
- Overall interpretation: Effective optimizer–activation pairings recur, but no single method has a uniform advantage across datasets and tasks.On most datasets, leading mean accuracies differ by less than run-to-run standard deviation; Tic-Tac-Toe Endgame is the main exception, where one configuration exceeded the next by more than two percentage points.
- Optimizer behaviour on fractally perturbed surfaces: SGD-type methods lead all Ackley and Himmelblau variants, while fractional and memory-based SGD variants match or exceed plain SGD.FSGD led both Ackley variants, FCSGD_GL solved standard Himmelblau exactly, and MemoryFSGD had the highest success rate on additive Ackley.
- Fractal activations in the classification experiments: Fractal activations dominate top-15 configurations on eight of ten datasets and provide the single best configuration on six, but activation-wide averages favour them on only four.The subcritical modified Weierstrass–tanh function is the most consistent, whereas the supercritical x+ sin variant is the weakest.
- Fractional optimizers in the classification experiments: Herrera-type local fractional scaling is the strongest optimizer design, taking six of ten dataset wins at negligible cost.No universally optimal order exists; ν ∈{0.75, 1.25, 1.50} covers all winning configurations, with fractal and standard activations succeeding on different datasets.
- Why the two experimental stages disagree about memory: Explicit gradient memory helps on deterministic multi-scale landscapes but less in minibatch training, where truncated GL memory amplifies sampling noise.Safeguards keep memory methods functional but not superior, while adaptive-memory methods consistently rank above plain GL-memory methods in group-level averages.
10. Conclusion
The study finds that fractional optimizers and fractal activations are compatible but yield benefits only in specific pairings. Results support joint configuration, repeated-run evaluation, and cautious use of explicit gradient memory, while broader architectures and systematic hyperparameter studies remain open.
- Main findings: SGD-type methods led every controlled-surface variant, with fractional and memory-based variants matching or exceeding classical SGD within that family.The controlled experiments used Ackley and Himmelblau surfaces with additive Weierstrass-type perturbations.
- Main findings: Fractal activations appeared among leading configurations on eight of ten datasets and produced the winner on six.Their advantage depended on specific optimizer pairings rather than grid-wide averages.
- Main findings: The subcritical modified Weierstrass–tanh was the most consistent fractal representative, while no universally optimal fractional order emerged.Most differences between leading configurations were small relative to run-to-run variation, supporting repeated-run evaluation.
- Memory effects: Gradient memory helped on deterministic multi-scale landscapes but amplified stochastic minibatch noise, whereas safeguards kept memory methods functional in the stochastic regime.The truncated Grünwald–Letnikov kernel extracted persistent oscillatory structure from deterministic gradients, but its differencing character amplified noise under stochastic gradients.
- Practical recommendations: The practical recommendations are to tune activation, optimizer formulation, and fractional order jointly; use Herrera-type scaling by default; and reserve explicit memory for low-noise settings.The proposed order set ν ∈{0.75, 1.25, 1.50} covered all winning configurations, while bounded adaptive memory was recommended as an alternative.
- Limitations and future work: Future work should test convolutional, recurrent, transformer, and graph architectures, regression and generative modelling, larger benchmarks, and systematically varied memory and trust-coefficient settings.The current experiments covered feed-forward networks on tabular classification tasks, with several methodological hyperparameters fixed.
Funding
The research was funded through SBA Research’s COMET Center affiliation, supported by Austrian federal and regional institutions and managed by the Austrian Research Promotion Agency.
- Funding: SBA Research is funded through the COMET Programme by BMIMI, BMWET, and the federal state of Vienna, with the programme managed by FFG.Additional support came from the Austrian Federal Ministry of Economy, Energy and Tourism, the National Foundation for Research, Technology and Development, and the Christian Doppler Research Association.
Institutional Review Board Statement … Appendix B.2. Adaptive Memory-Based Extensions
The appendices define the manuscript’s notation, abbreviations, perturbation symbols, and optimizer state conventions, then extend fixed- and adaptive-memory constructions to additional first-order optimizers. Adaptive memory keeps the fractional order and kernel fixed while varying only the mixing coefficient, interpolating between classical and safeguarded fractional-memory updates.
- Abbreviations: The abbreviation list distinguishes Grünwald–Letnikov derivatives, convergence and adaptive-order controls, standard optimizers, fractional variants, and fixed- or adaptive-memory families.The listed families include FSGD, MemoryFAdam, AdaptiveMemoryFAdagrad, FCSGD_GL, AdaGL, AOFGD, and FracM.
- Appendix A. Notation: The notation glossary defines scalar and vector conventions, iteration indices, componentwise powers, and symbols for fractional order, memory, adaptive mixing, and loss-aware control.It also catalogs notation for Weierstrass-type functions, discrete operators, network layers, and fractal activations.
- Appendix A.3. Adaptive Mixing and Related-Work Control Quantities: The appendices specify notation for loss-aware λ_t control, optimizer state variables, additive perturbations, and activation-specific constructions.For memory-based optimizers, squared-gradient accumulators use the raw gradient rather than the fractional or effective gradient.
- Appendix B.1. Fixed Memory-Based Extensions: The memory-based extensions apply the same finite-history Grünwald–Letnikov gradient, descent safeguard, and norm matching used in the main constructions.Coordinatewise scale estimates remain tied to the raw gradient, so fractional memory changes the update direction without replacing classical accumulators.
- Appendix B.1. Fixed Memory-Based Extensions: Fixed-memory variants extend the mechanism to Adagrad, AdamW, Nadam, and Adamax while preserving their characteristic accumulators, decay, acceleration, or infinity-norm scaling.MemoryFAdagrad retains decreasing coordinatewise learning rates; MemoryFAdamW keeps decoupled weight decay independent of fractional memory; MemoryFNadam uses memory in its first moment and look-ahead term.
- Appendix B.2. Adaptive Memory-Based Extensions: Adaptive-memory extensions use the same fractional gradient, safeguards, norm matching, and adaptive mixing coefficient while retaining raw-gradient accumulators.λ_t = 0 gives the exact base optimizer, and ν = 1 gives the classical fallback convention.
- Appendix B.2. Adaptive Memory-Based Extensions: AdaptiveMemoryFAdagrad increases memory contribution when recent gradients are stable and decreases it when gradients conflict or smoothed loss worsens.This makes its memory direction adaptive rather than fixed, while the Adagrad accumulator remains classical.
- Appendix B.2. Adaptive Memory-Based Extensions: Adaptive memory extends AdamW, Nadam, and Adamax without adapting ν: only bounded λ_t changes, interpolating between classical and safeguarded fractional-memory updates.Weight decay remains independent in AdaptiveMemoryFAdamW, Nadam is exact when λ_t = 0 or ν = 1, and Adamax can reduce memory during unstable phases.
Appendix B.3. Summary of the Optimizer-Family Extensions · Appendix C. Runtimes of the Experiments
Appendix B.3 shows that safeguarded fractional-memory mechanisms transfer beyond the four optimizer families used in the main experiments, while Appendix C defines how computational costs are measured and compared within each experimental block. The runtime appendix reports separate surface-step and neural-network training-time quantities, averaged over repeated runs on common hardware within each block.
- Appendix B.3. Summary of the Optimizer-Family Extensions: The two optimizer-extension groups apply fixed-memory and adaptive-memory principles to further first-order methods beyond the 21 main-comparison optimizers.Table B.25 summarizes these extensions rather than adding them to the main experimental comparison.
- Appendix B.3. Summary of the Optimizer-Family Extensions: Fixed-memory extensions use the safeguarded, norm-matched Grünwald–Letnikov gradient as the update direction while estimating scales from the raw gradient.At ν = 1, each extension recovers its corresponding base optimizer.
- Appendix B.3. Summary of the Optimizer-Family Extensions: Adaptive-memory extensions mix the raw gradient with the safeguarded memory gradient through the bounded coefficient λ_t.Their scale estimates remain based on the raw gradient, and the base optimizer is recovered at λ_t = 0 or ν = 1.
- Appendix B.3. Summary of the Optimizer-Family Extensions: The extensions demonstrate that the proposed memory mechanisms are not tied to the four optimizer families selected for the main experiments.The transfer rule applies the safeguarded fractional-memory direction where the base optimizer expects a descent direction.
- Appendix C. Runtimes of the Experiments: Appendix C makes explicit the computational cost accompanying accuracy and success-rate differences across both experimental blocks.Surface experiments report mean wall-clock time per optimization step, whereas neural-network experiments use mean training time.
- Appendix C. Runtimes of the Experiments: Neural-network runtime is the mean training time of each optimizer’s best-performing configuration per dataset, averaged over the ten datasets.All reported values average 40 repeated runs per configuration and use the same hardware within each experimental block.
- Appendix C. Runtimes of the Experiments: Runtime values are comparable within each table but should not be compared across tables or transferred to other hardware.The appendix attributes this comparability to shared hardware within each experimental block.
Appendix C.1. Per-Step Cost in the Surface Experiments · Appendix C.2. Training Time in the Classification Experiments · Appendix C.3. Summary
The appendices show that fractional mechanisms add predictable per-step cost, while classification training time depends strongly on the activation selected by each optimizer’s best configuration. Despite these overheads, all methods remained practical for the study’s experimental scale.
- Appendix C.1. Per-Step Cost in the Surface Experiments: The four baselines were cheapest on standard surfaces, requiring 2.3–3.7 ms per optimization step.Herrera-type optimizers added roughly −2% to +55% relative to their baselines.
- Appendix C.1. Per-Step Cost in the Surface Experiments: The additive fractal variant increased every optimizer’s per-step time by approximately 5.1–5.8 ms.This common offset came from evaluating the perturbation ladder PJ and its gradient at every step.
- Appendix C.1. Per-Step Cost in the Surface Experiments: Ackley and Himmelblau had nearly identical per-step times within each variant, with differences of only fractions of a millisecond.The measurements therefore mainly reflected update-rule and perturbation costs rather than the specific objective.
- Appendix C.2. Training Time in the Classification Experiments: The Herrera group combined the highest cross-dataset accuracy with the second-lowest mean training time, at 6.5 s.FAdam and FAdadelta were identified as its two strongest members.
- Appendix C.2. Training Time in the Classification Experiments: Classification training times reflected both optimizer cost and the activation function selected by each optimizer’s best per-dataset configuration.Within-group comparisons were more reliable when winning activations were similar, whereas cross-group comparisons required the corresponding accuracy tables.
- Appendix C.3. Summary: Per-step overhead was typically around 10% for Herrera-type scaling, 40–105% for safeguarded Grünwald–Letnikov memory, and 2–3× for adaptive-memory and related-work methods.These figures were measured against corresponding baselines on cheap objectives.
- Appendix C.3. Summary: The most expensive mean best-configuration training time was approximately 12 s, while the complete programme included 88 800 training runs and 2 × 1680 surface runs.Runtime comparisons should be interpreted alongside activation choices and the corresponding accuracy tables.