Source-linked AI summary
Universal statistical signatures of evolution in artificial intelligence architectures
Theodor Spiro
TL;DR
The paper asks whether evolutionary signatures generalize from biology to AI despite radically different substrates and search mechanisms. It compiles ablation evidence across machine-learning publications and finds that DFE shape, diversification dynamics, and convergent innovation quantitatively match biological evolution.
Problem
The paper examines whether evolutionary statistical signatures are conserved between AI architectures and biological organisms despite different physical scales and search mechanisms.
Method
The authors compile 935 single-component ablation experiments from 161 machine-learning publications and normalize performance changes as relative fitness effects.
Results
The DFE, diversification dynamics, and convergent innovation of AI architectures quantitatively match biological organisms, with directed search shifting the beneficial fraction to 13%.
Takeaways & Limitations
The findings support fitness-landscape topology, rather than search mechanism, as the determinant of these shared evolutionary signatures.
Takeaways & Limitations
The dataset may overrepresent successful architectures, while automated extraction differs systematically from manual curation and convergence criteria are weaker than in biology.
Abstract
from arXiv · showhide
We test whether artificial intelligence architectural evolution obeys the same statistical laws as biological evolution. Compiling 935 ablation experiments from 161 publications, we show that the distribution of fitness effects (DFE) of architectural modifications follows a heavy-tailed Student's t-distribution with proportions (68% deleterious, 19% neutral, 13% beneficial for major ablations, n=568) that place AI between compact viral genomes and simple eukaryotes. The DFE shape matches D. melanogaster (normalized KS=0.07) and S. cerevisiae (KS=0.09); the elevated beneficial fraction (13% vs. 1-6% in biology) quantifies the advantage of directed over blind search while preserving the distributional form. Architectural origination follows logistic dynamics (R^2=0.994) with punctuated equilibria and adaptive radiation into domain niches. Fourteen architectural traits were independently invented 3-5 times, paralleling biological convergences. These results demonstrate that the statistical structure of evolution is substrate-independent, determined by fitness landscape topology rather than the mechanism of selection.
Results · The distribution of fitness effects in AI architectures matches biological DFEs
Across 935 ablation experiments from 161 publications, AI architectural fitness effects show a negatively skewed, heavy-tailed Student’s t-distribution dominated by deleterious changes. Their distributional shape closely matches D. melanogaster and S. cerevisiae, while proportions place AI between compact-genome organisms and simple eukaryotes.
- The distribution of fitness effects in AI architectures matches biological DFEs: 935 ablation experiments from 161 publications measured normalized relative fitness effects across computer vision, natural language processing, audio, and other domains.Effects were defined relative to each full model, analogously to biological selection coefficients.
- The distribution of fitness effects in AI architectures matches biological DFEs: 568 major ablations formed the primary analysis because component removal is the most homogeneous class and closest to biological gene-knockout studies.The full dataset of 935 experiments served as a robustness check.
- The distribution of fitness effects in AI architectures matches biological DFEs: Skewness = −2.23 and kurtosis = 29.2 characterized major-ablation effects, best fit by a Student’s t-distribution over Laplace and Normal alternatives.The reported AIC values were Student’s t = −428, Laplace = +169, and Normal = +942.
- The distribution of fitness effects in AI architectures matches biological DFEs: 68.0% of major ablations were deleterious, 19.0% neutral, and 13.0% beneficial.The 95% confidence intervals were 64.1–71.8%, 15.8–22.2%, and 10.4–15.8%, respectively.
- The distribution of fitness effects in AI architectures matches biological DFEs: AI DFE comparisons with nine biological organisms used synthetic samples reconstructed from published summary statistics, preserving reported shapes and category proportions.This creates an empirical-AI versus parametric-biological comparison asymmetry.
- The distribution of fitness effects in AI architectures matches biological DFEs: 0.07 and 0.09 were the normalized KS distances for D. melanogaster and S. cerevisiae, respectively, whereas Bacteriophage φ6 and VSV virus had Euclidean proportion distances of 0.11 and 0.12.No organism was the best match across every metric; AI occupied an intermediate position between compact-genome organisms and simple eukaryotes.
- The distribution of fitness effects in AI architectures matches biological DFEs: β = 0.65 characterized AI’s deleterious effects, exceeding D. melanogaster at β ≈0.4, M. musculus at β ≈0.3, and H. sapiens at β ≈0.2.The AI confidence interval was 0.59–0.72; VSV comparison was complicated because its deleterious DFE favored a log-normal over a gamma fit.
- Stratification reveals conserved structure: 0.68, 0.32, and 0.51 were the deleterious fraction for major ablations, neutral fraction for minor ablations, and deleterious fraction for minor ablations, respectively.Major ablations were complete component removals, while minor ablations were hyperparameter changes; DFE structure was statistically invariant across computer vision, NLP, audio, and other domains.
The beneficial fraction quantifies directed selection
AI ablation studies show a 13.0% beneficial fraction, exceeding the 1–6% range observed across the biological comparison set. The elevated fraction quantifies directed selection while preserving the heavy-tailed, negatively skewed Student’s t DFE shape.
- The beneficial fraction quantifies directed selection: 13.0% (CI: 10.7–15.3%) is the AI beneficial fraction, exceeding the 1–6% range across the biological comparison set.This elevated fraction is presented as a quantitative measurement of the difference between directed and blind search.
- The beneficial fraction quantifies directed selection: ∼1–5% is the biological beneficial fraction, reflecting the rarity of improvement by random perturbation in systems already adapted to their environments.By contrast, AI experiments intentionally test modifications believed to be potentially useful, described as a “sighted mutagen.”
- The beneficial fraction quantifies directed selection: The DFE retains a heavy-tailed, negatively skewed Student’s t shape while the beneficial fraction is elevated in AI.The comparison links this preserved DFE shape and shifted beneficial fraction to differences between directed and blind search.
Diversification dynamics match paleontological radiation patterns
AI architectural diversity follows logistic growth toward saturation, with punctuated origination peaks and sequential adaptive radiation across domain niches. Its normalized radiation curve and paradigm transitions parallel paleontological radiation patterns.
- Diversification dynamics match paleontological radiation patterns: R^2 = 0.994 and K ≈142 indicate logistic growth toward saturation at approximately 88% of capacity across 125 architectures and six domain niches.Named AI architectures were tracked from 2012 to 2024.
- Diversification dynamics match paleontological radiation patterns: AI diversification filled niches sequentially: computer vision in 2012–2016, NLP from 2017+, and audio and multimodal domains from 2021+.This sequence mirrors ecological niche-filling patterns after mass extinctions.
- Diversification dynamics match paleontological radiation patterns: RNN decline preceded Transformer radiation by approximately two years, while GAN decline similarly preceded Diffusion radiation, paralleling mass-extinction recovery lags.The last new RNN variant appeared in 2015, before Transformer radiation from 2017+.
- Diversification dynamics match paleontological radiation patterns: The normalized, smoothed AI radiation curve falls between the Cambrian trilobite and post-K-Pg mammalian curves, indicating a shared functional form across substrates.Curves were normalized to peak time and peak rate and smoothed with Gaussian kernels.
Convergent evolution is pervasive and quantitatively comparable
AI architectures repeatedly evolve similar traits across application domains, with 14 traits independently invented at least three times. Their convergence is more tightly clustered than in biology, while several traits show functional analogies to biological adaptations.
- Convergent evolution is pervasive and quantitatively comparable: 14 architectural traits were independently invented at least three times by different research groups across application domains.Attention, feature normalization, gating, positional encoding, and contrastive self-supervised learning each had 5 independent inventions.
- Convergent evolution is pervasive and quantitatively comparable: 5 independent inventions each occurred for attention, normalization, gating, positional encoding, and contrastive self-supervised learning.
- Convergent evolution is pervasive and quantitatively comparable: 3–5 independent inventions per trait characterized AI convergences, significantly tighter than biology’s 3–100+ range (Mann–Whitney p = 0.035).The passage attributes this difference to roughly 20 major AI research groups versus millions of biological species.
- Convergent evolution is pervasive and quantitatively comparable: 4 independent origins remained for attention, normalization, and gating under stricter criteria requiring distinct application domains and no shared authors.Attention origins spanned NLP, CV, CV-video, and multimodal applications; normalization and gating likewise spanned multiple domains.
- Convergent evolution is pervasive and quantitatively comparable: Attention, normalization, and gating functionally parallel camera eyes, homeostatic regulation, and ion channels, respectively.Their analogous functions are selective information gathering, maintaining internal stability, and conditional information flow.
Lineage analysis reveals evolutionary maturation
Lineage analysis shows that architectural evolution undergoes maturation: descendants develop narrower DFEs, while increasing optimization is associated with stronger mutational constraint. AI and biological lineages occupy a shared optimization–constraint axis, although Mamba/SSM results provide a mixed test of the predicted neutral-space trend.
- Within-lineage maturation: Later Transformer descendants show narrower, more concentrated fitness effects than early descendants, consistent with progressive optimization reducing neutral space.The DFE shifts systematically with generational distance from the founding architecture.
- Optimization and constraint: More optimized systems show a higher fraction of deleterious mutations across AI lineages and biological organisms.CNN, Transformer, and Vision Transformer generations are compared with RNA viruses, bacteriophages, E. coli, yeast, Drosophila, and humans.
- Prediction test: 0.16 was the overall neutral fraction for Mamba/SSM architectures, below the dataset average of 0.19 for major ablations, producing a mixed prediction test.Mamba/SSM architectures comprised n = 38; stratification found Mamba major ablations (n = 23) were 100% deleterious.
Discussion
The study finds that evolution’s statistical structure is conserved from biological organisms to AI architectures despite differences in physical substrate, scale, timescale, and search mechanism. The DFE, diversification dynamics, and frequency of convergent innovation match quantitatively across both domains.
- Cross-domain conservation: Evolutionary statistical structure is conserved across carbon-based biology and silicon-based AI architectures.The comparison spans nanometer-to-centimeter scales, billion-year-to-decade timescales, and blind mutation-to-directed engineering.
- Cross-domain conservation: The distribution of fitness effects matches quantitatively between AI architectures and biological organisms.
- Cross-domain conservation: Diversification dynamics and the frequency of convergent innovation also match quantitatively across AI architectures and biological organisms.
Why do the patterns match despite directed selection?
The matching evolutionary patterns are attributed primarily to fitness-landscape topology rather than search mechanism: conserved DFE form coexists with a shifted beneficial fraction under directed selection.
- Fitness-landscape topology: Fitness-landscape topology, including hierarchical modularity, rugged local structure, and finite high-fitness basins, is proposed to generate heavy-tailed DFEs, logistic diversification, punctuated equilibria, and convergence regardless of search mechanism.The passage contrasts random mutation with intentional design as alternative search mechanisms producing the same statistical patterns.
- Conserved DFE form: AI DFE shape parameters β, skewness, and kurtosis fall within the biological range despite fundamentally different mutation processes.The conserved distributional form argues against search mechanism as the primary determinant of DFE shape.
- Directed selection: 13% beneficial effects in AI versus 1–6% in biology is the sole systematically shifted parameter, consistent with directed selection changing the fraction rather than the DFE form.The passage identifies this as a single-parameter shift while the broader distributional form remains conserved.
Alternative explanation: a property of modular systems?
The DFE shape alone may reflect modularity in engineered systems, but the combined signatures of stratification, logistic diversification, and convergent evolution more specifically support an evolutionary explanation.
- Alternative explanation: a property of modular systems?: Randomly removing components from aircraft or software could produce a heavy-tailed, negatively skewed DFE, making modularity an alternative explanation for the biological match.The passage frames this as a critical test of whether the findings are specific to evolution.
- Alternative explanation: a property of modular systems?: Software mutation testing also shows negatively skewed, heavy-tailed mutation effects, suggesting that heavy-tailed DFEs may be generic to modular engineered systems.Mutation testing applies systematic code modifications such as statement deletion and operator replacement.
- Alternative explanation: a property of modular systems?: The stratification pattern, logistic diversification dynamics, and convergent evolution are not predicted by generic modularity but are predicted by evolutionary theory.Their simultaneous agreement is difficult to explain without shared landscape structure.
- Alternative explanation: a property of modular systems?: The authors therefore conclude that DFE shape alone could reflect modularity, whereas its conjunction with diversification dynamics and convergence is more specifically evolutionary.This is the paper’s conservative framing of the alternative explanation.
Connection to thermodynamic theories of evolution
Thermodynamic frameworks explain the observed universality as adaptive optimization of dissipation on structured landscapes, balancing fitness and robustness through free energy. The framework also yields testable predictions about diversification, convergence, DFE stability, and beneficial effects.
- Thermodynamic framework: Free energy F = E −TS captures the trade-off between fitness, associated with low energy, and robustness, associated with high entropy.The framework treats biological organisms and neural networks as information-processing systems optimizing dissipation on structured landscapes.
- Testable predictions: K ≈142 and current ≈125 imply that architectural origination should continue declining as diversity approaches carrying capacity, with future radiation events smaller than the 2017 and 2021 peaks.This is the framework’s first testable prediction.
- Testable predictions: Attention and normalization mechanisms are predicted to be independently reinvented as AI expands into robotics, biological AI, and materials science.This prediction concerns adaptive radiation into new application domains.
- Testable predictions: The DFE shape parameter β is predicted to remain within [0.4, 0.7] regardless of which specific architectures dominate future data collection.The prediction concerns stability of DFE shape across changing architectural populations.
- Limitations: KS = 0.287 characterizes systematic differences between LLM extraction and manual curation, predominantly affecting the beneficial tail rather than distributional form.Additional limitations include publication bias, eight-order-of-magnitude timescale normalization, and a fundamental disanalogy in convergence analysis.
Methods
The study combined curated and automated extraction of architectural ablation experiments with comparative datasets on biological DFEs and architectural diversification. It evaluated distributional similarity, diversification dynamics, convergence, and lineage-specific DFE patterns using prespecified statistical procedures.
- Ablation data: 935?
- Ablation data: 140 experiments were manually curated from landmark papers, while 795 were extracted automatically from machine-learning publications published between 2014 and 2024.Ablations removed or modified one architectural component while holding all others constant and required a quantitative performance metric.
- DFE comparison: Eight organism-study combinations supplied published biological DFE estimates for comparison with AI architectures.The organisms included VSV virus, bacteriophage φ6, E. coli TEM-1 β-lactamase, S. cerevisiae, D. melanogaster, C. reinhardtii, M. musculus, and H. sapiens.
- DFE comparison: Synthetic DFE samples were generated from published summary statistics when individual-level mutation fitness data were unavailable.Individual-level data were available only for VSV virus and E. coli TEM-1; synthetic samples enabled KS distances and QQ plots for the other organisms.
- Diversification and lineage analyses: Diversification was analyzed from a 125-architecture catalog assigned to six domain niches, using logistic fits, annual-origination-rate CV, Mann–Whitney U testing, convergence intensity, and lineage-specific DFE tracking.The four tracked lineages were CNN, Transformer NLP, Vision Transformer, and Generative.
- DFE comparison: DFE comparisons used normalized Kolmogorov–Smirnov distance, QQ correlation, Euclidean distance in proportion space, moment matching, maximum-likelihood β fitting, AIC, and 2,000-resample bootstrap confidence intervals.Deleterious effects were fit using absolute values with |∆| < 3, and z-score normalization removed scale effects.