Paper deep dive
Universal statistical signatures of evolution in artificial intelligence architectures
Theodor Spiro
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/14/2026, 2:22:19 AM
Summary
This paper provides a quantitative analysis of artificial intelligence architectural evolution, demonstrating that it follows the same statistical laws as biological evolution. By analyzing 935 ablation experiments, the study shows that the distribution of fitness effects (DFE) in AI architectures follows a heavy-tailed Student's t-distribution, mirroring biological systems. The research identifies that AI evolution exhibits punctuated equilibrium, logistic diversification, and convergent evolution, suggesting that these evolutionary signatures are determined by fitness landscape topology rather than the specific substrate or search mechanism.
Entities (5)
Relation Signals (3)
Theodor Spiro â authored â Universal statistical signatures of evolution in artificial intelligence architectures
confidence 100% ¡ Universal statistical signatures of evolution in artificial intelligence architectures Theodor Spiro
Fitness Landscape Topology â determines â Evolutionary Statistical Structure
confidence 95% ¡ These results demonstrate that the statistical structure of evolution is substrate-independent, determined by fitness landscape topology
Artificial Intelligence Architectures â exhibitsstatisticalsignaturesof â Biological Evolution
confidence 90% ¡ We test whether artificial intelligence architectural evolution obeys the same statistical laws as biological evolution.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:We test whether artificial intelligence architectural evolution obeys the same statistical laws as biological evolution. Compiling 935 ablation experiments from 161 publications, we show that the distribution of fitness effects (DFE) of architectural modifications follows a heavy-tailed Student's t-distribution with proportions (68% deleterious, 19% neutral, 13% beneficial for major ablations, n=568) that place AI between compact viral genomes and simple eukaryotes. The DFE shape matches D. melanogaster (normalized KS=0.07) and S. cerevisiae (KS=0.09); the elevated beneficial fraction (13% vs. 1-6% in biology) quantifies the advantage of directed over blind search while preserving the distributional form. Architectural origination follows logistic dynamics (R^2=0.994) with punctuated equilibria and adaptive radiation into domain niches. Fourteen architectural traits were independently invented 3-5 times, paralleling biological convergences. These results demonstrate that the statistical structure of evolution is substrate-independent, determined by fitness landscape topology rather than the mechanism of selection.
Tags
Links
- Source: https://arxiv.org/abs/2604.10571v1
- Canonical: https://arxiv.org/abs/2604.10571v1
Trouble viewing inline? Open PDF directly â
Full Text
40,932 characters extracted from source content.
Expand or collapse full text
Universal statistical signatures of evolution in artificial intelligence architectures Theodor Spiro Independent researcher theospirin@gmail.com Abstract We test whether artificial intelligence architectural evolution obeys the same statistical laws as biolog- ical evolution. Compiling 935 ablation experiments from 161 publications, we show that the distribution of fitness effects (DFE) of architectural modifications follows a heavy-tailed Studentâs t-distribution with proportions (68% deleterious, 19% neutral, 13% ben- eficial for major ablations, n = 568) that place AI between compact viral genomes and simple eukary- otes. The DFE shape matches D. melanogaster (nor- malized KS = 0.07) and S. cerevisiae (KS = 0.09); the elevated beneficial fraction (13% vs. 1â6% in biology) quantifies the advantage of directed over blind search while preserving the distributional form. Architectural origination follows logistic dynamics (R 2 = 0.994) with punctuated equilibria and adap- tive radiation into domain niches. Fourteen architec- tural traits were independently invented 3â5 times, paralleling biological convergences. These results demonstrate that the statistical structure of evolu- tion is substrate-independent, determined by fitness landscape topology rather than the mechanism of se- lection. Keywords: evolution, distribution of fitness effects, neural network architecture, convergent evolution, fitness landscape The development of artificial intelligence architec- tures over the past decade presents a remarkable natural experiment in evolution. Like biological or- ganisms, AI architectures are subject to variation (novel design choices), selection (benchmark perfor- mance), and heredity (architectural components in- herited from predecessors through citation and code reuse). Unlike biological evolution, however, AI ar- chitectural evolution is directed by human engineers who intentionally seek improvementsâraising the question of whether the resulting evolutionary dy- namics are fundamentally different from their bio- logical counterpart or merely an accelerated version of the same process. Previous work has drawn qualitative parallels be- tween AI and biological evolution [19, 5]. Neural architecture search has been framed as an evolu- tionary algorithm. The historical progression from simple perceptrons to deep networks has been com- pared to the increase in organismal complexity. Scaling laws governing neural network performance [11, 10] have been likened to allometric relationships in biology. However, these comparisons have re- mained metaphorical. No study has quantitatively tested whether the statistical signatures of biological evolutionâthe shape of the distribution of fitness ef- fects, the dynamics of diversification, the frequency of convergent solutionsâare reproduced in AI archi- tectural evolution. Here we conduct this test. We compile the largest dataset of AI architectural ablation experiments to date (935 experiments from 161 publications) and compare the resulting evolutionary statistics against well-characterized biological systems spanning four orders of magnitude in genome complexity, from RNA viruses to humans. We test three specific hy- potheses: H1 The distribution of fitness effects (DFE) of ar- chitectural mutations matches biological DFEs in functional form and parameters. 1 arXiv:2604.10571v1 [q-bio.PE] 12 Apr 2026 H2 Architectural diversification dynamics exhibit punctuated equilibrium and logistic saturation, matching paleontological radiation patterns. H3 The frequency and intensity of convergent evo- lution in AI architectures is quantitatively com- parable to biological convergence. Our approach draws on an emerging theoretical framework that views both biological and artificial evolution as instances of adaptive search on rugged fitness landscapes. The DFE reflects the local topol- ogy of the fitness landscape; diversification dynamics reflect the global structure of the space of viable de- signs; convergent evolution reflects the existence of a limited number of high-fitness solutions to recurring functional challenges. If these statistical signatures match across substrates, it implies that the structure of evolution is determined primarily by the geometry of possibility spaceânot by the mechanism of search. Results The distribution of fitness effects in AI archi- tectures matches biological DFEs We compiled 935 ablation experiments from 161 ma- chine learning publications spanning computer vision (n = 345), natural language processing (n = 369), audio (n = 65), and other domains (n = 149). Each experiment reports performance change when a sin- gle architectural component is removed or modified; we normalized effects relative to the full model to obtain relative fitness effects (â), analogous to the selection coefficient s in biological DFE studies. Because ablation experiments encompass hetero- geneous operationsâfrom complete component re- moval to hyperparameter variationâwe present ma- jor ablations (component removal, n = 568) as the primary analysis, as these constitute the most ho- mogeneous class and the closest analogue to gene- knockout studies in biology. The full dataset (n = 935) including minor and intermediate ablations is presented as a robustness check. For major ablations, the DFE is characterized by negative skew (skewness = â2.23) and heavy tails (kurtosis = 29.2), best fit by a Studentâs t- distribution (AIC: Studentâs t = â428 vs. Laplace = +169, Normal = +942). The majority are dele- terious (68.0%; 95% CI: 64.1â71.8%), with a neutral fraction of 19.0% (CI: 15.8â22.2%) and a beneficial fraction of 13.0% (CI: 10.4â15.8%) (Fig. 1A). We compared the AI DFE against published DFE data for nine biological organisms (Table 1). Because individual-level DFE data are unavailable for most organisms, we generated synthetic samples from pub- lished summary statistics (gamma shape, category fractions) for distributional comparison. This ap- proach accurately preserves the shape and propor- tions reported in the original studies but means that our Q plots and KS distances compare an empiri- cal distribution (AI) against parametric reconstruc- tions (biology). We flag this asymmetry explicitly and base our primary conclusions on parameter com- parisons (β, category fractions) that do not depend on sample-level data. Table 1: Comparison of AI DFE (major abla- tions, n = 568) with biological DFEs. All biological comparisons use parametric reconstruc- tions from published summary statistics (cate- gory fractions, gamma shape where available); this asymmetry is discussed in Methods. KS norm : KolmogorovâSmirnov distance after z-score normal- ization. d Euclid : Euclidean distance in (del, neu, ben) proportion space. r Q : Pearson correlation of nor- malized Q plot. Organismf del f ben KS norm d Euclid r Q AI (major, n=568) 0.68 0.13â Bacteriophage Ď60.72 0.040.150.11 0.90 VSV virus0.69 0.040.190.12 0.90 S. cerevisiae0.60 0.030.090.22 0.87 D. melanogaster0.52 0.030.070.32 0.90 E. coli TEM-10.50 0.050.120.33 0.90 M. musculus0.55 0.030.140.28 0.90 C. reinhardtii0.45 0.050.120.39 0.87 H. sapiens0.55 0.050.200.26 0.91 Biological category fractions from: SanjuĂĄn et al. (2004), Burch & Chao (2003), Wloch et al. (2001), Keightley & Eyre-Walker (2007), Bank et al. (2014), BĂśndel et al. (2019), Halligan et al. (2009), Eyre-Walker et al. (2006). Where individual mutation data were unavailable, synthetic DFE samples were generated from published parameters. Organisms ordered byd Euclid (proportion similarity to AI). The closest matches depend on the metric used (Table 1). By normalized KS distance (distribu- 2 tional shape after z-score normalization), the best matches are D. melanogaster (0.07) and S. cerevisiae (0.09)âboth comparisons against synthetic DFEs reconstructed from published parameters. By Eu- clidean distance in DFE proportion space, the closest are Bacteriophage Ď6 (0.11) and VSV virus (0.12), which share AIâs high deleterious fraction. No sin- gle organism is the âbest matchâ across all metrics; rather, AI architectures occupy an intermediate po- sition between compact-genome organisms (viruses, high f del ) and simple eukaryotes (lower f del , higher f neu ). For the subset of organisms with published gamma shape parameter estimates from population-genetic inference, the AI value (β = 0.65; CI: 0.59â0.72) is higher than D. melanogaster (β â 0.4; [12]), M. mus- culus (β â 0.3; [9]), and H. sapiens (β â 0.2; [7]). For VSV virus, SanjuĂĄn et al. (2004) found that the deleterious DFE was better fit by a log-normal than a gamma distribution, complicating direct β compar- ison. The AI β value indicates a less leptokurtic dele- terious tail than multicellular eukaryotes, consistent with AIâs lower proportion of mildly deleterious mu- tations relative to strongly deleterious ones (Fig. 1D, 2D). Results are robust to dataset composition: the full dataset (n = 935, including minor and intermediate ablations) yields β = 0.63 and similar KS distances (Table S2). Stratification reveals conserved structure The DFE varies systematically with mutation mag- nitude in a manner that mirrors biology (Fig. 2A). Major ablations (complete component removal, n = 568) are predominantly deleterious (fraction delete- rious = 0.68) with a strongly negative mean effect, analogous to large-scale genomic deletions. Minor ablations (hyperparameter changes, n = 213) show a higher neutral fraction (0.32) and lower deleteri- ous fraction (0.51), analogous to point mutations. Intermediate modifications (component replacement, n = 154) fall between these extremes (fraction dele- terious = 0.72). This stratification is one of the most robust features of biological DFEs [8] and its repro- duction in AI architectures is not trivially expected. The DFE is invariant across ML domains (Fig. 2B): computer vision, NLP, audio, and other domains yield statistically similar distributions, sug- gesting that the DFE shape is a property of the de- sign space, not of the specific task. To control for potential selection bias, we com- pared manually curated data (n = 140) with LLM- extracted data (n = 795). The subsets differ sig- nificantly (KS = 0.287, p < 0.001; Fig. 2C), but the difference is systematic and interpretable: man- ual curation captured only component-removal abla- tions (0% beneficial), while LLM extraction captured the full range including substitutions that improve performance (15.3% beneficial). The LLM-extracted data is thus less biased than the manually curated data, and the combined dataset provides a more ac- curate estimate of the true DFE. The beneficial fraction quantifies directed se- lection The most informative departure from biological DFEs is the elevated beneficial fraction. At 13.0% (CI: 10.7â15.3%), the AI beneficial fraction exceeds all biological organisms in our comparison set (range: 1â6%). This is not a failure of the biological analogyâit is a quantitative measurement of the dif- ference between directed and blind search. In biology, the low beneficial fraction (âź1â5%) re- flects the rarity of improvement by random pertur- bation in a system already adapted to its environ- ment. In AI ablation studies, the experimenter in- tentionally tests modifications believed to be poten- tially usefulâa âsighted mutagen.â The fact that the DFE shape (Studentâs t, heavy-tailed, negatively skewed) is preserved while only the beneficial frac- tion is shifted demonstrates that the topology of the fitness landscape, not the search strategy, determines the distributional form. Diversification dynamics match paleontologi- cal radiation patterns We tracked the origination rate of named AI archi- tectures from 2012 to 2024, identifying 125 distinct architectures across six domain niches (Fig. 3A). Cu- mulative diversity follows a logistic growth curve 3 with R 2 = 0.994 and estimated carrying capacity K â 142, suggesting that AI architectural diversity is approaching saturation at approximately 88% of capacity (Fig. 3B). The origination rate shows two clear peaksâ 2017 (Transformer innovation, 16 new architec- tures) and 2021 (CLIP + Diffusion models, 19 new architectures)âseparated by periods of relative sta- sis (Fig. 3A). The coefficient of variation of annual origination rates (CV = 0.53) is consistent with punctuated equilibrium, falling between the Cam- brian trilobite radiation (CV = 0.79) and post-K-Pg mammalian radiation (CV = 2.19). Diversification proceeds by adaptive radiation into domain niches (Fig. 3C): computer vision is colo- nized first (2012â2016), followed by NLP (2017+), then audio and multimodal domains (2021+). This sequential niche-filling mirrors the ecological pattern following mass extinctions, where generalist niches are filled first and specialist niches later [16]. Paradigm transitions in AI exhibit a pattern anal- ogous to mass extinction followed by radiation: the decline of RNNs (last new variant: 2015) precedes the Transformer radiation (2017+) by approximately two years, comparable to the recovery lag observed in paleontological mass extinctions [4]. Similarly, the decline of GANs precedes the Diffusion model radi- ation. When normalized to peak time and peak rate and smoothed with Gaussian kernels, the AI radiation curve falls between the Cambrian trilobite curve and the post-K-Pg mammalian curve (Fig. 3D), suggest- ing a shared functional form for adaptive radiation across substrates. Convergent evolution is pervasive and quanti- tatively comparable We catalogued 14 architectural traits that were in- dependently invented three or more times by differ- ent research groups in different application domains (Fig. 4A). The most convergent traits include atten- tion mechanisms (5 independent inventions), feature normalization (5Ă), gating mechanisms (5Ă), posi- tional encoding (5Ă), and contrastive self-supervised learning (5Ă). The distribution of convergence counts differs significantly from biological convergences (Mannâ Whitney p = 0.035; Fig. 4B), with AI convergences more tightly clustered (3â5 independent inventions per trait) compared to biology (3â100+). However, this difference reflects the smaller number of inde- pendent AI lineages (âź20 major research groups) compared to biological lineages (millions of species). When normalized per lineage, AI convergence inten- sity is approximately 5Ă 10 4 times higher than bio- logical convergence, quantifying the degree to which directed search accelerates convergent discovery. Under a stricter criterion requiring different ap- plication domains with no shared authors, inven- tion counts decrease but remain substantial: atten- tion (4 independent origins: NLP, CV, CV-video, multimodal), normalization (4: CV, NLP, CV-style, CV-detection), and gating (4: NLP-recurrent, NLP- feedforward, CV, NLP-SSM). These counts are re- ported in Table S3. Functional analogies between AI and biological convergences are notable: attention mechanisms serve a function analogous to camera eyes (selective information gathering, 4â5 vs. 7 independent inven- tions); feature normalization parallels homeostatic regulation (maintaining internal stability, 4â5 vs. 3); gating mechanisms parallel ion channels (conditional information flow, 4â5 vs. 4). These parallels suggest that when computational challenges are similar, the number of viable solutions converges regardless of substrate. Lineage analysis reveals evolutionary matura- tion Within individual architectural lineages, the DFE shifts systematically with generational distance from the founding architecture (Fig. 4CâD). In the Transformer NLP lineage, early descendants retain broader fitness effects, while later descendants show a narrower, more concentrated DFEâconsistent with progressive optimization reducing neutral space. Across all lineages and biological organisms, a universal trend emerges (Fig. 4D): more optimized systems show a higher fraction of deleterious mu- tations. AI lineage points (CNN generations 1â 4 5, Transformer generations 1â5, Vision Transformer generations 1â4) are interspersed with biological data points (RNA virus, bacteriophage, E. coli, yeast, Drosophila, human) along a common axis, suggest- ing that the relationship between optimization level and mutational constraint is substrate-independent. We tested a within-study prediction: young lin- eages should retain more neutral space, analogous to genomes with lower functional density. Mamba/SSM architectures (n = 38) provide a mixed result. The overall neutral fraction (0.16) is lower than the dataset average (0.19 for major ablations), appear- ing to contradict the prediction. However, stratifi- cation reveals an important nuance: Mamba major ablations (n = 23) are 100% deleteriousâconsistent with a compact, highly optimized architecture where every component is criticalâwhile Mamba minor ab- lations (n = 10) show a neutral fraction of 0.60, substantially higher than the minor-ablation average of 0.32. This pattern suggests that Mambaâs core components are tightly constrained (few alternatives exist for a new paradigm), while its hyperparame- ter space remains underexploredâprecisely the sig- nature expected of a young lineage that has been optimized at the component level but not yet fully tuned. The prediction is thus partially confirmed but requires larger lineage-specific samples for definitive testing. Discussion Our results demonstrate that the statistical struc- ture of evolution is conserved across a transition that spans every physical parameter: from carbon to sili- con, from nanometer to centimeter, from billion-year to decade timescales, from blind mutation to directed engineering. The DFE, the dynamics of diversifica- tion, and the frequency of convergent innovation all match quantitatively between AI architectures and biological organisms. Why do the patterns match despite directed selection? The most parsimonious explanation is that the sta- tistical structure of evolution is determined primar- ily by the topology of the fitness landscape, not by the mechanism of search. A fitness landscape with hierarchical modularity, rugged local structure, and a finite number of high-fitness basins will pro- duce heavy-tailed DFEs, logistic diversification with punctuated equilibria, and convergent evolutionâ regardless of whether the search is conducted by ran- dom mutation or intentional design. This interpretation is supported by the observa- tion that AI DFE shape parameters (β, skewness, kurtosis) fall within the biological range despite the fundamentally different mutation process. If the search mechanism dominated, we would expect AI DFEs to differ systematicallyâshowing a different distributional form. Instead, the form is conserved and only one parameter is systematically shifted: the beneficial fraction (13% vs. 1â6% in biology). This single-parameter shift is precisely what the landscape hypothesis predicts: directed search reaches benefi- cial regions more efficiently but does not alter the landscape itself. Alternative explanation: a property of modu- lar systems? A critical question is whether our results are spe- cific to evolutionary processes or would arise from any complex modular system. If randomly removing components from a Boeing 737 or a software code- base also yields a heavy-tailed, negatively skewed DFE, then the match with biology reflects modu- larity rather than evolution per se. Evidence from software mutation testing is par- tially informative. In mutation testing, systematic code modifications (statement deletion, operator re- placement) are applied to software to evaluate test suite quality. The distribution of mutation effects in software is indeed negatively skewed with heavy tails [17], suggesting that heavy-tailed DFEs may be a generic property of modular engineered systems. However, two features distinguish the AI case. First, the gamma shape parameter (β = 0.65) falls specif- ically within the biological range (0.2â0.6), whereas software mutation effects have not been character- ized at this level of parametric detail. Second, the stratification pattern (major > intermediate > mi- 5 nor in deleteriousness), the logistic diversification dy- namics, and the convergent evolution are not predic- tions of a âgeneric modularityâ hypothesis but are specifically predicted by evolutionary theory. The coincidence of all three signatures matching simulta- neously is difficult to explain without invoking shared landscape structure. We therefore frame our results conservatively: the DFE shape match alone could reflect modularity; the conjunction of DFE shape, diversification dynamics, and convergence is more specifically evolutionary. Position in evolutionary parameter space AI architectures occupy a position between compact genomes (viruses, β â 0.5â0.6) and simple eukary- otes (yeast, Drosophila; β â 0.2â0.4) in DFE param- eter space. We propose that this reflects functional densityâthe fraction of components under selective constraint. Viral genomes have near-maximal func- tional density; eukaryotic genomes have substantial non-functional sequence. AI architectures, with their modular but non-redundant design, occupy an inter- mediate position. This is analogous to the DFE posi- tion of compact-genome organisms such as Bacterio- phage Ď6âan obligate parasite whose fitness is en- tirely determined by interaction with its host, much as a neural networkâs fitness is entirely determined by human evaluation. Coevolutionary dynamics The parallel to hostâparasite coevolution is deeper than analogy. In a companion study [18], we elicited probability forecasts from three major LLMs (GPT- 4o, Claude, Gemini) on 568 resolved prediction ques- tions from the Metaculus platform and found that their errors were highly correlated across models (r = 0.77), indicating that nominally independent AI systems share the same failure modes. More- over, the category-level pattern of LLM forecasting biases (which topics they over- or underestimate) al- ready closely matched the pattern of human forecast- ing biases measured before ChatGPT became avail- able, suggesting that LLMs inherited existing human cognitive biases from their training data rather than introducing novel ones. This establishes a human â AI transmission vector. The present study docu- ments the reverse vector: human researchers act as the selective environment for AI architectures, with their design intuitions, benchmark choices, and peer review constituting the selection pressure. Together, these findings describe a coevolutionary system. The human researcher is both the envi- ronment (selecting which architectures survive) and the substrate that AI modifies (through code sug- gestions, analysis tools, and cognitive assistance). As LLMs become increasingly integrated into the research workflow, the feedback loop tightens: AI- assisted researchers may converge on architectures that AI systems can best evaluate, creating a self- reinforcing cycle analogous to Red Queen dynam- ics [20]. The elevated beneficial fraction we observe (13%) may increase further as this coevolutionary feedback intensifiesâa testable prediction. Connection to thermodynamic theories of evolution The universality we observe is predicted by ther- modynamic frameworks that view evolution as op- timization of dissipation on structured landscapes [6]. Systems that process informationâwhether bio- logical organisms or neural networksâmust balance fitness (low energy) and robustness (high entropy). The free energy F = EâTS captures this trade-off. Heavy-tailed DFEs, logistic diversification, and con- vergent evolution can all be derived as consequences of adaptive search on hierarchical free energy land- scapes, suggesting that the statistical laws of evolu- tion are ultimately thermodynamic in origin. Predictions Our framework generates four testable predictions: 1. The architectural origination rate should con- tinue to decline as diversity approaches carrying capacity (K â 142, current â 125); future radi- ation events, if they occur, should be smaller in magnitude than the 2017 and 2021 peaks. 2. Attention and normalization mechanisms will be independently reinvented as AI expands into 6 new application domains (robotics, biological AI, materials science). 3. The DFE shape parameter β will remain within [0.4, 0.7] in future data collection, regardless of which specific architectures dominate. 4. The beneficial fraction will increase over time as AI-assisted research creates tighter coevolution- ary feedback between human designers and AI systems. Limitations Our dataset, while the largest of its kind, relies on published ablation studies that may overrepre- sent successful architectures. The LLM-extraction pipeline introduces systematic differences from man- ual curation (KS = 0.287), although these affect pre- dominantly the beneficial tail rather than the distri- butional form. The paleontological comparison re- quires normalization across timescales differing by eight orders of magnitude. The convergence anal- ysis faces a fundamental disanalogy with biology: biological lineages do not communicate, while ML researchers read across domains. Our âindependent inventionâ criterion (different research groups in dif- ferent domains) is therefore weaker than biological independence. Even under the strict criterion (no shared authors, different application domains), the most convergent traits retain ⼠3 independent ori- gins (Table S3), but we acknowledge that latent knowledge transfer through shared conferences and preprint culture may inflate apparent convergence. This limitation means that our convergence counts should be interpreted as upper bounds. At an ear- lier sample size (n = 594), a temporal DFE shift between early (2014â2018) and late (2019â2024) ar- chitectures was significant (p = 0.03); at n = 935, this significance was lost (p = 0.128), suggesting the effect was partly driven by small early-period sample size (see Supplementary Materials). Broader implications If the statistical laws of evolution are indeed substrate-independent, then biology and AI engi- neering are not merely analogousâthey are processes with shared statistical signatures, governed by the same landscape geometry despite fundamentally dif- ferent mechanisms of heredity and variation. This unification has practical implications: biological evo- lutionary theory can inform AI architecture design (predicting which modifications are likely neutral vs. deleterious), while AI evolution, with its complete and unbiased fossil record, can serve as a model sys- tem for testing evolutionary theory with a precision impossible in paleontology. Methods AI ablation data We collected ablation experiments from machine learning publications on arXiv and major venues (NeurIPS, ICML, ICLR, CVPR, ACL, EMNLP) published between 2014 and 2024. An ablation ex- periment was defined as the removal or modifica- tion of a single architectural component with all other components held constant, reporting a quan- titative performance metric. We collected 140 ex- periments through manual curation of landmark pa- pers and 795 through automated extraction using Claude Sonnet (Anthropic) with structured prompts designed to extract experiment metadata, baseline performance, ablated performance, and component description from PDF documents retrieved via the arXiv API. Candidate papers were identified through 49 search queries targeting âablation studyâ across six ML subcategories, yielding 3,040 unique papers filtered to 164 high-confidence candidates by key- word scoring. Each LLM-extracted entry was vali- dated against inclusion criteria; entries with |â| > 5 were excluded as likely extraction errors (5 entries removed). Fitness effects were computed as â = (ablatedâ baseline)/|baseline|, yielding relative fit- ness effects directly comparable to biological DFE conventions. Ablations were classified by type: ma- jor (complete component removal), minor (hyperpa- rameter or variant change), or intermediate (replace- ment with simpler alternative). 7 Biological DFE data We assembled published DFE estimates for eight organism-study combinations: VSV virus [15], Bac- teriophage Ď6 [3], E. coli TEM-1 β-lactamase [1], S. cerevisiae [22], D. melanogaster [13, 12], C. rein- hardtii [2], M. musculus [9], and H. sapiens [7]. Important methodological note. Individual- level mutation fitness data are available only for VSV virus [15] and E. coli TEM-1 [1]. For all other organisms, published DFE characterizations report summary statistics: category fractions (lethal, dele- terious, neutral, beneficial) and, for some, fitted distribution parameters. To enable distributional comparisons (KS distances, Q plots), we gener- ated synthetic DFE samples from these published summary statistics, drawing deleterious effects from gamma distributions with published shape parame- ters where available, and neutral/beneficial effects from appropriate distributions (see Table 1 foot- notes). This means that our KS distances and Q correlations compare an empirical distribution (AI) against parametric reconstructions (biology), which preserves shape and proportion information but may underestimate true biological variance. Our primary conclusions rest on parameter comparisons (category fractions, fitted β) that do not depend on sample- level reconstruction, and we present the distribu- tional comparisons as supporting evidence. Architectural diversity data We compiled a catalog of 125 named architectures from Papers With Code, arXiv, and primary publi- cations, recording: name, year of publication, parent architecture(s), domain(s) of application, and key ar- chitectural components. Architectures were assigned to domain niches (CV, NLP, Audio, Multimodal, RL, Graph). Paleontological comparison data were obtained from the Paleobiology Database (PBDB) via API queries for Cambrian trilobites (540â480 Ma, 5,831 taxa), post-K-Pg mammals (80â40 Ma, 7,391 taxa), and Cretaceous angiosperms (145â50 Ma, 2,988 taxa). Statistical analyses DFE comparison. AI and biological DFEs were compared using normalized KolmogorovâSmirnov distance (after z-score normalization to remove scale effects), Q correlation, Euclidean distance in pro- portion space (deleterious/neutral/beneficial), and moment matching. The gamma shape parameter β was estimated by maximum likelihood fit to the ab- solute values of deleterious effects (|â| < 3). Model comparison used AIC. Bootstrap confidence intervals (2,000 resamples) were computed for all key statis- tics. Diversification dynamics. Logistic curves were fit to cumulative diversity by nonlinear least squares. Punctuated equilibrium was quantified via CV of an- nual origination rates. Paleontological curves were smoothed with Gaussian kernels (Ď = 1.2 bins) and normalized to peak time and rate. Convergence analysis. Independent invention counts were compared using MannâWhitney U test. Per-lineage convergence intensity was computed as inventions per trait per independent lineage. Lineage analysis. For four architectural lineages (CNN, Transformer NLP, Vision Transformer, Gen- erative), we computed within-lineage DFE statistics at each generation and tracked fraction deleterious as a function of generational distance. Data and Code Availability All data, analysis scripts, and figure-generation code are available at https://github.com/mool32/ ai-evolution-universal-signatures.The dataset of 935 ablation experiments with metadata is provided as Supplementary Dataset S1. Acknowledgments The automated extraction pipeline used the Claude API (Anthropic). Paleontological data were ob- tained from the Paleobiology Database (paleo- biodb.org). We thank the contributors to Papers With Code and arXiv for enabling open-science in- frastructure. 8 References [1] Bank C, Hietpas RT, Jensen JD, Bolon DNA (2014) A systematic survey of an intragenic epistatic landscape. Mol Biol Evol 32:229â238. [2] BĂśndel KB et al. (2019) Inferring the distri- bution of fitness effects of spontaneous muta- tions in Chlamydomonas reinhardtii. PLoS Biol 17:e3000192. [3] Burch CL, Chao L (2003) Epistasis and its re- lationship to canalization in the RNA virus Ď6. Genetics 167:559â567. [4] Chen Z-Q, Benton MJ (2012) The timing and pattern of biotic recovery following the end- Permian mass extinction. Nat Geosci 5:375â383. [5] Elsken T, Metzen JH, Hutter F (2019) Neural architecture search: a survey. J Mach Learn Res 20:1â21. [6] England JL (2013) Statistical physics of self- replication. J Chem Phys 139:121923. [7] Eyre-Walker A, Woolfit M, Phelps T (2006) The distribution of fitness effects of new deleteri- ous amino acid mutations in humans. Genetics 173:891â900. [8] Eyre-Walker A, Keightley PD (2007) The distri- bution of fitness effects of new mutations. Nat Rev Genet 8:610â618. [9] Halligan DL et al. (2009) Evidence for pervasive adaptive protein evolution in wild mice. PLoS Genet 6:e1000825. [10] Hoffmann J et al. (2022) Training compute-optimal large language models. arXiv:2203.15556. [11] Kaplan J et al. (2020) Scaling laws for neural language models. arXiv:2001.08361. [12] Keightley PD, Eyre-Walker A (2007) Joint in- ference of the distribution of fitness effects of deleterious mutations and population demogra- phy. Genetics 177:2251â2261. [13] Loewe L, Charlesworth B (2006) Inferring the distribution of mutational effects on fitness in Drosophila. Biol Lett 2:426â430. [14] McGhee GR (2011) Convergent Evolution: Lim- ited Forms Most Beautiful (MIT Press). [15] SanjuĂĄn R, Moya A, Elena SF (2004) The dis- tribution of fitness effects caused by single- nucleotide substitutions in an RNA virus. Proc Natl Acad Sci USA 101:8396â8401. [16] Sepkoski J (1984) A kinetic model of Phanero- zoic taxonomic diversity. I. Post-Paleozoic families and mass extinctions. Paleobiology 10:246â267. [17] Andrews JH, Briand LC, Labiche Y (2005) Is mutation an adequate criterion of testing effec- tiveness? Proc 27th Int Conf on Software Engi- neering, p. 215â224. [18] Spiro T (2025) The oracleâs fingerprint: cor- related AI forecasting errors and the limits of bias transmission. Preprint available at https: //arxiv.org/abs/X.X. [19] Stanley KO, Clune J (2019) Designing neural networks through neuroevolution. Nat Mach In- tell 1:24â35. [20] Van Valen L (1973) A new evolutionary law. Evol Theory 1:1â30. [21] Vaswani A et al. (2017) Attention is all you need. Advances in Neural Information Process- ing Systems 30. [22] Wloch DM, Szafraniec K, Borts RH, Korona R (2001) Direct estimate of the mutation rate and the distribution of fitness effects in the yeast Saccharomyces cerevisiae. Genetics 159:441â 452. 9 1.00.80.60.40.20.00.2 Relative fitness effect () 0 5 10 15 20 25 30 35 Probability density A DFE: AI major ablations vs S. cerevisiae All ablations (n=935) Major ablations (n=568) S. cerevisiae (Wloch 2001, synthetic) 42024 VSV virus DFE quantiles (normalized) 4 2 0 2 4 AI DFE quantiles (normalized) r = 0.891 B Q plot: AI vs VSV virus (normalized) AI architectures VSV virus Bacteriophage 6 E. coli TEM-1 S. cerevisiae S. cerevisiae (Wloch) D. melanogaster C. reinhardtii M. musculus H. sapiens 0.0 0.2 0.4 0.6 0.8 1.0 Proportion C DFE categories across systems Deleterious Neutral Beneficial 6.05.55.04.54.03.53.02.5 Skewness 10 20 30 40 50 Excess kurtosis D Shape space: AI among biological DFEs AI (all, n=935) AI major AI minor AI intermediate Biological organisms Figure 1: Distribution of fitness effects in AI architectures matches biological DFEs. (A) His- togram of all AI ablation effects (n = 935) overlaid with synthetic DFE for S. cerevisiae (parametric recon- struction from published summary statistics). The primary analysis uses major ablations only (n = 568; see Table 1); the full dataset is shown here to illustrate the complete DFE shape. (B) Q plot comparing normalized AI DFE quantiles against synthetic VSV virus DFE (r = 0.89). (C) DFE category proportions across AI and biological organisms (all biological DFEs are parametric reconstructions). AI shows an el- evated beneficial fraction (13%) relative to all biological systems (1â6%). (D) Shape space (skewness vs. excess kurtosis). The AI DFE (star) falls within the cloud of biological DFEs (red circles). See Table 1 for comprehensive metric comparison. 10 42024 Fitness effect 0.0 0.2 0.4 0.6 0.8 1.0 Cumulative probability A DFE by mutation type major (n=568) minor (n=213) intermediate (n=154) 42024 Fitness effect 0.0 0.2 0.4 0.6 0.8 1.0 Cumulative probability B DFE by ML domain CV (n=345) NLP (n=369) Audio (n=65) Other (n=149) 42024 Fitness effect 0.0 0.2 0.4 0.6 0.8 1.0 Cumulative probability KS = 0.287 p = 0.0000 C Selection bias check: manual vs LLM Manual curation (n=140) LLM extraction (n=795) 0.00.10.20.30.40.50.6 Gamma shape parameter () AI (major) D. melanogaster M. musculus H. sapiens Probability density D DFE temporal evolution Figure 2: DFE stratification and universality. (A) CDF of fitness effects by mutation type: major ablations (n = 568, component removal) are more deleterious than minor ablations (n = 213, hyper- parameter changes), mirroring the biological pattern of deletions vs. point mutations. (B) DFE by ML domain: CV, NLP, Audio, and other domains show statistically similar distributions. (C) Methodological control: manually curated data (n = 140) vs. LLM-extracted data (n = 795). The systematic difference (KS = 0.287) reflects manual curationâs failure to capture beneficial mutations, confirming that automated extraction reduces rather than introduces bias. (D) Gamma shape parameter comparison: AI major abla- tions (β = 0.65) compared against the three organisms with published population-genetic β estimates (D. melanogaster â 0.4, M. musculus â 0.3, H. sapiens â 0.2). 11 2012201420162018202020222024 Year 0.0 2.5 5.0 7.5 10.0 12.5 15.0 17.5 New architectures AlexNet GAN + Seq2Seq ResNet Transformer GPT-3 / ViT CLIP + Diffusion LLaMA + Mamba A AI origination rate (red = key innovations) 2012201420162018202020222024 Year 0 20 40 60 80 100 120 Cumulative architectures B Cumulative diversity (logistic saturation) Observed Logistic (K=142, R²0.99) 2012201420162018202020222024 Year 0 2 4 6 8 10 12 14 16 18 New architectures C Adaptive radiation into domain niches CV NLP Audio Multimodal RL Graph 3210123 Normalized time (peak = 0) 0.0 0.2 0.4 0.6 0.8 1.0 1.2 1.4 Normalized origination rate D Universal radiation curve (smoothed) AI architectures Trilobites (Cambrian) Mammals (post K-Pg) Figure 3: Diversification dynamics. (A) Annual origination rate of AI architectures, 2012â2024. Red bars mark key innovations that triggered radiation events. (B) Cumulative diversity follows a logistic curve (R 2 = 0.994, K â 142), indicating approach to saturation. (C) Domain niche-filling: CV saturates first, followed by NLP, then Audio and Multimodalâanalogous to ecological succession. (D) Normalized and smoothed radiation curves for AI architectures, Cambrian trilobites, and post-K-Pg mammals. All three show the same qualitative pattern: accelerating rise, peak, and decline. 12 012345 Independent inventions Contrastive self-supervised learning Positional encoding Gating / multiplicative interaction Feature normalization Attention mechanism Patch-based tokenization Encoder-decoder architecture Learned/structured data augmentation Multi-scale feature processing Stochastic regularization Skip / residual connections Knowledge distillation Mixture of Experts / conditional computation Depthwise separable convolution A AI architectural convergences AI architectures Biology ( 20 inv.) 2.5 5.0 7.5 10.0 12.5 15.0 17.5 20.0 Independent inventions per trait B Convergence intensity (p = 0.035) 2.01.51.00.50.0 Fitness effect 0.0 0.2 0.4 0.6 0.8 1.0 CDF C Transformer lineage DFE evolution Transformer (2017) BERT (2019) RoBERTa (2019) T5 (2020) SwitchTransformer (2022) 12345 Generational distance / genome complexity 0.2 0.3 0.4 0.5 0.6 0.7 0.8 0.9 1.0 Fraction deleterious RNA virus Bacteriophage E. coli YeastDrosophila Human D Universal: more optimized more constrained AI: CNN AI: Transformer AI: ViT Biology (genome complexity) Figure 4: Convergent evolution and lineage maturation. (A) AI architectural convergences: 14 traits independently invented 3â5 times. Colors indicate functional category. (B) Convergence intensity: AI traits (3â5 inventions) vs. biological traits (⤠20 inventions, excluding outliers). MannâWhitney p = 0.035. (C) Transformer NLP lineage: DFE narrows from Transformer (2017) through BERT (2019) to Switch Transformer (2022), showing progressive optimization. (D) Universal trend: fraction deleterious increases with system maturity across both AI lineages (circles) and biological organisms (diamonds). AI and biology fall on a common continuum. 13