Paper deep dive
Evolving Many Worlds: Towards Open-Ended Discovery in Petri Dish NCA via Population-Based Training
Uljad Berdica, Jakob Foerster, Frank Hutter, Arber Zela
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/14/2026, 2:32:14 AM
Summary
PBT-NCA is a meta-evolutionary algorithm that evolves a population of Petri Dish Neural Cellular Automata (PD-NCA) to achieve open-ended discovery of complex, lifelike emergent phenomena. By utilizing a composite objective function that combines archive-based behavioral novelty and DINOv2-based visual diversity, the system avoids common failure modes like frozen equilibria or monocultures, sustaining effective complexity at the 'edge of chaos'.
Entities (4)
Relation Signals (3)
PBT-NCA â evolves â PD-NCA
confidence 100% ¡ PBT-NCA, a meta-evolutionary algorithm that evolves a population of PD-NCAs
PBT-NCA â uses â DinoV2
confidence 95% ¡ population-level visual diversity score derived from a frozen DINOv2 encoder
PD-NCA â exhibits â Emergent Phenomena
confidence 90% ¡ PD-NCA spontaneously generates a plethora of emergent lifelike phenomena
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The generation of sustained, open-ended complexity from local interactions remains a fundamental challenge in artificial life. Differentiable multi-agent systems, such as Petri Dish Neural Cellular Automata (PD-NCA), exhibit rich self-organization driven purely by spatial competition; however, they are highly sensitive to hyperparameters and frequently collapse into uninteresting patterns and dynamics, such as frozen equilibria or structureless noise. In this paper, we introduce PBT-NCA, a meta-evolutionary algorithm that evolves a population of PD-NCAs subject to a composite objective that rewards both historical behavioral novelty and contemporary visual diversity. Driven by this continuous evolutionary pressure, PBT-NCA spontaneously generates a plethora of emergent lifelike phenomena over extended horizons-a hallmark of true open-endedness. Strikingly, the substrate autonomously discovers diverse morphological survival and self-organization strategies. We observe highly regular, coordinated periodic waves; spore-like scattering where homogeneous groups eject cell-like clusters to colonize distant territories; and fluid, shape-shifting macro-structures that migrate across the substrate, maintaining stable outer boundaries that enclose highly active interiors. By actively penalizing monocultures and dead states, PBT-NCA sustains a state of effective complexity that is neither globally ordered nor globally random, operating persistently at the "edge of chaos".
Tags
Links
- Source: https://arxiv.org/abs/2604.11248v1
- Canonical: https://arxiv.org/abs/2604.11248v1
Trouble viewing inline? Open PDF directly â
Full Text
44,551 characters extracted from source content.
Expand or collapse full text
Evolving Many Worlds: Towards Open-Ended Discovery in Petri Dish NCA via Population-Based Training Uljad Berdica 1â Jakob Foerster 1 Frank Hutter 4,2,3 Arber Zela 2â 1 FLAIR, University of Oxford 2 ELLIS Institute TĂźbingen 3 University of Freiburg 4 Prior Labs uljadb@robots.ox.ac.ukarber.zela@tue.ellis.eu Abstract The generation of sustained, open-ended complexity from local interactions remains a fundamental challenge in artifi- cial life. Differentiable multi-agent systems, such as Petri Dish Neural Cellular Automata (PD-NCA), exhibit rich self- organization driven purely by spatial competition; however, they are highly sensitive to hyperparameters and frequently collapse into uninteresting patterns and dynamics, such as frozen equilibria or structureless noise. In this paper, we intro- duce PBT-NCA, a meta-evolutionary algorithm that evolves a population of PD-NCAs subject to a composite objective that rewards both historical behavioral novelty and contem- porary visual diversity. Driven by this continuous evolution- ary pressure, PBT-NCA spontaneously generates a plethora of emergent lifelike phenomena over extended horizonsâa hallmark of true open-endedness. Strikingly, the substrate autonomously discovers diverse morphological survival and self-organization strategies. We observe highly regular, coor- dinated periodic waves; spore-like scattering where homoge- neous groups eject cell-like clusters to colonize distant terri- tories; and fluid, shape-shifting macro-structures that migrate across the substrate, maintaining stable outer boundaries that enclose highly active interiors. By actively penalizing mono- cultures and dead states, PBT-NCA sustains a state of effective complexity that is neither globally ordered nor globally ran- dom, operating persistently at the âedge of chaosâ. Data/Code available at: arberzela/pbt-nca Introduction The quest to understand how simple, local interactions give rise to sustained, open-ended complexity is a central chal- lenge unifying evolutionary biology, artificial life, and non- equilibrium physics (Wolfram, 1983, 1984; Holland, 1992; Taylor et al., 2016). In nature, open-endedness is observed as the endless production of novel and non-trivial structures, allowing for the accumulation of behavioral and structural diversity (Bedau, 1996; Taylor et al., 2016; Lehman and Stanley, 2011; Stanley et al., 2019; Stanley, 2019). Replicat- ing this phenomenon for artificial systems entails a funda- â Equal Contribution. Meta-iteration t 1. Rollout & Score 2. Archive Update AâAâŞd 1 ,d 2 (top-m worlds, FIFO) 3. Exploit- Explore Create Offsprings Copy Crossover MutatePerturb Meta-iteration t+1 Legend Elite (top) Mid-rank Low-rank New child Discarded Figure 1: Overview of one PBT-NCA meta-iteration. Each colored square represents a PD-NCA world (weights, opti- mizer, world state, and hyperparameters). In Stage 1, all worlds are rolled out and scored by a fitness combining archive-based novelty and DINOv2 cross-world temporal di- versity. In Stage 2, the top-mbehavior descriptors are added to a FIFO archive. In Stage 3, the lowest-fitness worlds are replaced by mutated copies of elite parents. mental transition from static optimization to perpetual mor- phological and functional innovation, a property believed to be essential for achieving artificial superhuman intelligence (ASI) (Hughes et al., 2024). Many works suggest that this re- quires operating at the âedge of chaos,â a delicate phase tran- sition where dynamics are neither globally ordered and highly compressible, nor random and fully structureless (Langton, 1990; Kauffman, 1993; Zhang et al., 2025b). Cellular Automata (CA) (Wolfram, 1983, 1984, 2002) have long served as a powerful framework for modeling emergent phenomena from simple, local interactions. Building on this discrete foundation, continuous-state CA like Lenia (Chan, 2018, 2020, 2023; Plantec et al., 2023) demonstrated that arXiv:2604.11248v1 [cs.NE] 13 Apr 2026 060120180240300360420480 Meta-iteration 0.10 0.20 0.30 0.40 0.50 0.60 0.70 0.80 0.90 Mean Composite Score iter 20 iter 60 iter 110 iter 120 iter 230 iter 280 iter 295 iter 310 iter 350 iter 355 iter 395 iter 415 iter 460 rolling mean smoothed (window = 10) PBT-NCA ¡ Emergent Dynamics Throughout Meta-Iterations Figure 2: PBT-NCA evolving a population of 30 PD-NCA worlds, each with 7 NCA agents competing for territory. We plot the composite score function over meta-iterations and rollouts from the highest scoring world. Novel dynamics emerge from agentic competition in an open-ended progression. Videos available at: pbt-nca.github.io. hand-designed mathematical rules in continuous domains could produce a broad spectrum of lifelike patterns. More recently, this concept has been extended to deep learning through Neural Cellular Automata (NCA) (Mordvintsev et al., 2020). In NCAs, the local update rule is learned by training a neural network to grow, maintain, and regenerate target morphologies on a 2D (Niklasson et al., 2021; Randazzo et al., 2020, 2021) or 3D (Sudhakaran et al., 2021) grid. Petri Dish Neural Cellular Automata (PD-NCA) (Zhang et al., 2025a) extend conventional NCAs to a competitive, continual multi-agent setting, in which multiple NCA agents compete for territorial expansion within a shared substrate, hereafter referred to as a world. However, in the absence of carefully designed loss functions or well-tuned hyperparame- ters, the high-dimensional parameter space of PD-NCAs con- tains a lot of regions that lead to uninteresting, non-complex behaviors where system dynamics often freeze into static equilibria, devolve into high-entropy noise, or collapse to monocultures in which a single NCA agent uniformly domi- nates the substrate. To overcome these limitations, we introduce Population- Based Training of Petri Dish Neural Cellular Automata (PBT-NCA). As illustrated in Figure 1, PBT-NCA applies population-level exploit-and-explore cycles (Jaderberg et al., 2017) to evolve a parallel population of PD-NCA worlds under a novelty-driven selection pressure. By combining an archive-based score function (Lehman and Stanley, 2011) computed from handcrafted descriptors with a population- level visual diversity score derived from a frozen DINOv2 encoder (Oquab et al., 2023), we score each world by how different it is from both the history of past behaviors and the rest of the current population, rather than rewarding them for reaching a fixed target (Mordvintsev et al., 2020; Niklas- son et al., 2021; Randazzo et al., 2021). While previous methods like ASAL (Kumar et al., 2025) or ASAL++(Baid et al., 2025) use foundation models for outer-loop search over frozen world configurations, PBT-NCA directly applies selection pressure to a population where individual agents interact with each-other and adapt their local update rules. Emergent Symbiosis at the Edge of Chaos. Under our novelty-driven selection mechanism, PBT-NCA evolves worlds that spontaneously generate strikingly autopoietic lifelike entities (Figure 2), where structural regularities per- sist in the substrate, rather than collapsing to trivial equilibria or noise. We observe emergent phenomena including fluid, shape-shifting macro-structures with coordinated motion, dis- tributed territorial clusters self-replicating via long-range colonization, and the gradual formation of spatial organiza- tion. We further identify decentralized, trail-like locomotion patterns arising purely from local interactions, where agents start coordinating without centralized control (Figure 5). Cru- cially, these behaviors emerge without explicit objectives for cooperation or structure, but as a byproduct of continual pres- sure for novelty and adaptation. By evolving a population of worlds, PBT-NCA acts as an engine of open-endedness, enabling the ongoing accumulation of behavioral and struc- tural novelty over long time horizons. The resulting worlds reside at the âedge of chaosâ, a regime between order and randomness where rich, adaptive dynamics emerge (Langton, 1990; Kauffman, 1993; Mitchell et al., 1994). Population-Based Training of Petri Dish NCAs We extend the PD-NCA substrate (Zhang et al., 2025a) with a population-based meta-optimization loop designed to favor continual discovery of new ecological regimes (Jaderberg et al., 2017; Liu et al., 2019; Faldor and Cully, 2024b; Chan, 2023). The main motivation is that in single-world train- ing, once a particular interaction pattern becomes dominant, gradient descent tends to reinforce it, often leading to mono- cultures, frozen equilibria, or unstructured fluctuations. We therefore make selection explicitly non-stationary. Rather than rewarding one fixed target behavior, we score each world relative to both a memory of past behaviors and the diversity of the current population (Lehman and Stanley, 2011; Pugh et al., 2016). The goal is thus not to converge to one opti- mum, but to continually generate and maintain a population of worlds that remain behaviorally novel, visually distinct, and dynamically active. We name our method PBT-NCA and provide the pseudocode in Algorithm 1. The PD-NCA Substrate A PD-NCA 1 worldâŚis a differentiable multi-agent cellu- lar system whereNneural network (agents) act as distinct âspeciesâ competing for territory on a shared grid. At time t, the global stateX t â R HĂWĂC , whereHandWde- note the grid height and width, consists of a feature vector x t u,v â R C at each spatial position(u,v). This vector is ex- plicitly partitioned into attack channelsa t u,v â R C a , defense channelsd t u,v â R C d , and a hidden stateh t u,v â R C h , where C = C a +C d +C h , that allows agents to process information separately:x t u,v = [a t u,v ; d t u,v ; h t u,v ] . Alongside these cell- level traits, the system maintains a global âalivenessâ mask A t â R NĂHĂW that tracks exactly which agents occupy which cells. At every time step, each living agentkobserves its immediate Moore neighborhoodN u,v (X t )centered at (u,v). Based on this local view, its internal neural network parameterized byθ k decides how to act, proposing updates âx t,k u,v = f θ k (N u,v (X t )) . To ensure the agents never stop adapting, a background environment (k = 0) akin to a harsh climate or natural decay, constantly proposes its own ran- dom, static updates via aL 2 -normalized uniform noise tensor âx t,0 u,v = E u,v , drawn fromU (â1, 1), forcing agents to main- tain active defenses just to survive and effectively preventing evolutionary stagnation. Multiple agents can allocate resources, e.g. for attack or defense, to the same cell; the environment can also up- date the cellâs conditions at the same time. This compe- tition between agentskandlfor cell(u,v)is formalized via pairwise cosine similarities:Ď kl (u,v) =â¨a t,k u,v , d t,l u,v âŠâ â¨d t,k u,v , a t,l u,v âŠ. The total competitive strength of agentkat cell(u,v)is the sum of its pairwise interactions with ev- ery other agent and with the environment:Ψ k (u,v) = P l̸=k Ď kl (u,v) + Ď k,env (u,v). Crucially, this competition 1 https://pub.sakana.ai/pdnca/ Algorithm 1 PBT-NCA pseudocode Require:Population sizeP, meta-iterationsT, PD-NCA train epochsT world , exploit intervalK, replacement fractionĎ, archive incrementm, crossover and perturb probabilities P cross , P pert 1: Init. random population⌠i P i=1 , archiveAâ â 2: for t = 1 to T do 3:Train each ⌠i for T world steps; compute scoreF i 4:Add top-m descriptors toA 5:if t mod K = 0 then 6:for all ⌠c in bottom Ď fraction do 7:Sample parent ⌠p from top Ď fraction 8:Copy parent to child: ⌠c â ⌠p 9:Crossover each hparam w.p. P cross 10: Perturb each hparam byĂ(1Âą0.2)w.p.P pert 11:Add Gaussian noise to ⌠c weights 12:end for 13:end if 14: end for is not strictly winner-takes-all; strengths are normalized into contribution weightsw t,k u,v = softmax(Ψ k (u,v)) , meaning the cellâs new state is a weighted blend of the competing proposals:x t+1 u,v = clip x t u,v + P N k=0 w t,k u,v âx t,k u,v , with clipping enforcing valid ranges. These shares also dictate an agentâs ongoing aliveness, updated viaA t+1,k u,v = w t,k u,v if w t,k u,v > Îą , and0otherwise. If an agentâs share drops below Îą = 0.4, it dies off in that location and its remaining influ- ence is redistributed. This specific threshold allows up to two different agents to coexist in a single cell simultaneously, ensuring that territorial boundaries remain fluid and dynamic. A PD-NCA world evolves on two coupled levels: the grid state changes through local interactions, while the agents concurrently adapt their internal parameters. Ev- eryĎrollout steps, each agentkseeks territorial expan- sion by maximizing its log-aliveness across the entire grid: L k =â log P u,v A t+Ď,k u,v , and updating their parameters via gradient descent,θ k â θ k â Ρ k â θ k L k , whereΡ k is the current learning rate. This simple, gradient-based drive for growth naturally forces the substrate to self-organize into complex dynamics. Population State and World Rollouts We maintain a population ofPworlds. A world encompasses the full state of an ongoing adaptive ecosystem. Concretely, each world stores: (i) the parameters of all resident NCAs; (i) their optimizer states; (i) the mutable world state, including the grid contents, rollout counters, and any auxiliary state needed by the simulator; and (iv) learning hyperparameters controlling within-world adaptation, such as learning rate, batch size, and update frequency. At each meta-iteration, every world is rolled out for a fixed Trajectories t=1t=2 ¡ t=T ⌠1 ¡ ⌠2 ¡ ⌠3 ¡ one timestep count alive Species fractions env c 1 c 2 t=1 t=2 . . . t=T .55 .30.15 .40.35.25 .20.45.35 col Features Îź Ď Î´ .38.37.25 .14 .06.10 .18.08 .12 H win , ν .87.23 from argmax /|âa| cat â 2 Descriptor Îź Ď Î´ H,ν d i â R 3(N+1)+2 Archive novelty k-N archiveA, FIFO buffer N i = 1 k X jâkNN âĽd i âd j ⼠2 Figure 3: Handcrafted behavior descriptor and novelty score. Each world is rolled out; the firstN + 1channels of the trajectory (alive mass per entity) are spatially summed and row-normalised to produce species-fraction trajectories. The per-entity statistics are extracted: mean occupancyÎź, temporal standard deviationĎ, and turnoverδ. We use a k = 8 for the kNN. Trajectories t=1t=2 ¡ t=T ⌠1 ¡ ⌠2 ¡ ⌠3 ¡ one timestep Embed DINOv2 z t i â R 768 Distance d t ij = 1ââ¨z t i , z t j ⊠⌠1 ⌠2 ⌠3 ⌠1 ⌠2 ⌠3 â .82.77 .82 â .29 .77.29 â Row median med = .80 med = .56 med = .53 avgât D i = 1 T T X t=1 med j̸=i d t ij Figure 4: Visual diversity score. At each timestept, frames from allPworlds are embedded using a DINOv2 encoder. We compute the cosine distance matrix with masked diago- nals to score each world through its median distance to all other worlds, averaged across the rollout horizon. horizon ofT world inner steps. During this rollout, cellular updates and gradient-based learning remain interleaved, so the worldâs agents continue adapting while interacting on the grid. This produces a trajectoryT = (X 1 ,...,X T world ), which we utilize as the object for evaluating each worldâs novelty score. Scoring full trajectories rather than terminal frames is important because worlds with similar endpoints may differ greatly in coexistence ratios between different NCAs across time, transient competition or point of collapse. Composite Scoring Function We score each world⌠i ,i â 1,...,P, with a composite scalar functionF i =N i +D i , whereN i is an archive-based novelty score computed from a handcrafted behavioral de- scriptor, andD i is a population-level visual diversity score computed from image embeddings. These terms are com- plementary:N i captures how distinctly a world⌠i behaves relative to previously discovered ones, whileD i captures how much the worldâs individual timesteps differ from those of other worlds currently being trained in the population. Behavioral Descriptor and Archive Novelty.The archive novelty term is based on handcrafted features of substrate dynamics as illustrated in Figure 3. Instead of comparing high-dimensional grid states directly, we summarize each trajectory using statistics of fundamental concepts in ecology like species occupancy and change (Magurran, 2004). For each inner time step, we compute the fraction of total alive mass belonging to each entity, i.e., each of theNagents and the environment channel. We extract the following feature groups from the timeseries of each speciesâ fractions: 1. Mean occupancyÎź: the average fraction of alive mass assigned to each entity over the rollout. 2.Temporal standard deviationĎ: the variability of each entityâs occupancy over time. 3. Mean frame-to-frame occupancy changeδ: the average how quickly control of a cell changes from one agent (NCA) to another. 4. Winner-map entropyH: at each spatial location, we identify the dominant alive channel and compute the en- tropy of the winnerâs distributions across space and time to summarize how mixed or monopolized the world is. 5.Mean absolute alive-mass changeν: the average frame- to-frame alive mask change. These statistics are concatenated andâ 2 -normalized to form the feature vectord â R 3(N +1)+2 . This descriptor is more stable under small perturbations in the world state compared to the high-dimensional original frames and ensures that sep- arate ecological regimes such as collapse, oscillation, domi- nance, stable or volatile coexistence have distinct descriptors. We store past descriptors in a bounded FIFO archive A(Pugh et al., 2016; Lehman and Stanley, 2011). After each meta-iteration, the descriptors of the top-m(m = 2) worlds under the current novelty score are appended to the archive; if the archive exceeds its capacityA max , the oldest entries are removed first. Optionally, the archive may be cleared everyRmeta-iterations to reset the novelty signal once too many previously discovered regimes have accumu- lated. Given the archive, the novelty scoreN i of a world⌠i is the mean Euclidean distance from its descriptor to thek (k = 8) nearest neighbors inA(Lehman and Stanley, 2011), making it interesting when it differs from what has already been discovered. meta-iter 20meta-iter 125meta-iter 370 Figure 5: Shooters, Archipelago and Ant locomotion emerging from PBT-NCA (3 NCAs) at meta-iteration 20, 125 and 370. Figure 6: Self-replication emerging in PBT-NCA (3 NCA agents). A macroscopic entity (red agent) emits a small cluster of cells across the substrate, executing spatial colonization via a decentralized replicating strategy. Visual Diversity with DINOv2. The archive descriptor captures relevant ecological statistics but ignores many vi- sual properties of the rollout, such as morphology, spatial arrangement, and color pattern. We draw inspiration from ASAL (Kumar et al., 2025) to define a second term (Figure 4) based on a frozen DINOv2 vision encoder (Oquab et al., 2023) to complement it. For each world, we extract all the frames from its trajectory and encode them into a DINOv2 embedding. Letz t i denote the embedding of world⌠i at sampled framet. For each frame, we compare worldionly to the other worldsj ̸= iat the same sampled time point and compute cosine distances:d t ij = 1ââ¨z t i , z t j âŠ. To ensure robustness to outliers, the per-frame diversity of world⌠i is defined as the median pairwise distance across the population. Averaging over sampled frames gives the final diversity score D i = 1/T P T t=1 median j̸=i (d t ij ) , whereTis the number of sampled frames. Intuitively,D i is high when a world produces visual structures that the rest of the population is not currently producing. Since DINOv2 provides a rich per- ceptual representation without any task-specific labels, this term rewards novel spatial texture or morphology even when the handcrafted descriptor may remain similar across the population of worlds being selected. Table 1: PBT hyperparameter search space. â means that the hyperparameter is in the extended search space. HPTypeRangeLogDefault learning_ratefloat[10 â6 , 1.0]Yes3Ă 10 â4 batch_sizeint[1, 8]No8 steps_per_updateint[1, 64]No4 â softmax_tempfloat[0.05, 5.0]Yes1.0 â per_hid_updfloat[0.05, 1.0]No1.0 ExploitâExplore Once all worlds have been rolled out and scored, we rank the population by the composite score functionF. The top- mbehavioral descriptors are inserted into the archive as described above. EveryKmeta-iterations, we then perform a population-based exploitâexplore (Jaderberg et al., 2017; Liu et al., 2019). We set the number of worlds to replace to n replace = round(ĎP ), whereĎâ (0, 1)is the replacement fraction andPis the population size. Then replace lowest- scoring worlds are discarded and replaced by higher-scoring worlds. For each replacement, a parent is sampled uniformly from the elite subset, and the child is generated in four stages: 1.Copy. The parent world is deep-copied into the child. This includes neural-network parameters, optimizer state, world state, and hyperparameters. The inheritance is there- fore Lamarckian: descendants do not restart from random initialization but continue from a partially developed eco- logical and optimization state. 2. Crossover. Each hyperparameter is independently crossed over. With probabilityP cross (0.5in our experiments), the copied parent value is retained; otherwise, the child reverts to its own pre-copy value. Thus, the search mixes successful inherited settings with previously held settings. 3.Mutation. Each hyperparameter is perturbed indepen- dently with probabilityP pert (0.1in our experiments). The perturbation multiplies the current value by either1.2 or 0.8, chosen with equal probability, and clips the result to its allowed range. Integer hyperparameters are rounded after scaling. 4. Weight perturbation. Finally, independent Gaussian noise is added to all trainable neural network parameters. This injects local parametric variation without destroying the inherited world context. Steps 1 to 3 are directly adapted from Jaderberg et al. (2017) and Liu et al. (2019). The exploit phase (step 1) amplifies lineages that currently produce novel and diverse trajecto- ries, whereas the explore phase (steps 2-4) evaluates nearby dynamical possibilities by perturbing both hyperparameters and weights of the worlds in the population. Overall, the iter 100iter 50000iter 100000 Figure 7: Original PD-NCA (3 NCA agents) frames on itera- tions 100, 50000 and 100000. meta-iter 145meta-iter 485meta-iter 495 Figure 8: Random Search (3 NCA agents) frames on meta- iterations 145, 485 and 495. method couples two time scales. Within a rollout, gradient- based adaptation shapes the behavior of each world. Across meta-iterations, population-based selection decides which adaptive ecological regimes survive, diversify, and continue developing. This combination is what allows the system to search complex dynamics in an ever-diversifying population of PD-NCAs, which would be difficult to discover with a non-dynamical outer meta-optimization loop. Experiments and Simulations Our experiments aim to quantify the extent to which PBT- NCA sustains the continuous discovery of emerging complex dynamics and behaviours in the population. In all following experiments we use the hyperparameter space shown in Ta- ble 1. We set the population size toP = 30, meta-iterations toT = 500, world horizon toT world = 12, explore-exploit interval toK = 5, and the replacement factor toĎ = 0.25. We use Adam (Kingma and Ba, 2015) to optimize the NCA weights with the sampled/mutated learning rate. Unless stated otherwise, we use 3 NCA agents in each world. Fig- ure 5 and 6 demonstrate several complex dynamics, such as directed projectile emission and colonization from stable ter- ritorial clusters, as well as agents following learned territorial trailsâall emerging purely from local NCA competition. Comparison to Baselines. We compare our PBT-NCA against two other configurations: 1.Fixed-Hyperparameters PD-NCA. The original imple- mentation from Zhang et al. (2025a) using the default fixed hyperparameters from the original implementation. 2. Random Search (RS). Hyperparameters are sampled uni- formly from the PBT-NCA search space (Table 1). 80160240320400480 Meta-iteration 0.20 0.30 0.40 0.50 0.60 Mean Composite Score mean Âą std across 3 runs ¡ rolling mean smoothed (window = 10) 3 NCAs5 NCAs7 NCAscomposite score (left axis)novelty (right axis) 0.00 0.10 0.20 0.30 0.40 Mean Novelty PBT-NCA ¡ Mean Composite Score & Novelty over Meta-iterations 10 6 10 5 10 4 10 3 10 2 10 1 Learning Rate 2 4 6 8 Batch Size 060120180240300360420480 Meta-iteration 1 2 3 4 Steps per Update PBT-NCA ¡ Best-Member Hyperparameters over Meta-iterations Figure 9: Top figure: Mean novelty and composite score function over the number of meta-iterations. Bottom figure: Sampled hyperparameters over the number of meta-iterations for the best PBT-NCA (7 NCA agents). To match the compute budget of PBT-NCA, we use the same population size,P = 30, for RS. Each world was trained (rolled out) for a total of 6000 iterations (T = 500meta- iterationsĂ T world = 12iterations). RS discovers some interesting regular cyclic dynamics (Figure 8), whereas PD- NCA exhibits high entropy noise (Figure 7). Novelty, Open-Endedness, and Scalability To evaluate population-level performance, we track the mean composite score, Ě F = 1/P P P i=1 F i , across all members of the population. Figure 9 (top) illustrates the progression of Ě Fover the PBT-NCA meta-iterations (mean and standard deviation from three independent PBT-NCA runs) for worlds containing 3, 5, and 7 NCA agents. Across all configura- tions, an initial decline is followed by a steady increase in the composite score, suggesting the emergence of a diverse population of PBT-NCA substrates capable of continually generating novel dynamics (see Figures 2 and 10). Further- more, to isolate the individual components of the composite score, we plot the mean novelty, Ě N, in the same figure. We observe that worlds scaled to support a larger number of NCAs successfully retain higher levels of novelty in the later stages of training. We observe a strong selection pressure toward higher learning rates and smaller batch sizes (Figure 9, meta â iter 145 meta â iter 230 meta â iter 310 meta â iter 395 meta â iter 460 Figure 10: Emergent dynamics at meta-iteration 145, 230, 310, 395 and 495 of PBT-NCA (7 NCA agents), demonstrating the continuous novelty generation, a sign of open-endedness. We discover novel entities, such as spirals (145), amoeba (230), pirate world (archipelago+shooters) (310), alien (395) and glider (460). meta-iter 20 meta-iter 95meta-iter 125 Figure 11: Rigid, circuit-like geometric patterns discovered with PBT-NCA (3 NCA agents) on the extended search space. We notice the emergence of spaceships, gliders and information propagation, as discovered in continuous or discrete CA. bottom). By injecting greater gradient noise and accelerating parameter updates, this combination prevents worlds from collapsing into stagnant equilibria, directly supporting our novelty-seeking objective. Furthermore, when granted ac- cess to the extended search space (Figure 11), PBT-NCA discovers qualitatively distinct dynamics, abandoning fluid, organic morphologies in favor of rigid, circuit-like geometric expansions. This demonstrates that PBT-NCA can effectively scale also along the hyperparameter dimensions to unlock entirely new classes of emergent behavior. Edge-of-Chaos Analysis To quantify whether the emergent dynamics occupy the edge of chaosâthe narrow regime between frozen order and struc- tureless randomness where complex behaviour arises Lang- ton (1990); Kauffman (1993)âwe track two complementary metrics over PBT-NCA with 7 agents across 500 meta-steps. Ecological Persistence. LetH t be the per-pixel agentâs aliveness distribution entropy at timet. We defineEP =|t : H t > Îľ|/T withÎľ = 0.1 log 2 N, measuring the fraction of frames that sustain multi-agent coexistence. HighEPmeans the system avoids both extinction and monocultures. Effective Complexity. Following Gell-Mann and Lloyd (1996), we defineC t eff = H(grid t ) Ă (1 â L Z77 (grid t )), whereH(grid t )is the normalised spatial entropy andL Z77 is compressibility ratio. This product vanishes for both pure or- der (Hâ 0) and pure noise (L Z77 â 1), and peaks only when patterns are simultaneously varied and structurally compress- ible, the key characteristic of the edge of chaos. Results. Ecological Persistence remains atEP â 1.0 throughout training (Fig. 12, bottom), denoting that every 0 1 2 H(species) [bits] Species Entropy over Rollout Frames 10 0 10 1 10 2 Rollout frame 0.0 0.1 0.2 0.3 C eff Effective Complexity over Rollout Frames 100 200 300 400 Meta-training step 0100200300400500 Meta-training step 0.0 0.2 0.4 0.6 0.8 1.0 C eff & EP fraction Ecological Persistence & Effective Complexity 0.2 0.3 0.4 0.5 0.6 0.7 0.8 Mean Composite Score Figure 12: Ecological Persistence (EP) and Effective Com- plexity throughout meta-iterations and individual rollouts of PBT-NCA (7 NCA agents) worlds. computed rollout sustains multiple species above threshold at every frame, with mean species entropy Ě H â 2.3 bits (âź82% of the maximumlog 2 7) (Figure 12, top). Effective complex- ity averagesC eff â 0.21Âą 0.05, above the zero baselines of both ordered and chaotic regimes. Within-rollout trajectories (Figure 12, top) confirm that later runs maintain flat, elevated C t eff . This shows that PBT-NCAâs competitive pressure drives the population toward a stable attractor at the edge of chaos, sustaining robust coexistence and structurally rich dynamics. Discussion The quest for open-ended complexity remains a defining chal- lenge in artificial life (Stanley, 2019; Taylor et al., 2016). Our results demonstrate that by reframing spatial competition as a novelty-driven, meta-learning process, the PBT-NCA dis- covers and sustains a diverse manifold of emerging lifelike behaviors mirroring phenomena across computational and bi- ological domains. For instance, in Figure 5, 10 (meta-iter 310 and460) and 11, we observe the spontaneous genera- tion of mobile local structures akin to the âgunâ and âglidersâ, similarly as in traditional discrete and continuous CA (Gard- ner, 1970; Berto and Tagliabue, 2012; Chan, 2018). The system produces complex, macroscopic patterns be- yond the scope of localized agents. The traveling spiral waves (Figure 10,meta-iter 145) resemble the rotating spiral waves in stochastic multi-agent ecosystems (Reichenbach et al., 2008) and the nested territorial competition echoes the hypercycles of pre-biotic chemistry (Eigen and Schus- ter, 1977; Cronhjort, 1995). As noted by Crutchfield and Mitchell (1995), the persistence of such structurally stable âregular domainsâ with complex boundaries interacting be- tween them is a prerequisite for emergent computation. This is a promising sign that PBT-NCA may be able to imple- ment computational primitives. Crucially, our framework discovers these complex cycles autonomously, without the pattern-specific handcrafted loss functions required in prior work (Zhang et al., 2025a). Furthermore, the spatial segrega- tion into stable macro-structures and volatile frontiers (Figure 10,meta-iter 460) closely mimics the genetic drift and expanding boundaries of microbial populations (Hallatschek et al., 2007). The ejection of âcell-likeâ clusters to colonize new regions (Figure 6) further highlights this biological re- alism, matching spatial migration observed in nature (Baym et al., 2016; Nadell et al., 2016). Most interestingly, atmeta-iter 230of Figure 10, we observe the emergence of fluid, shape-shifting macro- structures that migrate across the substrate with coordinated, differentiated cell behavior, mirroring motility-induced phase separation (Cates and Tailleur, 2015). The system does not merely visit these complex states transiently; it sustains cel- lular differentiation, self-replication, and coordination at the âedge of chaosâ (Langton, 1990; Kauffman, 1993). We hypothesize that the sustained morphological diversity in PBT-NCA is driven by dynamics akin to the Red Queen hypothesis (Van Valen, 1973; Cliff and Miller, 1995; Ku- mar et al., 2026). The landscape never settles as individual agents must survive an evolving competitive environment. Our composite novelty score rewards lineages that break away from both historical and contemporary norms. PBTâs exploit-explore mechanism ensures that descendants inherit partially developed ecologies, allowing them to compound complex survival strategies rather than being reinitialized. Limitations and Future Work. While PBT-NCA repre- sents a concrete step toward artificial open-endedness, sev- eral limitations remain. The substrate is based on a 2D grid that oversimplifies the thermodynamic constraints driving evolutionary innovation in natural ecosystems (Schneider and Sagan, 2006; SchrĂśdinger, 1992; England, 2013). Addi- tionally, leveraging DINOv2 for visual diversity evaluation may inject anthropocentric, natural-image biases into the systemâs definition of ânoveltyâ. In the future, we aim to leverage hardware-accelerated frameworks like CAX (Fal- dor and Cully, 2024a) and hyperscale evolution strategies (Sarkar et al., 2025) to support larger grids and more agents (Chan, 2023). This will allow us to study the scaling laws of emergent dynamics in PD-NCA, as well as co-evolving other components such as architectures (Stanley and Miikkulainen, 2002; Zoph and Le, 2017; White et al., 2023), environments (Wang et al., 2019; Ferreira et al., 2022; Zhang et al., 2024; Bruce et al., 2024), and the fundamental update strategies (Andrychowicz et al., 2016; Chen et al., 2022). Acknowledgements Uljad Berdica is funded by the Rhodes Scholarship and the EPSRC Centre for Doctoral Training in Autonomous Intel- ligent Machines and Systems. Jakob Foerster is partially funded by the UKRI grant EP/Y028481/1, which was origi- nally selected for funding by the ERC. Jakob Foerster is also supported by the JPMC Research Award and the Amazon Re- search Award. This project received compute resources from a gracious grant provided by the Isambard-AI National AI Research Resource. Frank Hutter acknowledges the financial support of the Hector Foundation. The authors thank Lukas Seier, Hannah Erlebach, Antonio LeĂłn Villares, Michael Beukman and Philipp Bordne for their valuable feedback and discussions. References Andrychowicz, M., Denil, M., Colmenarejo, S. G., Hoffman, M. W., Pfau, D., Schaul, T., Shillingford, B., and de Freitas, N. (2016). Learning to learn by gradient descent by gradient descent. In Proceedings of the 30th International Conference on Neural In- formation Processing Systems, NeurIPSâ16, page 3988â3996, Red Hook, NY, USA. Curran Associates Inc. Baid, N., Erlebach, H., Hellegouarch, P., and Wieser, F. (2025). Guiding evolution of artificial life using vision-language mod- els. In Artificial Life Conference Proceedings 37, volume 2025, page 22. MIT Press One Rogers Street, Cambridge, MA 02142-1209, USA journals-info . . . . Baym, M., Lieberman, T. D., Kelsic, E. D., Chait, R., Gross, R., Yelin, I., and Kishony, R. (2016). Spatiotemporal microbial evolution on antibiotic landscapes. Science, 353(6304):1147â 1151. Bedau, M. A. (1996). Measurement of evolutionary activity, teleol- ogy, and life. Berto, F. and Tagliabue, J. (2012). Cellular automata. Stanford Encyclopedia of Philosophy. Bruce, J., Dennis, M., Edwards, A., Parker-Holder, J., Shi, Y., Hughes, E., Lai, M., Mavalankar, A., Steigerwald, R., Apps, C., et al. (2024). Genie: Generative interactive environments. arXiv preprint arXiv:2402.15391. Cates, M. E. and Tailleur, J. (2015). Motility-induced phase separa- tion. Annu. Rev. Condens. Matter Phys., 6(1):219â244. Chan, B. W.-C. (2018). Lenia - biology of artificial life. Complex Syst., 28. Chan, B. W.-C. (2020). Lenia and expanded universe. Artificial Life, ALIFE 2020: The 2020 Conference on Artificial Life:221â 229. Chan, B. W.-C. (2023). Towards large-scale simulations of open- ended evolution in continuous cellular automata. In Proceed- ings of the Companion Conference on Genetic and Evolution- ary Computation, GECCO â23 Companion, page 127â130, New York, NY, USA. Association for Computing Machinery. Chen, T., Chen, X., Chen, W., Wang, Z., Heaton, H., Liu, J., and Yin, W. (2022). Learning to optimize: a primer and a benchmark. J. Mach. Learn. Res., 23(1). Cliff, D. and Miller, G. F. (1995). Tracking the red queen: Measure- ments of adaptive progress in co-evolutionary simulations. In Proceedings of the Third European Conference on Advances in Artificial Life, page 200â218, Berlin, Heidelberg. Springer- Verlag. Cronhjort, M. B. (1995). Hypercycles versus parasites in the ori- gin of life: model dependence in spatial hypercycle systems. Origins of Life and Evolution of the Biosphere, 25(1):227â233. Crutchfield, J. P. and Mitchell, M. (1995). The evolution of emer- gent computation. Proceedings of the National Academy of Sciences, 92(23):10742â10746. Eigen, M. and Schuster, P. (1977). The hypercycle: A principle of natural self-organization. Part A: Emergence of the hypercycle. Naturwissenschaften, 64(11):541â565. England, J. L. (2013). Statistical physics of self-replication. The Journal of Chemical Physics, 139(12):121923. Faldor, M. and Cully, A. (2024a). Cax: Cellular automata acceler- ated in jax. arXiv preprint arXiv:2410.02651. Faldor, M. and Cully, A. (2024b). Toward artificial open-ended evolution within lenia using quality-diversity. Artificial Life, ALIFE 2024: Proceedings of the 2024 Artificial Life Confer- ence:85. Ferreira, F., Nierhoff, T., Sälinger, A., and Hutter, F. (2022). Learn- ing synthetic environments and reward networks for reinforce- ment learning. In International Conference on Learning Rep- resentations. Gardner, M. (1970). Mathematical games. Scientific American, 223(4):120â123. Gell-Mann, M. and Lloyd, S. (1996). Information measures, effec- tive complexity, and total information. Complex., 2(1):44â52. Hallatschek, O., Hersen, P., Ramanathan, S., and Nelson, D. R. (2007). Genetic drift at expanding frontiers promotes gene segregation. Proceedings of the National Academy of Sciences, 104(50):19926â19930. Holland, J. H. (1992). Adaptation in Natural and Artificial Systems: An Introductory Analysis with Applications to Biology, Control, and Artificial Intelligence. The MIT Press. Hughes, E., Dennis, M., Parker-Holder, J., Behbahani, F., Mavalankar, A., Shi, Y., Schaul, T., and Rocktaschel, T. (2024). Open-endedness is essential for artificial superhuman intelli- gence. arXiv preprint arXiv:2406.04268. Jaderberg, M., Dalibard, V., Osindero, S., et al. (2017). Popu- lation based training of neural networks. In arXiv preprint arXiv:1711.09846. Kauffman, S. A. (1993). The Origins of Order: Self-Organization and Selection in Evolution. Oxford University Press. Kingma, D. P. and Ba, J. (2015). Adam: A method for stochas- tic optimization. In International Conference on Learning Representations (ICLR). Kumar, A., Bahlous-Boldi, R., Sharma, P., Isola, P., Risi, S., Tang, Y., and Ha, D. (2026). Digital red queen: Adversarial program evolution in core war with llms. Kumar, A., Lu, C., Kirsch, L., Tang, Y., Stanley, K. O., Isola, P., and Ha, D. (2025). Automating the search for artificial life with foundation models. arXiv preprint arXiv:2412.17799. Langton, C. G. (1990). Computation at the edge of chaos: Phase transitions and emergent computation. Physica D: Nonlinear Phenomena, 42(1):12â37. Lehman, J. and Stanley, K. O. (2011). Abandoning objectives: Evolution through the search for novelty alone. Evolutionary Computation, 19(2):189â223. Liu, S., Lever, G., Heess, N., Merel, J., Tunyasuvunakool, S., and Graepel, T. (2019). Emergent coordination through compe- tition. In International Conference on Learning Representa- tions. Magurran, A. (2004). Measuring Biological Diversity. Wiley. Mitchell, M., Hraber, P. T., and Crutchfield, J. P. (1994). Revisiting the edge of chaos: Evolving cellular automata to perform computations. Complex Systems, 7(2):89â130. Mordvintsev, A., Randazzo, E., Niklasson, E., and Levin, M. (2020). Growing neural cellular automata. Distill, 5(2):e23. Nadell, C. D., Drescher, K., and Foster, K. R. (2016). Spatial structure, cooperation and competition in biofilms. Nature Reviews Microbiology, 14(9):589â600. Niklasson,E.,Mordvintsev,A.,Randazzo,E.,and Levin, M. (2021).Self-organising textures.Distill. https://distill.pub/selforg/2021/textures. Oquab, M., Darcet, T., Moutakanni, T., Vo, H., Szafraniec, M., Khalidov, V., Fernandez, P., Haziza, D., Massa, F., El-Nouby, A., et al. (2023). Dinov2: Learning robust visual features without supervision. arXiv preprint arXiv:2304.07193. Plantec, E., Hamon, G., Etcheverry, M., Oudeyer, P.-Y., Moulin- Frier, C., and Chan, B. W.-C. (2023). Flow-lenia: Towards open-ended evolution in cellular automata through mass con- servation and parameter localization. Artificial Life, ALIFE 2023: Ghost in the Machine: Proceedings of the 2023 Artifi- cial Life Conference:131. Pugh, J. K., Soros, L. B., and Stanley, K. O. (2016). Quality diversity: A new frontier for evolutionary computation. In Frontiers in Robotics and AI. Randazzo, E., Mordvintsev, A., Niklasson, E., and Levin, M. (2021). Adversarial reprogramming of neural cellular automata. Distill. https://distill.pub/selforg/2021/adversarial. Randazzo, E., Mordvintsev, A., Niklasson, E., Levin, M., and Grey- danus, S. (2020). Self-classifying mnist digits. Distill. Reichenbach, T., Mobilia, M., and Frey, E. (2008). Self-organization of mobile populations in cyclic competition. Journal of Theo- retical Biology, 254(2):368â383. Sarkar, B., Fellows, M., Duque, J. A., Letcher, A., Villares, A. L., Sims, A., Cope, D., Liesen, J., Seier, L., Wolf, T., Berdica, U., Goldie, A. D., Courville, A., Sevegnani, K., Whiteson, S., and Foerster, J. N. (2025). Evolution strategies at the hyperscale. Schneider, E. D. and Sagan, D. (2006). Into the cool : energy flow, thermodynamics, and life. SchrĂśdinger, E. (1992). What is Life?: With Mind and Matter and Autobiographical Sketches. Cambridge paperback library. Cambridge University Press. Stanley, K. O. (2019). Why open-endedness matters. Artificial life, 25(3):232â235. Stanley, K. O., Lehman, J., and Soros, L. B. (2019). Designing open-ended intelligence. Artificial Life, 25(3):224â246. Stanley, K. O. and Miikkulainen, R. (2002). Evolving neural net- works through augmenting topologies. Evolutionary Compu- tation, 10:99â127. Sudhakaran, S., Grbic, D., Li, S., Katona, A., Najarro, E., Glanois, C., and Risi, S. (2021). Growing 3d artefacts and func- tional machines with neural cellular automata. arXiv preprint arXiv:2103.08737. Taylor, T., Bedau, M., Channon, A., Ackley, D., Banzhaf, W., Beslon, G., Dolson, E., Froese, T., Hickinbotham, S., Ikegami, T., et al. (2016). Open-ended evolution: Perspectives from the oee workshop in york. Artificial life, 22(3):408â423. Van Valen, L. (1973). A new evolutionary law. Evolutionary theory, 1. Wang, R., Lehman, J., Clune, J., and Stanley, K. O. (2019). Paired open-ended trailblazer (poet): Evolvable evolutionary chal- lenges and their solutions through a generative environment model. In Proceedings of the Genetic and Evolutionary Com- putation Conference, pages 142â151. White, C., Safari, M., Sukthanker, R. S., Ru, B., Elsken, T., Zela, A., Dey, D., and Hutter, F. (2023). Neural architecture search: Insights from 1000 papers. ArXiv, abs/2301.08727. Wolfram, S. (1983). Statistical mechanics of cellular automata. Rev. Mod. Phys., 55:601â644. Wolfram, S. (1984). Cellular automata as models of complexity. Nature, 311(5985):419â424. Wolfram, S. (2002). A New Kind of Science. Wolfram Media. Zhang, I., Risi, S., and Darlow, L. (2025a). Petri dish neural cellular automata. Zhang, J., Lehman, J., Stanley, K., and Clune, J. (2024). OMNI: Open-endedness via models of human notions of interesting- ness. In The Twelfth International Conference on Learning Representations. Zhang, S., Patel, A., Rizvi, S. A., Liu, N., He, S., Karbasi, A., Zappala, E., and van Dijk, D. (2025b). Intelligence at the edge of chaos. In The Thirteenth International Conference on Learning Representations. Zoph, B. and Le, Q. (2017). Neural architecture search with rein- forcement learning. In International Conference on Learning Representations.