Paper deep dive
Delayed Repression and Emergent Instability in Adaptive Multi-Agent Systems
Igor Itkin
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 7/8/2026, 10:22:32 PM
Summary
This paper investigates how institutional processing delays alone can destabilize adaptive multi-agent systems without external shocks. Using a delayed replicator equation and networked simulations with 240 agents, the author demonstrates that a critical delay threshold triggers a supercritical Hopf bifurcation, leading to bounded runaway oscillations. Counterintuitively, reactive agents collapse under delay, fixed-policy agents remain immune, and Q-learning agents show partial resilience by encoding punishment history.
Entities (8)
Relation Signals (8)
Delayed Replicator Equation → exhibits → Hopf Bifurcation
confidence 95% · We derive a closed-form critical delay beyond which the unique interior equilibrium loses stability through a Hopf bifurcation
Institutional Delay → triggers → Hopf Bifurcation
confidence 95% · When institutional processing introduces delay, the alarm signal is stale... Above a critical threshold, this stale feedback generates oscillations or runaway.
Institutional Delay → destabilizes → Reactive Agents
confidence 90% · Reactive agents are perfectly stable without delay yet collapse once delay is introduced (96% runaway by delay >= 8)
Institutional Delay → doesnotdestabilize → Fixed-Policy Agents
confidence 90% · fixed-policy agents are immune (0% at all delays)
Hopf Bifurcation → istype → Supercritical Bifurcation
confidence 90% · prove via center manifold reduction that the bifurcation is supercritical (bounded oscillations, not explosive growth) for the entire sigmoid response family.
Institutional Delay → partiallydestabilizes → Q-Learning Agents
confidence 90% · Q-learning agents are only partially resilient (66% at delay 20)
Q-learning → buffers → Institutional Delay
confidence 85% · learning buffers this through punishment memory encoded in value functions.
Sigmoid Response Function → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Regulatory institutions (from content moderation platforms to financial supervisors) observe, deliberate, and intervene only after a characteristic delay. We ask whether this processing lag alone can destabilize a multi-agent system that would otherwise remain stable, without exogenous shocks, coordination among agents, or malicious actors. We study this in two stages. First, we analyze a delayed replicator equation in which autonomous agents benefit from radical behavior but face punishment based on a lagged institutional alarm signal. We derive a closed-form critical delay beyond which the unique interior equilibrium loses stability through a Hopf bifurcation, and prove via center manifold reduction that the bifurcation is supercritical (bounded oscillations, not explosive growth) for the entire sigmoid response family. Second, we embed N=240 agents on a network with reinforcement learning (tabular Q-learning) and cross institutional delay with three decision architectures: fixed-policy, reactive (a memoryless threshold heuristic), and Q-learning. The hierarchy is opposite to the naive expectation that learning amplifies instability. Reactive agents are perfectly stable without delay yet collapse once delay is introduced (96% runaway by delay >= 8); fixed-policy agents are immune (0% at all delays); Q-learning agents are only partially resilient (66% at delay 20). The destabilizing ingredient is reactivity to delayed signals, not learning: agents that immediately exploit low-alarm windows trigger oscillatory feedback loops, while learning buffers this through punishment memory encoded in value functions. Throughout, "runaway" denotes bounded large-amplitude oscillation crossing a radical-fraction threshold, consistent with the supercritical bifurcation, not unbounded growth.
Tags
Links
- Source: https://arxiv.org/abs/2605.30392v2
- Canonical: https://arxiv.org/abs/2605.30392v2
Trouble viewing inline? Open PDF directly →
Full Text
95,268 characters extracted from source content.
Expand or collapse full text
Delayed Repression and Emergent Instability in Adaptive Multi-Agent Systems Igor Itkin Independent Researcher, Tel Aviv, Israel ig.itkin@gmail.com ORCID: 0009-0004-9513-8463 (May 2026) Abstract Regulatory institutions (from content moderation platforms to financial supervisors) observe, deliberate, and intervene only after a characteristic delay. We ask whether this processing lag alone can destabilize a multi-agent system that would otherwise remain stable, without exogenous shocks, coordination among agents, or malicious actors. We study this in two stages. First, we analyze a delayed replicator equation in which autonomous agents benefit from radical behavior but face punishment based on a lagged institutional alarm signal. We derive a closed-form critical delay beyond which the unique interior equilibrium loses stability through a Hopf bifurcation, and prove via center manifold reduction that the bifurcation is supercritical (bounded oscillations, not explosive growth) for the entire sigmoid response family. Second, we embed N=240N=240 agents on a network with reinforcement learning (tabular Q-learning) and cross institutional delay with three decision architectures: fixed-policy, reactive (a memoryless threshold heuristic), and Q-learning. The hierarchy is opposite to the naive expectation that learning amplifies instability. Reactive agents are perfectly stable without delay yet collapse once delay is introduced (96% runaway by delay≥ 8\,≥\,8); fixed-policy agents are immune (0% at all delays); Q-learning agents are only partially resilient (66% at delay= 20\,=\,20). The destabilizing ingredient is reactivity to delayed signals, not learning: agents that immediately exploit low-alarm windows trigger oscillatory feedback loops, while learning buffers this through punishment memory encoded in value functions. Throughout, “runaway” denotes bounded large-amplitude oscillation crossing a radical-fraction threshold, consistent with the supercritical bifurcation, not unbounded growth. 1 Introduction Regulatory bodies observe aggregate statistics, deliberate, and intervene only after a characteristic delay. This paper asks what happens in an adaptive multi-agent system (a population of agents that modify their behavior in response to observed rewards and punishments) when institutional feedback arrives with such a delay. The answer, at least in our model, is that delay alone can destabilize an otherwise stable system: no exogenous shock is needed. The replicator equation (Taylor and Jonker, 1978; Hofbauer and Sigmund, 1998) provides the analytical framework, and network structure shapes the conditions under which autonomy persists (Nowak, 2006; Szabó and Fáth, 2007; Lieberman et al., 2005), but the core mechanism we study is temporal: a lag between what agents do and when the institution responds. (a) No delay (Δ=0 =0)Agentschoose actionsAlarmA(t)A(t)Institutionsets p(t)p(t)Punishmentξi(t) _i(t)radical fractioncurrent staterepressioncost → retreat⟶ Stable equilibrium x∗x^*(b) With delay (Δ>Δc > _c)Agentschoose actionsAlarmA(t−Δ)A(t- )Institutionsets p(t)p(t)Punishmentξi(t) _i(t)radical fractionstale staterepressioncost → retreat⟶ Oscillations / runaway Figure 1: The delay-destabilization mechanism. (a) Without delay, the institution observes the current state and adjusts repression in real time; negative feedback drives the system to a stable equilibrium x∗x^*. (b) When institutional processing introduces delay Δ , the alarm signal is stale: agents have already changed behavior by the time repression arrives. Above a critical threshold Δc _c, this stale feedback generates oscillations or runaway. The paper derives Δc _c analytically (Section 2) and tests which agent architectures are most vulnerable (Section 4). Several lines of work address time delays in replicator dynamics. Kuang (1993) established the mathematical framework for delay differential equations in population dynamics. In evolutionary game theory, Alboszta and Miekisz (2004) showed that time delay can destabilize evolutionarily stable strategies in discrete replicator dynamics, while Wesson et al. (2016) proved the existence of Hopf bifurcations in two-strategy delayed replicator equations and characterized the critical delay at which limit cycles emerge. The effect is not universal: Iijima (2012) found that in a class of discrete delayed dynamics the stability of the mixed equilibrium is preserved independent of the delay, showing that whether delay destabilizes is model-dependent. Mittal et al. (2020) analyzed the delayed replicator-mutator equation, showing how delay generates limit cycles even in cooperative regimes. The upshot is that delay can qualitatively change stable equilibria into oscillatory or unstable ones, but the outcome depends on the specific response structure. We build on this line of work because it provides the analytical foundation (the Hopf bifurcation framework for delayed replicator equations) that we extend to nonlinear institutional response functions and validate in an agent-based simulation. Despite this analytical foundation, the interaction between delayed institutional feedback and adaptive learning agents on structured networks has not been studied. Most analytical results assume well-mixed populations with fixed strategy revision protocols (Miekisz, 2008). Real multi-agent systems, however, involve agents that learn from experience, interact on heterogeneous graphs, and face punishment signals that propagate through institutional channels with variable latency. The literature on evolutionary dynamics on graphs (Ohtsuki et al., 2006; Santos and Pacheco, 2005; Perc et al., 2013, 2017; McAvoy and Allen, 2021) and on multi-agent reinforcement learning (Sutton and Barto, 2018; Zhang et al., 2021; Gronauer and Diepold, 2022) have proceeded largely separately. Recent work on reward delays in single-agent and multi-agent reinforcement learning (RL) (Bouteiller et al., 2021; Zhang et al., 2023) addresses convergence properties but not the evolutionary stability questions that arise when delayed feedback drives population-level regime transitions. We draw on both traditions because the gap lies at their intersection: delayed institutional feedback (from evolutionary game theory) has not been combined with adaptive learning agents (from multi-agent RL) on structured networks. How does the delay-instability mechanism manifest when agents learn from experience rather than following fixed revision protocols? Our central finding is counterintuitive: adaptive learning does not amplify delay-induced instability—it partially buffers it. The destabilizing ingredient is not learning but reactivity to delayed signals. Agents that react immediately to low-alarm windows exploit them and trigger oscillatory feedback loops; agents that learn from cumulative punishment history resist this trap. We reach this conclusion through a two-stage approach. First, we analyze the delayed replicator equation with a nonlinear sigmoid response function, specializing the Hopf bifurcation framework of Wesson et al. (2016) to derive a closed-form critical delay Δc _c and prove via center manifold reduction that the bifurcation is supercritical for the entire admissible sigmoid parameter class (Proposition 2). This guarantees bounded limit cycles rather than explosive instability above the critical threshold. We validate the analytical result through numerical ordinary differential equation (ODE) integration and a discrete mean-field bridge connecting the continuous theory to the agent-based simulation. Second, we construct a networked multi-agent simulation with N=240N=240 reinforcement learning agents on a modular graph and run a factorial experiment crossing delay with agent decision architecture. Three architectures are compared: non-reactive agents (fixed policy), reactive agents (threshold heuristic without memory), and Q-learning agents (tabular reinforcement learning with cumulative value estimates). The results reveal a clear hierarchy. Reactive agents are the cleanest demonstration of the central claim: with no delay they are fully stable (100% stable at Δ=0 =0), and delay alone drives them to catastrophic runaway (96% by delay≥ 8\,≥\,8). Non-reactive (fixed-policy) agents are immune (0% runaway at all tested values), confirming that the feedback loop is necessary. Q-learning agents achieve partial resilience (66% at delay= 20\,=\,20) by encoding historical punishment into their value functions, though they carry an intrinsic oscillatory tendency at Δ=0 =0 (only ≈42%≈42\% stable) that the memoryless reactive baseline lacks. The reduced ODE is a caricature, not a forecast. It identifies a local instability mechanism; the simulations then test whether that mechanism survives contact with a realistic population of adaptive agents on a heterogeneous network. The key empirical question is how different agent architectures interact with delay, and the two-factor design reveals that reactivity, not learning per se, is the destabilizing ingredient. Section 2 develops the analytical theory (delayed replicator equation, critical delay, Hopf bifurcation proof). Section 3 describes the networked simulation model, including agent learning dynamics and network topologies. Section 4 presents the experiments and results. Section 5 discusses implications and limitations. The companion paper extends the setting to noisy selective control on modular networks. Code and data: https://github.com/YehudaItkin/delayed-repression-instability. 2 Theory: Delayed Repression Dynamics 2.1 Problem formulation We consider a population of N agents that choose between conformist and autonomous (radical) behavior over discrete time steps. An institution observes the population state, but with a processing delay of Δ steps. Based on the delayed observation, the institution sets a repression probability that imposes costs on radical agents. Formally: • Input: system parameters (N,Δ,k,agent architecture)(N, ,k,agent architecture), where k is the institutional response sharpness. • State: the radical fraction x(t)∈[0,1]x(t)∈[0,1] at each time step t. • Dynamics: the delayed replicator equation x˙(t)=x(t)(1−x(t))[a−Cp(x(t−Δ))] x(t)=x(t)(1-x(t))[a-C\,p(x(t- ))], where p(⋅)p(·) is the institutional response function and a=B−Sa=B-S is the net autonomy advantage. • Question: for what values of Δ does the unique interior equilibrium x∗x^* lose stability? We formalize this question in a reduced mean-field model (this section), derive the stability boundary Δc _c in closed form, and then test whether the mechanism survives in a full agent-based simulation where agents maximize cumulative discounted reward ∑tγtri(t) _tγ^tr_i(t) via Q-learning (Section 3). 2.2 Reduced mean-field model The analytical model builds on the replicator equation framework introduced by Taylor and Jonker (1978) and developed extensively in Hofbauer and Sigmund (1998), incorporating the delay mechanisms analyzed in the general setting by Kuang (1993) and in evolutionary games specifically by Wesson et al. (2016). Let x(t)∈[0,1]x(t)∈[0,1] denote the fraction of agents adopting an autonomous (radical) strategy at time t. Non-autonomous agents receive a baseline payoff S. Autonomous agents receive benefit B, but incur punishment cost C>0C>0 with probability determined by a delayed aggregate signal. Define the net autonomy advantage as a:=B−S.a:=B-S. We study the delayed replicator equation x˙(t)=x(t)(1−x(t))[a−Cp(x(t−Δ))], x(t)=x(t)(1-x(t)) [a-Cp(x(t- )) ], (1) where Δ≥0 ≥ 0 is the repression delay and p:[0,1]→[0,1]p [0,1]→[0,1] is a smooth increasing response function representing the probability that the institution activates repression given the observed population state. The logistic growth factor x(1−x)x(1-x) ensures that the dynamics vanish at the population boundaries, as is standard in replicator equations. The key modeling assumption is that the institution observes and responds to the population state at time t−Δt- rather than time t, reflecting bureaucratic processing time, information aggregation delays, or deliberation periods. We model the response function as a sigmoid (logistic function): p(x)=σ(k(x−xc))=11+e−k(x−xc),p(x)=σ(k(x-x_c))= 11+e^-k(x-x_c), (2) where k>0k>0 controls the sharpness of the institutional response and xc∈(0,1)x_c∈(0,1) is the threshold at which the institution transitions from low to high repression probability. This choice is motivated on three grounds. First, the sigmoid is the standard Fermi update function in stochastic evolutionary game theory (Traulsen et al., 2006), where it parameterizes the sensitivity of strategy revision to payoff differences; here it plays an analogous role for institutional sensitivity to the observed radical fraction. Second, threshold-based responses are empirically characteristic of institutions that tolerate low-level deviation but activate enforcement above a critical level, a pattern common to regulatory agencies, content moderation systems, and immune responses (Scheffer, 2009). Third, given an inflection point xcx_c, the sigmoid is the unique smooth monotone solution of p′(x)=kp(x)(1−p(x))p (x)=kp(x)(1-p(x)) with p(xc)=1/2p(x_c)=1/2, making it the natural one-parameter family indexed by sharpness k for any system with logistic-type threshold activation. Larger k corresponds to a more decisive institution that switches abruptly between tolerance and punishment, while smaller k represents a gradual, proportional response. 2.3 Interior equilibrium and stability without delay Proposition 1 (Interior equilibrium). The boundary points x=0x=0 and x=1x=1 are stationary. An interior stationary point x∗∈(0,1)x^*∈(0,1) satisfies p(x∗)=aC.p(x^*)= aC. (3) If p is strictly increasing on [0,1][0,1], such a point exists and is unique if and only if p(0)<aC<p(1).p(0)< aC<p(1). (4) For a steep sigmoid with p(0)≈0p(0)≈ 0 and p(1)≈1p(1)≈ 1, this reduces approximately to 0<a<C0<a<C. Proof. At a stationary point of (1), either x=0x=0, x=1x=1, or a−Cp(x∗)=0a-Cp(x^*)=0. The latter condition is equivalent to p(x∗)=a/Cp(x^*)=a/C. Strict monotonicity of p gives existence and uniqueness precisely when a/Ca/C lies in the image of p restricted to the open interval (0,1)(0,1). ∎ At the interior equilibrium, institutional punishment exactly offsets the net benefit of autonomy. For the sigmoid response function, the equilibrium is given explicitly by x∗=xc+1klog(aC−a),x^*=x_c+ 1k ( aC-a ), (5) provided this value lies in (0,1)(0,1). Having located the interior equilibrium, we next establish that it is locally asymptotically stable whenever repression is instantaneous. Delay is therefore necessary for the instabilities studied in subsequent sections. Theorem 1 (Local stability without delay). Assume an interior equilibrium x∗∈(0,1)x^*∈(0,1) exists and p′(x∗)>0p (x^*)>0. For Δ=0 =0, x∗x^* is locally asymptotically stable. Proof. Let F(x,y)=x(1−x)[a−Cp(y)].F(x,y)=x(1-x)[a-Cp(y)]. With no delay, the system is x˙=F(x,x) x=F(x,x). Linearize around x∗x^* by writing x(t)=x∗+u(t)x(t)=x^*+u(t). The linear coefficient is A=∂xF(x∗,x∗)+∂yF(x∗,x∗).A= _xF(x^*,x^*)+ _yF(x^*,x^*). At the interior equilibrium, a−Cp(x∗)=0a-Cp(x^*)=0, so ∂xF(x∗,x∗)=(1−2x∗)[a−Cp(x∗)]=0 _xF(x^*,x^*)=(1-2x^*)[a-Cp(x^*)]=0. Hence A=−Cx∗(1−x∗)p′(x∗)<0,A=-Cx^*(1-x^*)p (x^*)<0, which implies local asymptotic stability since the linearized equation u˙=Au u=Au decays exponentially. ∎ 2.4 Delay-induced instability When Δ>0 >0, the linearization around x∗x^* takes the form of a scalar delay differential equation. Since ∂xF(x∗,x∗)=0 _xF(x^*,x^*)=0 as shown above, the linearized dynamics involve only the delayed term. u˙(t)=βu(t−Δ), u(t)=β\,u(t- ), (6) where β=−Cx∗(1−x∗)p′(x∗)<0.β=-Cx^*(1-x^*)p (x^*)<0. (7) This is a standard scalar delay equation of the form studied in Kuang (1993). Let b=−β>0b=-β>0. Then (6) becomes u˙(t)=−bu(t−Δ) u(t)=-b\,u(t- ), which is the Hayes equation whose stability boundary is classically known. Theorem 2 (Critical delay and Hopf crossing). For the scalar linearized delay equation (6), the interior equilibrium is locally asymptotically stable for 0≤bΔ<π2.0≤ b < π2. At Δc=π2Cx∗(1−x∗)p′(x∗), _c= π2Cx^*(1-x^*)p (x^*), (8) the characteristic equation has a simple conjugate pair of purely imaginary roots λ=±ibλ=± ib that cross the imaginary axis with positive speed as Δ increases through Δc _c. Thus Δc _c is the first local stability boundary and the reduced nonlinear delay differential equation (DDE) satisfies the local Hopf crossing conditions. The criticality of this bifurcation is established in Proposition 2. Proof. Seek solutions of the characteristic equation by substituting u(t)=eλtu(t)=e^λ t into (6), obtaining λ=βe−λΔ.λ=β e^-λ . At the first stability crossing, let λ=iωλ=iω with ω>0ω>0. Since β=−bβ=-b, iω=−be−iωΔ=−b(cos(ωΔ)−isin(ωΔ)).iω=-be^-iω =-b( (ω )-i (ω )). Equating real and imaginary parts gives 0=−bcos(ωΔ),ω=bsin(ωΔ).0=-b (ω ), ω=b (ω ). The first condition requires cos(ωΔ)=0 (ω )=0, so ωΔ=π/2+nπω =π/2+nπ for non-negative integer n. The first crossing corresponds to n=0n=0, giving ωΔ=π/2ω =π/2 and hence ω=bω=b from the second equation. Substituting b=Cx∗(1−x∗)p′(x∗)b=Cx^*(1-x^*)p (x^*) yields (8). It remains to verify the transversality condition and the absence of other imaginary-axis roots. Write the characteristic equation as λ+be−λΔ=0λ+be^-λ =0. Differentiating with respect to Δ : dλdΔ=bλe−λΔ1−bΔe−λΔ. dλd = bλ e^-λ 1-b e^-λ . At λ=ibλ=ib, Δ=Δc=π/(2b) = _c=π/(2b), we have e−ibΔc=e−iπ/2=−ie^-ib _c=e^-iπ/2=-i, so be−λΔ=−ibbe^-λ =-ib. Substituting: dλdΔ|Δ=Δc=b21+iπ/2, dλd |_ = _c= b^21+iπ/2, and therefore Re(dλ/dΔ)=b2/(1+π2/4)>0Re(dλ/d )=b^2/(1+π^2/4)>0. The eigenvalue crosses the imaginary axis with positive speed. To verify no other purely imaginary roots exist at Δc _c, note that if λ=iωλ=iω is a root, then taking moduli in iω=−be−iωΔciω=-be^-iω _c gives |ω|=b|ω|=b, so the only purely imaginary roots are λ=±ibλ=± ib. By the Hopf bifurcation theorem for delay differential equations (Kuang, 1993), a periodic orbit bifurcates from the equilibrium at Δ=Δc = _c. The direction and stability of this bifurcation are determined by the first Lyapunov coefficient, computed via center manifold reduction in Proposition 2. ∎ Proposition 2 (Supercritical Hopf bifurcation). For the sigmoid response p(x)=σ(k(x−xc))p(x)=σ(k(x-x_c)) and any admissible parameters such that the interior equilibrium x∗∈(0,1)x^*∈(0,1) exists (equivalently, p(0)<a/C<p(1)p(0)<a/C<p(1); for steep sigmoids this reduces to 0<a<C0<a<C), the Hopf bifurcation at Δ=Δc = _c is supercritical: the first Lyapunov coefficient satisfies Re(c1(0))<0Re(c_1(0))<0, the bifurcating periodic orbits are orbitally stable, and their amplitude grows continuously from zero as Δ increases through Δc _c. The proof proceeds by center manifold reduction following the Hassard–Kazarinoff–Wan formalism (Hassard et al., 1981) applied to the infinite-dimensional DDE phase space (Hale and Verduyn Lunel, 1993). The key steps are: (i) expansion of the nonlinearity to cubic order around x∗x^*; (i) projection onto the center eigenspace via the Hale–Verduyn Lunel bilinear form; (i) computation of the center manifold corrections W20W_20, W11W_11; and (iv) extraction of the first Lyapunov coefficient c1(0)c_1(0) from the resulting normal form. For the sigmoid response, the closed-form expression for Re(c1(0))Re(c_1(0)) factors into a positive prefactor times (ρ−1)ℬ(ρ-1)B, where ρ=p(x∗)=a/C∈(0,1)ρ=p(x^*)=a/C∈(0,1) is the equilibrium repression probability and ℬB is shown to be strictly positive for all ρ∈(0,1)ρ∈(0,1) and k>0k>0 by a positive-definiteness argument on a quadratic form. Since ρ<1ρ<1, it follows that Re(c1(0))<0Re(c_1(0))<0 universally. The full derivation, including symbolic verification and numerical validation over 54,00054,000 parameter combinations, is given in Appendix A. For the sigmoid response function, the derivative at equilibrium takes the form p′(x∗)=kaC(1−aC),p (x^*)=k\, aC (1- aC ), (9) Substituting into (8) yields the explicit critical delay Δc=π2kax∗(1−x∗)(1−aC). _c= π2k\,a\,x^*(1-x^*) (1- aC ). (10) This expression reveals that the critical delay decreases with increasing sharpness k, with increasing autonomy advantage a (provided a<Ca<C), and is minimized near mixed-population states where x∗(1−x∗)x^*(1-x^*) is large. 2.5 Summary and testable hypotheses Prior work established the existence of Hopf bifurcations in delayed replicator equations with symmetric, linear payoffs (Wesson et al., 2016; Wesson and Rand, 2016). Ben-Khalifa et al. (2018) showed that the shape of the delay distribution (Dirac, uniform, Gamma) affects the Hopf threshold, and Wettergren (2023) demonstrated that asymmetric delays on costs versus benefits produce alternating stability–instability windows absent in the single-delay case. In feedback-evolving games, Yan et al. (2021) found that delayed environmental feedback from cooperation, combined with punishment of defectors, generates oscillations through the same Hopf mechanism, a parallel finding where the delay acts on agent contributions rather than institutional observation. More recently, Mohamadichamgavi and Bodnar (2025) combined logistic population growth with strategy-dependent delays, producing richer bifurcation structure than the constant-population case. In multi-agent reinforcement learning, Chen et al. (2020) formalized delay-aware Markov games and showed that delay from one agent propagates to others’ learning, while Derman et al. (2021) proved that deterministic non-stationary Markov policies suffice under execution delay without state augmentation, which provides theoretical backing for why Q-learning handles delayed environments. The present analysis extends this body of work in two directions. First, we replace the symmetric payoff structure with an asymmetric institutional response: a nonlinear sigmoid function p(x)p(x) that models threshold-based repression (Section 2.2). This changes both the equilibrium condition (p(x∗)=a/Cp(x^*)=a/C rather than a symmetric Nash condition) and the form of the critical delay (Equation 10), which now depends on the sharpness parameter k controlling institutional decisiveness. Second, we prove that the Hopf bifurcation is supercritical universally across the sigmoid parameter class (Proposition 2), not just for isolated parameter values. Prior results left the bifurcation direction as a numerical observation; we provide a closed-form proof via center manifold reduction, validated symbolically and numerically over 54,00054,000 parameter combinations (Appendix A). Together, these results yield a complete analytical characterization: for any sigmoid response, the delayed system transitions from a stable equilibrium to bounded limit cycles at a critical delay Δc _c that decreases with sharpness k and is maximally fragile near the institutional threshold. The analytical results generate three testable hypotheses for the networked agent-based simulation (Section 3). We state them here to motivate the experimental design in Section 4. 1. Delay destabilizes. Increasing the institutional delay Δ should move the system from stable equilibrium to oscillatory dynamics and, beyond the local bifurcation, potentially to runaway-crackdown regimes. This follows directly from Theorem 2: the stability margin shrinks as Δ grows. 2. Sharpness amplifies. Increasing the response sharpness k should reduce the stability margin by lowering Δc _c (Equation 10), so that at fixed delay the system crosses from stable to unstable as k increases. 3. Mixed populations are most fragile. Fragility should be greatest where the product x∗(1−x∗)p′(x∗)x^*(1-x^*)p (x^*) is large. For a sigmoid response, this maximum occurs near the response midpoint xcx_c when the population is neither overwhelmingly conformist nor overwhelmingly autonomous. These hypotheses concern local stability. Theorem 2 proves that the equilibrium loses stability via a Hopf bifurcation, and Proposition 2 establishes that this bifurcation is supercritical: the resulting periodic orbit is orbitally stable with amplitude growing continuously from zero. The theory does not, by itself, characterize global dynamics such as large-amplitude runaway or crackdown. The transition from continuous ODE to discrete networked simulation introduces three additional complexity layers: (i) finite populations with stochastic action selection, (i) heterogeneous network structure, and (i) adaptive agents that learn from experience. The experiments in Section 4 test whether the delay-destabilization mechanism survives these extensions, and which agent architecture is most vulnerable. 3 Networked Simulation Model The mean-field ODE (Section 2) proves that delay can destabilize an equilibrium, but it assumes a well-mixed population of identical agents with a fixed strategy-revision rule. Real multi-agent systems violate all three assumptions: agents are heterogeneous, interact on structured networks, and adapt their behavior through learning. This section constructs a simulation that relaxes these assumptions. The goal is not to replicate the ODE dynamics quantitatively (the integer delay steps in the simulation do not map onto the continuous time units of the ODE) but to test whether the three hypotheses from Section 2.5 survive in a substantially richer system. 3.1 Agents and actions Each agent i∈1,…,Ni∈\1,…,N\ chooses an action at each discrete time step t: ai(t)∈L,M,R,a_i(t)∈\L,M,R\, corresponding to loyal, moderate, and radical behavior. These actions are mapped to a numerical activity level f(L)=0f(L)=0, f(M)=1f(M)=1, f(R)=2f(R)=2. Loyal agents conform fully to institutional expectations, moderate agents exercise limited autonomy, and radical agents pursue maximum autonomy at the risk of attracting punishment. Each agent is characterized by a learning rate αi _i, an exploration parameter ϵi _i, and a detectability score did_i that modulates how visible the agent’s behavior is to the institutional observer. 3.2 Network topologies Agents interact on a directed graph G=(V,E)G=(V,E) where edges represent influence relationships. The simulation supports three topology classes. Erdős–Rényi random graphs provide a homogeneous-mixing baseline with approximately Poisson degree distribution (Szabó and Fáth, 2007). Scale-free networks generated via preferential attachment produce heavy-tailed degree distributions in which hub nodes exert disproportionate influence (Santos and Pacheco, 2005). Stochastic block model graphs create modular community structure with dense intra-community connectivity and sparse inter-community links (Holland et al., 1983; Girvan and Newman, 2002). In the modular topology, bridge nodes are identified using betweenness centrality (Freeman, 1977), which measures the fraction of shortest paths between all node pairs that pass through a given node. Nodes with high betweenness centrality serve as conduits for information and influence between communities. This definition ensures that bridge status reflects actual inter-community connectivity rather than merely high intra-community degree. 3.3 Aggregate alarm and delayed repression The institutional observer computes an aggregate alarm signal by averaging weighted activity levels across the population: A(t)=1N∑i=1Ndif(ai(t)).A(t)= 1N _i=1^Nd_i\,f(a_i(t)). (11) The detectability weights did_i allow heterogeneous visibility: agents in exposed positions or with high-profile behavior contribute more to the institutional alarm than those operating in obscurity. The repression probability for agent i at time t depends on the delayed alarm: pi(t)=utσ(k(A(t−Δ)−Ac))di,p_i(t)=u_t\,σ (k(A(t- )-A_c) )\,d_i, (12) where σ(⋅)σ(·) is the logistic sigmoid from Section 2, Δ is the institutional delay measured in discrete time steps, k is the response sharpness, AcA_c is the alarm threshold, and ut∈[0,2.5]u_t∈[0,2.5] is the regulator force (the institutional control intensity). The regulator force utu_t may be fixed at a constant value, adjusted by a heuristic rule, or selected by a learning regulator agent. The product σ(⋅)diσ(·)\,d_i ensures that more detectable agents face higher punishment probability when the system is under repression. 3.4 Local influence and rewards Each agent observes its local neighborhood through an influence signal: Ii(t)=∑j∈−(i)wjif(aj(t)),I_i(t)= _j ^-(i)w_ji\,f(a_j(t)), (13) where −(i)N^-(i) denotes the set of agents with directed edges into i and wjiw_ji are influence weights. The agent’s immediate reward combines a benefit from its action, a social influence component, and a punishment cost: ri(t)=B(ai(t),χi)+λIi(t)f(ai(t))−ξi(t)C(ai(t),κi),r_i(t)=B(a_i(t), _i)+λ\,I_i(t)\,f(a_i(t))- _i(t)\,C(a_i(t), _i), (14) where ξi(t)∼Bernoulli(pi(t)) _i(t) (p_i(t)) is a stochastic punishment indicator (the agent either receives the full cost or nothing at each step), χi _i is a charisma parameter scaling the benefit of radical behavior, λ is the influence coupling strength, and κi _i scales the punishment cost. The benefit function is B(a,χ)=χ⋅f(a)B(a,χ)=χ· f(a) (linear in activity level, so radical actions yield the highest benefit), and the cost function is C(a,κ)=κ⋅f(a)C(a,κ)=κ· f(a) (punishment is proportional to activity level). This incentive gradient toward radicalism is counterbalanced by the repression mechanism. 3.5 Learning dynamics In learning variants, each agent employs tabular Q-learning (Sutton and Barto, 2018), a standard reinforcement learning algorithm in which the agent maintains a table of estimated long-run values Q(s,a)Q(s,a) for each state–action pair and updates them toward observed rewards. The agent uses a compact state representation. The observation state for agent i at time t is a tuple comprising a discretized local influence bucket, a discretized delayed alarm bucket, a binary recent-punishment indicator, and a binary bridge-membership flag. This yields a state space of manageable size (3×3×2×2=363× 3× 2× 2=36 states) that agents can explore within the simulation horizon. The Q-update rule is standard: Qi(s,a)←Qi(s,a)+αi[ri(t)+γmaxa′Qi(s′,a′)−Qi(s,a)],Q_i(s,a)← Q_i(s,a)+ _i [r_i(t)+γ _a Q_i(s ,a )-Q_i(s,a) ], where s is the current observation state, s′s is the next observation state (computed from the next time step’s delayed alarm and local influence), and γ is the discount factor. Action selection follows an ϵε-greedy policy with exploration rate ϵi=0.08 _i=0.08 (a standard choice in tabular RL ensuring sufficient exploration without excessive noise; Sutton and Barto, 2018, Chapter 2). The learning rate is αi=0.10 _i=0.10 and the discount factor γ=0.95γ=0.95, both standard defaults for tabular Q-learning in environments with moderate horizon length (Sutton and Barto, 2018). We verified robustness with a 3×33× 3 sweep over α∈0.05,0.1,0.2α∈\0.05,0.1,0.2\ and γ∈0.9,0.95,0.99γ∈\0.9,0.95,0.99\: across all nine combinations the Q-learning runaway rate is ≈15%≈15\% at Δ=0 =0, 35%35\% at Δ=8 =8, and 6565–75%75\% at Δ=20 =20, so both the monotonic delay–runaway relationship and the architecture hierarchy are preserved (results under experiments/results/paper1_v2/alpha_gamma_robust). The learning model is deliberately minimal: it provides agents with the capacity to detect and exploit temporal patterns in repression without pretending to model human cognition. In fixed-policy variants, agents do not learn. Instead, they sample actions from a stationary probability distribution over L,M,R\L,M,R\ that does not depend on the environment state. This provides a baseline against which to measure the destabilizing potential of adaptive behavior under delayed feedback. 3.6 Simulation loop Algorithm 1 summarizes the complete simulation loop, showing how the components defined above connect into a training procedure. Each agent’s objective is to maximize cumulative discounted reward ∑t=0T−1γtri(t) _t=0^T-1γ^tr_i(t); there is no explicit loss function to minimize, because the Q-learning update rule (above) approximates the optimal policy through temporal-difference bootstrapping rather than gradient descent. 1:Input: graph G, delay Δ , sharpness k, horizon T, seeds 2:Initialize Q-tables Qi(s,a)←0Q_i(s,a)← 0 for all agents i, states s, actions a 3:Initialize alarm history buffer H←[0,…,0]H←[0,…,0] of length Δ+1 +1 4:for t=0,1,…,T−1t=0,1,…,T-1 do 5: Compute delayed alarm: Adelayed←H[t−Δ]A_delayed← H[t- ] (stale observation) 6: for each agent i do 7: Observe state si←(influence bucket,alarm bucket,punished?,bridge?)s_i←(influence bucket,alarm bucket,punished?,bridge?) 8: Select action ai←ϵ-greedy(Qi,si,ϵi)a_i←ε-greedy(Q_i,s_i, _i) 9: end for 10: Compute current alarm: A(t)←1N∑idi⋅f(ai)A(t)← 1N _id_i· f(a_i); append to H 11: Compute repression: pi(t)←ut⋅σ(k(Adelayed−Ac))⋅dip_i(t)← u_t·σ(k(A_delayed-A_c))· d_i 12: Sample punishment: ξi(t)∼Bernoulli(pi(t)) _i(t) (p_i(t)) 13: Compute rewards: ri(t)←B(ai,χi)+λIi(t)⋅f(ai)−ξi(t)⋅C(ai,κi)r_i(t)← B(a_i, _i)+λ\,I_i(t)· f(a_i)- _i(t)· C(a_i, _i) 14: for each agent i (learning variants only) do 15: Observe next state si′s _i from step t+1t+1 observations 16: Update: Qi(si,ai)←Qi(si,ai)+αi[ri(t)+γmaxa′Qi(si′,a′)−Qi(si,ai)]Q_i(s_i,a_i)← Q_i(s_i,a_i)+ _i[r_i(t)+γ _a Q_i(s _i,a )-Q_i(s_i,a_i)] 17: end for 18:end for 19:Output: time-series history A(t),xR(t),regime label\A(t),x_R(t),regime label\ Algorithm 1 Simulation loop for one run. The key feature is that repression at step t is computed from the alarm at step t−Δt- (line 5), creating the delayed feedback loop that drives instability. 4 Experiments Stage 1:Delayed Replicator ODEStage 2:Networked SimulationRegimeClassificationH1: Delay destabilizesH2: Sharpness amplifiesH3: Mixed pop. fragileExp. 1: ODE valid.Exp. 2: Discrete bridgeExp. 3–4: SweepsExp. 5: Crossed designExp. 6: RL regulatorQ1: Survives? ✓Q2: Sharpness? ✓Q3: Learning buffersQ4: RL regulator?hypotheses50 seeds/cond.Section 2Sections 3–4Section 4 Figure 2: Solution pipeline. Stage 1 derives analytical predictions from the delayed replicator ODE. Stage 2 tests these predictions in a networked agent-based simulation with three decision architectures. Regime classification assigns each run to stable, oscillatory, or runaway. The mapping from hypotheses (H1–H3) through experiments (1–6) to research questions (Q1–Q4) is detailed in Table 1. The analytical theory (Section 2) establishes that delay can destabilize a well-mixed population with a fixed strategy-revision rule. The simulation model (Section 3) introduces three layers of realism: finite heterogeneous populations, network structure, and adaptive learning via reinforcement learning. The experiments address the following research questions: 1. Does the delay-destabilization mechanism survive in the networked simulation? The ODE predicts instability above Δc _c; the simulation adds noise, discretization, and network effects that could either amplify or suppress this mechanism. 2. How does sigmoid sharpness k interact with delay? The theory predicts that higher k lowers Δc _c (Equation 10). We test whether the simulation reproduces this monotonic relationship. 3. Does adaptive learning amplify or buffer delay-induced instability? The naive expectation is that learning agents exploit delayed low-alarm windows more effectively, amplifying instability. The alternative is that cumulative learning provides implicit memory that dampens oscillations. 4. Can an adaptive regulator compensate for the destabilizing effect of adaptive agents? If agents learn to exploit delay, can an RL-trained regulator learn to counteract them? Table 1 maps each experiment to the research question it addresses and the key controlled variables. Experiment Research question Varied factor Controlled factors Experiment 1: ODE validation Baseline Δ/Δc / _c Continuous ODE Experiment 2: Discrete mean-field Q1 (discretization) η, Δ Mean-field, no network Experiment 3: Delay sweep Q1 Δ k=20k=20, Q-learning Experiment 4: Sharpness sweep Q2 k Δ=15 =15, Q-learning Experiment 5: Crossed design Q3 Δ× × architecture k=10k=10, 3 architectures Experiment 6: RL regulator Q4 Regulator type Δ=6 =6, Q-learning Table 1: Mapping of experiments to research questions. Each experiment varies one factor while controlling others to isolate the effect. Experiment 1 and Experiment 2 validate the analytical theory; Experiment 3 through Experiment 6 test the hypotheses in the full simulation. The progression is deliberate: Experiment 1 confirms the ODE theory is mathematically correct; Experiment 2 bridges continuous and discrete dynamics; Experiment 3 and Experiment 4 establish dose-response curves for delay and sharpness; Experiment 5 is the central experiment testing the interaction between delay and agent architecture; Experiment 6 is exploratory. All results reported in this section use 50 independent seeds per condition. Raw time-series histories, summary statistics, and configuration manifests are stored under experiments/results/paper1_v2/ with per-experiment subdirectories. The progression from ODE validation through discrete mean-field to full network simulation reveals both the robustness of the core delay-instability mechanism and the quantitative gaps introduced by agent heterogeneity and adaptive learning. 4.1 Experimental setup Shared settings. All networked experiments use N=240N=240 agents on a modular stochastic block model graph (6 communities of 40 nodes, intra-community edge probability pin=0.08p_in=0.08, inter-community pout=0.004p_out=0.004). The population size N=240N=240 is large enough for stable mean-field-like alarm statistics but small enough for complete Q-table exploration within 500 steps. The 6-community modular structure follows standard stochastic block model designs (Holland et al., 1983) and produces well-defined bridge nodes for the companion paper’s analysis; the intra/inter edge probabilities (0.08/0.0040.08/0.004) yield a modularity ratio of pin/pout=20p_in/p_out=20, ensuring clear community separation. The simulation horizon is 500 time steps, sufficient for Q-learning agents to reach stable policies (convergence diagnostic below confirms Q-values equilibrate by step 100–200). We use 50 independent seeds per condition throughout, with results frozen after completion to ensure reproducibility. The sharpness parameter k is set high (k=10k=10 or k=20k=20) to ensure the system operates above the critical threshold Δc _c at the tested delay values; lower k would require longer delays to produce measurable effects. The alarm threshold AcA_c is set to the theoretical equilibrium level. These parameter choices follow from the analytical predictions: we choose settings where the theory predicts instability and test whether the simulation confirms it. Baselines and comparison strategy. This paper studies a mechanism (delay-induced instability), not a method that competes against alternatives on a benchmark. The natural baselines are therefore ablation baselines: conditions that remove specific components of the mechanism to isolate their contribution. Fixed-policy agents serve as the null model (no feedback loop, so delay cannot destabilize). Reactive agents isolate the effect of memoryless reactivity. Q-learning agents add adaptive memory. The comparison across these three architectures answers Q3. No prior work has studied the specific interaction between institutional delay, agent learning, and network structure that we examine; accordingly, we compare agent architectures rather than competing models from the literature. The ODE validation (Experiment 1) and discrete mean-field (Experiment 2) serve as additional baselines connecting the simulation to the analytical theory. Regime classification. Each simulation run is classified into one of three regimes based on tail-half statistics (the final 50% of the time series, after transient dynamics decay). Stable: the radical fraction remains below the runaway threshold Rmax=0.40R_ =0.40 and oscillation amplitude is small. Oscillatory: substantial periodic variation in radical fraction (detected by spectral analysis: peak-to-total power ratio >0.25>0.25 with amplitude >0.08>0.08, or a fallback test of tail amplitude >0.12>0.12 or tail standard deviation >0.04>0.04) without exceeding RmaxR_ . Runaway: the maximum radical fraction exceeds Rmax=0.40R_ =0.40. We emphasize that “runaway” denotes crossing this threshold, not unbounded growth: across all runs so labeled, the tail-period mean radical fraction never exceeds 0.300.30, consistent with the supercritical-Hopf prediction of bounded limit cycles rather than literal radical domination. The threshold Rmax=0.40R_ =0.40 is chosen as a substantively meaningful level at which radical agents dominate; we verify that the central ordering (fixed ≤ Q-learning ≤ reactive in delay sensitivity) is preserved across thresholds from 0.30 to 0.50 (Section 5). 4.2 Theory validation (Experiment 1, Experiment 2) Experiment 1: ODE validation. This experiment validates the analytical theory by numerically integrating the delayed replicator equation (1) directly, sweeping Δ from 0 to 4Δc4 _c. It is the ground-truth baseline: if the ODE does not bifurcate at Δc _c, the theory is wrong and the simulation experiments are moot. The integration uses parameter values a=2.0a=2.0, C=7.0C=7.0, k=10k=10, xc=0.72x_c=0.72, representative magnitudes chosen so that the interior equilibrium x∗≈0.63x^*≈0.63 sits at a non-degenerate point away from the boundaries, with critical delay Δc≈0.47 _c≈0.47. The supercriticality proof (Appendix A) shows the qualitative bifurcation is insensitive to these specific values across all admissible parameters. Direct numerical integration of Equation (1) confirms Theorem 2. Figure 3 shows the bifurcation diagram: below Δc _c, trajectories converge to x∗x^*; at Δ/Δc=1 / _c=1, stable limit cycles emerge with amplitude growing continuously from zero, consistent with the supercritical Hopf bifurcation. Figure 4 shows representative trajectories, and Figure 5 shows that increasing sharpness k at fixed Δ/Δc=1.5 / _c=1.5 sharpens the nonlinear response and increases limit-cycle amplitude. Figure 3: Bifurcation diagram for the delayed replicator ODE (Experiment 1). Horizontal axis: normalized delay Δ/Δc / _c. Below Δc _c, the equilibrium is stable. At Δ/Δc=1 / _c=1, a Hopf bifurcation produces stable limit cycles with continuously growing amplitude (solid: numerical envelope; dashed: theoretical threshold). Conclusion: the analytical theory (Theorem 2) is confirmed exactly—the instability occurs precisely at the predicted threshold. Figure 4: Representative ODE trajectories at six delay levels (Experiment 1). Below Δc _c: damped convergence to equilibrium. Above: sustained limit cycles with amplitude growing with Δ . The transition is sharp, consistent with the supercritical Hopf bifurcation (Proposition 2). Figure 5: Effect of sharpness k on ODE dynamics at fixed Δ/Δc=1.5 / _c=1.5 (Experiment 1). All panels are at the same distance beyond the stability boundary, yet higher k produces sharper, larger-amplitude limit cycles saturating near the population boundaries. Conclusion: sharpness amplifies the nonlinear consequences of delay-induced instability, confirming Hypothesis 2. Experiment 2: Discrete mean-field. This experiment bridges continuous ODE dynamics and discrete multi-agent simulation. Discretization itself can introduce instabilities absent in the continuous limit (Alboszta and Miekisz, 2004). We study a discrete-time mean-field system with step size η and verify that the continuous bifurcation structure is recovered as η→0η→ 0. This isolates discretization effects from network structure, agent heterogeneity, and learning. Figure 6 presents the phase diagram in (Δ,η)( ,η) space. For small η, the discrete system recovers the ODE bifurcation structure. As η increases, the stability margin narrows and the dynamics overshoot the ODE limit-cycle amplitude (Figure 7). Figure 8 quantifies this: the stability margin decreases monotonically with η. This establishes that discretization introduces additional instability beyond delay. The result provides a baseline for interpreting the full simulation, where discrete time steps are inherent. Figure 6: Phase diagram of the discrete mean-field system in (Δ,η)( ,η) space (Experiment 2). Color: regime classification. Conclusion: the stability region contracts with increasing η, confirming that discretization itself is an additional instability source—the full simulation’s discrete time steps make the system more fragile than the ODE predicts. Figure 7: Discrete mean-field trajectories at Δ=1.2Δc =1.2 _c for several η values (Experiment 2). Small η: bounded oscillations matching the ODE limit cycle. Large η: overshooting dynamics exceeding the ODE envelope. Conclusion: the full simulation’s discrete time steps (η≈1η≈ 1) introduce additional instability beyond the continuous theory. Figure 8: Stability margin vs. discrete step size η (Experiment 2). As η→0η→ 0, the margin recovers Δc _c from the continuous theory. Conclusion: discretization reduces the stability margin monotonically. This provides a quantitative bridge between the ODE prediction and the simulation’s inherent discrete dynamics. 4.3 Dose-response (Experiment 3, Experiment 4) Experiment 3: Delay sweep. This experiment tests Hypothesis 1 (delay destabilizes) in the full networked simulation with Q-learning agents. We sweep repression delay from 0 to 30 steps using k=20k=20 (placing the system well above Δc _c) and 500 time steps (sufficient for Q-value convergence). If the ODE mechanism survives, runaway frequency should increase monotonically with delay. Table 2 reports runaway rate vs. delay (k=20k=20, Q-learning agents). At delay=0=0 the system is already only 46% stable (44% oscillatory, 10% runaway): the non-delayed dynamics sit near the oscillatory boundary, and the 10% runaway reflects stochastic noise rather than delay (Theorem 1). Increasing delay shifts this distribution decisively into runaway, which rises to 72% at delay=30=30 (95% confidence interval (CI) [58%,83%][58\%,83\%]), an excess of +62+62 percentage points over the delay-zero runaway floor. The monotonic trend confirms the qualitative prediction: delay converts the marginally-stable baseline into large-amplitude runaway. Sixty-two percentage points is not a subtle effect. This sweep uses Q-learning agents, whose adaptive exploration leaves them near the oscillatory boundary even at Δ=0 =0; the cleanest demonstration of destabilization from a fully stable baseline is the reactive architecture (Experiment 5 and Figure 12), which is 100% stable at Δ=0 =0 and collapses entirely once delay is introduced. Delay (steps) Runaway rate 95% CI Excess over Δ=0 =0 (p) 0 0.10 [0.04,0.21][0.04,0.21] — 2 0.12 [0.06,0.24][0.06,0.24] +2+2 p 4 0.20 [0.11,0.33][0.11,0.33] +10+10 p 6 0.30 [0.19,0.44][0.19,0.44] +20+20 p 8 0.36 [0.24,0.50][0.24,0.50] +26+26 p 10 0.44 [0.31,0.58][0.31,0.58] +34+34 p 14 0.56 [0.42,0.69][0.42,0.69] +46+46 p 20 0.62 [0.48,0.74][0.48,0.74] +52+52 p 30 0.72 [0.58,0.83][0.58,0.83] +62+62 p Table 2: Runaway rate vs. repression delay (Experiment 3; k=20k=20, N=240N=240, 500 steps, Q-learning agents, 50 seeds). The 10% baseline at Δ=0 =0 is stochastic noise; the “Excess” column isolates the delay contribution. Conclusion: delay increases runaway by 62 percentage points (from 10% to 72%), confirming Hypothesis 1—delay alone destabilizes the networked system (question 1). p = percentage points. Wilson score 95% CIs reported. Figure 9: Delay sweep (Experiment 3; k=20k=20, Q-learning, 50 seeds). Left: runaway rate with 95% Wilson band. Right: excess runaway over the Δ=0 =0 baseline. Conclusion: monotonic increase confirms Hypothesis 1: delay alone produces a large, dose-dependent destabilization in the networked simulation. Experiment 4: Sharpness sweep. This experiment tests Hypothesis 2 by holding delay fixed at 15 steps and sweeping k∈3,5,7,10,15,20,30,40k∈\3,5,7,10,15,20,30,40\. Higher k should reduce Δc _c (Equation 10), so the system should cross from stable to unstable as k increases. Delay of 15 is chosen so that low-k conditions are below threshold while high-k conditions are above it. At fixed delay=15=15 (Table 3), runaway rate increases from 16% (k=3k=3) to 64% (k=40k=40). The transition is sharpest between k=7k=7 (28%) and k=10k=10 (52%), after which the system saturates: once operating far beyond Δc _c, additional sharpening produces only marginal instability gains. k Stable Oscillatory Runaway 3 70% 14% 16% 5 56% 24% 20% 7 40% 32% 28% 10 26% 22% 52% 15 26% 14% 60% 20 24% 16% 60% 30 20% 18% 62% 40 24% 12% 64% Table 3: Regime distribution by sharpness k at fixed delay=15=15 (Experiment 4; N=240N=240, 500 steps, 50 seeds). Conclusion: runaway increases from 16% (k=3k=3) to 64% (k=40k=40), with a sharp transition at k=7k=7–1010 corresponding to the system crossing Δc _c. This confirms Hypothesis 2 (question 2): sharpness amplifies delay vulnerability. Figure 10: Runaway rate vs. sharpness k at fixed delay=15=15 (Experiment 4; 50 seeds). Vertical lines: k=10k=10 (used in Experiment 5) and k=20k=20 (Experiment 3). Conclusion: the sharp transition at k=7k=7–1010 marks the system crossing Δc _c (Equation 10). Above this threshold, further sharpening produces diminishing returns, a ceiling effect. 4.4 Central experiment: crossed delay × architecture (Experiment 5) This is the central experiment. A two-factor design crosses delay (Δ∈0,4,8,14,20 ∈\0,4,8,14,20\) with agent architecture: fixed-policy (no environmental response, the null model), reactive (threshold heuristic without memory), and Q-learning (tabular RL with cumulative value estimates). We use k=10k=10 to ensure fixed-policy agents remain stable as a clean baseline. The naive expectation is that learning amplifies instability; the alternative is that cumulative learning provides implicit memory that dampens oscillations. The two-factor crossing reveals a clear hierarchy (Figure 11 and Table 4). Figure 11: Runaway rate vs. delay for three agent architectures (Experiment 5, the central experiment; k=10k=10, N=240N=240, 500 steps, 50 seeds per cell). Shaded: 95% Wilson CIs. Conclusion: the hierarchy is reactive (96%) >> Q-learning (66%) >> fixed (0%). Learning buffers rather than amplifies instability (question 3), contradicting the naive expectation. The destabilizing ingredient is memoryless reactivity to delayed signals. This experiment uses k=10k=10 (lower than the aggressive k=20k=20 of Experiment 3) and 500 time steps with N=240N=240 agents. The results reveal three distinct regimes: 1. Non-reactive (fixed): 0% runaway at all delays (92% stable, 8% oscillatory from finite-sample binomial noise in the action draws, occasionally triggering the spectral classifier). These agents cannot exploit temporal structure in the repression signal. 2. Reactive (threshold heuristic): fully stable without delay (100% stable at delay=0=0), then a sharp transition to large-amplitude runaway as delay grows: 80% runaway at delay=4=4, rising to 96% at delay=8=8 (95% CI [87%,99%][87\%,99\%]). A fine-resolution sweep (Figure 12) locates the threshold sharply between delay=2=2 (90% stable) and delay=3=3 (62% runaway), the discrete-time analog of the critical delay Δc _c. This is the cleanest instance of the central claim: an otherwise fully stable population destabilized by delay alone. The reactive policy immediately exploits low-alarm windows but has no memory of past punishment to dampen the resulting oscillations. 3. Q-learning: 10% runaway at delay=0=0, rising to 66% at delay=20=20 (95% CI [52%,78%][52\%,78\%]). Q-values encode cumulative punishment experience, which creates inertia against full exploitation of every low-alarm window. The key finding is the ordering: reactive >> Q-learning >> fixed in delay sensitivity. The immunity of fixed-policy agents is structural, not contingent on k: because their action distribution does not depend on the alarm signal, the feedback loop required for delay-induced instability is broken regardless of parameter values. Reactivity to delayed signals is the destabilizing mechanism; adaptive learning partially buffers this through implicit memory rather than amplifying it. Delay (steps) Fixed Reactive Q-learning 95% CI (Q-learning) 0 0% 0% 10% [4%,21%][4\%,21\%] 4 0% 80% 18% [10%,31%][10\%,31\%] 8 0% 96% 32% [21%,46%][21\%,46\%] 14 0% 96% 50% [37%,63%][37\%,63\%] 20 0% 96% 66% [52%,78%][52\%,78\%] Table 4: Runaway rate by delay and agent architecture (Experiment 5; k=10k=10, N=240N=240, 500 steps, 50 seeds per cell). Conclusion: the ordering fixed (0%) << Q-learning (up to 66%) << reactive (up to 96%) holds at every delay Δ≥4 ≥ 4; at Δ=0 =0 all three are at the floor (fixed 0%, reactive 0%, Q-learning 10% from exploration noise). Q-learning’s implicit punishment memory provides partial protection; memoryless reactivity is catastrophically fragile. This answers question 3: learning buffers, not amplifies, delay-induced instability. Figure 12: Fine-resolution delay sweep for reactive agents (k=10k=10, N=240N=240, 50 seeds per delay). The population is fully stable at Δ=0 =0 and through Δ=2 =2 (90% stable), then transitions sharply to runaway between Δ=2 =2 and Δ=3 =3—the discrete-time analog of the critical delay Δc _c. Conclusion: delay alone, applied to an otherwise stable population, drives the destabilization (question 1). 4.5 Exploratory: RL regulator (Experiment 6) This experiment compares three governance configurations: fixed-policy agents with a static regulator, Q-learning agents with a static regulator, and Q-learning agents with an RL-trained regulator. The RL regulator selects force ut∈0,0.5,1.0,1.5,2.0,2.5u_t∈\0,0.5,1.0,1.5,2.0,2.5\ to minimize rreg=−xR−0.1utr_reg=-x_R-0.1\,u_t, where xRx_R is the radical fraction and utu_t is the control intensity. The regulator shares the institution’s information delay, so it cannot circumvent the destabilizing lag. This experiment is exploratory: 50 seeds provide limited statistical power, and we report results as suggestive. Figure 13: Regime distribution for three governance configurations (Experiment 6; 500 steps, 50 seeds). Conclusion: the RL regulator reduces runaway from 32% to 20% but increases oscillations from 32% to 40%, suggesting mode conversion (disaster into bounded oscillation) rather than full stabilization. The effect is suggestive but not statistically conclusive at 50 seeds (question 4). The learning ablation compares three governance configurations (Table 5). Fixed-policy agents achieve 92% stability, establishing the baseline. Q-learning agents with a static regulator produce 36% stable, 32% oscillatory, and 32% runaway. Introducing an RL-trained regulator (a Q-learning agent that observes the delayed alarm and radical fraction, selects force ut∈0,0.5,1.0,1.5,2.0,2.5u_t∈\0,0.5,1.0,1.5,2.0,2.5\, and receives reward rreg=−xR−0.1utr_reg=-x_R-0.1\,u_t, where xRx_R is the radical fraction) shifts the distribution to 40% stable, 40% oscillatory, and 20% runaway. This result is suggestive rather than conclusive: the 12 percentage-point reduction in runaway (32% to 20%) on 50 seeds yields a confidence interval that includes zero effect. The simultaneous increase in oscillatory outcomes from 32% to 40% suggests that if the regulator does help, it converts catastrophic runaway into bounded oscillations rather than restoring stability. We treat this as exploratory evidence that adaptive governance can transform instability modes, not as a confirmed finding. Condition Stable Oscillatory Runaway Fixed agents (static regulator) 92% 8% 0% Q-learning (static regulator) 36% 32% 32% Q-learning (RL regulator) 40% 40% 20% Table 5: Regime distribution by governance configuration (Experiment 6; 50 seeds each). Conclusion: adaptive governance converts runaway into oscillation (32%→ 20% runaway, 32%→ 40% oscillatory) rather than restoring stability. This mode conversion is the realistic ceiling for a controller sharing the institution’s delay (question 4, exploratory). 4.6 Summary of findings The central result is the two-factor crossing (Experiment 5): reactive agents collapse catastrophically under delay (96%), Q-learning agents achieve partial resilience (66%), and non-reactive agents are immune (0%). Reactivity is the destabilizing mechanism; learning buffers it. The story is simpler than one might expect. The ODE validation (Experiment 1) and discrete mean-field (Experiment 2) confirm that the analytical mechanism is correct and survives discretization. The delay sweep (Experiment 3) and sharpness sweep (Experiment 4) map the dose-response and activation boundary. The RL regulator (Experiment 6) is exploratory: suggestive of mode conversion but not statistically conclusive. All regime classifications use the adaptive classifier with runaway threshold Rmax>0.40R_ >0.40. The quantitative gap between mean-field thresholds and simulation thresholds reflects the expected difference between an analytically tractable mechanism and a system with agent heterogeneity, stochastic exploration, and network structure. 5 Discussion The results confirm the core analytical prediction (delay destabilizes) and reveal a surprising empirical finding (learning buffers rather than amplifies). This section discusses three qualitative insights that go beyond the numerical results: why the ODE-to-simulation gap exists and what it teaches us about the mechanism, what the architecture hierarchy reveals about the nature of delay vulnerability, and what practical implications follow for institutional design. 5.1 Why the quantitative gap between theory and simulation matters (Q1) The ODE predicts a sharp bifurcation at Δc _c; the simulation shows a gradual dose-response curve. This gap is not a failure of the theory—it is informative about the mechanism. Three stabilizing buffers are absent from the mean-field model and present in the simulation: stochastic exploration (ϵε-greedy noise damps incipient oscillations), the three-action space (agents can retreat to moderate behavior rather than oscillating between extremes), and spatial heterogeneity (network structure disperses the synchronized oscillations required for large-amplitude instability). These are not defects of the simulation; they are what the mean-field theory trades away for a closed-form answer. The qualitative lesson is that delay-induced instability is robust to realistic complications, but the sharp threshold of the ODE becomes a gradual erosion of stability in a heterogeneous population. Institutions operating near the theoretical boundary should not expect a clean phase transition; they should expect slowly worsening oscillatory behavior that may be difficult to distinguish from normal fluctuations until runaway is already underway. 5.2 What the architecture hierarchy reveals about the mechanism (Q3) The ordering (reactive agents most fragile, Q-learning agents partially resilient, fixed-policy agents immune) points to a specific causal mechanism rather than a generic correlation between complexity and instability. The mechanism is a feedback loop: when the alarm drops (due to institutional processing lag), reactive agents immediately escalate. Punishment arrives too late; by then the alarm spike triggers the next oscillation. The critical feature is not that reactive agents are “simple” but that they are memoryless: each decision is based entirely on the current (stale) signal, with no integration over past experience. Q-learning agents break this loop because their Q-values carry forward the memory of past punishment. The “soft brake” is not a design feature—it is an emergent consequence of temporal-difference learning applied to a delayed-feedback environment. This distinction has a practical implication: the risk factor for delay vulnerability is not whether agents are adaptive, but whether they integrate information over time. Any decision process that reacts only to current signals (whether a simple threshold, a rule-based policy, or a sophisticated model with no memory state) will be vulnerable. Institutions seeking stability should therefore prioritize reducing processing delay over restricting agent adaptiveness. The problem is the lag, not the learning. 5.3 Policy levers: response sharpness and adaptive governance (Q2, Q4) The theory predicts that gradual institutional responses (low k) expand the stability margin. The simulation confirms this, but with an important qualification: the protective effect of gradual response is conditional on agent architecture. Fixed-policy agents are stable regardless of sharpness, because they do not close the feedback loop. For adaptive agents, reducing sharpness helps but cannot eliminate delay vulnerability—it only raises the threshold. The policy implication is that institutions face a genuine dilemma: sharp responses are decisive but fragile; gradual responses are resilient but slow to act. Every parameter trades one vulnerability for another. The second lever is the regulator itself. The RL regulator experiment is exploratory, and the specific numbers should be interpreted cautiously. The qualitative pattern, however, is suggestive: the regulator appears to convert catastrophic runaway into bounded oscillations rather than restoring full stability. This mode conversion—disaster into nuisance—may be the realistic ceiling for any controller that shares the institution’s information delay. A regulator that observes the same stale signal as the repression mechanism cannot anticipate the population’s response; it can only react to the consequences of its own delayed actions, introducing a secondary feedback loop that sustains limit cycles while preventing unbounded growth. We report this as a pattern worth investigating, not a confirmed finding: 50 seeds is enough to see a pattern but not enough to bet on it. 5.4 Limitations This model is deliberately simple, and the simplifications cost something. The Q-learning model is deliberately minimal and does not capture sophisticated strategic reasoning, communication, or coordination among agents. Real adaptive agents in social or technical systems may exhibit richer behavioral repertoires. The three-action space is a simplification; continuous action spaces might produce different stability characteristics. The regime classifier uses fixed thresholds, and the precise quantitative rates depend on these choices. We verified that the central ordering (fixed ≤ Q-learning ≤ reactive in delay sensitivity) is preserved across runaway thresholds from Rmax>0.30R_ >0.30 through Rmax>0.50R_ >0.50, though the absolute rates shift substantially (e.g., reactive at delay=20=20 ranges from 100% at threshold 0.30 to 64% at threshold 0.50). A convergence diagnostic confirms that Q-learning agents reach stable behavior by step 100–200: the mean radical fraction and its standard deviation change by less than 0.001 between the last two 100-step windows at all tested delays, indicating that 500 steps is sufficient for the Q-values to equilibrate. The simulation uses finite populations (N=240N=240) on specific graph realizations; larger populations or different graph families might shift the quantitative boundaries. The architecture hierarchy itself, however, is robust to the network parameterization: re-running the crossed delay × architecture design on a denser modular graph (4 communities, pin=0.15p_in=0.15, pout=0.02p_out=0.02) reproduces the same ordering—fixed agents immune (0% at all delays), reactive agents catastrophic (100% by delay 8), Q-learning agents partially resilient (77% at delay 20)—with the denser network somewhat more unstable (reactive reaches 27% runaway even at delay 0), consistent with its stronger influence coupling. The reactive baseline uses a single threshold heuristic; other reactive architectures (imitation dynamics, best-response with noise) might show different levels of delay sensitivity, though the qualitative finding (that memoryless reactivity is more fragile than learning) should hold for any policy that lacks temporal integration. The RL regulator’s Q-update bootstraps toward a next-state computed from the current (undelayed) alarm rather than the next delayed observation, introducing a minor state-transition inconsistency; since Experiment 6 is presented as exploratory evidence, this does not affect the paper’s confirmed findings. Finally, the theory assumes a single aggregate alarm signal; systems with multiple observation channels or local feedback might exhibit different instability structures. Reality, of course, does not present its instabilities one at a time. 5.5 Relationship to prior work Our analytical results connect directly to the Hopf bifurcation analysis of Wesson et al. (2016) for two-strategy delayed replicator dynamics, extending their framework to include an asymmetric institutional response function rather than symmetric frequency-dependent payoffs. The discrete-time instability amplification we observe in Experiment 2 echoes the finding of Alboszta and Miekisz (2004) that delay can destabilize the ESS in discrete replicator dynamics; Iijima (2012) documents the complementary regime in which stability is instead preserved, underscoring that the destabilizing effect of delay is model-dependent. The network simulation builds on the evolutionary-dynamics-on-graphs tradition (Lieberman et al., 2005; Ohtsuki et al., 2006; Szabó and Fáth, 2007), adding delayed institutional feedback as a novel destabilizing mechanism distinct from the structural effects studied in that literature. The interaction between learning and delay connects to recent work on reward delays in multi-agent reinforcement learning (Zhang et al., 2023), though our focus on population-level regime transitions rather than individual convergence is novel. The companion paper extends the present mechanism to noisy selective control on modular networks, addressing the question of how imperfect observation and targeted governance alter stability conditions. That extension is necessary because real institutions rarely observe true system state directly, and the consequences of classification errors depend strongly on network position, a consideration we deliberately set aside here so that delay gets a fair hearing on its own. 6 Conclusion We studied how institutional processing delay affects the stability of a multi-agent system in which agents adapt their behavior in response to delayed punishment signals. We derived a closed-form critical delay Δc _c for the delayed replicator equation with a sigmoid response function and proved that the resulting Hopf bifurcation is supercritical for the entire admissible parameter class. We then tested three agent architectures (fixed-policy, reactive, and Q-learning) in a networked simulation with 240 agents across 50 seeds per condition. The central result answers question 3: learning does not amplify delay-induced instability—it partially buffers it. Non-reactive agents are immune to delay (0% runaway), reactive agents collapse catastrophically (96%), and Q-learning agents achieve partial resilience (66%). The delay sweep (question 1) confirms a monotonic dose-response reaching +62+62 percentage points of excess runaway, the sharpness sweep (question 2) confirms that sharper institutional responses lower the stability threshold, and the RL regulator experiment (question 4) suggests that adaptive governance converts runaway into bounded oscillations rather than restoring stability. The destabilizing ingredient, then, is memoryless reactivity to delayed signals, not learning: agents that immediately exploit low-alarm windows trigger oscillatory feedback loops, while agents with cumulative memory resist this trap. The practical implication is twofold. Institutions should prefer gradual responses over sharp thresholds, because sharpness amplifies the instability that reactivity exploits; and adaptive memory is not the disease but the closest thing to a cure this system has. Three extensions are natural: scaling to larger populations and alternative network families to test how far the architecture hierarchy holds; continuous action spaces, which may produce different stability characteristics; and combining delay with noisy observation, which the companion paper addresses for selective governance on modular networks. Acknowledgments The author thanks Ilya Makarov for valuable feedback on the manuscript. References Alboszta and Miekisz [2004] Jan Alboszta and Jacek Miekisz. Stability of evolutionarily stable strategies in discrete replicator dynamics with time delay. Journal of Theoretical Biology, 231(2):175–179, 2004. doi: 10.1016/j.jtbi.2004.06.012. Ben-Khalifa et al. [2018] Nesrine Ben-Khalifa, Rachid El-Azouzi, and Yezekael Hayel. Discrete and continuous distributed delays in replicator dynamics. Dynamic Games and Applications, 8(4):713–732, 2018. doi: 10.1007/s13235-017-0225-7. Bouteiller et al. [2021] Yann Bouteiller, Simon Ramstedt, Giovanni Beltrame, Christopher Pal, and Jonathan Binas. Reinforcement learning with random delays. In International Conference on Learning Representations (ICLR), 2021. Chen et al. [2020] Baiming Chen, Mengdi Xu, Zuxin Liu, Liang Li, and Ding Zhao. Delay-aware multi-agent reinforcement learning for cooperative and competitive environments. arXiv preprint arXiv:2005.05441, 2020. Derman et al. [2021] Esther Derman, Gal Dalal, and Shie Mannor. Acting in delayed environments with non-stationary markov policies. In International Conference on Learning Representations (ICLR), 2021. Freeman [1977] Linton C. Freeman. A set of measures of centrality based on betweenness. Sociometry, 40(1):35–41, 1977. doi: 10.2307/3033543. Girvan and Newman [2002] Michelle Girvan and Mark E. J. Newman. Community structure in social and biological networks. Proceedings of the National Academy of Sciences, 99(12):7821–7826, 2002. doi: 10.1073/pnas.122653799. Gronauer and Diepold [2022] Sven Gronauer and Klaus Diepold. Multi-agent deep reinforcement learning: A survey. Artificial Intelligence Review, 55:895–943, 2022. doi: 10.1007/s10462-021-09996-w. Hale and Verduyn Lunel [1993] Jack K Hale and Sjoerd M Verduyn Lunel. Introduction to Functional Differential Equations, volume 99 of Applied Mathematical Sciences. Springer, 1993. Hassard et al. [1981] Brian D Hassard, Nicholas D Kazarinoff, and Yieh-Hei Wan. Theory and Applications of Hopf Bifurcation, volume 41 of London Mathematical Society Lecture Note Series. Cambridge University Press, 1981. Hofbauer and Sigmund [1998] Josef Hofbauer and Karl Sigmund. Evolutionary Games and Population Dynamics. Cambridge University Press, 1998. Holland et al. [1983] Paul W. Holland, Kathryn Blackmond Laskey, and Samuel Leinhardt. Stochastic blockmodels: First steps. Social Networks, 5(2):109–137, 1983. doi: 10.1016/0378-8733(83)90021-7. Iijima [2012] Ryota Iijima. On delayed discrete evolutionary dynamics. Journal of Theoretical Biology, 300:1–6, 2012. doi: 10.1016/j.jtbi.2012.01.001. Kuang [1993] Yang Kuang. Delay Differential Equations with Applications in Population Dynamics. Academic Press, 1993. Lieberman et al. [2005] Erez Lieberman, Christoph Hauert, and Martin A. Nowak. Evolutionary dynamics on graphs. Nature, 433(7023):312–316, 2005. doi: 10.1038/nature03204. McAvoy and Allen [2021] Alex McAvoy and Benjamin Allen. Fixation probabilities in evolutionary dynamics under weak selection. Journal of Mathematical Biology, 82:14, 2021. doi: 10.1007/s00285-021-01568-4. Miekisz [2008] Jacek Miekisz. Evolutionary game theory and population dynamics. In Multiscale Problems in the Life Sciences, volume 1940 of Lecture Notes in Mathematics, pages 269–316. Springer, 2008. doi: 10.1007/978-3-540-78362-6_5. Mittal et al. [2020] Sourabh Mittal, Archan Mukhopadhyay, and Sagar Chakraborty. Evolutionary dynamics of the delayed replicator–mutator equation: Limit cycle and cooperation. Physical Review E, 101(4):042410, 2020. doi: 10.1103/PhysRevE.101.042410. Mohamadichamgavi and Bodnar [2025] Javad Mohamadichamgavi and Marek Bodnar. Bifurcation analysis of replicator dynamics with logistic growth and strategy-dependent time delays in snowdrift game. Dynamic Games and Applications, 2025. doi: 10.1007/s13235-025-00671-1. Nowak [2006] Martin A. Nowak. Five rules for the evolution of cooperation. Science, 314(5805):1560–1563, 2006. doi: 10.1126/science.1133755. Ohtsuki et al. [2006] Hisashi Ohtsuki, Christoph Hauert, Erez Lieberman, and Martin A. Nowak. A simple rule for the evolution of cooperation on graphs and social networks. Nature, 441(7092):502–505, 2006. doi: 10.1038/nature04605. Perc et al. [2013] Matjaž Perc, Jesús Gómez-Gardeñes, Attila Szolnoki, Luis M. Floría, and Yamir Moreno. Evolutionary dynamics of group interactions on structured populations: A review. Journal of the Royal Society Interface, 10(80):20120997, 2013. doi: 10.1098/rsif.2012.0997. Perc et al. [2017] Matjaž Perc, Jillian J. Jordan, David G. Rand, Zhen Wang, Stefano Boccaletti, and Attila Szolnoki. Statistical physics of human cooperation. Physics Reports, 687:1–51, 2017. doi: 10.1016/j.physrep.2017.05.004. Santos and Pacheco [2005] Francisco C. Santos and Jorge M. Pacheco. Scale-free networks provide a unifying framework for the emergence of cooperation. Physical Review Letters, 95(9):098104, 2005. doi: 10.1103/PhysRevLett.95.098104. Scheffer [2009] Marten Scheffer. Critical Transitions in Nature and Society. Princeton University Press, 2009. Sutton and Barto [2018] Richard S. Sutton and Andrew G. Barto. Reinforcement Learning: An Introduction. MIT Press, 2018. Szabó and Fáth [2007] György Szabó and Gábor Fáth. Evolutionary games on graphs. Physics Reports, 446(4–6):97–216, 2007. doi: 10.1016/j.physrep.2007.04.004. Taylor and Jonker [1978] Peter D. Taylor and Leo B. Jonker. Evolutionary stable strategies and game dynamics. Mathematical Biosciences, 40(1–2):145–156, 1978. doi: 10.1016/0025-5564(78)90077-9. Traulsen et al. [2006] Arne Traulsen, Martin A Nowak, and Jorge M Pacheco. Stochastic dynamics of invasion and fixation. Physical Review E, 74(1):011909, 2006. doi: 10.1103/PhysRevE.74.011909. Wesson and Rand [2016] Elizabeth Wesson and Richard Rand. Hopf bifurcations in delayed rock–paper–scissors replicator dynamics. Dynamic Games and Applications, 6(1):139–156, 2016. doi: 10.1007/s13235-015-0138-2. Wesson et al. [2016] Elizabeth Wesson, Richard H. Rand, and David G. Rand. Hopf bifurcations in two-strategy delayed replicator dynamics. International Journal of Bifurcation and Chaos, 26(1):1650006, 2016. doi: 10.1142/S0218127416500061. Wettergren [2023] Thomas A. Wettergren. Replicator dynamics of evolutionary games with different delays on costs and benefits. Applied Mathematics and Computation, 458:128228, 2023. doi: 10.1016/j.amc.2023.128228. Yan et al. [2021] Fang Yan, Xiaojie Chen, Zhipeng Qiu, and Attila Szolnoki. Cooperator driven oscillation in a time-delayed feedback-evolving game. New Journal of Physics, 23:053017, 2021. doi: 10.1088/1367-2630/abf205. Zhang et al. [2021] Kaiqing Zhang, Zhuoran Yang, and Tamer Başar. Multi-agent reinforcement learning: A selective overview of theories and algorithms. In Handbook of Reinforcement Learning and Control, Studies in Systems, Decision and Control, pages 321–384. Springer, 2021. Zhang et al. [2023] Yuyang Zhang, Runyu Zhang, Yuantao Gu, and Na Li. Multi-agent reinforcement learning with reward delays. In Proceedings of the 5th Annual Learning for Dynamics and Control Conference (L4DC), volume 211 of PMLR, pages 692–704, 2023. Appendix A Proof of Supercritical Hopf Bifurcation This appendix establishes Proposition 2: the Hopf bifurcation at Δ=Δc = _c is supercritical for all admissible sigmoid parameters. The proof follows the Hassard–Kazarinoff–Wan center manifold reduction [Hassard et al., 1981] applied to the DDE phase space [Hale and Verduyn Lunel, 1993]. We proceed in five steps, summarized in Table 6. Step Goal Key result Equation 1 Linearize and set up phase space Hayes equation, β=−bβ=-b (1) 2 Eigenvectors and adjoint normalization D¯=1/(1+iπ/2) D=1/(1+iπ/2) 3 Nonlinear expansion to cubic order f2f_2, f3f_3 (15) 4 Center manifold and normal form c1(0)c_1(0) (17) 5 Sign analysis of Re(c1)Re(c_1) Re(c1)<0Re(c_1)<0 universally (18) Table 6: Roadmap of the center-manifold reduction for Proposition 2. Conclusion: the five steps culminate in Re(c1)<0Re(c_1)<0 for all admissible sigmoid parameters, establishing that the Hopf bifurcation is supercritical (stable bounded limit cycles rather than explosive growth). Notation. Throughout this appendix: ρ=p(x∗)=a/C∈(0,1)ρ=p(x^*)=a/C∈(0,1) is the equilibrium repression probability, α0=x∗(1−x∗) _0=x^*(1-x^*), α1=1−2x∗ _1=1-2x^*, and b=Cα0p′(x∗)b=C _0p (x^*). All derivatives of p are evaluated at x∗x^*. Proof of Proposition 2. Step 1: Linearization and phase space. Write u(t)=x(t)−x∗u(t)=x(t)-x^* and define F(u,ud)F(u,u_d) as the full nonlinearity, where ud=u(t−Δ)u_d=u(t- ). Since a−Cp(x∗)=0a-Cp(x^*)=0 at equilibrium, the instantaneous linearization vanishes: ∂uF(0,0)=0 _uF(0,0)=0. The linearized equation is u˙(t)=βu(t−Δ),β=−b<0, u(t)=β\,u(t- ), β=-b<0, the Hayes equation, whose phase space is C=C([−Δc,0],ℝ)C=C([- _c,0],R). Step 2: Eigenvectors and adjoint normalization. At Δ=Δc = _c, the characteristic equation λ+be−λΔ=0λ+be^-λ =0 has roots λ=±iωλ=± iω with ω=bω=b. Define: Eigenvector: φ(θ)=eiωθ,θ∈[−Δc,0]. (θ)=e^iωθ, θ∈[- _c,0]. Adjoint eigenvector: ψ(s)=D¯eiωs,s∈[0,Δc]. ψ(s)= D\,e^iω s, s∈[0, _c]. Normalizing via the Hale–Verduyn Lunel bilinear form ⟨ψ,φ⟩=1 ψ, =1 yields D¯=11+iπ/2. D= 11+iπ/2. Step 3: Nonlinear expansion to cubic order. The full nonlinearity is F(u,ud)=(α0+α1u−u2)[−Cp′ud−12Cp′ud2−16Cp′ud3+⋯],F(u,u_d)=( _0+ _1u-u^2) [-Cp u_d- 12Cp u_d^2- 16Cp u_d^3+·s ], where α0=x∗(1−x∗) _0=x^*(1-x^*) and α1=1−2x∗ _1=1-2x^*. Collecting by order: f2(u,ud) f_2(u,u_d) =−Cα1p′uud =-C _1p \;u\,u_d −12Cα0p′ud2, -\; 12\,C _0p \;u_d^2, (15) f3(u,ud) f_3(u,u_d) =Cp′u2ud = -Cp \;u^2u_d −12Cα1p′uud2−16Cα0p′ud3. -\; 12\,C _1p \;u\,u_d^2\;-\; 16\,C _0p \;u_d^3. For the sigmoid response, the derivatives at x∗x^* evaluate to: p′ p =kρ(1−ρ), =k\,ρ(1-ρ), p′ p =k2ρ(1−ρ)(1−2ρ), =k^2\,ρ(1-ρ)(1-2ρ), (16) p′ p =k3ρ(1−ρ)(1−6ρ(1−ρ)). =k^3\,ρ(1-ρ) (1-6ρ(1-ρ) ). Step 4: Center manifold reduction and normal form. On the center manifold, u≈z+z¯u≈ z+ z and ud≈i(z¯−z)u_d≈ i( z-z) (since e−iωΔc=−ie^-iω _c=-i). Projecting f2f_2 onto the center eigenspace gives the second-order normal form coefficients g20g_20, g11g_11, g02g_02. The center manifold corrections W20(θ)W_20(θ) and W11(θ)W_11(θ) satisfy the boundary value problems (2iω−)W20 (2iω-A)\,W_20 =H20, =H_20, −W11 -A\,W_11 =H11, =H_11, where A is the infinitesimal generator of the linearized semigroup. The right-hand sides H20H_20, H11H_11 are determined by projecting the quadratic nonlinearity onto the stable complement. The resulting values at θ=0θ=0 and θ=−Δcθ=- _c are verified symbolically (supplementary code) to satisfy their boundary conditions identically for all parameter values. These corrections yield the third-order coefficient g21g_21 and the first Lyapunov coefficient: c1(0)=i2ω(g20g11−2|g11|2−13|g02|2)+g212.c_1(0)\;=\; i2ω (\,g_20\,g_11-2\,|g_11|^2- 13\,|g_02|^2\, )\;+\; g_212. (17) Step 5: Sign analysis of Re(c1)Re(c_1). Carrying out the full computation and rationalizing all complex denominators (supplementary code provides the symbolic derivation), we obtain: Re(c1(0))=Ckρ5α0(4+π2)⏟>0⋅(k,ρ,α0,α1)⏟sign?,Re(c_1(0))= Ckρ5 _0(4+π^2)_>0·\; N(k,ρ, _0, _1)_sign?, (18) where the numerator factors as =(ρ−1)⋅ℬN=(ρ-1)·B, with ρ−1<0ρ-1<0 for all ρ∈(0,1)ρ∈(0,1). It remains to show ℬ>0B>0. ∎ Lemma 1 (Positivity of ℬB). For all ρ∈(0,1)ρ∈(0,1), k>0k>0, and x∗∈(0,1)x^*∈(0,1), define ℬ= \;=\; 2Q(ρ)(α0k)2⏟term A−(7π−8)(2ρ−1)(α0k)α1⏟cross-term B 2\,Q(ρ)\;( _0k)^2_term A\;-\; (7π-8)(2ρ-1)\;( _0k)\, _1_cross-term B +2(3π−2)α12⏟term C+10πα0⏟remainder, +\; 2(3π-2)\; _1^2_term C\;+\; 10π\, _0_remainder, where Q(ρ)=−(7π−8)ρ(1−ρ)+(3π−2)Q(ρ)=-(7π-8)ρ(1-ρ)+(3π-2). Then ℬ>0B>0. Proof. Since ρ(1−ρ)≤1/4ρ(1-ρ)≤ 1/4, we have Q(ρ)≥(3π−2)−(7π−8)/4=5π/4>0Q(ρ)≥(3π-2)-(7π-8)/4=5π/4>0. The first three terms form a quadratic form in the variables v1=α0kv_1= _0k and v2=α1v_2= _1: ℬ0(v1,v2)=2Qv12−(7π−8)(2ρ−1)v1v2+2(3π−2)v22.B_0(v_1,v_2)=2Q\,v_1^2-(7π-8)(2ρ-1)\,v_1v_2+2(3π-2)\,v_2^2. This is positive definite whenever the discriminant condition 4AC>B24AC>B^2 holds, i.e., 16Q(3π−2)>(7π−8)2(2ρ−1)2.16\,Q\,(3π-2)>(7π-8)^2(2ρ-1)^2. Substituting μ=ρ(1−ρ)μ=ρ(1-ρ) so that (2ρ−1)2=1−4μ(2ρ-1)^2=1-4μ, the condition becomes [−20π(7π−8)]μ+95π2−80π>0.[-20π(7π-8)]\,μ+95π^2-80π>0. Since the coefficient of μ is negative, the worst case is μ=1/4μ=1/4 (i.e., ρ=1/2ρ=1/2), giving 20π(3π−2)>020π(3π-2)>0. Hence ℬ0B_0 is positive definite for all ρ∈(0,1)ρ∈(0,1). Adding the strictly positive remainder 10πα0>010π _0>0 gives ℬ=ℬ0+10πα0>0B=B_0+10π _0>0. ∎ Conclusion of the proof. Since ρ−1<0ρ-1<0 and ℬ>0B>0 (Lemma 1), we have =(ρ−1)ℬ<0N=(ρ-1)B<0. The prefactor in (18) is strictly positive. Therefore Re(c1(0))<0Re(c_1(0))<0 for all admissible parameters. Bifurcation direction and orbital stability. The standard Hassard–Kazarinoff–Wan quantities are μ2=−Re(c1(0))Re(dλ/dΔ)|Δc>0,β2=2Re(c1(0))<0. _2=- Re(c_1(0))Re(dλ/d ) |_ _c>0, _2=2Re(c_1(0))<0. Since μ2>0 _2>0, the bifurcation is supercritical (periodic orbits exist for Δ>Δc > _c). Since β2<0 _2<0, the bifurcating periodic orbits are orbitally stable. Numerical validation. Quantity Value Interpretation Re(c1(0))Re(c_1(0)) −40.85-40.85 <0<0: supercritical μ2 _2 15.9515.95 >0>0: orbits exist above Δc _c β2 _2 −81.70-81.70 <0<0: orbits are stable Table 7: Bifurcation quantities at representative parameters (a,C,k,xc)=(2,5,10,0.5)(a,C,k,x_c)=(2,5,10,0.5). Conclusion: Re(c1)<0Re(c_1)<0, μ2>0 _2>0, and β2<0 _2<0 jointly confirm a supercritical Hopf bifurcation, with stable limit cycles emerging for Δ>Δc > _c. Three independent checks validate the algebraic reduction: 1. Grid evaluation: on a 30×60×3030× 60× 30 grid (30 values of ρ uniform in (0,1)(0,1), 60 values of k log-uniform in [0.1,104][0.1,10^4], 30 values of xcx_c uniform in [0.02,0.98][0.02,0.98]), Re(c1)Re(c_1) is strictly negative at every admissible point—those with an interior equilibrium x∗∈(0,1)x^*∈(0,1), which is the hypothesis of the proposition. Grid points with no interior equilibrium fall outside the proposition’s scope and are excluded. 2. DDE integration: Direct numerical integration confirms continuous amplitude growth from zero above Δc _c for all tested parameter sets. 3. BVP residuals: Boundary condition residuals for W20W_20 and W11W_11 remain below 10−1210^-12 across 10310^3 random parameter draws. Supplementary scripts (Python/SymPy and Wolfram Language) reproduce all symbolic derivations and numerical checks. Appendix B Experiment Details This appendix provides implementation details for all experiments reported in the main text. Full experiment configurations are stored as JSON manifests in experiments/configs/ and each run produces a deterministic output bundle. B.1 Output structure Each experimental run produces three artifacts: a history.csv file containing the full time series of behavioral fractions, alarm, repression probability, and punishment counts at each time step; a summary.json file containing aggregate statistics computed over the final 50% of the simulation horizon and a regime label; and a config_resolved.json file recording the exact parameter values used. Each sweep directory additionally produces a summary.csv aggregating all per-run summaries. The regime label is assigned by the adaptive classifier (spectral analysis with runaway threshold Rmax>0.40R_ >0.40); the old fixed-threshold classifier is also recorded as regime_fixed for comparison. B.2 Common parameters Unless otherwise noted, all networked simulations use N=240N=240 agents on a stochastic block model graph with 6 communities, intra-community edge probability pin=0.08p_in=0.08, inter-community edge probability pout=0.004p_out=0.004, and bridge fraction 12% (defined by betweenness centrality). Agent Q-learning parameters: α=0.10α=0.10, ϵ=0.08ε=0.08, γ=0.95γ=0.95. Influence coupling strength λ=0.22λ=0.22; alarm threshold Ac=0.72A_c=0.72. Content types are disabled for all Paper 1 experiments (enable_content=false; agents receive neutral payoffs with no content-type bonuses). All experiments use 50 seeds (1–50). B.3 Experiment catalog Table 8 summarizes all six experiments at a glance; the paragraphs that follow give the full configuration for each. Exp. System Varied factor Fixed parameters Runs Headline finding 1 Delayed replicator ODE Δ/Δc∈[0,4] / _c∈[0,4] a=2a=2, C=7C=7, k=10k=10, xc=0.72x_c=0.72 80 Hopf bifurcation at Δc _c exactly; supercritical 2 Discrete mean-field step size η, Δ same as Exp. 1 grid discretization shrinks the stability margin monotonically 3 Network, Q-learning Δ∈0–30 ∈\0--30\ k=20k=20, N=240N=240 450 runaway 10%→72%10\%→ 72\% (+62+62 p) 4 Network, Q-learning k∈3–40k∈\3--40\ Δ=15 =15, N=240N=240 400 runaway 16%→64%16\%→ 64\%; knee at k=7–10k=7--10 5 Network, crossed Δ× × architecture k=10k=10, N=240N=240 750 reactive 96%>96\%> Q-learning 66%>66\%> fixed 0%0\% 6 Network, governance regulator type Δ=6 =6, k=10k=10 150 RL regulator converts runaway to oscillation (exploratory) Table 8: Summary of all Paper 1 experiments. Experiments 1–2 validate the analytical theory (the ODE and its discretization); Experiments 3–6 test it in the networked simulation with N=240N=240 agents and 50 seeds per condition. “Runs” counts independent simulations: Experiment 1 counts ODE integrations and Experiment 2 is a deterministic Δ×η ×η grid, while Experiments 3–6 are seeded (9×509×50, 8×508×50, 3×5×503×5×50, 3×503×50). Conclusion: the delay-instability mechanism survives every added layer of realism, and the central crossing (Experiment 5) isolates reactivity—not learning—as the destabilizing ingredient. All headline numbers reproduce from the frozen summary.csv files under experiments/results/paper1_v2/ (for Experiment 5, group by the learning_mode column). Experiment 1: ODE validation. Numerical integration of the delayed replicator equation (1) using Euler stepping for the DDE (global truncation error O(dt)O(dt); at dt=0.01dt=0.01 the bifurcation point is resolved to within 1% of Δc _c, verified by halving the step size). Parameters: a=2.0a=2.0, C=7.0C=7.0, k=10k=10, xc=0.72x_c=0.72. Delay swept from 0 to 4Δc4 _c in 80 values. Integration horizon: 120 time units with step size dt=0.01dt=0.01. Output: bifurcation diagram, representative trajectories, and sharpness sweep. Experiment 2: Discrete mean-field. Euler discretization of the delayed replicator equation with step size η swept over 0.01,0.02,0.05,0.1,0.2,0.5,1.0\0.01,0.02,0.05,0.1,0.2,0.5,1.0\. Same base parameters as Experiment 1. Delay swept from 0 to 60 steps jointly with η to produce a phase diagram. Each condition run for 2000 steps with deterministic initial conditions at x∗+0.05x^*+0.05. Experiment 3: Delay sweep. Full networked simulation with common parameters above plus k=20k=20, 500 time steps. Delay swept over 0,2,4,6,8,10,14,20,30\0,2,4,6,8,10,14,20,30\ steps. Static regulator with force ut=1.0u_t=1.0. Seeds: 50 per delay level. Experiment 4: Sharpness sweep. Same base configuration as Experiment 3 with delay fixed at 15 steps. Sharpness k swept over 3,5,7,10,15,20,30,40\3,5,7,10,15,20,30,40\. Seeds: 50 per sharpness level. Experiment 5: Crossed delay × architecture. Two-factor design crossing delay ∈0,4,8,14,20∈\0,4,8,14,20\ with agent architecture ∈fixed,reactive,Q-learning∈\fixed,reactive,Q-learning\. Fixed agents use stationary action weights [0.65,0.27,0.08][0.65,0.27,0.08] for [L,M,R][L,M,R], chosen to approximate the empirical radical fraction at Q-learning convergence without delay and thus provide a matched non-adaptive baseline. Reactive agents use a threshold heuristic that responds to the delayed alarm, local influence signal, and punishment state but maintains no memory across steps. Q-learning agents use α=0.10α=0.10, ϵ=0.08ε=0.08. Default sharpness k=10k=10. Simulation horizon: 500 steps. Seeds: 50 per cell. Experiment 6: RL regulator. Three conditions: (1) fixed agents with static regulator (ut=1.0u_t=1.0), (2) Q-learning agents with static regulator, (3) Q-learning agents with RL regulator. The RL regulator is a tabular Q-learning agent (αreg=0.05 _reg=0.05, ϵreg=0.10 _reg=0.10, γreg=0.95 _reg=0.95) observing a discretized state (4 alarm buckets × 3 radical-fraction buckets) and selecting utu_t from 0.0,0.5,1.0,1.5,2.0,2.5\0.0,0.5,1.0,1.5,2.0,2.5\. Regulator reward: rreg=−xR−0.1utr_reg=-x_R-0.1u_t. Default delay (Δ=6 =6 steps) and sharpness (k=10k=10). Simulation horizon: 500 steps. Seeds: 50 per condition. B.4 Reproducibility All results are generated from frozen random seeds (seeds 1–50 for each condition). The simulation code is deterministic given a seed, configuration, and Python/NumPy version. Results reported in the main text are stored under experiments/results/paper1_v2/. The regime classifier thresholds are documented in experiments/src/autonomy_lab/metrics.py and were fixed prior to running the final experiments.