Paper deep dive
The Tragedy of the Commons in Multi-Population Resource Games
Yamin Vahmian, Keith Paarporn
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/20/2026, 2:26:09 PM
Summary
This paper analyzes the 'tragedy of the commons' in a multi-population resource game featuring a bi-level decision-making hierarchy. It models strategic interactions where high-level 'greedy' agents control extraction rates for their respective populations, while a single 'responsible' agent implements pro-environmental policies. The study characterizes a unique symmetric Nash equilibrium in the high-level resource extraction game, identifying conditions under which the common environmental resource is sustained or depleted as the number of greedy populations increases.
Entities (6)
Relation Signals (5)
Resource Extraction Game → hasequilibrium → Nash Equilibrium
confidence 95% · We characterize a unique symmetric Nash equilibrium in the high-level game
Tragedy of the Commons → iscausedby → Self-optimizing behaviors
confidence 95% · Self-optimizing behaviors can lead to outcomes where collective benefits are ultimately destroyed, a well-known phenomenon known as the ``tragedy of the commons
Greedy Agents → degrades → Common Environmental Resource
confidence 90% · populations benefit from a common environmental resource that degrades with higher extractive efforts made by high-level agents
Responsible Agent → sustains → Common Environmental Resource
confidence 90% · one “responsible” agent implements a pro-environmental policy for its population such that the common resource can be sustained
Number of Populations → influences → Resource Depletion
confidence 85% · While the equilibrium resource level degrades as the number of populations grows large, there are instances where it does not become depleted.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Self-optimizing behaviors can lead to outcomes where collective benefits are ultimately destroyed, a well-known phenomenon known as the ``tragedy of the commons". These scenarios are widely studied using game-theoretic approaches to analyze strategic agent decision-making. In this paper, we examine this phenomenon in a bi-level decision-making hierarchy, where low-level agents belong to multiple distinct populations, and high-level agents make decisions that impact the choices of the local populations they represent. We study strategic interactions in a context where the populations benefit from a common environmental resource that degrades with higher extractive efforts made by high-level agents. We characterize a unique symmetric Nash equilibrium in the high-level game, and investigate its consequences on the common resource. While the equilibrium resource level degrades as the number of populations grows large, there are instances where it does not become depleted. We identify such regions, as well as the regions where the resource does deplete.
Tags
Links
- Source: https://arxiv.org/abs/2602.20603v1
- Canonical: https://arxiv.org/abs/2602.20603v1
Trouble viewing inline? Open PDF directly →
Full Text
51,974 characters extracted from source content.
Expand or collapse full text
The Tragedy of the Commons in Multi-Population Resource Games †thanks: Yamin Vahmian, Keith Paarporn Y. Vahmian and K. Paarporn are with the Department of Computer Science, University of Colorado, Colorado Springs. Contact: yvahmian,kpaarpor@uccs.edu. This paper is based on previous work that studied a two-population setting [11], but with no strategic game interactions between high-level decision-makers. This work is supported by ONR grant #N000142612120, and in part by NSF grant #ECCS-2346791. Abstract Self-optimizing behaviors can lead to outcomes where collective benefits are ultimately destroyed, a well-known phenomenon known as the “tragedy of the commons”. These scenarios are widely studied using game-theoretic approaches to analyze strategic agent decision-making. In this paper, we examine this phenomenon in a bi-level decision-making hierarchy, where low-level agents belong to multiple distinct populations, and high-level agents make decisions that impact the choices of the local populations they represent. We study strategic interactions in a context where the populations benefit from a common environmental resource that degrades with higher extractive efforts made by high-level agents. We characterize a unique symmetric Nash equilibrium in the high-level game, and investigate its consequences on the common resource. While the equilibrium resource level degrades as the number of populations grows large, there are instances where it does not become depleted. We identify such regions, as well as the regions where the resource does deplete. I Introduction Population games are well-suited for analyzing the emergent collective behaviors among a large number of individuals [15]. A central topic in this area concerns whether the self-interested behaviors lead to outcomes where common benefits are destroyed. This is the well-known phenomenon called the “tragedy of the commons”. A primary research goal is to identify, and ultimately influence, the factors that can avert such tragedies. Feedback-evolving games are a class of models that directly incorporates an environmental state that co-evolves with population behaviors [20, 18]. They have been extensively applied to examine the utilization of common resources [17, 21], social distancing in epidemics [16, 14], and ecological dynamics [18]. This research has extended our understanding of which types of learning dynamics [2, 17, 12, 13] and control policies [19, 8] are able to avert tragic outcomes. However, much of this research is in the context of a single population, where everyone is independently interacting with a common environment. In many scenarios, individuals’ behaviors are influenced by higher-level decision-making agents. For example, neighboring states or countries set different environmental policies, yet they all consume resources from a shared environment (e.g. fish in the ocean, drinking water, and clean air). In this paper, we build on this framework by considering multiple populations and a bi-level hierarchy of decision-making. Each population is represented by a high-level agent that is able to set the incentive policies for their local population. We focus on a particular setting where one “responsible” agent implements a pro-environmental policy for its population such that the common resource can be sustained in the absence of other populations. All other high-level agents are considered “greedy”, who implement policies that induce total consumption behavior in their populations. Our study investigates the strategic decision-making among the greedy high-level agents, where they are able to control the allowable magnitude of consumption in their populations. The goal in this paper then is to characterize the emergent strategic behavior in this hierarchical setting, and identify conditions for which the multi-population behavior induces or averts a tragedy of the commons. Specifically, we focus on a normal-form game between the high-level greedy agents with coupling constraints on their strategies, which we call the resource extraction game. The utilities in this game are based on the stable long-term outcomes from the evolutionary behavior from the low-level individuals. Each greedy agent seeks to extract as much of the resource as possible – however, higher extractive efforts degrade the resource level. Our primary contribution is the characterization of Nash equilibria in the resource extraction game (Theorem 4.1). In particular, our results show there is a unique symmetric Nash equilibrium for any instance of the game, i.e. all greedy agents choose the same extractive effort. We then identify the parameter regimes for which the common resource can be sustained in the equilibrium, even when the number of greedy populations grows large (Proposition 5.1). Related works: There is an emerging body of work that focuses on strategic interactions in multi-population settings. Evolutionary dynamics over a network of populations is considered in [9], where convergence to equilibria is established. Externalities between two interacting populations have been studied in contexts such as epidemic spreading [4] and shared environmental states [10, 3]. Several recent papers propose hierarchical decision-making structures: [5] proposes new extensions for population games, [7] studies mean-field games with minor and major agents in the regulation of carbon emissions, and [6, 1] consider competing service providers as high-level agents that strategically set prices. Section I provides preliminaries on single-population feedback-evolving games. Section I formulates our multi-population setup and resource extraction game. Section IV presents an analysis and main results. Section V examines consequences of the main results with numerical studies. I Preliminary: single population Consider a single population of interacting agents that have access to a common pool resource, where their choices are to cooperate (low consumption, L) or defect (high consumption, H). At time t≥0t≥ 0, let x(t)∈[0,1]x(t)∈[0,1] denote the fraction of low consumers and n(t)∈[0,1]n(t)∈[0,1] the resource level, with n=1n=1 corresponding to a fully abundant state and n=0n=0 to a depleted state. In a standard feedback-evolving game, these quantities obey the following coupled dynamics: x˙ x =x(1−x)g(x,n) =x(1-x)g(x,n) (1) n˙ n =n(1−n)(θx−α(1−x)) =n(1-n)(θ x-α(1-x)) where g(x,n):=πL(x,n)−πH(x,n)g(x,n):= _L(x,n)- _H(x,n) is the payoff difference between low and high consumers where πL(x,n) _L(x,n), πH(x,n) _H(x,n) denote their payoffs respectively (to be defined soon), θ>0θ>0 is the restoration rate due to low consumption, and α>0α>0 is the resource extraction rate due to high consumption activity. The x˙ x equation above is the standard replicator dynamics, and the n˙ n equation describes logistic growth for the resource level. This system exhibits a feedback mechanism: agents’ strategies influence n which in turn affects the payoffs. We note that the system (1) is forward-invariant in the interior set, (x,n)∈(0,1)2(x,n)∈(0,1)^2. The agent payoffs are described by the environment-dependent 2×22× 2 payoff matrix, An=n[R1S1T1P1]+(1−n)[R0S0T0P0]A_n=n bmatrixR_1&S_1\\ T_1&P_1 bmatrix+(1-n) bmatrixR_0&S_0\\ T_0&P_0 bmatrix (2) Here, the first row and column corresponds to a low consumer, and the second row and column corresponds to a high consumer. Entry ijij (i,j∈ℒ,ℋi,j∈\L,H\) indicates the experienced payoff to an agent using strategy i when encountering an agent using strategy j. We denote x∈[0,1]x∈[0,1] as the fraction (or frequency) of agents in the population using strategy ℒL. The payoff experienced by each type of agent is then given by πℒ(x,n)=[An[x,1−x]⊤]1,πℋ(x,n)=[An[x,1−x]⊤]2 _L(x,n)=[A_n[x,1-x] ]_1, _H(x,n)=[A_n[x,1-x] ]_2 (3) The payoffs are determined by the parameters in the A0A_0 and A1A_1 matrices. The A1A_1 matrix is the payoff matrix when the environment is abundant. Following the literature on feedback-evolving games, we make the following assumption about the A1A_1 matrix. Assumption 1. High consumption is the dominant strategy in A1A_1, i.e. δTR1:=T1−R1>0 _TR1:=T_1-R_1>0 and δPS1:=P1−S1>0 _PS1:=P_1-S_1>0 On the other hand, the A0A_0 matrix describes incentives when the environment is near depletion, and can be interpreted to be an “environmental policy” that is implemented to reduce pressure on the common resource (e.g. subsidies for using electric vehicles). The payoff structure of the A0A_0 matrix is completely determined from the parameters δSP0:=S0−P0 _SP0:=S_0-P_0 and δRT0:=R0−T0 _RT0:=R_0-T_0. In this paper, we will also make the following assumption about the parameters (δSP0,δRT0)( _SP0, _RT0). Assumption 2. We assume that ∂g∂n(x)<0 ∂ g∂ n(x)<0 for all x∈[0,1]x∈[0,1]. This is equivalent to δSP0>−δPS1 _SP0>- _PS1 and δRT0>−δTR1 _RT0>- _TR1. Assumption 2 asserts that the relative payoff to low consumers in the responsible population monotonically decreases as the resource level improves. In other words, high consumption becomes more incentivized as resource become more available. The result below summarizes the findings of the originating work [20], which identified conditions on the policies (δSP0,δRT0)( _SP0, _RT0) that enable long-term behavior of (1) to sustain the environmental resource. Theorem 2.1 (adapted from [20]). The environmental policy (δSP0,δRT0)( _SP0, _RT0) determines the asymptotic properties of (1) as follows. 1) Resource sustained: If (δSP0,δRT0)( _SP0, _RT0) satisfies δSP0>0 and −θαδSP0≤δRT0<δTR1δPS1δSP0, _SP0>0 and - θα _SP0≤ _RT0< _TR1 _PS1 _SP0, (4) then the fixed point (x∗,n∗)=(α+θ,g(x∗,0)−∂g∂n(x∗))∈(0,1)×[0,1)(x^*,n^*)=( α+θ, g(x^*,0)- ∂ g∂ n(x^*))∈(0,1)×[0,1) is the only asymptotically stable fixed point in the system. Here, n∗=0n^*=0 if and only if −θαδSP0=δRT0- θα _SP0= _RT0. 2) Oscillating tragedy of the commons: If δSP0>0 _SP0>0 and δTR1δPS1δSP0<δRT0 _TR1 _PS1 _SP0< _RT0, then the system exhibits convergence to the heteroclinic cycle composed of the boundary of [0,1]2[0,1]^2. If δTR1δPS1δSP0=δRT0 _TR1 _PS1 _SP0= _RT0, then all trajectories of the system are closed orbits centered around the neutrally stable interior fixed point (x∗,n∗)=(α+θ,g(x∗,0)−∂g∂n(x∗))∈(0,1)2(x^*,n^*)=( α+θ, g(x^*,0)- ∂ g∂ n(x^*))∈(0,1)^2. 3) Resource collapse: If (δSP0,δRT0)( _SP0, _RT0) does not satisfy the conditions of items 1) or 2), then the only asymptotically stable fixed point has n=0n=0. The policies described in item 1 above induces a stable and non-zero resource level, and the policies in items 2 and 3 induce undesirable environmental outcomes – either the resource oscillates between bad and good states indefinitely, or it collapses entirely. We consequently obtain a natural definition for a set of “responsible” policies. Figure 1: The shaded green region depicts the set of all policies (δSP0,δRT0)( _SP0, _RT0) that are responsible, as defined as in Definition 1. Definition 1. An environmental policy (δSP0,δRT0)( _SP0, _RT0) is responsible if it satisfies (4) and Assumption 2. Specifically, max−θαδSP0,−δTR1≤δRT0<δTR1δPS1δSP0. \ -θα _SP0,- _TR1\≤ _RT0< _TR1 _PS1 _SP0. (5) In other words, a responsible policy for a population averts a “tragedy of the commons” (TOC), which we refer to as instances where n(t)→0n(t)→ 0. Figure 1 shows a diagram of the region of responsible policies. I Model: Extraction game with multiple populations Consider an extension of (1) to multiple populations, all of whom share the same common-pool resource. In this paper, we focus on a scenario where one of the populations has a responsible policy, and all other populations do not, which are termed “greedy”. The state of this system is specified as (x,x1,…,xM,n)∈[0,1]M+2(x,x_1,…,x_M,n)∈[0,1]^M+2, where x denotes the fraction of cooperators in the responsible population, xix_i the fraction of cooperators in greedy population i∈:=1,…,Mi :=\1,…,M\, and n the common resource level. The state evolves according to the dynamics x˙ x =x(1−x)g(x,n) =x(1-x)g(x,n) (6) x˙i x_i =xi(1−xi)gi(x,n),∀i∈ =x_i(1-x_i)g_i(x,n),\ ∀ i n˙ n =ϵn(1−n)(θx+∑i∈θixi =ε n(1-n) (θ x+ _i _ix_i . −(α(1−x)+∑i∈αi(1−xi))). .- (α(1-x)+ _i _i(1-x_i) ) ). Here, we have introduced corresponding rate parameters αi,θi>0 _i, _i>0 for the greedy populations. In line with Assumption 1, we also assume corresponding payoff parameters in the abundant state satisfy δTR1i,δPS1i>0 _TR1^i, _PS1^i>0. However, for all greedy populations, we will assume that their environmental policies all satisfy δSP0i,δRT0i<0 _SP0^i, _RT0^i<0. This implies that the payoff differences for greedy populations satisfy gi(xi,n):=πLi−πHi<0g_i(x_i,n):= _L^i- _H^i<0 for all states (xi,n)(x_i,n), where πLi,πHi _L^i, _H^i are defined analogously to (3). This means that low consumption is never incentivized in greedy populations. We reserve the un-indexed parameters α, θ, δTR1,δPS1 _TR1, _PS1, and δSP0,δRT0 _SP0, _RT0 for those associated with the responsible population. By construction, system (6) is forward-invariant on the set (0,1)M+2(0,1)^M+2, and can only admit stable fixed points for which xi=0x_i=0 for all i∈i . Lemma 3.1. Denote α¯:=∑i∈αi α:= _i _i. The asymptotic dynamics of the multi-population system (6) are summarized below. 1) Suppose α¯>θ α>θ. Then a tragedy of the commons is asymptotically stable, i.e. limt→∞n(t)=0 _t→∞n(t)=0. 2) Suppose α¯<θ α<θ. (a) If α¯−θα+α¯δSP0≤δRT0<δTR1δPS1δSP0 α-θα+ α _SP0≤ _RT0< _TR1 _PS1 _SP0, then the only asymptotically stable fixed point is (x1∗,M,n∗)(x_1^*,0_M,n^*), where x∗:=α+α¯α+θ,n∗:=−g(x∗,0)∂g∂n(x∗)∈(0,1).x^*:= α+ α+θ, n^*:=- g(x^*,0) ∂ g∂ n(x^*)∈(0,1). (7) (b) If δRT0<α¯−θα+α¯δSP0 _RT0< α-θα+ α _SP0, then a tragedy of the commons is the only asymptotically stable outcome. 3) Suppose α¯=θ α=θ. If δRT0>0 _RT0>0, then there is a locally stable line segment of fixed points given by (1,M,n):n∈[0,δRT0δRT0+δTR1) \(1,0_M,n):n∈ [0, _RT0 _RT0+ _TR1 ) \ (8) All other fixed points are isolated and unstable. If δRT0≤0 _RT0≤ 0, then a tragedy of the commons is asymptotically stable. Proof. The main arguments for this result are fully detailed in [11], which established the case for a single greedy population. Because all greedy states xi(t)→0x_i(t)→ 0 by construction, the extension to M greedy populations easily generalizes because it only requires using the total consumption rate α¯ α in place of the consumption rate for the single greedy population. We therefore omit these details, as no new fundamental arguments are needed. ∎ In item 1 above, the greedy populations induce a tragedy of the commons if their total consumption rate, α¯ α, exceeds the responsible population’s restoration rate θ. Item 2a provides a region of sustainable policies for population 1. This region gets smaller as α¯ α increases while remaining less than θ. For notational convenience, let us denote this region as (α¯):=(y1,y2):maxα¯−θα+α¯y1,−δTR1≤y2<δTR1δPS1y1 ( α):= \(y_1,y_2): \ α-θα+ αy_1,- _TR1 \≤ y_2< _TR1 _PS1y_1 \ (9) Observe that (α¯)V( α) is a sub-region of the set of responsible policies (Definition 1), in which the lower bound increases as the total extraction α¯ α increases. Item 2b provides a region where the responsible population fails to sustain the resource even though α¯<θ α<θ. Figure 2: The resource extraction game is a strategic-form game among M players who each represent one of the greedy populations. Player i∈i chooses extraction rate αi≥0 _i≥ 0 for its population, where higher rates degrade the resource. Hierarchical resource extraction game We now pose a high-level decision-making problem that is the central focus of this paper. We model the resource extraction game as a complete-information game in which each high-level agent observes the aggregate extraction level and the individual extraction choices of the other agents. This assumption enables agents to compute best responses, and reflects settings where extraction decisions or their aggregate effects are publicly observable or effectively shared among decision-makers. Consider M high-level agents that represent each of the greedy populations. The agents are strategic in selecting its population’s resource extraction rate, αi≥0 _i≥ 0. This is representative of, for example, the intensity of a corporation’s extraction of a natural resource in the presence of competitors, as well as analogous settings such as competing jurisdictions regulating shared environmental resources or firms influencing common infrastructure capacity. The utility of each agent i∈i is defined as Ui():=αiR()U_i( α):= _iR( α) (10) where :=(α1,…,αM) α:=( _1,…, _M) is the strategy profile of the extraction decisions, and R()R( α) is the steady-state resource level induced by the constituent low-level agents, characterized in Lemma 3.1: R(α¯):=−g(α+α¯α+θ,0)∂g∂n(α+α¯α+θ), if α¯≤θ and (δSP0,δRT0)∈(α¯)0, elseR( α):= cases- g( α+ α+θ,0) ∂ g∂ n( α+ α+θ), if α≤θ and ( _SP0, _RT0) ( α)\\ 0, else cases (11) Lemma 3.1 (item 3) states that when α¯=θ α=θ, there is a range of possible asymptotic resource levels. In our definition of R(⋅)R(·) above, we have elected to assign the highest stable resource level 111This is done for two reasons. First, it makes the resource function well-defined and left-continuous at θ. Thus, R attains a maximum value in the interval [0,θ][0,θ]. Second, since no particular initial condition is prescribed in the model, we may view the irresponsible populations to be “optimistic” regarding the best case among all possible outcomes., δRT0δRT0+δTR1 _RT0 _RT0+ _TR1. The utilities (10) define a strategic-form game between the M local authorities, with continuous strategy spaces i:=ℝ≥0A_i:=R_≥ 0. Figure 2 provides an illustration. We refer to this game as the resource extraction game, and denote an instance of it as the tuple ℰ(δSP0,δRT0,α,θ)E( _SP0, _RT0,α,θ) of fixed parameters. Table I summarizes the key system parameters and decision variables in the resource extraction game. In our forthcoming analysis, we seek to characterize Nash equilibria of ℰE, i.e. any strategy profile α that satisfies Ui(αi,α−i)≥Ui(αi′,α−i)U_i( _i, _-i)≥ U_i( _i , _-i) for all i∈i and for all αi′≠αi _i ≠ _i. In particular, we wish to draw attention to the extractive behavior and resource level R()R( α) under equilibria of the game, and how these quantities change as the number of greedy populations increases. TABLE I: Parameters and variables of the multi-population system System Parameters δSP0,δRT0 _SP0, _RT0 Responsible policy in deplete state (see Figure 1) δTR1,δPS1 _TR1, _PS1 Responsible policy in replete state (fixed, both positive) α,θα,θ Consumption, restoration rates of responsible population Decision Variables αi _i Consumption rate of greedy population i=1,…,Mi=1,…,M α¯ α Total consumption rate of greedy populations α¯−i α_-i Total consumption rate of greedy populations, excluding population i IV Analysis and Main Results In this section, we seek to characterize Nash Equilibria of the extraction game ℰE. Notice from (11) that the shared resource will be destroyed if the total extraction α¯>θ α>θ, upon where all players will receive zero utility. Indeed, there are an infinite number of equilibria under which the resource is destroyed. Lemma 4.1. An equilibrium α of ℰE satisfies R()=0R( α)=0 if and only if for all i∈i , we have α¯−αi>θ α- _i>θ. Under the equilibria described above, the total extraction will still exceed θ regardless of any player’s unilateral deviation. We thus focus our attention on finding equilibria for which the resource level is sustained, i.e. R()>0R( α)>0. Moving forward, we consider extraction profiles that do not destroy the resource, i.e. ones that obey the following coupling constraint, ∈:=:α¯≤θ and (δSP0,δRT0)∈(α¯). α :=\ α: α≤θ and ( _SP0, _RT0) ( α)\. (12) This causes player i’s available extraction choices to be coupled with the profile of extraction rates α−i _-i from the other players. Let us define a player’s restricted strategy set as i(α−i) _i( _-i) := = (13) [0,θ−α¯−i], if δRT0>0[0,αδRT0+θδSP0δSP0−δRT0−α¯−i], if α¯−i−θα+α¯−iδSP0≤δRT0≤0. -22.76219pt. Let us now denote ℰM∗(δSP0,δRT0,α,θ)E^*_M( _SP0, _RT0,α,θ) to mean the extraction game ℰE under the coupling constraint C. In the following result, we demonstrate that it is a concave game. Lemma 4.2. The game ℰM∗(δSP0,δRT0,α,θ)E^*_M( _SP0, _RT0,α,θ) is a concave game. That is, for any i∈i and any α−i _-i, the function Ui(αi,α−i)U_i( _i, _-i) is concave in αi _i. Proof. To establish concavity, we need to show that Ui′(αi,α−i)<0U_i ( _i, _-i)<0 for all αi∈i(α−i) _i _i( _-i) (here, U′U denotes partial derivative w.r.t. αi _i). For all αi∈i(α−i) _i _i( _-i), it holds that R(αi,α−i)=−g(x,0)∂g∂n(x)R( _i, _-i)=- g(x,0) ∂ g∂ n(x), where x∗=α+α¯α+θx^*= α+ α+θ (7). We have Ui′(αi,α−i)=2R′(αi,α−i)+αiR′(αi,α−i).U _i( _i, _-i)=2R ( _i, _-i)+ _iR ( _i, _-i). (14) We introduce the following parameters: a a :=δSP0−δRT0+δPS1−δTR1 = _SP0- _RT0+ _PS1- _TR1 (15) b b :=δRT0−δSP0 = _RT0- _SP0 c c :=−(δPS1+δSP0)<0 =-( _PS1+ _SP0)<0 d d :=δSP0>0 = _SP0>0 Because these are payoff parameters of the responsible population, we know that c<0c<0, d>0d>0, but the signs of a,ba,b can be positive or negative depending on the policy. Using this notation, we can write g(x,n) g(x,n) =axn+bx+cn+d =axn+bx+cn+d (16) ∂g∂n(x) ∂ g∂ n(x) =ax+c =ax+c for any x,n∈[0,1]x,n∈[0,1]. After algebraic computations, we can write the second derivative in the form Ui′(αi,α−i)=2Y(α+θ)(∂g∂n(x∗))2(−1+αiα+θa∂g∂n(x∗)),U _i( _i, _-i)= 2Y(α+θ)( ∂ g∂ n(x^*))^2 (-1+ _iα+θ a ∂ g∂ n(x^*) ), (17) where Y:=bc−ad=δTR1δSP0−δRT0δPS1>0.Y:=bc-ad= _TR1 _SP0- _RT0 _PS1>0. (18) Hence, the sign of Ui′U_i is the sign of the term in the parentheses, which we can show is negative. The condition for which this is negative is: −1+aα+θαi∂g∂n(x)<0⇔0>∂g∂n(α+α¯−iα+θ). -1+ aα+θ _i ∂ g∂ n(x)<0 0> ∂ g∂ n( α+ α_-iα+θ). (19) Here, we have used the fact that ∂g∂n(⋅)<0 ∂ g∂ n(·)<0 (Assumption 2), and that one can write ∂g∂n(x)=ax+c ∂ g∂ n(x)=ax+c with c:=−(δPS1+δSP0)c:=-( _PS1+ _SP0). This concludes the proof. ∎ In the next result, we leverage the concavity of the game ℰM∗E^*_M to establish the best-response functions, BRi(α−i):=argmaxαi∈i(α−i)Ui(αi,α−i).BR_i( _-i):= _ _i _i( _-i)U_i( _i, _-i). (20) Lemma 4.3. Consider any ∈ α and any player i. Define Fi(α¯−i):= F_i( α_-i)= (21) α+θa(−∂g∂n(α+α¯−iα+θ)+1bb∂g∂n(α+α¯−iα+θ)Y). α+θa (- ∂ g∂ n ( α+ α_-iα+θ )+ 1b b ∂ g∂ n ( α+ α_-iα+θ )Y ). where a,ba,b are defined in (15), and Y is defined in (18). Define C(α¯−i;δSP0):=12[−(δTR1+θ−α¯−iα+θδPS1)+ C( α_-i; _SP0)= 12 [- ( _TR1+ θ- α_-iα+θ _PS1 )+ . (22) (δTR1+θ−α¯−iα+θδPS1)2+4θ−α¯−iα+θδTR1δSP0]. . ( _TR1+ θ- α_-iα+θ _PS1 )^2+4 θ- α_-iα+θ _TR1 _SP0 ]. Then player i’s best response is given as follows. If C(α¯−i;δSP0)≤δRT0≤δTR1δPS1δSP0C( α_-i; _SP0)≤ _RT0≤ _TR1 _PS1 _SP0, then BRi(α−i)=θ−α¯−i.BR_i( _-i)=θ- α_-i. (23) If maxα¯−i−θα−α¯−iδSP0,−δTR1≤δRT0<C(α¯−i;δSP0) \ α_-i-θα- α_-i _SP0,- _TR1\≤ _RT0<C( α_-i; _SP0), then BRi(α−i)=Fi(α¯−i),if a≠012(θδSP0+αδRT0δSP0−δRT0−α¯−i),if a=0.BR_i( _-i)= casesF_i( α_-i),&if a≠ 0\\ 12 ( θ _SP0+α _RT0 _SP0- _RT0- α_-i ),&if a=0 cases. (24) Proof. From Lemma 4.2, the function D(αi)=Ui(αi,α−i)D( _i)=U_i( _i, _-i) is concave and continuous in its restricted strategy set αi∈i(α−i) _i _i( _-i) (13). D′(αi) D ( _i) =R()+αiR′(αi,α−i) =R( α)+ _iR ( _i, _-i) (25) =−g(x,0)∂g∂n(x)−αiY(α+θ)(∂g∂n(x))2. =- g(x,0) ∂ g∂ n(x)- _i Y(α+θ)( ∂ g∂ n(x))^2. where x=α+αi+∑j≠iαjα+θx= α+ _i+ _j≠ i _jα+θ. We first verify that it is increasing at αi=0 _i=0. Writing α^−i=α+α¯−iα+θ α_-i= α+ α_-iα+θ for compactness, and invoking Assumption 2, the condition D′(0)>0D (0)>0 is equivalent to the condition g(α^−i,0)=δSP0(θ−α¯−iα+θ)+δRT0(α+α¯−iα+θ)>0g( α_-i,0)= _SP0( θ- α_-iα+θ)+ _RT0( α+ α_-iα+θ)>0. From the constraint (δSP0,δRT0)∈(α¯i)( _SP0, _RT0) ( α_i) (9), we use the fact that δRT0>α¯−i−θα+θδSP0 _RT0> α_-i-θα+θ _SP0 to obtain g(α^−i,0)>0g( α_-i,0)>0. Case 1: C(α¯−i;δSP0)≤δRT0≤δTR1δPS1δSP0C( α_-i; _SP0)≤ _RT0≤ _TR1 _PS1 _SP0. In this regime, it holds that δRT0>0 _RT0>0 so that i(α−i)=[0,θ−α¯−i]A_i( _-i)=[0,θ- α_-i]. The expression D′(θ−α¯−i)D (θ- α_-i) is a convex quadratic function in δRT0 _RT0, whose only positive root is given precisely by C(α¯−i;δSP0)C( α_-i; _SP0). Therefore for all δRT0≥Ci(α¯−i;δSP0) _RT0≥ C_i( α_-i; _SP0), D′(θ−α¯−i)≥0D (θ- α_-i)≥ 0. Because D′(0)>0D (0)>0, this means that D is monotonically increasing on αi∈i(α−i) _i _i( _-i), so that BRi(α−i)=θ−α¯−iBR_i( _-i)=θ- α_-i. Case 2a: 0<δRT0<C(α¯−i;δSP0)0< _RT0<C( α_-i; _SP0). Here, i(α−i)=[0,θ−α¯i]A_i( _-i)=[0,θ- α_i], and D must have a critical point in (0,θ−α¯i)(0,θ- α_i) because D′(θ−α¯i)<0D (θ- α_i)<0. The solution to the equation D′(αi)=0D ( _i)=0 satisfies: ∂g∂n(x)g(x,0)+Yαiα+θ=0 ∂ g∂ n(x)g(x,0)+Y _iα+θ=0 (26) Using (16), we can write ∂g∂n(x)g(x,0)=(aαiα+θ+∂g∂n(α^−i))(bαiα+θ+g(α^−i,0)) ∂ g∂ n(x)g(x,0)=(a _iα+θ+ ∂ g∂ n( α_-i))(b _iα+θ+g( α_-i,0)). The equation D′(αi)=0D ( _i)=0 then becomes ab(αiα+θ)2+2b∂g∂n(α^−i)αiα+θ+∂g∂n(α^−i)g(α^−i,0)=0. ab ( _iα+θ )^2+2b ∂ g∂ n ( α_-i ) _iα+θ+ ∂ g∂ n ( α_-i )g ( α_-i,0 )=0. (27) In the case that a=0a=0, then we obtain the solution (24) (second entry). In the case that a≠0a≠ 0, the roots of this equation can be reduced to be: αi±=α+θa(−∂g∂n(α^−i)±1bb∂g∂n(α^−i)Y)=Fi(α¯−i). _i^±= α+θa (- ∂ g∂ n( α_-i)± 1b b ∂ g∂ n( α_-i)Y )=F_i( α_-i). (28) We are interested only in solutions αi≥0 _i≥ 0. We know Y is a positive constant (proof of Lemma 4.2, and ∂g∂n(⋅)<0 ∂ g∂ n(·)<0 (Assumption 2). It also holds here that b<0b<0 because δRT0<Ci(α−i;δSP0) _RT0<C_i( _-i; _SP0). The sign of the radicand in (28) is positive, and we have that αi+≥0 _i^+≥ 0. This is because the sign of the parentheses term in (28) is the same sign as the parameter a, thus establishing that αi>0 _i>0. To see this, the condition that −∂g∂n(α^−i)+1bb∂g∂n(α^−i)Y>0- ∂ g∂ n( α_-i)+ 1b b ∂ g∂ n( α_-i)Y>0 is equivalent to ag(α^−i,0)>0ag( α_-i,0)>0. Since g(α^−i,0)>0g( α_-i,0)>0, the condition holds if and only if a>0a>0. Now, consider αi− _i^-. Here, the parentheses term is positive because it is the sum of two positive terms. If a<0a<0, then αi−<0 _i^-<0 and thus αi+≥0 _i^+≥ 0 is the only non-negative root. If a>0a>0, it holds that both roots are non-negative. However, we can deduce that αi+ _i^+ is the root that lies in i(α−i)A_i( _-i) for the following reasons. First, it holds that αi+<αi− _i^+< _i^-. Because D(αi)D( _i) is continuous and concave for all αi∈i(α−i) _i _i( _-i), it can have at most one critical point on this interval. Case 2b: max−δTR1,α¯−i−θα−α¯−iδSP0≤δRT0≤0 \- _TR1, α_-i-θα- α_-i _SP0\≤ _RT0≤ 0. Here, i(α−i)=[0,αrδRT0+θδSP0δSP0−δRT0−α¯−i]A_i( _-i)=[0, _r _RT0+θ _SP0 _SP0- _RT0- α_-i]. Also in this regime, we have that b<0b<0. At αi=αrδRT0+θδSP0δSP0−δRT0−α¯−i _i= _r _RT0+θ _SP0 _SP0- _RT0- α_-i, it holds that g(α+α¯α+θ,0)=0g( α+ α+θ,0)=0, and thus D(αrδRT0+θδSP0δSP0−δRT0−α¯−i)=0D( _r _RT0+θ _SP0 _SP0- _RT0- α_-i)=0. Since D′(0)>0D (0)>0, there is a critical point in the interior of i(α−i)A_i( _-i). This must be given precisely by the expression αi+ _i^+ (28). ∎ With the best-response functions established, we now state our main result, which characterizes the equilibrium extraction rate for all instances of ℰM∗E^*_M. Theorem 4.1. Define E± E_± :=12M2[−α+θa(2M∂g∂n(α+θ)−M−1bY) = 12M^2 [- α+θa (2M ∂ g∂ n( α+θ)- M-1bY ) . (29) ±(α+θ)2Ya2b(4M2∂g∂n(α+θ)+(M−1)2Yb)]. .± (α+θ)^2Ya^2b (4M^2 ∂ g∂ n( α+θ)+(M-1)^2 Yb ) ]. The game ℰM∗(δSP0,δRT0,α,θ)E^*_M( _SP0, _RT0,α,θ) admits a unique symmetric Nash equilibrium (α∗,…,α∗)(α^*,…,α^*), given as follows. If C(M−1Mθ;δSP0)≤δRT0≤δTR1δPS1δSP0C( M-1Mθ; _SP0)≤ _RT0≤ _TR1 _PS1 _SP0, then α∗=θM.α^*= θM. (30) If max−θαδSP0,−δTR1≤δRT0<C(M−1Mθ;δSP0) \ -θα _SP0,- _TR1\≤ _RT0<C( M-1Mθ; _SP0), then α∗=E+,if a<0E−,if a>01M+1θδSP0+αδRT0δSP0−δRT0,if a=0,α^*= casesE_+,&if a<0\\ E_-,&if a>0\\ 1M+1 θ _SP0+α _RT0 _SP0- _RT0,&if a=0 cases, (31) where a is defined in (15). Proof. The approach to proving this result is to consider symmetric profiles (γ,…,γ)(γ,…,γ), and leverage the best-response function from Lemma 4.3 by solving the equation γ=BR((M−1)γ)γ=BR((M-1)γ) under each of its conditions (23), (24). Case 1: In order for a symmetric profile (γ,…,γ)(γ,…,γ) to be an equilibrium and satisfy C((M−1)γ;δSP0)≤δRT0≤δTR1δPS1δSP0C((M-1)γ; _SP0)≤ _RT0≤ _TR1 _PS1 _SP0, it must hold from Lemma 4.3 that γ=θ−γ¯−i=θ−(M−1)γ=θ- γ_-i=θ-(M-1)γ, which yields the unique equilibrium (30). Case 2: In order for a symmetric profile γ to be an equilibrium and satisfy max(M−1)γ−θα−(M−1)γ,−δTR1≤δRT0<C((M−1)γ;δSP0) \ (M-1)γ-θα-(M-1)γ,- _TR1\≤ _RT0<C((M-1)γ; _SP0) (with a≠0a≠ 0), the equation γ=F((M−1)γ)γ=F((M-1)γ) must hold. Its solutions satisfy the quadratic equation Q(γ):=M2γ2+K1γ+K0=0 Q(γ)=M^2γ^2+K_1γ+K_0=0 (32) where K1 K_1 :=α+θa(2M∂g∂n(α+θ)−M−1bY) = α+θa (2M ∂ g∂ n( α+θ)- M-1bY ) (33) K0 K_0 :=(α+θ)2a2∂g∂n(α+θ)(∂g∂n(α+θ)−Yb). = (α+θ)^2a^2 ∂ g∂ n( α+θ) ( ∂ g∂ n( α+θ)- Yb ). The solutions to this equation are given by (29). Observe that the radicand is positive, due to ∂g∂n(⋅)<0 ∂ g∂ n(·)<0, and b<0b<0 because δRT0<C((M−1)γ;δSP0)<δSP0 _RT0<C((M-1)γ; _SP0)< _SP0. This means that the solutions are real-valued. We can also verify that when a<0a<0, we have K1>0K_1>0 and K0<0K_0<0, and when a>0a>0, we have K1<0K_1<0 and K0>0K_0>0. In the case that a<0a<0, it holds that Q(0)<0Q(0)<0 and Q′(0)>0Q (0)>0, which implies there is a unique positive solution given by E+E_+. In the case that a>0a>0, it holds that Q(0)>0Q(0)>0 and Q′(0)<0Q (0)<0, which implies that there are two positive solutions. However, only the solution E−E_- provides a symmetric profile that is feasible in the restricted strategy sets of all players, i.e. (E−,…,E−)∈(E_-,…,E_-) . It holds that ME+<θME_+<θ (identical argument applies for E−E_-), since otherwise we would be in Case 1. It then follows that C((M−1)E+;δSP0)>C(M−1Mθ;δSP0)C((M-1)E_+; _SP0)>C( M-1Mθ; _SP0), because C(⋅;δSP0)C(·; _SP0) is a decreasing function in its first argument. We are then able to characterize the equilibria in the first two entries of (31). When a=0a=0, the equation γ=12(θδSP0+αδRT0δSP0−δRT0−(M−1)γ)γ= 12 ( θ _SP0+α _RT0 _SP0- _RT0-(M-1)γ ) must hold, which yields the equilibrium in the last entry of (31). ∎ In the symmetric equilibrium α∗α^* of ℰM∗(δSP0,δRT0,α,θ)E^*_M( _SP0, _RT0,α,θ), the players choose positive extraction rates, though never large enough to destroy the resource. The equilibrium characterization in Theorem 4.1 has two primary regimes with respect to the responsible policy (δSP0,δRT0)( _SP0, _RT0). In the case of (30) where mutual cooperation is highly incentivized (high δRT0 _RT0), the total extraction Mα∗Mα^* reaches its upper limit θ. Here, we note that the boundary C(M−1Mθ;δSP0)C( M-1Mθ; _SP0) is monotonically decreasing in M and reaches 0 in the limit. In the case of (31), mutual cooperation is not as highly incentivized. In this regime, the total extraction rate is less than the upper limit of αδRT0+θδSP0δSP0−δRT0 α _RT0+θ _SP0 _SP0- _RT0. In the next section, we investigate the impact of equilibrium behavior on the system’s resource level, the agents’ overall utilities, and the total extraction rate. (a) (b) (c) (d) Figure 3: (a) Resource level RM∗R^*_M under the symmetric equilibrium (Theorem 4.1) over the responsible policy space for M=6M=6 greedy populations, with parameters θ=1.0,α=0.40,δTR1=2.1,δPS1=2.0θ=1.0,\ α=0.40,\ _TR1=2.1,\ _PS1=2.0. The red dashed line indicates the curve C(0;δSP0)C(0; _SP0). Figures (b), (c), and (d) show the resource levels RM∗R^*_M vs. the number of greedy populations M for selected parameter values (shown as dots in subplot (a)). In (b), δRT0=0.8 _RT0=0.8 is above C(0;δSP0)C(0; _SP0), meaning that α∗=θ/Mα^*=θ/M for all M≥1M≥ 1, leading to a constant value for RM∗R^*_M. In (c), δRT0=0.2>0 _RT0=0.2>0 is below C(0;δSP0)C(0; _SP0). RM∗R^*_M decreases initially then stabilizes at a positive constant (Proposition 5.1). In (d) δRT0=−1.0<0 _RT0=-1.0<0. RM∗R^*_M monotonically decreases toward zero, indicating complete resource depletion. V Impact of equilibrium extraction We are interested in assessing the quality of the symmetric equilibrium α∗α^* characterized in Theorem 4.1. We will examine equilibrium outcomes as the number of greedy populations becomes large. In particular, we are interested in determining in which cases the common resource could be preserved even with a growing number of greedy populations. Let us denote αM∗ _M^* as the equilibrium (individual) extraction rate of any of the M≥1M≥ 1 greedy populations. The result below characterizes the limiting values of the equilibrium total consumption rate, α¯M∗:=MαM∗ α^*_M:=M _M^*, and the equilibrium resource level, RM∗:=R(α¯M∗)R^*_M:=R( α^*_M). Proposition 5.1. Consider any game ℰM∗(δSP0,δRT0,α,θ)E^*_M( _SP0, _RT0,α,θ). If 0<δRT0≤δTR1δPS1δSP00< _RT0≤ _TR1 _PS1 _SP0, then 1. α¯∞∗:=limM→∞α¯M∗=θ α^*_∞:= _M→∞ α^*_M=θ. 2. R∞∗:=limM→∞RM∗=δRT0δRT0+δTR1>0R^*_∞:= _M→∞R^*_M= _RT0 _RT0+ _TR1>0. If max−θαδSP0,−δTR1≤δRT0≤0 \ -θα _SP0,- _TR1\≤ _RT0≤ 0, then 3. α¯∞∗=αδRT0+θδSP0δSP0−δRT0∈(0,θ) α^*_∞= α _RT0+θ _SP0 _SP0- _RT0∈(0,θ). 4. R∞∗=0R^*_∞=0. Proof. In the limit of large M, C(M−1Mθ;δSP0)→C(θ;δSP0)=0C( M-1Mθ; _SP0)→ C(θ; _SP0)=0. Thus, if δRT0>0 _RT0>0, the equilibrium consumption α∗=θMα^*= θM for sufficiently large M (30). We therefore obtain the limits for items 1 and 2. Now, suppose max−θαδSP0,−δTR1≤δRT0≤0 \ -θα _SP0,- _TR1\≤ _RT0≤ 0, and suppose a<0a<0. By (31), we have α¯∗(M)=M⋅E+ α^*(M)=M· E_+. Applying L’Hopital’s rule, we obtain α¯∞∗=−α+θa(∂g∂n(α+θ)−Y2b)+(α+θ)Y2ab α^*_∞=- α+θa( ∂ g∂ n( α+θ)- Y2b)+ (α+θ)Y2ab. This expression, after substituting Y=bc−adY=bc-ad (18), becomes αδRT0+θδSP0δSP0−δRT0 α _RT0+θ _SP0 _SP0- _RT0, which is the total consumption capacity in this regime. Similar calculations can be done when considering a≤0a≤ 0, by applying the other cases of (31). This yields item 3. To calculate R∞∗=−g(α+α¯∞∗α+θ,0)∂g∂n(α+α¯∞∗α+θ)R^*_∞=- g( α+ α^*_∞α+θ,0) ∂ g∂ n( α+ α^*_∞α+θ), we simply observe that g(α+α¯∞∗α+θ,0)=bδSP0δSP0−δRT0+δSP0=0g( α+ α^*_∞α+θ,0)=b _SP0 _SP0- _RT0+ _SP0=0 (after substituting b from (15)). This yields item 4. ∎ The significant observation in Proposition 5.1 is that the responsible population’s policy (δSP0,δRT0)( _SP0, _RT0) will determine whether the common resource can be sustained, or will become depleted with a growing number of greedy populations. The multi-population system averts a tragedy of the commons in the regime 0<δRT0≤δTR1δPS1δSP00< _RT0≤ _TR1 _PS1 _SP0, and induces a tragedy of the commons in the regime max−θαδSP0,−δTR1≤δRT0≤0 \ -θα _SP0,- _TR1\≤ _RT0≤ 0. In both regions, each greedy population’s individual utility, α∗RM∗α^*R_M^*, diminishes to zero as M increases. Moreover, the total equilibrium consumption α¯M∗ α_M^* approaches the resource’s capacity (13). Figure 3 provides numerical illustrations of the equilibrium resource levels as they depend on the responsible policy (δSP0,δRT0)( _SP0, _RT0) and the number of populations M. Here, we observe three distinct regimes. For large values of δRT0 _RT0, the total equilibrium consumption is the upper limit θ for any number of populations M, and thus the resource level remains a positive constant value (no tragedy, see Figure 3(b)). For intermediate values of δRT0 _RT0, the equilibrium resource level is decreasing in M, but it settles to a positive constant value (no tragedy, see Figure 3(c)). For low values of δRT0 _RT0, the equilibrium resource level is decreasing in M and it converges to zero (a tragedy, see Figure 3(d)). VI Conclusion In this paper, we investigate a resource extraction game whose players are high-level decision-makers who can determine the consumption rates of the local populations they represent. The players seek to extract as much of the resource for their populations as possible, however, higher total extraction degrades the resource. Our primary results establish a unique symmetric Nash equilibrium. We then identified conditions for which the equilibrium extraction behavior can lead to the destruction of the resource as the number of players increases, a scenario referred to as a tragedy of the commons. The hierarchical extraction game studied here captures real-world settings such as multiple countries, firms, or regional authorities extracting from a shared environmental resource (e.g., fisheries, groundwater basins, or atmospheric carbon sinks), where each decision-maker controls aggregate consumption behavior within its jurisdiction while collectively impacting a common resource. This work may be extended along a few directions. The uniqueness of the symmetric equilibrium could be rigorously established using the theory of M-player concave games. Also, we seek to characterize equilibria in scenarios where populations have different masses. An important direction for future work is to relax the symmetry assumptions by incorporating heterogeneity across populations or considering asymptotic regimes with many players, thereby improving the model’s relevance to large-scale real-world systems. References [1] A. Aghaei, F. Cai, and T. Wu (2025) Game-theoretic analysis of policy impacts in competition between reverse supply chains involving traditional and e-channels. Smart Cities 8 (1). External Links: ISSN 2624-6511, Document Cited by: §I. [2] M. R. Arefin and J. Tanimoto (2021) Imitation and aspiration dynamics bring different evolutionary outcomes in feedback-evolving games. Proceedings of the Royal Society A 477 (2251), p. 20210240. Cited by: §I. [3] K. Betz, F. Fu, and N. Masuda (2024) Evolutionary game dynamics with environmental feedback in a network with two communities. Bulletin of Mathematical Biology 86 (7), p. 84. Cited by: §I. [4] J. Certório, R. J. La, and N. C. Martins (2023) Epidemic population games for policy design: two populations with viral reservoir case study. In 2023 62nd IEEE Conference on Decision and Control (CDC), p. 7667–7674. Cited by: §I. [5] Y. Chen, N. C. Martins, and M. Arcak (2025) Hierarchical decision-making in population games. arXiv preprint arXiv:2509.05808. Cited by: §I. [6] M. Datar, E. Altman, and H. Le Cadre (2022) Strategic resource pricing and allocation in a 5g network slicing stackelberg game. IEEE Transactions on Network and Service Management 20 (1), p. 502–520. Cited by: §I. [7] G. Dayanikli and M. Lauriere (2024) Multi-population mean field games with multiple major players: application to carbon emission regulations. In 2024 American Control Conference (ACC), p. 5075–5081. Cited by: §I. [8] L. Gong, W. Yao, J. Gao, and M. Cao (2022) Limit cycles analysis and control of evolutionary game dynamics with environmental feedback. Automatica 145, p. 110536. Cited by: §I. [9] A. Govaert, L. Zino, and E. Tegling (2022) Population games on dynamic community networks. IEEE Control Systems Letters 6, p. 2695–2700. Cited by: §I. [10] Y. Kawano, L. Gong, B. D. Anderson, and M. Cao (2018) Evolutionary dynamics of two communities under environmental feedback. IEEE Control Systems Letters 3 (2), p. 254–259. Cited by: §I. [11] K. Paarporn and J. Nelson (2024) Two competing populations with a common environmental resource. In 2024 IEEE 63rd Conference on Decision and Control (CDC), p. 3197–3202. External Links: Document Cited by: The Tragedy of the Commons in Multi-Population Resource Games †thanks: , §I. [12] K. Paarporn (2023) Non-myopic agents can stabilize cooperation in feedback-evolving games. In 2023 59th Annual Allerton Conference on Communication, Control, and Computing (Allerton), p. 1–7. Cited by: §I. [13] K. Paarporn (2024) The madness of people: rational learning in feedback-evolving games. In 2024 European Control Conference (ECC), p. 311–316. Cited by: §I. [14] C. M. Saad-Roy and A. Traulsen (2023) Dynamics in a behavioral–epidemiological model for individual adherence to a nonpharmaceutical intervention. Proceedings of the National Academy of Sciences 120 (44), p. e2311584120. Cited by: §I. [15] W. H. Sandholm (2010) Population games and evolutionary dynamics. MIT Press. Cited by: §I. [16] A. Satapathi, N. K. Dhar, A. R. Hota, and V. Srivastava (2023) Coupled evolutionary behavioral and disease dynamics under reinfection risk. IEEE Transactions on Control of Network Systems 11 (2), p. 795–807. Cited by: §I. [17] L. Stella and D. Bauso (2023) The impact of irrational behaviors in the optional prisoner’s dilemma with game-environment feedback. International Journal of Robust and Nonlinear Control 33 (9), p. 5145–5158. Cited by: §I. [18] A. R. Tilman, J. B. Plotkin, and E. Akçay (2020) Evolutionary games with environmental feedbacks. Nature communications 11 (1), p. 915. Cited by: §I. [19] X. Wang, Z. Zheng, and F. Fu (2020) Steering eco-evolutionary game dynamics with manifold control. Proceedings of the Royal Society A 476 (2233), p. 20190643. Cited by: §I. [20] J. S. Weitz, C. Eksin, K. Paarporn, S. P. Brown, and W. C. Ratcliff (2016) An oscillating tragedy of the commons in replicator dynamics with game-environment feedback. Proceedings of the National Academy of Sciences 113 (47), p. E7518–E7525. Cited by: §I, Theorem 2.1, §I. [21] T. Zhang, H. Gupta, K. Suprabhat, and L. Stella (2023) A multi-agent reinforcement learning approach to promote cooperation in evolutionary games on networks with environmental feedback. In 2023 62nd IEEE Conference on Decision and Control (CDC), p. 2196–2201. Cited by: §I.