Paper deep dive
Social welfare optimisation under institutional reward and punishment
Van An Nguyen, Vuong Khang Huynh, Huu Loi Bui, Hai Anh Ha, Quang Dung Le, Tan Dat Nguyen, Ngoc Ngu Nguyen, Zhao Song, Manh Hong Duong, Le Hong Trang, The Anh Han
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/8/2026, 11:04:11 PM
Summary
This paper develops a welfare-centric framework for institutional incentives in finite, well-mixed populations playing social dilemmas (Donation Game and Public Goods Game). It contrasts traditional approaches that minimize cost and maximize cooperation frequency with a framework that maximizes total population payoff net of institutional expenditure. The authors derive explicit expressions for expected social welfare under reward and punishment mechanisms, analyzing the impact of incentive efficiency and selection intensity. Key findings include identifying parameter regimes with single optima or non-monotonic phase transitions, proving optimal incentives are either zero or concentrated around closed-form targets, and establishing conditions where rewards systematically outperform punishments in maximizing social welfare. The study reveals a systematic gap between cost/cooperation-focused optimization and true welfare maximization.
Entities (10)
Relation Signals (10)
Social Welfare → isoptimizedby → Institutional Punishment
confidence 95% · We develop a welfare-centric framework for institutional incentives... considering both rewards for cooperators and punishments for defectors.
Social Welfare → isoptimizedby → Institutional Reward
confidence 95% · We develop a welfare-centric framework for institutional incentives... considering both rewards for cooperators and punishments for defectors.
Institutional Reward → appliesto → Donation Game
confidence 90% · We adopt the well-established social dilemmas games, namely the Donation Game and Public Goods Game
Institutional Punishment → appliesto → Public Goods Game
confidence 90% · We adopt the well-established social dilemmas games, namely the Donation Game and Public Goods Game
Social Welfare → dependson → Selection Intensity
confidence 90% · characterise how it depends on incentive efficiency and selection intensity.
Social Welfare → dependson → Incentive Efficiency
confidence 90% · characterise how it depends on incentive efficiency and selection intensity.
Social Welfare → divergesfrom → Cooperation Frequency
confidence 90% · reveal a systematic gap between incentives optimised for cost or cooperation frequency and those that maximise welfare.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Institutional incentives are widely used to promote cooperation among autonomous, self-regarding agents, from human societies to multi-agent and AI systems. Existing work typically treats incentive design as a bi-objective problem: minimise institutional cost while achieving a high long-run frequency of cooperation. Whether such schemes also maximise social welfare - total population payoff net of institutional expenditure - has remained largely unexplored. We develop a welfare-centric framework for institutional incentives in finite, well-mixed populations playing a social dilemma (Donation Game and Public Goods Game), considering both rewards for cooperators and punishments for defectors. For each mechanism, we derive explicit expressions for expected social welfare and characterise how it depends on incentive efficiency and selection intensity. Analytically, we identify parameter regimes where social welfare has a single optimal incentive level and regimes with qualitative phase transitions, in which welfare becomes non-monotonic with multiple local optima. We prove that any welfare-maximising incentive is either zero or concentrated around a simple closed-form target, and we provide an efficient algorithm to compute these optima. Comparing reward and punishment, we further derive close-formed conditions under which reward outperform punishment in terms of social welfare for any given budget. Overall, our results reveal a systematic gap between incentives optimised for cost or cooperation frequency and those that maximise welfare.
Tags
Links
- Source: https://arxiv.org/abs/2605.31330v1
- Canonical: https://arxiv.org/abs/2605.31330v1
PDF not stored locally. Use the link above to view on the source site.
Full Text
105,896 characters extracted from source content.
Expand or collapse full text
Social welfare optimisation under institutional reward and punishment Van An Nguyen Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT), Vietnam Vietnam National University - Ho Chi Minh City (VNU-HCM), Vietnam Vuong Khang Huynh Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT), Vietnam Vietnam National University - Ho Chi Minh City (VNU-HCM), Vietnam Huu Loi Bui Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT), Vietnam Vietnam National University - Ho Chi Minh City (VNU-HCM), Vietnam Hai Anh Ha Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT), Vietnam Vietnam National University - Ho Chi Minh City (VNU-HCM), Vietnam Quang Dung Le Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT), Vietnam Vietnam National University - Ho Chi Minh City (VNU-HCM), Vietnam Tan Dat Nguyen Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT), Vietnam Vietnam National University - Ho Chi Minh City (VNU-HCM), Vietnam Ngoc Ngu Nguyen Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT), Vietnam Vietnam National University - Ho Chi Minh City (VNU-HCM), Vietnam Zhao Song School of Computing, Engineering and Digital Technologies, Teesside University, United Kingdom Manh Hong Duong School of Mathematics, University of Birmingham, Birmingham, United Kingdom Le Hong Trang Faculty of Computer Science and Engineering, Ho Chi Minh City University of Technology (HCMUT), Vietnam Vietnam National University - Ho Chi Minh City (VNU-HCM), Vietnam The Anh Han School of Computing, Engineering and Digital Technologies, Teesside University, United Kingdom Abstract Institutional incentives are widely used to promote cooperation among autonomous, self-regarding agents, from human societies to multi-agent and AI systems. Existing work typically treats incentive design as a bi-objective problem: minimise institutional cost while achieving a high long-run frequency of cooperation. Whether such schemes also maximise social welfare—total population payoff net of institutional expenditure—has remained largely unexplored. We develop a welfare-centric framework for institutional incentives in finite, well-mixed populations playing a social dilemma (Donation Game and Public Goods Game), considering both rewards for cooperators and punishments for defectors. For each mechanism, we derive explicit expressions for expected social welfare and characterise how it depends on incentive efficiency and selection intensity. Analytically, we identify parameter regimes where social welfare has a single optimal incentive level and regimes with qualitative phase transitions, in which welfare becomes non-monotonic with multiple local optima. We prove that any welfare-maximising incentive is either zero or concentrated around a simple closed-form target, and we provide an efficient algorithm to compute these optima. Comparing reward and punishment, we further derive close-formed conditions under which reward outperform punishment in terms of social welfare for any given budget. Overall, our results reveal a systematic gap between incentives optimised for cost or cooperation frequency and those that maximise welfare. Contents 1 Introduction 2 Models and Methods 2.1 Evolutionary game models 2.2 Social dilemmas: Donation and Public Goods games 2.2.1 Donation Game (DG) 2.2.2 Public Goods Game (PGG) 2.3 Cost optimisation under institutional incentives 2.4 Social welfare optimisation under institutional incentives 2.4.1 Institutional reward 2.4.2 Institutional punishment 3 Main results 3.1 Reward 3.2 Institutional punishment vs institutional reward 4 Numerical analyses and validation 4.1 Impact of varying incentive efficiency a and selection intensity β 4.2 Linearity under neutral and strong selection limits 4.3 Convergence of the optimal incentive 4.4 Social welfare dynamics under reward vs punishment 4.5 Social welfare vs institutional cost optimisation: optimal incentives 5 Conclusions and Future Work References A Proof of the main results: reward A.1 Zero Sum Transfer (a=1a=1) A.2 For a>1a>1 A.3 For a<1a<1 A.4 Influence of Selection Intensity (β) B Proof of the main results: punishment B.1 Social welfare in institutional punishment B.2 Algorithm 1 for punishment B.3 Comparison between reward and punishment B.4 Additional numerical simulations 1 Introduction The evolution of cooperation has long been a central puzzle in evolutionary biology, the social sciences, and multi-agent systems [43, 24, 35, 49, 36]. In many strategic settings, such as one-shot interactions in the Prisoner’s Dilemma and the Public Goods Game [2, 3, 9], classical evolutionary theory predicts that selection on individual fitness typically favours selfish behaviour [43, 27]. Yet, cooperative behaviour remains widespread in both humans and other animals [13, 33]. This apparent tension between individual and collective benefits has motivated extensive research into mechanisms that enable, stabilise, and shape cooperation in social dilemmas [36, 21, 46, 53, 30]. To address this puzzle [36, 39], numerous mechanisms have been proposed, including spatial structure [46, 47], kin and group selection [18, 48], direct and indirect reciprocity [32, 53], and institutional incentives [14, 40, 5, 45]. A particularly important line of work studies external institutional influence, where a central authority invests resources to steer populations towards cooperation [5, 10, 42]. For example, an external institution, such as an international body or a local authority, can conditionally reward cooperative individuals or punish defective ones based either on global statistics or local spatial neighbourhood information [7, 20, 50, 41, 4, 51, 45]. In well-mixed populations, the analysis of these institutional interventions has traditionally been framed as a bi-objective optimisation problem [23, 10]. In this framework, a decision-maker conditionally rewards cooperators to guarantee a desired cooperation level while simultaneously minimising the interference cost. Subsequent work has developed rigorous stochastic analyses of such institutional incentives, characterising optimal schemes across different selection intensities and revealing sharp phase transitions in cost efficiency [12, 11]. Consequently, prior models predominantly evaluate the success of external interventions through these two specific performance criteria: maximising the frequency of cooperation and minimising the institutional investment [50, 17]. Despite these analytical advances, existing models evaluating external interventions largely overlook a crucial holistic metric: social welfare. Social welfare captures the system-level impact of these interventions, broadly defined as the total population payoff minus the external institutional investment [19]. From a societal perspective, cooperation is valuable only insofar as it results in net benefits once the costs of promoting it are taken into account [29, 34, 28, 44]. Focusing solely on the dual metrics of cooperation levels or budgetary savings risks endorsing interventions that are formally successful by those measures yet ultimately detrimental to overall welfare [21]. Addressing this gap is essential not only for theoretical completeness but also for practical relevance, especially in contexts such as public policy, distributed systems, and resource allocation [35, 34, 21]. Therefore, in this paper, we analytically study the optimisation of social welfare under costly institutional incentives in well-mixed populations. We examine whether schemes that minimise cost and maximise cooperation also maximise social welfare, and whether welfare‑maximisation leads to smaller, larger, or qualitatively different investments than cost‑focused approaches.. We explicitly embed social welfare into the analytical framework of institutional incentives in finite well-mixed populations. This enables a systematic comparison between cost-based and welfare-based optima and reveals when they coincide and when they diverge, including cases where maximising cooperation or minimising cost leads away from welfare-maximising policies. We adopt the well-established social dilemmas games, namely the Donation Game and the Public Goods Game [43], studying how social welfare optimisation responds to variations in institutional incentives over a broad parameter range. By treating social welfare as the central optimisation objective, this work provides a more realistic and policy-relevant framework for understanding how cooperative behaviour should be engineered in biological, social, and artificial systems [21]. Organisation of the paper The rest of the paper is organized as follows. In Section 2, we introduce the evolutionary game models and mathematically formulate the problem of optimising the social welfare for both institutional reward and punishment. Section 3 presents the main analytical results while Section 4 provides numerical findings and validations. Further discussion and outlook is given in Section 5. Detailed proofs of the main analytical results are provided in Appendix, namely Appendix A (for reward) and Appendix B (for punishment). Additional numerical simulations with different parameters’ values to show the robustness of our results are also provided in Appendix B. 2 Models and Methods This section describes the evolutionary game model used to formulate the bi-objective optimisation of institutional incentives [12, 23]. We then derive the social welfare functions for both institutional reward and punishment for a well-mixed population setting. 2.1 Evolutionary game models We consider a finite, well-mixed population of N players, who interact with each other using a cooperation dilemma, namely the Donation Game or Public Goods Game. The population evolves according to the Fermi strategy update rule [41]: a player X with fitness fXf_X adopts the strategy of another player Y with fitness fYf_Y with probability PX,Y=(1+e−β(fY−fX))−1,P_X,Y= (1+e^-β(f_Y-f_X) )^-1, where β>0β>0 denotes the selection intensity. The population dynamics are modelled as an absorbing Markov chain over the state space S0,S1,…,SN\S_0,S_1,…,S_N\, where SiS_i represents the state with i cooperators. The homogeneous states S0S_0 and SNS_N are absorbing, while S1,…,SN−1S_1,…,S_N-1 are transient. Let ΠC(i) and ΠD(i) _C(i) and _D(i) be the expected payoffs of a cooperative player (C-player) and a defective player (D-player) in state SiS_i of the population, respectively. Let U=uiji,j=1N−1U=\u_ij\_i,j=1^N-1 be the transition matrix among transient states. In the absence of institutional incentives, for 1≤i≤N−11≤ i≤ N-1, the transition probabilities are ui,i±k u_i,i± k =0,k≥2, =0, k≥ 2, (1) ui,i+1 u_i,i+1 =N−iNiN(1+e−β[ΠC(i)−ΠD(i)])−1, = N-iN iN (1+e^-β[ _C(i)- _D(i)] )^-1, ui,i−1 u_i,i-1 =N−iNiN(1+eβ[ΠC(i)−ΠD(i)])−1, = N-iN iN (1+e^β[ _C(i)- _D(i)] )^-1, ui,i u_i,i =1−ui,i+1−ui,i−1. =1-u_i,i+1-u_i,i-1. 2.2 Social dilemmas: Donation and Public Goods games Our analysis will be carried out for cooperation dilemmas in both pairwise and multi-player settings, described below. 2.2.1 Donation Game (DG) The Donation Game is a special case of the Prisoners’ Dilemma [43], where cooperation corresponds to providing the co-player with a benefit b at a personal cost c, with b>cb>c, while defection yields no benefit and incurs no cost. The payoff matrix of the game (for the row player) is given by CDC( b−c−c) Db0. ~&C&D C&b-c&-c D&b&0 . Let πX,Y _X,Y denote the payoff of a player using strategy X∈C,DX∈\C,D\ when interacting with a player using strategy Y∈C,DY∈\C,D\. In a well-mixed population of size N, at state SiS_i where there are i cooperators, the expected payoffs of a C-player and a D-player are given by ΠC(i) _C(i) =(i−1)πC,C+(N−i)πC,DN−1=(i−1)(b−c)+(N−i)(−c)N−1, = (i-1) _C,C+(N-i) _C,DN-1= (i-1)(b-c)+(N-i)(-c)N-1, ΠD(i) _D(i) =iπD,C+(N−i−1)πD,DN−1=ibN−1. = i _D,C+(N-i-1) _D,DN-1= ibN-1. Therefore, the payoff difference between cooperation and defection is δ=ΠC(i)−ΠD(i)=−(c+bN−1),δ= _C(i)- _D(i)=- (c+ bN-1 ), which is negative and independent of the population state SiS_i, in accordance with the general assumption introduced earlier. 2.2.2 Public Goods Game (PGG) In the Public Goods Game, individuals interact in groups of size n [25]. Each player can either cooperate by contributing an amount c>0c>0 to a common pool, or defect by contributing nothing. The total contribution within a group is multiplied by an enhancement factor r, with 1<r<n1<r<n, and the resulting amount is equally shared among all group members, independently of their strategies. Since defectors benefit from the public good without paying the cost, the game constitutes a social dilemma. In a well-mixed population of size N, at state SiS_i where there are i cooperators, groups are formed by multivariate hypergeometric sampling. Hence, the expected payoffs of a C-player and a D-player are given by ΠC(i) _C(i) =∑k=0n−1(i−1k)(N−in−1−k)(N−1n−1)((k+1)rcn−c)=rcn(1+(i−1)n−1N−1)−c, =Σ^n-1_k=0 i-1k N-i\,n-1-k\, N-1\,n-1\, ( (k+1)rcn-c )= rcn (1+(i-1) n-1N-1 )-c, ΠD(i) _D(i) =∑k=0n−1(ik)(N−1−in−1−k)(N−1n−1)krcn=rc(n−1)n(N−1)i. =Σ^n-1_k=0 ik N-1-i\,n-1-k\, N-1\,n-1\, krcn= rc(n-1)n(N-1)\,i. Therefore, the payoff difference between cooperation and defection is δ=ΠC(i)−ΠD(i)=−c(1−r(N−n)n(N−1)),δ= _C(i)- _D(i)=-c (1- r(N-n)n(N-1) ), which is negative and independent of the population state SiS_i. 2.3 Cost optimisation under institutional incentives To reward a cooperator (respectively, punish a defector), the institution has to spend an amount θ (per-capita incentive), such that the payoff of the targeted individual increases by aθaθ (respectively, a decrease by a^θ aθ), where a and a a denote the efficiencies of reward and punishment, respectively. We derive the expected cost of providing institutional incentives [22, 12]. Under institutional incentives, for 1≤i≤N−11≤ i≤ N-1, the transition probabilities in (1) are modified as follows for reward (and similarly for punishment): ui,i±k u_i,i± k =0,k≥2, =0, k≥ 2, (2) ui,i+1 u_i,i+1 =N−iNiN(1+e−β[ΠC(i)−ΠD(i)+aθ])−1, = N-iN iN (1+e^-β[ _C(i)- _D(i)+aθ] )^-1, ui,i−1 u_i,i-1 =N−iNiN(1+eβ[ΠC(i)−ΠD(i)+aθ])−1, = N-iN iN (1+e^β[ _C(i)- _D(i)+aθ] )^-1, ui,i u_i,i =1−ui,i+1−ui,i−1. =1-u_i,i+1-u_i,i-1. Let =(I−U)−1=(nik)i,k=1N−1N=(I-U)^-1=(n_ik)_i,k=1^N-1 denote the fundamental matrix of this chain. The entry nikn_ik gives the expected number of visits to state SkS_k when starting from state SiS_i. Since mutants can appear with equal probability in S0S_0 and SNS_N, the expected number of visits to SjS_j is, 12(n1j+nN−1,j) 12(n_1j+n_N-1,j). Thus, the expected costs of using only reward and only punishment are given by Er(θ)=θ2∑j=1N−1(n1j+nN−1,j)j,Ep(θ)=θ2∑j=1N−1(n1j+nN−1,j)(N−j).E_r(θ)= θ2 _j=1^N-1(n_1j+n_N-1,j)j, E_p(θ)= θ2 _j=1^N-1(n_1j+n_N-1,j)(N-j). (3) Now, we derive the cooperation frequency under institutional incentives [22]. Since the population consists of two strategies, the fixation probabilities of a single cooperator in a population of defectors and vice versa are given by ρD,C _D,C =(1+∑i=1N−1∏k=1i1+eβ[ΠC(k)−ΠD(k)+aθ]1+e−β[ΠC(k)−ΠD(k)+aθ])−1, = (1+ _i=1^N-1 _k=1^i 1+e^β[ _C(k)- _D(k)+aθ]1+e^-β[ _C(k)- _D(k)+aθ] )^-1, ρC,D _C,D =(1+∑i=1N−1∏k=1i1+eβ[ΠD(k)−ΠC(k)−aθ]1+e−β[ΠD(k)−ΠC(k)−aθ])−1. = (1+ _i=1^N-1 _k=1^i 1+e^β[ _D(k)- _C(k)-aθ]1+e^-β[ _D(k)- _C(k)-aθ] )^-1. As the stationary frequency of cooperation is given by ρD,CρD,C+ρC,D _D,C _D,C+ _C,D, maximising this frequency is equivalent to maximising maxθ(ρD,CρC,D). _θ ( _D,C _C,D ). (4) This ratio simplifies as ρD,CρC,D _D,C _C,D = = ∏k=1N−1uk,k−1uk,k+1=∏k=1N−11+eβ[ΠC(k)−ΠD(k)+aθ]1+e−β[ΠC(k)−ΠD(k)+aθ] _k=1^N-1 u_k,k-1u_k,k+1= _k=1^N-1 1+e^β[ _C(k)- _D(k)+aθ]1+e^-β[ _C(k)- _D(k)+aθ] (5) = = eβ∑k=1N−1(ΠC(k)−ΠD(k)+aθ) e^β _k=1^N-1( _C(k)- _D(k)+aθ) = = eβ(N−1)(δ+aθ) e^β(N-1)(δ+aθ) Given that a minimal level of population cooperation ω∈[0,1]ω∈[0,1] is required. That is, the following inequality must hold, ρD,CρD,C+ρC,D≥ω _D,C _D,C+ _C,D≥ω. It follows from (5) that θ≥θω=1a(N−1)βlog(ω1−ω)−δ.θ≥ _ω= 1a(N-1)β \! ( ω1-ω )-δ. (6) 2.4 Social welfare optimisation under institutional incentives Below we derive the population social welfare under institutional reward and punishment. 2.4.1 Institutional reward Recall that a∈[0,+∞)a∈[0,+∞) represents the efficiency of the reward mechanism, i.e. a (per-capita) institutional cost of θ yields a payoff increase of aθaθ for a targeted cooperator. The reward mechanism is deemed cost-efficient when a≥1a≥ 1. Our analysis examines the mathematical properties of social welfare under three distinct efficiency regimes: a=1a=1, a<1a<1, and a>1a>1. We begin by deriving SW(θ)SW(θ), the expected total social welfare across all population states. For a specific state SiS_i containing i cooperators, social welfare is calculated as the total payoff of all players in that state (denoted by PiP_i) strictly net of the total institutional cost (denoted by θi _i). We define Δ:=ΠD(i)/i := _D(i)/i, that is Δ=bN−1in Donation Game,rc(n−1)n(N−1)in Public Goods Game. = cases bN-1 Donation Game,\\ rc(n-1)n(N-1) Public Goods Game. cases Note that both δ and Δ are independent of the states of the population. We obtain, Pi=i[ΠC(i)+aθ]+(N−i)ΠD(i)P_i=i [ _C(i)+aθ ]+(N-i)\, _D(i). As θi=iθ _i=iθ, the population social welfare in state SiS_i is: SWi(θ) SW_i(θ) =Pi−θi =P_i- _i =i[ΠC(i)+aθ]+(N−i)ΠD(i)−iθ =i [ _C(i)+aθ ]+(N-i)\, _D(i)-iθ =iΠC(i)+(N−i)ΠD(i)+iθ(a−1) =i _C(i)+(N-i) _D(i)+iθ(a-1) =i(ΠC(i)−ΠD(i))+NΠD(i)+i(a−1)θ =i ( _C(i)- _D(i) )+N _D(i)+i(a-1)θ =iδ+N(iΔ)+i(a−1)θ =iδ+N(i )+i(a-1)θ =i(δ+NΔ+(a−1)θ). =i (δ+N +(a-1)θ ). To evaluate the expected population-level social welfare, we aggregate SWi(θ)SW_i(θ) over all transient states SiS_i (i=1i=1 to N−1N-1), weighted by the expected visitation rates (n1,i+nN−1,i2 n_1,i+n_N-1,i2) derived from the Fundamental Matrix (N) defined in [12]. SW(θ) SW(θ) =12∑iSWi(θ)(n1,i+nN−1,i) = 12 _iSW_i(θ)(n_1,i+n_N-1,i) =12∑i(δ+NΔ+(a−1)θ)(n1,i+nN−1,i) = 12 _ii (δ+N +(a-1)θ )(n_1,i+n_N-1,i) =12(δ+NΔ+(a−1)θ)∑i(n1,i+nN−1,i) = 12(δ+N +(a-1)θ) _ii(n_1,i+n_N-1,i) =N22f(x)g(x)(δ+NΔ+(a−1)θ), = N^22 f(x)g(x) (δ+N +(a-1)θ ), (7) where x:=β(aθ+δ)x:=β(aθ+δ), f(x)f(x) and g(x)g(x) are two functions defined below [12]: f(x) f(x) =(1+ex)[(1+ex+⋯+e(N−2)x)HN+e(N−1)x∑j=1N−1e−jxj], =(1+e^x) [ (1+e^x+·s+e^(N-2)x )H_N+e^(N-1)x _j=1^N-1 e^-jxj ], (8) g(x) g(x) =1+ex+⋯+e(N−1)x. =1+e^x+·s+e^(N-1)x. (9) The main objective of this paper is to study the mathematical problem of optimising the total expected social welfare (under institutional reward) maxθ>0SW(θ). _θ>0SW(θ). (10) 2.4.2 Institutional punishment For punishment, the institution incurs a cost θ, resulting in a reduction of a^θ aθ to the defector’s payoff. In this case, P^i=iΠC(i)+(N−i)[ΠD(i)−a^θ]. P_i=i\, _C(i)+(N-i) [ _D(i)- aθ ]. Combining with the total institutional cost θ^i=(N−i)θ θ_i=(N-i)θ yields aggregate the social welfare in state SiS_i as: SW^i(θ)=P^i−θ^i=iΠC(i)+(N−i)[ΠD(i)−(1+a^)θ] SW_i(θ)= P_i- θ_i=i\, _C(i)+(N-i) [ _D(i)-(1+ a)θ ] By similar computations as in the reward case, we obtain the following expression for the total expected social welfare in the punishment incentive (see Appendix B.1 for the detailed calculations) SW^(θ)=N22f(x)g(x)(δ+NΔ)−N22f^(x)g(x)(1+a^)θ, SW(θ)= N^22 f(x)g(x) (δ+N )- N^22 f(x)g(x)(1+ a)θ, (11) where f^(x)=(1+ex)[(1+ex+⋯+e(N−2)x)HN+∑j=1N−1e(j−1)xj]. f(x)=(1+e^x) [ (1+e^x+·s+e^(N-2)x )H_N+ _j=1^N-1 e^(j-1)xj ]. The corresponding social welfare optimisation problem (under institutional punishment) is defined as maxθ>0SW^(θ). _θ>0 SW(θ). (12) 3 Main results The aim of this paper is to provide a rigorous analysis of the total expected social welfare SWSW and SW SW and the associated optimisation problems (10) and (12). From a mathematical point of view, these are non-trivial optimisation problems since the objective functions are generally non-convex and depend on several parameters, namely the population and group sizes N and n (which can be arbitrarily large), the payoff entries, the strength of selection as well as the efficiency of the institutional incentives. We obtain both analytical results, analysing qualitative properties of the social welfare functions, and numerical results, offering algorithms to practically compute the optimal institutional cost. 3.1 Reward The first result of our paper is the following theorem, which shows that both the efficiency of the reward mechanism (a) and the strength of selection (β) exert non-trivial effects on the qualitative behaviour of SW(θ)SW(θ). When a=1a=1 (zero-sum transfer), SWSW is increasing and then decreasing for all β. However, when a≠1a≠ 1 (non-zero sum transfer), the monotonicity of SWSW undergoes an intriguing phase-transition phenomenon as a and β vary. We provide analytical formula for computing the critical thresholds and the optimiser in each scenario. Theorem 1. (Behaviour and optimisation of the total expected social welfare) (1) Zero-sum transfer (a=1a=1) When a=1a=1, there exists a unique θ∗>0θ^*>0 such that SW(θ)SW(θ) is increasing on (0,θ∗)(0,θ^*) and decreasing on (θ∗,+∞)(θ^*,+∞). Thus, SWSW has a unique global maximiser at θ=θ∗θ=θ^*. (2) Efficient institutional reward (a>1a>1). We define a threshold: β∗=−F∗>0.β^*=- F^*K>0. (13) (i) For β≤β∗β≤β^*, there exists a threshold θ0>0 _0>0 such that SW(θ)SW(θ) is non-decreasing on (θ0,+∞)( _0,+∞). (i) (behaviour above the threshold value) For β>β∗β>β^*, the number of changes of the sign of dSW(θ)/dθdSW(θ)/dθ is at least two for all N and there exists an N0N_0 such that the number of changes is exactly two for N≤N0N≤ N_0. As a consequence, for N≤N0N≤ N_0, there exists θ1<θ2 _1< _2 such that, for β>β∗β>β^*, SW(θ)SW(θ) is increasing when θ<θ1θ< _1, decreasing when θ1<θ<θ2 _1<θ< _2 and increasing when θ>θ2θ> _2. (3) Inefficient institutional reward (a<1a<1) (i) There exist a threshold a∗∈(0,1)a^*∈(0,1) such that SW(θ)SW(θ) is strictly decreasing on (θ0,+∞)( _0,+∞) when a∗<a<1a^*<a<1. (i) (behaviour under the threshold value) For 0<a<a∗0<a<a^* and β<β∗β<β^*, SW(θ)SW(θ) is non-decreasing on (θ0,+∞)( _0,+∞). Consequently: maxθ≥θ0SW(θ)=SW(θ0) _θ≥ _0SW(θ)=SW( _0) (i) (behaviour above the threshold value) For 0<a<a∗0<a<a^* and β>β∗β>β^*, the number of changes of the sign of dSW(θ)/dθdSW(θ)/dθ is at least two for all N and there exists an N0N_0 such that the number of changes is exactly two for N≤N0N≤ N_0. As a consequence, for N≤N0N≤ N_0, there exist θ1<θ2 _1< _2 such that, for β>β∗β>β^*, SW(θ)SW(θ) is decreasing when θ<θ1θ< _1, increasing when θ1<θ<θ2 _1<θ< _2 and decreasing when θ>θ2θ> _2. Thus, for N≤N0N≤ N_0: maxθ≥θ0SW(θ)=maxSW(θ0),SW(θ2) _θ≥ _0SW(θ)= \SW( _0),SW( _2)\ (iv) Moreover, for sufficiently large β and small θ, SW(θ)SW(θ) is increasing as θ→0+θ→ 0^+ (see proof in Lemma 3 in Appendix). Plots of the qualitative behaviour of the social welfare SWSW as a function of θ for various parameters as well as numerical calculations of the critical thresholds θ∗θ^* and β∗β^* are presented in Figure 1 for Donation Game and Figure 2 for Public Goods Games. Our second result is the following theorem which shows that the maximiser, θ∗θ^*, of the optimisation problem (10) is either zero or located around a specific value, θ∞=−δa _∞=- δa. Theorem 2 (localsation around θ∞ _∞). Let θ⋆θ be a maximiser of the social welfare objective SW(θ)SW(θ) over θ≥0θ≥ 0, and define θ∞=−δa. _∞=- δa. Then either θ⋆=0θ =0, or θ⋆>0θ >0 and |θ⋆−θ∞|=(1aβ). |θ - _∞ |=O\! ( 1aβ ). Based on Theorem 2, we develop the following algorithm (Algorithm 1) to approximate the optimal incentive θ∗θ^*. It is applicable to both the Donation Game and the Public Goods Game described. Input: Game parameters b,cb,c; Parameters N∈ℤ+N ^+, a,β∈ℝ+a,β ^+ (a<1a<1). Search radius r>0r>0 and number of steps Nsteps∈ℤ+N_steps ^+ Output: θ∗θ^* maximising SWSW within valid bounds Step 1: Compute Bound and Starting Point Compute values of δ and Δ\;\;\;\;\;\;Compute values of δ and θlimit←δ+NΔ1−a\;\;\;\;\;\; _limit← δ+N 1-a ; // Define the hard upper limit θstart←min(−δa,θlimit)\;\;\;\;\;\; _start← ( -δa, _limit ); Step 2: Initialize Grid Search θbest←0\;\;\;\;\;\; _best← 0; SWbest←SW(0)\;\;\;\;\;\;SW_best← SW(0); Δθ←2r/Nsteps\;\;\;\;\;\; θ← 2r/N_steps ; // Calculate step size Step 3: Evaluate Interval for k←0k← 0 to NstepsN_steps do θcurr←max(0,(θstart−r)+k⋅Δθ); _curr← (0,( _start-r)+k· θ); if θcurr>θlimit _curr> _limit then break ; // Terminate search if limit is exceeded val←SW(θcurr)val← SW( _curr); if val>SWbestval>SW_best then SWbest←valSW_best← val; θbest←θcurr _best← _curr; Step 4: Return Optimal Incentive return θbest _best; Algorithm 1 Restricted Interval Grid Search with Upper and Lower Bounds Cutoff Theorem 1 has shown the influence of the strength of selection to the total expected social welfare. The following theorem further characterises the asymptotic behaviour of SWSW in the neutral and strong selection limits, namely when β tends to zero and infinity respectively. The convergence of the optimiser θ∗θ^* to θ∞ _∞ is numerically presented in Figure 5. Theorem 3. (Neutral and strong selection limits) SW(θ)SW(θ) exhibits linear behaviour in the limits β→0+β→ 0^+ and β→+∞β→+∞. More precisely, we have limβ→0+SW(θ) _β→ 0^+SW(θ) =N2HN(δ+NΔ+(a−1)θ), =N^2H_N(δ+N +(a-1)θ), limβ→+∞SW(θ) _β→+∞SW(θ) =N2HN(δ+NΔ+(a−1)θ)forθ=−δa,N22(HN+1)(δ+NΔ+(a−1)θ)forθ>−δa,N22(HN+1N−1)(δ+NΔ+(a−1)θ)forθ<−δa. = casesN^2H_N(δ+N +(a-1)θ) θ=- δa,\\ N^22(H_N+1)(δ+N +(a-1)θ) θ>- δa,\\ N^22 (H_N+ 1N-1 )(δ+N +(a-1)θ) θ<- δa. cases In the above formulae, HNH_N denotes the harmonic number HN=∑j=1N−11j.H_N= _j=1^N-1 1j. In Figures 3 and 4 we numerically demonstrate the asymptotic limits of the social welfare described in Theorem 3 respectively for Donation Game and Public Goods Game. 3.2 Institutional punishment vs institutional reward We now consider institutional punishment 111A counterpart of Theorem 1 for the punishment could be obtained, but we omit it and only present Theorem 4 since it provides an algorithm to compute the optimiser in practice.. The following theorem is the counterpart of Theorem 2 for this type of incentive. It shows the localisation property of the optimal of SW(θ)SW(θ) for the punishment. Theorem 4. Let θ⋆θ be a maximiser of the social welfare function SW^(θ) SW(θ) over θ≥0θ≥ 0, and define θ^∞=−δa^. θ_∞=- δ a. Then either θ⋆=0θ =0, or θ⋆>0θ >0 and |θ⋆−θ^∞|=(logβa^β). |θ - θ_∞ |=O\! ( β aβ ). Compared with the convergence rate for institutional reward obtained in Theorem 2, the convergence for the punishment one is slower because of the extra logarithmic factor. The following theorem shows that if the reward efficiency is sufficiently high compared to that of the punishment, then the institution can always adjust the investment cost so that the reward incentive outperforms the punishment one. Theorem 5. Let a and a a be the efficiencies of institutional reward and punishment, respectively. For N≥3N≥ 3, it satisfies that SW^(θ)≤SW(a^θ/a) SW(θ)≤ SW( aθ/a) for all θ≥0θ≥ 0 if and only if a≥a^ηN−1η0+a^(η0+ηN−1),a≥ a _N-1 _0+ a( _0+ _N-1), where η0=1N−1+HNandηN−1=1+HN, _0= 1N-1+H_N _N-1=1+H_N, Specifically, when the two types of institutional incentives are equally efficient, i.e. a=a^a= a, reward leads to a higher level of social welfare than punishment whenever a≥ηN−1−η0ηN−1+η0=N−2N+2(N−1)HN.a≥ _N-1- _0 _N-1+ _0= N-2N+2(N-1)H_N. Comparisons between the reward and punishment incentives are numerically demonstrated in Figure 6. Idea of the proofs The technically detailed proofs of the above theorems are deferred to Appendices A and B. Here we provide the underlying ideas of these proofs. The proofs of Theorems 1-3-5 rely on the analytically explicit formulas of the objective functions, which are the total expected social welfare SWSW and SW SW given in (7)-(11). We are able to derive explicitly these formulas because of the crucial fact mentioned earlier that in Donation Game and Public Goods Game the payoff difference between cooperation and defection is independent of the population states. Although these expressions are still complicated, they enable us to compute the derivative of the objective functions. Finding the roots of the derivative functions then boils down to finding the roots of certain polynomials, see Equations (A.1) and (A5) below, which has been studied in [12]. The qualitative behaviour, especially the monotonic properties, of the objective functions is then followed by a thorough analysis of the sign of the derivative functions. The later step is also the key step in the proofs of Theorems 2-4. By characterising the sign of the derivative functions in appropriate intervals, we are able to provide lower and upper estimates for the location of their roots, which are the optimisers of the corresponding optimisation problems. 4 Numerical analyses and validation This section analyses the behaviour of the expected social welfare (SWSW) with respect to the incentive impact θ, under different incentive regimes and game parameters. To further verify the analytical results (Theorems 1-5) presented above, we focus on: 1. The effect of efficiency parameter a and selection intensity β on the variation of the SWSW curve for the reward case, notably on [θ0,+∞) [ _0,+∞ ). 2. Approximate linearity of SW(θ)SW(θ) under neutral and strong selection limits. 3. Convergence of the optimal value of θ, i.e. θ∗θ^*, with respect to selection intensity β. 4. Comparison between the social welfare functions, i.e. SW(θ)SW(θ) and SW^(θ) SW(θ), under institutional reward and punishment policies. 5. Comparison of the optimal incentive levels θ∗θ^* obtained under social welfare maximisation, institutional cost minimisation, and minimum frequency of cooperation constraints. 4.1 Impact of varying incentive efficiency a and selection intensity β Figure 1: In the Donation Game, depending on the selection intensity, the relationship between social welfare and the institutional incentive transitions is either monotonic behaviour or exhibits a clear extremum. Social welfare SW(θ)SW(θ) as a function of the per-capita institutional cost θ, for reward in the Donation Game (DG). The shape of the social welfare function undergoes qualitative transitions with incentive efficiency (a) and selection intensity (β). As a increases, SWSW changes from predominantly decreasing (a<1a<1), to nearly flat (a=1a=1), and eventually to increasing (a>1a>1). Increasing β further reveals a threshold β∗β^*: below β∗β^*, SWSW is monotonic, whereas above β∗β^* the curve develops additional extrema, indicating a phase transition in welfare-maximising incentive levels. Figure 2: In the Public Goods Game, depending on the selection intensity, the relationship between social welfare and the institutional incentive transitions is either monotonic behaviour or exhibits a clear extremum. Shown are the numerical results for the Public Goods Game (PGG). The figures demonstrate the changes in the overall tendency of the SWSW curve with varying efficiency parameter (from downward-sloping for a<1a<1, to nearly level at a=1a=1 and eventually upward-sloping when a>1a>1), and the behaviour of the curve around threshold β∗β^*. Increasing a progressively reshapes the SWSW curves, changing their overall tendency with respect to θ from decreasing to nearly flat and eventually increasing. This reflects a transition from a dissipative regime (a<1a<1), where incentives reduce the overall population welfare, to an amplifying regime (a>1a>1), where incentives enhance this outcome. Increasing β sharpens the structure of the SWSW curve and changes its behaviour. As shown in Figures 1 and 2, when β is below the threshold defined in Theorem 1 (i.e. β<β∗β<β^*), it results in a monotonic SWSW curve on [θ0,+∞)[ _0,+∞). For β>β∗β>β^*, the local extrema at θ0 _0 can be observed clearly. The SWSW function develops a distinct local extrema beyond θ0 _0, indicating a phase transition. With β≈β∗β≈β^*, it is unclear whether the SWSW will exhibit a phase transition or a monotonic behaviour for all cases. This suggests that the change in the curve’s behaviour as β approaches and surpasses the threshold is gradual and continuous, rather than abrupt. The resulting observations in the behaviour of the SWSW curve with respect to a and β further supports the analytical findings from Theorem 1. 4.2 Linearity under neutral and strong selection limits Figure 3: In the Donation Game, social welfare undergoes a sharp phase transition at a critical incentive threshold under extreme selection intensities. Shown are the numerical results for the Donation Game (DG). The figures illustrate how the SWSW curve approaches a near-linear form under regimes β→0+β→ 0^+ and β→+∞β→+∞, compared with its behaviour at an intermediate selection intensity. Figure 4: In the Public Goods Game, social welfare undergoes a sharp phase transition at a critical incentive threshold under extreme selection intensities. Shown are the numerical results for the Public Goods Game (PGG). The figures illustrate how the SWSW curve approaches a near-linear form under regimes β→0+β→ 0^+ and β→+∞β→+∞, compared with its behaviour at an intermediate selection intensity. The SWSW behaviour observed in Figures 3 and 4 are consistent with Theorem 3. In the limiting regimes β→0+β→ 0^+ and β→+∞β→+∞, SWSW approaches a piecewise-linear profile, with approximately linear behaviour on [0,θ0)[0, _0) and (θ0,+∞)( _0,+∞). However, it exhibits a steep, near-vertical transition around θ0 _0 in several cases. 4.3 Convergence of the optimal incentive Figure 5: The optimal institutional incentive undergoes a sharp phase transition at a critical selection threshold, rapidly converging to its theoretical limit under strong selection. Convergence of the optimal incentive θ∗θ^* to theoretical limits as selection intensity β increases on a logarithmic scale. (Left) Theorem 2: Optimal reward incentive converging to θ∞ _∞. Simulated using the Donation Game (DG) with population size N=100N=100, benefit b=2.0b=2.0, cost c=1.0c=1.0, and a reward transfer a=0.8a=0.8. (Right) Theorem 4: Optimal punishment incentive converging to θ^∞ θ_∞. Simulated using the DG with N=100N=100, benefit b=5.0b=5.0, cost c=0.2c=0.2, and a punishment efficiency a^=0.6 a=0.6. To validate the localisation properties of the optimal incentive θ∗θ^*, we evaluate the distance between the numerical maximiser and the theoretical limits: θ∞ _∞ for reward and θ^∞ θ_∞ for punishment, as the selection intensity β increases. Figure 5 illustrates this relationship for both mechanisms. Notably, the convergence dynamics exhibit a distinct phase transition: under weak selection, the optimal incentive θ∗θ^* remains at 0, but once a critical threshold β∗β^* is exceeded, θ∗θ^* jumps away from zero and rapidly approaches its theoretical limiting value. For institutional reward, Theorem 2 states that the optimal incentive θ∗θ^* satisfies |θ∗−θ∞|=(1aβ)|θ^*- _∞|=O( 1aβ). The left panel of Figure 5 visually confirms this bound for the Donation Game (N=100,b=2.0,c=1.0,a=0.8N=100,b=2.0,c=1.0,a=0.8), demonstrating that as β→+∞β→+∞, the distance between θ∗θ^* and θ∞ _∞ collapses to zero. Similarly, for the punishment mechanism, Theorem 4 establishes the bound |θ∗−θ^∞|=(logβa^β)|θ^*- θ_∞|=O( β aβ). The right panel validates this constraint using a modified parameter configuration (N=100,b=5.0,c=0.2,a^=0.6N=100,b=5.0,c=0.2, a=0.6), showing a matching pattern of convergence. It is critical to note that for the punishment mechanism, achieving a non-trivial optimal incentive (where θ∗>0θ^*>0) is highly sensitive to the game’s payoff structure. In our observations, this typically occurs only when the benefit-to-cost ratio (b/cb/c) is exceptionally large. We utilised a ratio of b/c=25b/c=25 in this simulation specifically to capture this dynamic. Under less extreme conditions, the optimal punishment incentive frequently collapses to zero, underscoring the limitations of punishment compared to reward. Regarding the convergence speeds, the theoretical bounds provide vital context for interpreting the numerical results. Theorem 2 dictates that the reward mechanism converges at a rate of (1/(aβ))O(1/(aβ)), while Theorem 4 establishes a slightly slower theoretical convergence rate for punishment at (logβ/(a^β))O( β/( aβ)). Despite this mathematical distinction, the logarithmic scale of the x-axes in both figures reveals that once the phase transition threshold is crossed, both incentive mechanisms exhibit a rapid collapse toward their respective limits. Consequently, from an applied institutional perspective, both optimal incentives stabilise almost immediately once selection intensity becomes sufficiently strong. 4.4 Social welfare dynamics under reward vs punishment Figure 6: The efficiency threshold dictates the absolute dominance of institutional reward over punishment. Comparison of reward and punishment in the Donation Game (N=100N=100, b=5.0b=5.0, c=0.2c=0.2, a=0.3a=0.3). The left panel (a^=0.45 a=0.45) shows strict dominance of the shifted reward when the theoretical efficiency condition is met (see Theorem 5), while the right panel (a^=0.6 a=0.6) shows punishment outperforming reward when the condition is violated. To properly observe the regime where punishment could theoretically dominate, we utilize the same extreme payoff configuration from our convergence analysis: the Donation Game with a population size N=100N=100, benefit b=5.0b=5.0, and cost c=0.2c=0.2. This very high benefit-to-cost ratio (b/c=25b/c=25) is necessary because, under standard payoff conditions, the reward mechanism almost universally outperforms in terms of social welfare maximisation. In the left panel (a^=0.45 a=0.45), the parameters satisfy the condition established in Theorem 5. Consequently, the shifted reward curve, SW(a^θ/a)SW( aθ/a) (shown in green), acts as a strict upper bound to the punishment curve (SW^(θ) SW(θ), shown in red) across all incentive levels. Not only does reward dominate, but the absolute gap in total social welfare between the two mechanisms remains visible even as the incentive θ increases. Conversely, the right panel illustrates a regime where this efficiency condition is explicitly violated by increasing the punishment efficiency to a^=0.6 a=0.6. Here, the red punishment curve eclipses the green shifted reward curve for intermediate incentive values, demonstrating that the reward dominance fails to hold under these specific parameters. 4.5 Social welfare vs institutional cost optimisation: optimal incentives Figure 7: Efficient rewards can conflict with cost: multi-objective comparison of optimal incentive levels for the DG (reward case, β=10.0β=10.0). The panels show that the social-welfare–maximising incentive often differs substantially from the cost-minimising incentive assuming minimum cooperation targets. The figures illustrate how the incentive values that maximise SW(θ)SW(θ), minimise Er(θ)E_r(θ), and satisfy cooperation-frequency thresholds vary across three reward-efficiency regimes—ineffective reward (a<1a<1), zero-sum transfer (a=1a=1), and effective reward (a>1a>1). Vertical lines mark the optimal incentives and the threshold values for target cooperation frequencies. Figure 8: Efficient rewards can conflict with cost: multi-objective comparison of optimal incentive levels for the PGG (reward case, β=10.0β=10.0). The panels show that the social-welfare–maximising incentive can differ substantially from the cost-minimising incentive. The figures illustrate how the incentive values that maximise SW(θ)SW(θ), minimise Er(θ)E_r(θ), and satisfy cooperation-frequency thresholds vary across three reward-efficiency regimes—ineffective reward (a<1a<1), zero-sum transfer (a=1a=1), and effective reward (a>1a>1). Vertical lines mark the optimal incentives and the threshold values for target cooperation frequencies. This section compares the optimal incentive levels to three inter-related optimisation objectives: social welfare maximisation, institutional cost minimisation, and institutional cost minimisation with a minimum target cooperation frequency. It is crucial to focus on whether the social welfare optimum remains feasible and compatible with the institutional cost, once the frequency of cooperation threshold is imposed. From Figures 7 and 8, for DG and PGG, respectively, we show that the optimal incentives θ∗θ^* for achieving the objectives can be significantly different, for varying game parameters (see additional figures A3-A8 in Appendix for other parameter settings). (i) Institutional cost minimisation under the constraints of the frequency of cooperation. The frequency of cooperation imposes a lower bound θω _ω on the incentive level (see Equation 6). Since Er(θ)E_r(θ) increases with the incentive level in the relevant regimes, institutional cost does not favour incentives above this threshold. Hence, once a target cooperation frequency is fixed, the cost-minimising feasible incentive is naturally located at θω _ω. This shows that the cost objective and the frequency of cooperation should be interpreted jointly: the threshold defines feasibility, while Er(θ)E_r(θ) selects the lowest feasible incentive. (i) Comparison with the social welfare maximisation objective. The central comparison is between the social welfare optimum and the feasible incentive to minimise institutional cost. If θSW∗θ^*_SW lies close to θω _ω, then social welfare maximisation, cost minimisation, and cooperation enforcement are largely compatible. If θSW∗>θωθ^*_SW> _ω, then maximising social welfare requires incentives above the minimum level needed to satisfy the cooperation target, creating a trade-off with institutional cost. Conversely, if θSW∗<θωθ^*_SW< _ω, the frequency of cooperation forces the institution to choose an incentive level beyond the welfare optimum. In the inefficient and zero-transfer regimes, a≤1a≤ 1, the interaction between the objectives depends on the shape of SW(θ)SW(θ). When a≤a∗, SW(θ)a≤ a^*, SW(θ) is non-increasing over the feasible interval, or only admits a local minimiser. As a result, social welfare maximisation and institutional cost minimisation are aligned, since both favour the smallest feasible incentive. When SW(θ)SW(θ) admits an interior local maximiser (a>a∗a>a^*), the welfare optimum may lie above the cooperation threshold, producing a partial trade-off with the institutional cost. However, if the frequency of cooperation threshold exceeds this local maximiser, the feasible solution is instead dictated by the cooperation requirement. In the efficient-transfer regime, a>1a>1, the social welfare objective tends to favour larger incentive levels, whereas institutional cost minimisation favours the smallest feasible incentive. The frequency of cooperation threshold determines the minimum admissible value of θ, but the welfare optimum may lie above this threshold. This creates a stronger conflict between increasing welfare and controlling institutional cost. The neutral and strong selection limits provide simplified benchmark cases for the multi-objective comparison. Under both limits, SW(θ)SW(θ) approaches a linear limiting profile, while the thresholds for cooperation either become extremely large or converge to a common value. These limiting behaviours reduce the complexity of the trade-off among the objectives. 5 Conclusions and Future Work In this work, we extend and generalise the classical framework of institutional incentives by putting social welfare at the centre of the analysis, rather than focusing only on minimising institutional costs or maximising cooperation levels [21, 10, 5, 50]. For institutional reward, we derive an explicit expression for population-level social welfare in finite, well-mixed populations and show that both how efficiently rewards are transferred and how strongly selection acts are crucial for whether rewarding cooperation produces a net societal benefit. We identify parameter regimes where social welfare has a single best incentive level (where it first increases and then decreases as incentives grow), and other regimes where it exhibits more complex phase transitions, with multiple local optima as efficiency and selection change. We also show that any welfare-maximising incentive is either zero or very close to a simple, analytically determined target level. For institutional punishment, we find qualitatively similar localisation properties and derive explicit conditions under which sufficiently efficient rewards always outperform punishment in terms of social welfare for any given budget. Taken together, these results demonstrate that incentive schemes optimised for cost efficiency or cooperation frequency can be markedly suboptimal or even harmful when social welfare is considered. This challenges much of the existing literature on the evolution of cooperation and institutional incentives, which typically uses cooperation levels or institutional spending cost as the main performance criteria [26, 42, 41, 6, 5, 51, 40, 16, 10, 15, 52, 30]. Moreover, our findings provide a welfare-based analytical framework for evaluating mechanisms of cooperation and clarify when higher cooperation genuinely improves collective outcomes [21]. As important implications for institutional design and public policy, our findings underscore the importance of welfare-centric optimisation and offer structural guidance on how to calibrate reward and sanction systems. Our framework suggests that centralised controllers and incentive designers should be judged not just on how much cooperation they induce among agents, but on how well they maximise net social welfare under realistic resource constraints. Several important gaps remain open for future work. Firstly, as our work develops the reward-based and punishment-based social welfare frameworks separately, a natural next step is to investigate adaptive and hybrid combinations [5, 31, 10, 30] with reward and punishment for enhanced social welfare. Secondly, our numerical results suggest clear structural behaviours of social welfare on several key intervals, but a rigorous global characterisation remains open. Addressing this gap would strengthen the proposed framework and provide a more complete understanding of optimal incentive design in finite populations. Thirdly, our current formulation assumes a uniform incentive level across all targeted individuals, whereas a more realistic institutional design would allow incentives to depend on the specific individual, their role, or their local contribution to the population state [37, 8, 7, 1]. The generalisation to a heterogenous incentives could reveal richer optimal policies and better reflect real social systems. Finally, the model currently assumes symmetric mutant emergence from the two absorbing states. Allowing asymmetric mutation probabilities, or more generally state-dependent mutation structures, would make the framework more realistic and may substantially change both the welfare landscape and the resulting optimal intervention strategies. Together, these extensions would help move the present framework from a stylised welfare analysis toward a more general and policy-relevant theory of institutional incentive design in finite populations. Acknowledgements Z.S. and TAH are supported by EPSRC (grant EP/Y00857X/1). MHD was supported by EPSRC grant EP/Y008561/1. T.A.H. acknowledges travel support from the HCMUT-VNUHCM (Adjunct Professorship scheme HCMUT-VNUHCM). Competing interest Authors declare that they have no conflict of interest. References [1] Z. Alalawi, P. Bova, T. Cimpeanu, A. Di Stefano, M. H. Duong, E. F. Domingos, T. A. Han, M. Krellner, N. B. Ogbo, S. T. Powers, et al. (2026) Trust ai regulation? discerning users are vital to build trust and effective ai regulation. Applied Mathematics and Computation 508, p. 129627. Cited by: §5. [2] M. Archetti and I. Scheuring (2012) Game theory of public goods in one-shot social dilemmas without assortment. Journal of theoretical biology 299, p. 9–20. Cited by: §1. [3] R. Axelrod and W. D. Hamilton (1981) The evolution of cooperation. science 211 (4489), p. 1390–1396. Cited by: §1. [4] H. Brandt, C. Hauert, and K. Sigmund (2006) Punishing and abstaining for public goods. Proceedings of the National Academy of Sciences of the United States of America 103 (2), p. 495. Cited by: §1. [5] X. Chen, T. Sasaki, Å. Brännström, and U. Dieckmann (2015) First carrot, then stick: how the adaptive hybridization of incentives promotes cooperation. Journal of The Royal Society Interface 12 (102), p. 20140935. Cited by: §1, §5, §5, §5. [6] T. Cimpeanu, T. A. Han, and F. C. Santos (2019) Exogenous Rewards for Promoting Cooperation in Scale-Free Networks. In ALIFE 2019, p. 316–323. Cited by: §5. [7] T. Cimpeanu, C. Perret, and T. A. Han (2021) Cost-efficient interventions for promoting fairness in the ultimatum game. Knowledge-Based Systems 233, p. 107545. Cited by: §1, §5. [8] T. Cimpeanu, Z. Song, and T. A. Han (2025) The hidden price of cooperation: incentives and social welfare on networks. In Artificial Life Conference Proceedings 37, Vol. 2025, p. 85. Cited by: §5. [9] M. Doebeli and C. Hauert (2005) Models of cooperation based on the prisoner’s dilemma and the snowdrift game. Ecology letters 8 (7), p. 748–766. Cited by: §1. [10] M. H. Duong, C. Durbac, and T. Han (2023) Cost optimisation of hybrid institutional incentives for promoting cooperation in finite populations. Journal of Mathematical Biology 87 (5), p. 77. Cited by: §1, §1, §5, §5, §5. [11] M. H. Duong, C. Durbac, and T. Han (2026) Cost of institutional incentives for promoting cooperation in games and collective risk games. Dynamic Games and Applications, p. 1–34. Cited by: §1. [12] M. H. Duong and T. A. Han (2021-10) Cost efficiency of institutional incentives for promoting cooperation in finite populations. Proceedings of the Royal Society A: Mathematical, Physical and Engineering Sciences 477 (2254), p. 20210568. External Links: ISSN 1364-5021, Document, Link, https://royalsocietypublishing.org/rspa/article-pdf/doi/10.1098/rspa.2021.0568/360899/rspa.2021.0568.pdf Cited by: §A.1, §A.1, §A.2, Appendix A, §B.1, §B.2, §B.2, §1, §2.3, §2.4.1, §2.4.1, §2, §3. [13] E. Fehr and U. Fischbacher (2004) Social norms and human cooperation. Trends in cognitive sciences 8 (4), p. 185–190. Cited by: §1. [14] E. Fehr and S. Gächter (2000) Cooperation and punishment in public goods experiments. American Economic Review 90 (4), p. 980–994. Cited by: §1. [15] J. García and A. Traulsen (2019) Evolution of coordinated punishment to enforce cooperation from an unbiased strategy space. Journal of the Royal Society Interface 16 (156). Cited by: §5. [16] A. R. Góis, F. P. Santos, J. M. Pacheco, and F. C. Santos (2019) Reward and punishment in climate change dilemmas. Scientific reports 9 (1), p. 16193. Cited by: §5. [17] J. Gross, C. Graf, and C. S. Rossetti (2025) The hidden costs of human cooperation. Trends in Cognitive Sciences. Cited by: §1. [18] W. D. Hamilton (1964) The genetical evolution of social behaviour. i. Journal of theoretical biology 7 (1), p. 17–52. Cited by: §1. [19] T. A. Han, M. H. Duong, and M. Perc (2024) Evolutionary mechanisms that promote cooperation may not promote social welfare. J R Soc Interface 21 (220), p. 20240547. External Links: Document Cited by: §1. [20] T. A. Han, S. Lynch, L. Tran-Thanh, and F. C. Santos (2018) Fostering cooperation in structured populations through local and global interference strategies. In Proceedings of the 27th international joint conference on artificial intelligence, p. 289–295. Cited by: §1. [21] T. A. Han, Z. Song, T. Cimpeanu, M. H. Duong, M. Krellner, V. Capraro, and M. Perc (2026) Cooperation versus social welfare. Physics of Life Reviews 56, p. 33–60. External Links: ISSN 1571-0645, Document, Link Cited by: §1, §1, §1, §5, §5. [22] T. A. Han and L. Tran-Thanh (2018) Cost-effective external interference for promoting the evolution of cooperation. Scientific reports 8 (1), p. 15997. Cited by: §2.3, §2.3. [23] T. A. Han and L. Tran-Thanh (2018) Cost-effective external interference for promoting the evolution of cooperation. In Scientific Reports, Cited by: §1, §2. [24] T. A. Han (2022) Emergent behaviours in multi-agent systems with evolutionary game theory. AI Communications 35 (4), p. 327–337. Cited by: §1. [25] C. Hauert, A. Traulsen, H. Brandt, M. A. Nowak, and K. Sigmund (2007) Via freedom to coercion: the emergence of costly punishment. science 316 (5833), p. 1905–1907. Cited by: §2.2.2. [26] C. Hilbe and A. Traulsen (2012) Emergence of responsible sanctions without second order free riders, antisocial punishment or spite. Scientific reports 2 (1), p. 458. Cited by: §5. [27] J. Hofbauer and K. Sigmund (1998) Evolutionary games and population dynamics. Cambridge university press. Cited by: §1. [28] M. Kaneko and K. Nakamura (1979) The nash social welfare function. Econometrica: Journal of the Econometric Society, p. 423–435. Cited by: §1. [29] Ö. Karsu and A. Morton (2015) Inequity averse optimization in operational research. Eur J Oper Res 245 (2), p. 343–359. Cited by: §1. [30] L. Liu, L. Wang, W. Niu, and S. Hua (2026) Dynamic sanctioning mechanism for cooperative multi-agent systems. Expert Systems with Applications 296, p. 128873. Cited by: §1, §5, §5. [31] Y. Liu, L. Wang, R. Guo, S. Hua, L. Liu, L. Zhang, et al. (2025) Evolution of trust in the n-player trust game with transformation incentive mechanism. Journal of the Royal Society Interface 22 (224). Cited by: §5. [32] M. A. Nowak and K. Sigmund (2005) Evolution of indirect reciprocity. Nature 437 (1291-1298). Cited by: §1. [33] M. A. Nowak (2006) Five rules for the evolution of cooperation. science 314 (5805), p. 1560–1563. Cited by: §1. [34] J. A. Nyman (2006) The efficiency of equity. Cumb L Rev 37, p. 461. Cited by: §1. [35] A. Paiva, F. Santos, and F. Santos (2018) Engineering pro-sociality with autonomous agents. In Proceedings of the AAAI conference on artificial intelligence, Vol. 32. Cited by: §1, §1. [36] M. Perc, J. J. Jordan, D. G. Rand, Z. Wang, S. Boccaletti, and A. Szolnoki (2017) Statistical physics of human cooperation. Physics Reports 687, p. 1–51. Cited by: §1, §1. [37] M. Perc and A. Szolnoki (2015) A double-edged sword: benefits and pitfalls of heterogeneous punishment in evolutionary inspection games. Scientific reports 5 (1), p. 11027. Cited by: §5. [38] W. H. Press, S. A. Teukolsky, W. T. Vetterling, and B. P. Flannery (2007) Numerical recipes: the art of scientific computing. Cambridge University Press. Cited by: Appendix A. [39] D. G. Rand and M. A. Nowak (2011) The evolution of antisocial punishment in optional public goods games. Nature communications 2 (1), p. 434. Cited by: §1. [40] T. Sasaki, Å. Brännström, U. Dieckmann, and K. Sigmund (2012) The take-it-or-leave-it option allows small penalties to overcome social dilemmas. Proceedings of the National Academy of Sciences 109 (4), p. 1165–1169. Cited by: §1, §5. [41] K. Sigmund, H. De Silva, A. Traulsen, and C. Hauert (2010) Social learning promotes institutions for governing the commons. Nature 466 (7308), p. 861–863. Cited by: §1, §2.1, §5. [42] K. Sigmund, C. Hauert, and M. A. Nowak (2001) Reward and punishment. Proceedings of the National Academy of Sciences 98 (19), p. 10757–10762. Cited by: §1, §5. [43] K. Sigmund (2010) The calculus of selfishness. Princeton University Press. Cited by: §1, §1, §2.2.1. [44] Z. Song and T. A. Han (2026) Emergence of cooperation and commitment in optional prisoner’s dilemma. Applied Mathematical Modelling 155, p. 116603. Cited by: §1. [45] Z. Sun, X. Chen, and A. Szolnoki (2023) State-dependent optimal incentive allocation protocols for cooperation in public goods games on regular networks. IEEE Transactions on Network Science and Engineering 10 (6), p. 3975–3988. Cited by: §1. [46] G. Szabó and G. Fath (2007) Evolutionary games on graphs. Physics reports 446 (4-6), p. 97–216. Cited by: §1, §1. [47] A. Szolnoki, G. Szabó, and M. Perc (2011) Phase diagrams for the spatial public goods game with pool punishment. Physical Review E—Statistical, Nonlinear, and Soft Matter Physics 83 (3), p. 036101. Cited by: §1. [48] A. Traulsen and M. A. Nowak (2006) Evolution of cooperation by multilevel selection. Proceedings of the National Academy of Sciences 103 (29), p. 10952–10955. External Links: Document, Link, https://w.pnas.org/doi/pdf/10.1073/pnas.0602530103 Cited by: §1. [49] K. Tuyls and S. Parsons (2007) What evolutionary game theory tells us about multiagent learning. Artificial Intelligence 171 (7), p. 406–416. Cited by: §1. [50] S. Wang, X. Chen, and A. Szolnoki (2019) Exploring optimal institutional incentives for public cooperation. Commun Nonlinear Sci Numer Simul 79, p. 104914. Cited by: §1, §1, §5. [51] S. Wang, L. Liu, and X. Chen (2021) Incentive strategies for the evolution of cooperation: analysis and optimization. Europhysics Letters 136 (6), p. 68002. Cited by: §1, §5. [52] Y. Wu, B. Zhang, and S. Zhang (2017) Probabilistic reward or punishment promotes cooperation in evolutionary games. Chaos, Solitons & Fractals 103, p. 289–293. Cited by: §5. [53] C. Xia, J. Wang, M. Perc, and Z. Wang (2023) Reputation and reciprocity. Physics of life reviews 46, p. 8–45. Cited by: §1, §1. Appendix A Proof of the main results: reward A.1 Zero Sum Transfer (a=1a=1) Proof of the case of zero-sum transfer (a=1a=1). Under the condition a=1a=1, the term (a−1)θ(a-1)θ inside the bracket in the right-hand side of (7) vanishes. Thus the expected total social welfare, SW(θ)SW(θ), is simplified to: SW(θ)=Kf(x)g(x),SW(θ)=K f(x)g(x), where we recall that x=β(aθ+δ)=β(θ+δ)x=β(aθ+δ)=β(θ+δ) and K:=N22(δ+NΔ)K:= N^22(δ+N ). For the DG game: K=N22(δ+NΔ)=N22[−(c+bN−1)+NbN−1]=N22(b−c)>0sinceb>c.K= N^22(δ+N )= N^22 [- (c+ bN-1 )+ NbN-1 ]= N^22(b-c)>0 b>c. For the PGG game: K=δ+NΔ=N22[−c(1−r(N−n)n(N−1))+Nrc(n−1)n(N−1)]=N22c(r−1)>0sincec>0,r>1.K=δ+N = N^22 [-c (1- r(N-n)n(N-1) )+ Nrc(n-1)n(N-1) ]= N^22c(r-1)>0 c>0,~r>1. Thus K>0K>0 for both games. Therefore, the problem of maximising SW(θ)a=1SW(θ)_a=1 reduces to the following optimisation problem maxθ≥θωΨ(θ),whereΨ(θ):=f(x)g(x)=f(β(θ+δ))g(β(θ+δ)). _θ≥ _ω (θ), (θ):= f(x)g(x)= f(β(θ+δ))g(β(θ+δ)). Using the chain rule, we compute the derivative of Ψ : Ψ′(θ)=x′(θ)dx(f(x)g(x))=βf′(x)g(x)−f(x)g′(x)g(x)2:=βuP(u)g2(x), (θ)=x (θ) ddx ( f(x)g(x) )=β f (x)g(x)-f(x)g (x)g(x)^2:=β uP(u)g^2(x), (A1) where following [12] we have defined u=exu=e^x and P(u)=f(x)g′(x)−f′(x)g(x)u.P(u)= f(x)g (x)-f (x)g(x)u. (A2) More precisely, P(u) P(u) :=(1+u)[(∑j=0N−2(HN+1N−1−j)uj)(∑j=1N−1juj−1)−(∑j=1N−2(HN+1N−1−j)juj−1)(∑j=0N−1uj)] :=(1+u) [ ( _j=0^N-2(H_N+ 1N-1-j)u^j ) ( _j=1^N-1ju^j-1 )- ( _j=1^N-2 (H_N+ 1N-1-j )ju^j-1 ) ( _j=0^N-1u^j ) ] −(∑j=0N−2(HN+1N−1−j)uj)(∑j=0N−1uj), - ( _j=0^N-2 (H_N+ 1N-1-j )u^j ) ( _j=0^N-1u^j ), (A3) where HNH_N denotes the harmonic number HN=∑j=1N−11j.H_N= _j=1^N-1 1j. We will use the following properties of P which is proved in [12, Proposition 1.6, supplementary document] (P1) P(u)P(u) is a polynomial of order 2N−42N-4. (P2) For u>0u>0, P(u)P(u) has exactly one solution u0>1u_0>1 and sign(P(u))=sign(u−u0) sign(P(u))= sign(u-u_0). From (A1) and the second property above, we have Ψ′(θ) (θ) has a unique positive root θ0=log(u0)β−δ _0= (u_0)β-δ. In addition, Ψ(θ) (θ) is increasing on (0,θ0)(0, _0) and decreasing on (θ0,+∞)( _0,+∞). This guarantees that Ψ has a global maximum at θ0 _0. ∎ Numerical calculation of the maximiser We have established the existence of a unique global maximiser for the expected total welfare SW(θ)SW(θ) in the case a=1a=1. The maximiser θ0 _0 which is computed from the unique positive root u0>1u_0>1 of the polynomial P defined in (A.1). However, for large N, according to Abel’s impossibility theorem, finding the analytical value for u0u_0 (thus θ0 _0) is analytically intractable since P(u)P(u) is a polynomial of degree 2N−42N-4. Therefore, we compute u0u_0 numerically using Brent’s method, which combines bisection, secant, and inverse quadratic interpolation to ensure both reliability and fast convergence [38]. Since P(u)P(u) is a polynomial admitting a unique positive root, Brent’s method can be used to compute u0u_0 accurately and efficiently once a valid bracketing interval is established. To this end, we first determine an interval [umin,umax][u_ ,u_ ] such that P(umin)P(umax)<0P(u_ )P(u_ )<0. Noting that P(1)<0P(1)<0, we construct the upper bound by iteratively doubling: umax=2p,where p=mink∈ℤ+:P(2k)>0,u_ =2^p, p= \k ^+:P(2^k)>0\, thereby ensuring the existence of a sign change within the interval. The maximum overall social welfare is achieved when the incentive θ forces the system into the state defined by x0x_0. Since x=β(aθ+δ)x=β(aθ+δ), the Optimal Social Welfare Incentive (θ0SW(θ) _0^SW(θ)) is derived as: θ0SW(θ)=x0aβ−δa _0^SW(θ)= x_0aβ- δa A.2 For a>1a>1 Proof of the case of sufficient transfer (a>1a>1). To determine the optimal incentive θ that maximises social welfare, we analyse the derivative of SW(θ)SW(θ) with respect to θ. Substituting the expressions for f(x)f(x), g(x)g(x), and their derivatives, the derivative dSW(θ)/dθdSW(θ)/dθ can be computed explicitly as follows: We take the derivative of Social Welfare objective function from 7, for θ such that u>u0u>u_0, as follows: dSW(θ)dθ dSW(θ)dθ =N22[aβf′(x)g(x)−f(x)g′(x)g2(x)(δ+NΔ+(a−1)θ)+(a−1)f(x)g(x)] = N^22 [aβ f (x)g(x)-f(x)g (x)g^2(x) (δ+N +(a-1)θ )+(a-1) f(x)g(x) ] =N2(a−1)2g2(x)[aβ(f′(x)g(x)−f(x)g′(x))(δ+NΔa−1+θ)+f(x)g(x)] = N^2(a-1)2g^2(x) [aβ(f (x)g(x)-f(x)g (x)) ( δ+N a-1+θ )+f(x)g(x) ] =N2(a−1)uP(u)2g2(x)[−aβ(δ+NΔa−1+θ)+f(x)g(x)uP(u)] = N^2(a-1)uP(u)2g^2(x) [-aβ ( δ+N a-1+θ )+ f(x)g(x)uP(u) ] =N2(a−1)uP(u)2g2(x)[f(x)g(x)uP(u)−aβθ+aβδ+NΔ1−a] = N^2(a-1)uP(u)2g^2(x) [ f(x)g(x)uP(u)-aβθ+aβ δ+N 1-a ] =N2uP(u)2g2(x)(a−1)[F(u)+β], = N^2uP(u)2g^2(x)(a-1) [F(u)+ ], (A4) where we have defined F(u):=f(x)g(x)uP(u)−xand :=δ+aNΔ1−a.F(u):= f(x)g(x)uP(u)-x K:= δ+aN 1-a. (A5) From [12] we have important properties of F(u)F(u) on (u0,+∞)(u_0,+∞), where u0u_0 is the unique positive root of the polynomial P, as follows: (F1) F(u)>0F(u)>0, (F2) F(u)>1+u−log(u)F(u)>1+u- (u). In particular, limF(u)=+∞ F(u)=+∞. (F3) There u∗u^* such that F∗:=F(u∗)=minu>u0F(u)F^*:=F(u^*)= _u>u_0F(u). (F4) When N≤N0=100N≤ N_0=100, u∗u^* is unique and sign(dF/du)=sign(u−u∗) sign(dF/du)= sign(u-u^*). We consider the sign of K. For Donation Game: =11−a(−c−bN−1+aNbN−1) = 11-a (-c- bN-1+ aNbN-1 ) =Nb(1−a)(N−1)[a−c(N−1)Nb−1N] = Nb(1-a)(N-1) [a- c(N-1)Nb- 1N ] =Nb(1−a)(N−1)[a−c(N−1)+bNb]. = Nb(1-a)(N-1) [a- c(N-1)+bNb ]. For Public Good Game: =11−a[−c+cr(N−n)n(N−1)+aNrc(n−1)n(N−1)] = 11-a [-c+c r(N-n)n(N-1)+aN rc(n-1)n(N-1) ] =c(1−a)n(N−1)[−n(N−1)+r(N−n)+aNr(n−1)] = c(1-a)n(N-1) [-n(N-1)+r(N-n)+aNr(n-1) ] =cNr(n−1)(1−a)n(N−1)[a−n(N−1)−r(N−n)Nr(n−1)]. = cNr(n-1)(1-a)n(N-1) [a- n(N-1)-r(N-n)Nr(n-1) ]. We define the threshold value of a a∗:=c(N−1)+bNbin Donation Game,n(N−1)−r(N−n)Nr(n−1)in Public Goods Game.a^*:= cases c(N-1)+bNb Donation Game,\\ n(N-1)-r(N-n)Nr(n-1) Public Goods Game. cases Then it follows that in both games sign()=sign(1−a)sign(a−a∗) sign(K)= sign(1-a) sign(a-a^*). Note further that since c(N−1)+b>0[c(N−1)+b]−Nb=(c−b)(N−1)<0n(N−1)−r(N−n)=(n−r)(N−1)+r(n−1)>0[n(N−1)−r(N−n)]−Nr(n−1)=n(N−1)(1−r)<0 casesc(N-1)+b>0\\ [c(N-1)+b]-Nb=(c-b)(N-1)<0\\ n(N-1)-r(N-n)=(n-r)(N-1)+r(n-1)>0\\ [n(N-1)-r(N-n)]-Nr(n-1)=n(N-1)(1-r)<0 cases we deduce that 0<a∗<10<a^*<1. In the case of sufficient transfer, since a>1>a∗a>1>a^*, K is always negative. We recall the threshold value β∗β^* β∗=−F∗>0.β^*=- F^*K>0. (i) Then for β≤β∗β≤β^*, SW(θ)SW(θ) is non-decreasing on (θ0,+∞)( _0,+∞), where we recall that θ0=logu0−βδβa _0= u_0-βδβ a. (i) For β>β∗β>β^*, the number of changes of the sign of dSW(θ)/dθdSW(θ)/dθ is at least two for all N and there exists an N0N_0 such that the number of changes is exactly two for N≤N0N≤ N_0. As a consequence, for N≤N0N≤ N_0, there exist θ1<θ2 _1< _2 such that, for β>β∗β>β^*, SW(θ)SW(θ) is increasing when θ<θ1θ< _1, decreasing when θ1<θ<θ2 _1<θ< _2 and increasing when θ>θ2θ> _2. ∎ A.3 For a<1a<1 Proof of the case of insufficient transfer (a<1a<1). (i) In the case of insufficient transfer a<1a<1, then by the definition of a∗a^*, we have >0K>0 when a>a∗a>a^* and <0K<0 when a<a∗a<a^*. For >0K>0, SW(θ)SW(θ) is strictly decreasing on (θ0,+∞)( _0,+∞). (i) With the threshold β∗β^* from 13, for β<β∗β<β^*, SW(θ)SW(θ) is non-decreasing on (θ0,+∞)( _0,+∞). Consequently: maxθ≥θ0SW(θ)=SW(θ0) _θ≥ _0SW(θ)=SW( _0) (i) For β>β∗β>β^*, the number of changes of the sign of dSW(θ)/dθdSW(θ)/dθ is at least two for all N and there exists an N0N_0 such that the number of changes is exactly two for N≤N0N≤ N_0. As a consequence, for N≤N0N≤ N_0, there exist θ1<θ2 _1< _2 such that, for β>β∗β>β^*, SW(θ)SW(θ) is decreasing when θ<θ1θ< _1, increasing when θ1<θ<θ2 _1<θ< _2 and decreasing when θ>θ2θ> _2. Thus, for N≤N0N≤ N_0: maxθ≥θ0SW(θ)=maxSW(θ0),SW(θ2) _θ≥ _0SW(θ)= \SW( _0),SW( _2)\ (iv) Moreover, for sufficiently large β and small θ, SW(θ)SW(θ) is increasing as θ→0+θ→ 0^+ (shown in Lemma 3). ∎ Proof of Theorem 2 In this section, we will prove Theorem 2. To this end, we first need some axillary results. The following proposition presents some properties of the expected total social welfare SWSW. Proposition 1 (Basic properties of SW(θ)SW(θ)). Recall the Social Welfare objective SW(θ)=N22f(x)g(x)(δ+NΔ+(a−1)θ)SW(θ)= N^22 f(x)g(x) (δ+N +(a-1)θ ) where x=β(aθ+δ)x=β(aθ+δ). With a<1a<1, the following properties hold: 1. Positive interval: SW(θ)≥0 iff θ∈I:=[0,δ+NΔ1−a]SW(θ)≥ 0 ~ iff ~ θ∈ I:= [0, δ+N 1-a ]. 2. Local extrema: SW′(θ)=0SW (θ)=0 if and only if: 1−aδ+NΔ+(a−1)θ=−aβuP(u)f(x)g(x). 1-aδ+N +(a-1)θ=-aβ uP(u)f(x)g(x). (A6) Proof of Proposition 1. The first statement follows directly from the formula of SW(θ)SW(θ). In fact, since f(x)f(x) and g(x)g(x) are positive polynomials, the sign of SW(θ)SW(θ) is the same as the sign of factor δ+NΔ+(a−1)θδ+N +(a-1)θ. For the second statement, we recall from (A4) that dSW(θ)dθ=N2(a−1)2g2(x)[aβuP(u)(δ+NΔ1−a−θ)+f(x)g(x)], dSW(θ)dθ= N^2(a-1)2g^2(x) [aβ uP(u) ( δ+N 1-a-θ )+f(x)g(x) ], where u=exu=e^x. From this we deduce (A6). ∎ In the following proposition, we give an alternative representation for f(x)f(x) and g(x)g(x). Proposition 2. The functions f(x)f(x) and g(x)g(x) can be expressed in the following forms: f(x)=∑j=0N−1ηjuj,andg(x)=∑j=0N−1uj,f(x)= _j=0^N-1 _ju^j, g(x)= _j=0^N-1u^j, where u(θ)=eβ(aθ+δ)=exu(θ)=e^β(aθ+δ)=e^x and η0=1N−1+HN,ηj=2HN+1N−j+1N−j−1for 1≤j≤N−2,andηN−1=1+HN. _0= 1N-1+H_N, _j=2H_N+ 1N-j+ 1N-j-1 1≤ j≤ N-2, _N-1=1+H_N. Proof of Proposition 2. This follows directly from the formula of f(x)f(x) and g(x)g(x) in (8) and (9) respectively. ∎ Let S(u)=f(x)g(x)S(u)= f(x)g(x) and for convenience, R(θ)=S(u(θ))R(θ)=S(u(θ)). Specifically, let: S(u)=∑j=0N−1ηjuj∑j=0N−1uj and R(θ)=∑j=0N−1ηjejβ(aθ+δ)∑j=0N−1ejβ(aθ+δ).S(u)= _j=0^N-1 _ju^j _j=0^N-1u^j and R(θ)= _j=0^N-1 _je^jβ(aθ+δ) _j=0^N-1e^jβ(aθ+δ). Lemma 1. S(u)>η0S(u)> _0 for all u>0u>0. Proof of Lemma 1. It follows from the formula of ηj _j that ηj>η0 _j> _0 for all 1≤j≤N−11≤ j≤ N-1. Therefore, we have f(u)>∑j=0N−1η0uj=η0g(u)f(u)> _j=0^N-1 _0u^j= _0g(u), which implies S(u)>η0S(u)> _0. ∎ Lemma 2. The function S(u)S(u) is Lipschitz continuous and strictly increasing for all u∈(0,1)u∈(0,1). Proof of Lemma 2. By the chain rule, we have: dSdu dSdu =dSdxdxdu=−uP(u)g(x)21u=−P(u)g(x)2 = dSdx dxdu= -uP(u)g(x)^2 1u= -P(u)g(x)^2 Thus, dS/dudS/du is continuous on (0,1)(0,1). Furthermore, the sign of dS/dudS/du matches the sign of −P(u)-P(u). Since P(u)<0P(u)<0 for all u∈(0,u0)u∈(0,u_0) and u0>1u_0>1, S(u)S(u) is strictly increasing for all u∈(0,1)u∈(0,1). ∎ Let L be Lipschitz bound for S(u)S(u) in the interval [0,1][0,1]. The property stated in the following Lemma yields the lower bound for the search interval of Algorithm 1 Lemma 3 (Left-side dominance by SW(0)SW(0)). There exists a constant rl>1r_l>1, independent of β, such that, for all sufficiently large β, SW(θ)≤SW(0)SW(θ)≤ SW(0) for all θ∈[0,μ]θ∈[0,μ] where μ=−δa−logrlaβμ= -δa- r_laβ. To prove lemma 3, we divide [0,μ][0,μ] into two parts and use the following two claims. Claim 1. For sufficiently large β, SW(θ)SW(θ) is strictly decreasing on the interval I0:=[0,1aβ]I_0:= [0, 1aβ ]. Proof of Claim 1. First, we make sure that β is large enough such that u(1aβ)=eβδ+1<1u( 1aβ)=e^βδ+1<1, so that S′(u(θ))>0S (u(θ))>0, and so is R′(θ)=u′(θ)S′(u(θ))R (θ)=u (θ)S (u(θ)), for all θ∈I0θ∈ I_0. We proceed by contradiction. Suppose there exists a point θ~∈I0 θ∈ I_0 such that SW′(θ~)≥0SW ( θ)≥ 0. Differentiating the objective function SW(θ)SW(θ) yields: SW′(θ)=N22((δ+NΔ)R′(θ)−(1−a)θR′(θ)−(1−a)R(θ)).SW (θ)= N^22 ((δ+N )R (θ)-(1-a)θ R (θ)-(1-a)R(θ) ). For SW′(θ~)≥0SW ( θ)≥ 0 to hold, we require: (δ+NΔ)R′(θ~)−(1−a)θ~R′(θ~)≥(1−a)R(θ~)(δ+N )R ( θ)-(1-a) θR ( θ)≥(1-a)R( θ) We know a<1a<1, and u′(θ)=aβeβ(aθ+δ)≤aβeβδ+1u (θ)=aβ e^β(aθ+δ)≤ aβ e^βδ+1 for all θ∈I0θ∈ I_0. Combining these facts with Lemma 1 and Lemma 2, the inequality implies: (δ+NΔ)aβeβδ+1L≥(δ+NΔ)R′(θ~)−(1−a)θ~R′(θ~)≥(1−a)R(θ~)>(1−a)η0.(δ+N )aβ e^βδ+1L≥(δ+N )R ( θ)-(1-a) θR ( θ)≥(1-a)R( θ)>(1-a) _0. Since δ<0δ<0, we have limβ→∞βeβδ=0 _β→∞β e^βδ=0. Thus, for sufficiently large β, the above inequality fails. Thus, we conclude that for sufficiently large β, it holds that SW′(θ)<0SW (θ)<0 for all θ∈I0θ∈ I_0. This implies SW(θ)<SW(0)SW(θ)<SW(0) for all θ∈I0θ∈ I_0 for sufficiently large β. ∎ Claim 2. There exists an ε0>1 _0>1, independent of β, such that for all θ∈I1:=[1aβ,−δa−logε0aβ]θ∈ I_1:= [ 1aβ, -δa- _0aβ ], SW(θ)≤SW(0)SW(θ)≤ SW(0). Proof of Claim 2. First, note that this interval is valid for sufficiently large β since 1aβ≤−δa−logε0aβ 1aβ≤ -δa- _0aβ, for β≥1+logε0−δβ≥ 1+ _0-δ. Suppose by contradiction that there exists a θ~∈I1 θ∈ I_1 such that SW(θ~)>SW(0)SW( θ)>SW(0). This is equivalent to: (δ+NΔ)R(θ~)−(1−a)θ~R(θ~)>(δ+NΔ)R(0).(δ+N )R( θ)-(1-a) θR( θ)>(δ+N )R(0). Rearranging the terms yields: (δ+NΔ)(R(θ~)−R(0))>(1−a)θ~R(θ~), (δ+N )(R( θ)-R(0))>(1-a) θR( θ), (A7) which implies R(θ~)−R(0)>0R( θ)-R(0)>0. Applying Lemma 1 and Lemma 2 yields: (1−a)θ~η0<(1−a)θ~R(θ~) (1-a) θ _0<(1-a) θR( θ) <(δ+NΔ)(R(θ~)−R(0)) <(δ+N )(R( θ)-R(0)) <(δ+NΔ)L(u(θ~)−u(0))<(δ+NΔ)Lu(θ~). <(δ+N )L(u( θ)-u(0))<(δ+N )Lu( θ). (A8) By making the substitution θ~=−δa−logεaβ θ= -δa- aβ (where ε0≤ε≤e−βδ−1 _0≤ ≤ e^-βδ-1), we have u(θ~)=ε−1u( θ)= ^-1. Rearranging the left-most and right-most sides of the inequality, we have: ε(−δ−logεβ)<(δ+NΔ)La(1−a)η0. (-δ- β )< (δ+N )La(1-a) _0. Observe that the left-hand side is an increasing function of ε , for all 1≤ε≤e−βδ−11≤ ≤ e^-βδ-1. Therefore, by choosing an ε0 _0 strictly independent of β such that −δε0≥(δ+NΔ)La(1−a)η0-δ _0≥ (δ+N )La(1-a) _0, and a β large enough such that ε0 _0 is valid (i.e. 1≤ε0≤e−βδ−11≤ _0≤ e^-βδ-1) the inequality fails for all valid ε≥ε0 ≥ _0, a contradiction. Hence, proving our claim. ∎ Proof of Lemma 3. Lemma 3 follows by combining Claim 1 and Claim 2 and setting rl=ε0r_l= _0. ∎ We are now in the position to prove Theorem 2. Proof of Theorem 2. We establish the upper and lower bound of the search interval. For the upper bound: Since the polynomial P(u)P(u) has a unique root u0>1u_0>1, from (A6), for every θ∈Iθ∈ I, the LHS is strictly positive and strictly increasing with respect to θ, while the RHS is strictly negative for all u>u0u>u_0. Therefore, any extrema θ∗∈Iθ^*∈ I (i.e. solution to equation (A6)) must yield u∈(0,u0)u∈(0,u_0). This implies that eβ(aθ∗+δ)<u0e^β(aθ^*+δ)<u_0 which is equivalent to θ∗−(−δa)=θ∗−θ∞<logu0aβ.θ^*- ( -δa )=θ^*- _∞< u_0aβ. This gives the upper bound. The lower bound is a consequence of Lemma 3. In fact, since on [0,μ][0,μ], SW(θ)≤SW(0)SW(θ)≤ SW(0)), any meaningful extrema θ∗θ^* must satisfy θ∗>μθ^*>μ, or equivalently θ∗−δa>−logrlaβ.θ^*- -δa>- r_laβ. This establishes the lower bound of θ∗θ^* and completes the proof of Theorem 2. ∎ A.4 Influence of Selection Intensity (β) The intensity of selection β also plays a crucial role in deciding the overall stability and long-term optimal structure of the system’s Social Welfare (SWSW). We recall from (7) that SW(θ)=N22f(x)g(x)(δ+NΔ+(a−1)θ),SW(θ)= N^22 f(x)g(x)(δ+N +(a-1)θ), where x=β(aθ+δ)x=β(aθ+δ). Neutral Selection Limit (β→0+β→ 0^+) The weak selection (neutral) limit corresponds to β→0β→ 0, where payoff differences have only a small effect on strategy adoption probabilities. In this regime, updates are nearly random and the dynamics are dominated by neutral drift, with selection acting only as a weak perturbation. It follows from the above formula for the total expected social welfare and Proposition 2 that limβ→0SW(θ)=N22f(0)g(0)(δ+NΔ+(a−1)θ)=N22∑j=0N−1ηjN(δ+NΔ+(a−1)θ)=N2HN(δ+NΔ+(a−1)θ). _β→ 0SW(θ)= N^22 f(0)g(0)(δ+N +(a-1)θ)= N^22 _j=0^N-1 _jN(δ+N +(a-1)θ)=N^2H_N(δ+N +(a-1)θ). Strong Selection Limit (β→+∞β→+∞) By a straightforward adaption of the proof of [12, Proposition 1.12] we have limβ→+∞f(x)g(x)=2HNforθ=−δa,HN+1forθ>−δa,HN+1N−1forθ<−δa. _β→+∞ f(x)g(x)= cases2H_N θ=- δa,\\ H_N+1 θ>- δa,\\ H_N+ 1N-1 θ<- δa. cases It immediately follows that limβ→+∞SW(θ)=N2HN(δ+NΔ+(a−1)θ)forθ=−δa,N22(HN+1)(δ+NΔ+(a−1)θ)forθ>−δa,N22(HN+1N−1)(δ+NΔ+(a−1)θ)forθ<−δa. _β→+∞SW(θ)= casesN^2H_N(δ+N +(a-1)θ) θ=- δa,\\ N^22(H_N+1)(δ+N +(a-1)θ) θ>- δa,\\ N^22 (H_N+ 1N-1 )(δ+N +(a-1)θ) θ<- δa. cases Appendix B Proof of the main results: punishment B.1 Social welfare in institutional punishment In our model where the institution punishes Defectors (instead of rewarding Cooperators in above sections), the total payoff received in the population with i cooperators and incentive efficiency a a, is: P^i=iΠC(i)+(N−i)[ΠD(i)−a^θ]. P_i=i\, _C(i)+(N-i) [ _D(i)- aθ ]. Combining with the total institutional cost θ^i=(N−i)θ θ_i=(N-i)θ yields aggregate the social welfare in state SiS_i as: SW^i(θ)=P^i−θ^i=iΠC(i)+(N−i)[ΠD(i)−(1+a^)θ] SW_i(θ)= P_i- θ_i=i\, _C(i)+(N-i) [ _D(i)-(1+ a)θ ] Which then by carrying out the same algebraic process as the reward case, yields the Social Welfare function: SW^(θ) SW(θ) =12∑iSW^i(θ)(n1,i+nN−1,i) = 12 _i SW_i(θ)(n_1,i+n_N-1,i) =12∑i(i(δ+NΔ)−(N−i)(1+a^)θ)(n1,i+nN−1,i) = 12 _i(i(δ+N )-(N-i)(1+ a)θ)(n_1,i+n_N-1,i) =12(δ+NΔ)∑i(n1,i+nN−1,i)−12(1+a^)θ∑i(N−i)(n1,i+nN−1,i) = 12(δ+N ) _ii(n_1,i+n_N-1,i)- 12(1+ a)θ _i(N-i)(n_1,i+n_N-1,i) =N22f(x)g(x)(δ+NΔ)−N22f^(x)g(x)(1+a^)θ, = N^22 f(x)g(x) (δ+N )- N^22 f(x)g(x)(1+ a)θ, (A9) where as derived in [12] f^(x)=(1+ex)[(1+ex+⋯+e(N−2)x)HN+∑j=1N−1e(j−1)xj] f(x)=(1+e^x) [ (1+e^x+·s+e^(N-2)x )H_N+ _j=1^N-1 e^(j-1)xj ] Moreover, one can write f^(x) f(x) in the following form: f^(x)=∑j=0N−1η^juj, f(x)= _j=0^N-1 η_ju^j, where u=eβ(a^θ+δ)=exu=e^β( aθ+δ)=e^x, and η^j=ηN−1−j η_j= _N-1-j for all 0≤j≤N−10≤ j≤ N-1. B.2 Algorithm 1 for punishment We claim that Algorithm 1 also works for the punishment case. We will justify the our algorithm via the following theorem, which is the counterpart of Theorem 2 in the reward case. The proof will be carried out in a similar manner to the reward case. That is, by providing an upper and lower bound for the quantity |θ∗−θ^∞||θ^*- θ_∞|. Firstly, by applying the chain rule, we have: dSW^dθ=dSW^dxdxdθ d SWdθ= d SWdx dxdθ Since x=β(a^θ+δ)x=β( aθ+δ), we have dxdθ=a^β dxdθ= aβ. On the other hand, from equation A9, we have: dSW^dx d SWdx =N22(δ+NΔ)f′(x)g(x)−f(x)g′(x)g(x)2−N22(1+a^)(f^′(x)g(x)−f^(x)g′(x)g(x)2θ+f^(x)g(x)1a^β) = N^22(δ+N ) f (x)g(x)-f(x)g (x)g(x)^2- N^22(1+ a) ( f (x)g(x)- f(x)g (x)g(x)^2θ+ f(x)g(x) 1 aβ ) =N2(1+a^)2g(x)2(−δ+NΔ1+buP(u)+θuP^(u)−Q^(u)a^β) = N^2(1+ a)2g(x)^2 (- δ+N 1+buP(u)+θ u P(u)- Q(u) aβ ) where uP^(u)=f^(x)g′(x)−f^′(x)g(x)u P(u)= f(x)g (x)- f (x)g(x) and Q^(u)=f^(x)g(x) Q(u)= f(x)g(x) are polynomials defined in [12]. Therefore, we have: dSW^dθ=N2a^β(1+a^)2g(x)2(θuP^(u)−δ+NΔ1+a^uP(u)−Q^(u)a^β) d SWdθ= N^2 aβ(1+ a)2g(x)^2 (θ u P(u)- δ+N 1+ auP(u)- Q(u) aβ ) (A10) Upper Bound We will prove the following lemma: Lemma 4. There exists a positive number pup_u, independent of β, such that for large enough β, if θ is an extrema of SW SW then θ<−δa^+logpuβa^βθ< -δ a+ p_uβ aβ Proof. Two results from [12] showed that: P^(u)=−u2N−4P(1u) P(u)=-u^2N-4P ( 1u ) and Q^(u)≥43(N−1)(1+u)uP^(u)∀u>0 Q(u)≥ 43(N-1)(1+u)u P(u) ∀ u>0 And since P(u)P(u) has a unique positive root u0>1u_0>1, P^(u) P(u) also has a unique positive root u0∗=1u0<1u^*_0= 1u_0<1. Furthermore, since P(u)P(u) is strictly negative for all u∈(0,u0)u∈(0,u_0) and strictly positive for all u∈(u0,+∞)u∈(u_0,+∞), it follows that P^(u) P(u) is strictly negative for all u∈(0,1u0)u∈(0, 1u_0) and strictly positive for all u∈(1u0,+∞)u∈( 1u_0,+∞). Consider θ such that u(θ)∈(u0,+∞)u(θ)∈(u_0,+∞) where we define u(θ)=eβ(a^θ+δ)u(θ)=e^β( aθ+δ), a bit different from the reward case. From (A10), it holds that if θ~ θ (corresponding to u~ u) is an extrema of SW^(θ) SW(θ) then: θ~u~P^(u~)=δ+NΔ1+a^u~P(u~)+Q^(u~)a^β>Q^(u~)a^β≥43a^β(N−1)(1+u~)u~P^(u~) θ u P( u)= δ+N 1+ a uP( u)+ Q( u) aβ> Q( u) aβ≥ 43 aβ(N-1)(1+ u) u P( u) (A11) Comparing left-most and right-most side of the above inequality yields: θ~>43a^β(N−1)(1+u~)>4u~3a^β(N−1) θ> 43 aβ(N-1)(1+ u)> 4 u3 aβ(N-1) Let u~=a^βv u= aβ v, by substituting θ~=−δa^+logu~a^β θ= -δ a+ u aβ into the above inequality, we have: −δa^+loga^βa^β+logva^β>4v3(N−1) -δ a+ aβ aβ+ v aβ> 4v3(N-1) Using the inequality x≥elogx≥ e x for all x>0x>0 gives us: −δa^+1e+logva^β≥−δa^+loga^βa^β+logva^β>4v3(N−1) -δ a+ 1e+ v aβ≥ -δ a+ aβ aβ+ v aβ> 4v3(N-1) Observe that there exists a constant v∗v^* independent of β such that for all large enough β, the above inequality fails for all v≥v∗v≥ v^*, which means that if θ~ θ is an extrema of SW SW, and therefore a root of the expression in (A10), it holds that u(θ~)<a^βv∗u( θ)< aβ v^*, and therefore θ~<−δa^+loga^βv∗a^β θ< -δ a+ aβ v^* aβ By setting pu=a^v∗p_u= av^*, we have proven our lemma. ∎ Lower Bound Let S^(u)=f^(x)g(x) S(u)= f(x)g(x) and for convenience, R^(θ)=S^(u(θ)) R(θ)= S(u(θ)). Specifically, let: S^(u)=∑j=0N−1ηj^uj∑j=0N−1uj and R^(θ)=∑j=0N−1η^jejβ(a^θ+δ)∑j=0N−1ejβ(a^θ+δ) S(u)= _j=0^N-1 _ju^j _j=0^N-1u^j and R(θ)= _j=0^N-1 η_je^jβ( aθ+δ) _j=0^N-1e^jβ( aθ+δ) First, notice that Lemma 1 also applies for S^(u) S(u), that is S^(u)>η0 S(u)> _0 for all u>0u>0. Furthermore, we have: dS^du=dS^dxdxdu=−uP^(u)g(x)21u=−P^(u)g(x)2 d Sdu= d Sdx dxdu= -u P(u)g(x)^2 1u= - P(u)g(x)^2 which, from the analyses in the Upper Bound section, is positive for all u∈(0,1u0)u∈(0, 1u_0) and negative for all u∈(1u0,+∞)u∈( 1u_0,+∞). With that, we present a similar Lemma as in the reward case: Lemma 5 (Left-side dominance by SW^(0) SW(0)). There exists a constant pl>1p_l>1, independent of β, such that for all sufficiently large β, the following holds: If μ=−δa^−logpla^βμ= -δ a- p_l aβ, then SW^(θ)≤SW^(0) SW(θ)≤ SW(0) for all θ∈[0,μ]θ∈[0,μ]. Proof. The proof is carried out similarly like in the reward case, with the only difference being that the quantity (1−a)R(θ)(1-a)R(θ) is replaced with (1+a^)R^(θ)(1+ a) R(θ). ∎ B.3 Comparison between reward and punishment In this section, we demonstrate that rewarding cooperators frequently yields better outcomes than punishing defectors. Specifically, we observed that if the transfer efficiency of rewards is sufficiently high relative to punishments, any budget allocated to punishment can be replaced by a corresponding reward level that achieves equal or greater Social Welfare. We define the following functions for all θ≥0θ≥ 0 and incentive efficiency a>0a>0 and a^>0 a>0: F(a,θ)=∑j=0N−1ηjejβ(aθ+δ),F^(a^,θ)=∑j=0N−1ηj^ejβ(a^θ+δ) andG(a,θ)=∑j=0N−1ejβ(aθ+δ).F(a,θ)= _j=0^N-1 _je^jβ(aθ+δ), F( a,θ)= _j=0^N-1 _je^jβ( aθ+δ) and G(a,θ)= _j=0^N-1e^jβ(aθ+δ). (A12) With these definitions, we can write the Social Welfare functions as follow: SW(θ) SW(θ) =N22(δ+NΔ−(1−a)θ)F(a,θ)G(a,θ), = N^22 (δ+N -(1-a)θ ) F(a,θ)G(a,θ), (A13a) SW^(θ) SW(θ) =N22(δ+NΔ)F(a^,θ)G(a^,θ)−N22(1+a^)θF^(a^,θ)G(a^,θ). = N^22 (δ+N ) F( a,θ)G( a,θ)- N^22(1+ a)θ F( a,θ)G( a,θ). (A13b) To prove Theorem 5, we will need two auxiliary lemmas. The first one expresses the difference between the two (rescaled) objective functions. Lemma 6. For all θ≥0θ≥ 0, we have SW^(θa^)−SW(θa)=N2θ2G(1,θ)(1−aF(1,θ)−1+a^a^F^(1,θ)) SW ( θ a )-SW ( θa )= N^2θ2G(1,θ) ( 1-aaF(1,θ)- 1+ a a F(1,θ) ) (A14) Proof. From (A13a)-(A13b), we have: SW(θa)=N22(δ+NΔ−1−aθ)F(1,θ)G(1,θ),SW ( θa )= N^22 (δ+N - 1-aaθ ) F(1,θ)G(1,θ), and SW^(θa^)=N22(δ+NΔ)F(1,θ)G(1,θ)−N221+a^a^θF^(1,θ)G(1,θ). SW ( θ a )= N^22 (δ+N ) F(1,θ)G(1,θ)- N^22 1+ a aθ F(1,θ)G(1,θ). From these expressions, by rearranging terms, then we obtain (A14) as claimed. ∎ The second lemma estimates the ratio between F and F F that appear in (A14). Lemma 7. For all a,θ∈ℝa,θ , we have: η0ηN−1≤F(a,θ)F^(a,θ)≤ηN−1η0, _0 _N-1≤ F(a,θ) F(a,θ)≤ _N-1 _0, or, equivalently: 1−N−2(N−1)(HN+1)≤F(a,θ)F^(a,θ)≤1+N−2(N−1)HN+11- N-2(N-1)(H_N+1)≤ F(a,θ) F(a,θ)≤ 1+ N-2(N-1)H_N+1 Proof. It follows from the formulas of F and F F in (A12) that F(a,θ)F^(a,θ)=∑j=0N−1ηjejβ(aθ+δ)∑j=0N−1η^jejβ(aθ+δ). F(a,θ) F(a,θ)= _j=0^N-1 _je^jβ(aθ+δ) _j=0^N-1 η_je^jβ(aθ+δ). Therefore, by the Generalized Mediant Inequality, the value of this ratio lies between the smallest and largest component fraction of the mediant minjηjejβ(aθ+δ)η^jejβ(aθ+δ)≤F(a,θ)F^(a,θ)≤maxjηjejβ(aθ+δ)η^jejβ(aθ+δ), _j _je^jβ(aθ+δ) η_je^jβ(aθ+δ)≤ F(a,θ) F(a,θ)≤ _j _je^jβ(aθ+δ) η_je^jβ(aθ+δ), that is minjηjη^j≤F(a,θ)F^(a,θ)≤maxjηjη^j. _j _j η_j≤ F(a,θ) F(a,θ)≤ _j _j η_j. A direct pairwise comparison shows that: η0ηN−1≤ηjη^j≤ηN−1η0 _0 _N-1≤ _j η_j≤ _N-1 _0 for all 0≤j≤N−10≤ j≤ N-1. Therefore: η0ηN−1≤F(a,θ)F^(a,θ)≤ηN−1η0, _0 _N-1≤ F(a,θ) F(a,θ)≤ _N-1 _0, which, after an algebraic manipulation step, yields the second desired inequality. Furthermore, these bounds are tight, as the ratio asymptotically approaches the lower and upper bounds as x→−∞x→-∞ and x→+∞x→+∞, respectively. ∎ We are now in the position to prove Theorem 5. Proof of Theorem 5. For a≥1a≥ 1, the right-hand side of equation (A14) is strictly non-positive, implying that SW^(θ)≤SW(a^θ/a) SW(θ)≤ SW( aθ/a) is trivially true for all θ≥0θ≥ 0. Therefore, we restrict our focus to the case where a<1a<1. The condition for the (shifted) punishment to strictly outperform the (shifted) reward, SW^(θ/a^)>SW(θ/a) SW(θ/ a)>SW(θ/a), is equivalent to: 1−aF(1,θ)−1+a^a^F^(1,θ)>0 1-aaF(1,θ)- 1+ a a F(1,θ)>0 (A15) which, by letting a∗=1/a^*=1/a and a^∗=1/a a^*=1/ a, can be equivalently rewritten as a ratio of the polynomials F(1,θ)F^(1,θ) F(1,θ) F(1,θ) >a^∗+1a∗−1. > a^*+1a^*-1. (A16) By combining inequality (A16) with Lemma 7, we know that the ratio F(a,θ)F^(a,θ) F(a,θ) F(a,θ) is bounded above by ηN−1η0 _N-1 _0. Therefore, if we enforce: a^∗+1a∗−1≥ηN−1η0 a^*+1a^*-1≥ _N-1 _0 or equivalently: η0a^∗+η0≥ηN−1a∗−ηN−1 _0 a^*+ _0≥ _N-1a^*- _N-1 (A17) then inequality (A16) can never be satisfied for any θ. This guarantees SW^(θ/a^)≤SW(θ/a) SW(θ/ a)≤ SW(θ/a), or SW^(θ)≤SW(a^θ/a) SW(θ)≤ SW( aθ/a) globally. By some algebraic manipulation and substitution a=1/a∗a=1/a^* and a^=1/a^∗ a=1/ a^*, inequality (A17) is equivalent to: a≥a^ηN−1η0+a^(η0+ηN−1)a≥ a _N-1 _0+ a( _0+ _N-1) which is exactly the claimed condition in Theorem 5. ∎ From Theorem 5, we have the following Corollary: Corollary 1. If it holds true that: a≥ηN−1η0+ηN−1a≥ _N-1 _0+ _N-1 then SW^(θ)≤SW(a^θ/a) SW(θ)≤ SW( aθ/a) for all θ≥0θ≥ 0, regardless of a a. Proof. Notice that the function f(a^)=a^ηN−1η0+a^(η0+ηN−1)f( a)= a _N-1 _0+ a( _0+ _N-1) is strictly increasing for all a^>0 a>0. Therefore, taking the limit a^→+∞ a→+∞ yields the desired result. ∎ B.4 Additional numerical simulations This section presents additional figures for the punishment case and the multi-objective comparison, to further illustrate the behaviour of the social welfare function. The supplementary results provide visual evidence for the key structural properties of the SWSW curve across different game parameter settings. Figure A1: Numerical results for the Donation Game (DG) for the punishment-based policy. The SWSW curves of varying selection intensity seemingly exhibit phase transitions and monotonic behaviours similar to the reward case, although with different thresholds for a a and β∗β^*. Figure A2: Numerical results for the Public Goods Game (PGG) for the punishment-based policy. The SWSW curves of varying selection intensity seemingly exhibit phase transitions and monotonic behaviours similar to the reward case, although with different thresholds for a a and β∗β^*. Figure A3: Multi-objective comparison of optimal incentive levels for the DG (reward case, β=1.0β=1.0). The figures illustrate how the incentive values associated with maximising SW(θ)SW(θ), minimising Er(θ)E_r(θ), and satisfying frequency of cooperation thresholds vary across the three reward-efficiency regimes, namely ineffective reward (a<1)(a<1), zero-sum transfer (a=1)(a=1), and effective reward (a>1)(a>1). Vertical lines indicate the relevant optimal incentives and threshold values for target cooperation frequencies. Figure A4: Multi-objective comparison of optimal incentive levels for the PGG (reward case, β=1.0β=1.0). The figures illustrate how the incentive values associated with maximising SW(θ)SW(θ), minimising Er(θ)E_r(θ), and satisfying frequency of cooperation thresholds vary across the three reward-efficiency regimes, namely ineffective reward (a<1)(a<1), zero-sum transfer (a=1)(a=1), and effective reward (a>1)(a>1). Vertical lines indicate the relevant optimal incentives and threshold values for target cooperation frequencies. Figure A5: Multi-objective comparison of optimal incentive levels for the DG (reward case). Under the limit β→0+β→ 0^+, the minimum incentive levels required to achieve the target cooperation frequencies take extremely large values, causing the cooperation frequency thresholds to dominate the comparison with the social welfare and institutional cost objectives. As a result, the threshold requirements make the other objective-specific optima practically unattainable, except in the efficient-transfer regime (a>1)(a>1), where the social welfare optimum remains compatible with the required incentive level. Figure A6: Multi-objective comparison of optimal incentive levels for the PGG (reward case). Under the limit β→0+β→ 0^+, the minimum incentive levels required to achieve the target cooperation frequencies take extremely large values, causing the cooperation frequency thresholds to dominate the comparison with the social welfare and institutional cost objectives. As a result, the threshold requirements make the other objective-specific optima practically unattainable, except in the efficient-transfer regime (a>1)(a>1), where the social welfare optimum remains compatible with the required incentive level. Figure A7: Multi-objective comparison of optimal incentive levels for the DG (reward case) under the strong-selection limit β→+∞β→+∞. In this limit, the minimum incentive levels required to attain the target cooperation frequencies converge to −δ-δ. Consequently, the threshold requirements reduce the effective comparison to the two objective-specific optima: maximising SW(θ)SW(θ) and minimising Er(θ)E_r(θ) over the feasible interval [−δ,+∞)[-δ,+∞). For a≤1a≤ 1, the two objectives are largely aligned, since the SW(θ)SW(θ) curve is expected to be non-increasing over the feasible range. However, because an extremely large value of β is still finite, the sharp near-vertical transition in SW(θ)SW(θ) can create a small discrepancy between the solutions for the social welfare and the institutional cost objectives. For a>1a>1, the objectives become strongly conflicting: social welfare favours larger incentive levels, whereas institutional cost minimisation favours the smallest feasible incentive. Figure A8: Multi-objective comparison of optimal incentive levels for the PGG (reward case) under the strong-selection limit β→+∞β→+∞. In this limit, the minimum incentive levels required to attain the target cooperation frequencies converge to −δ-δ. Consequently, the threshold requirements reduce the effective comparison to the two objective-specific optima: maximising SW(θ)SW(θ) and minimising Er(θ)E_r(θ) over the feasible interval [−δ,+∞)[-δ,+∞). For a≤1a≤ 1, the two objectives are largely aligned, since the SW(θ)SW(θ) curve is expected to be non-increasing over the feasible range. However, because an extremely large value of β is still finite, the sharp near-vertical transition in SW(θ)SW(θ) can create a small discrepancy between the solutions for the social welfare and the institutional cost objectives. For a>1a>1, the objectives become strongly conflicting: social welfare favours larger incentive levels, whereas institutional cost minimisation favours the smallest feasible incentive.