Paper deep dive
Can We Volunteer Out of the Peer Review Crisis?
Theo Tang, Toby Handfield, Julian Garcia
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 7/8/2026, 7:37:32 AM
Summary
The paper addresses the strain on the peer review system caused by rapidly growing manuscript volumes outpacing reviewer capacity. It proposes a game-theoretic model of a voluntary pre-review lottery where authors accept a chance of random rejection to reduce reviewer burden. The analysis demonstrates that a Nash equilibrium emerges where self-interested scientists voluntarily participate if they possess sufficient epistemic motivation to value collective literature quality. This mechanism reduces review noise and improves the average quality of published science, offering a decentralized solution to the peer review crisis exacerbated by AI-driven submission growth.
Entities (13)
Relation Signals (13)
Theo Tang ā affiliatedwith ā Monash University
confidence 99% Ā· Department of Data Science and Artificial Intelligence, Monash University, Melbourne, Australia
Toby Handfield ā affiliatedwith ā Monash University
confidence 99% Ā· SOPHIS, Monash University, Melbourne, Australia
Julian Garcia ā affiliatedwith ā Monash University
confidence 99% Ā· Department of Data Science and Artificial Intelligence, Monash University, Melbourne, Australia
Voluntary Lottery ā reduces ā Reviewer Burden
confidence 96% Ā· a voluntary lottery in which authors accept a chance of random pre-review rejection, reducing reviewer burden and improving the quality of surviving evaluations.
Voluntary Lottery ā improves ā Average Accepted Quality
confidence 95% Ā· reducing reviewer burden and improving the quality of surviving evaluations. We show that a Nash equilibrium emerges in which authors voluntarily enter the lottery.
Peer Review System ā suffersfrom ā Review Noise
confidence 95% Ā· The result is a widening mismatch: reviewer scarcity, noisier assessments, and declining confidence in editorial decisions.
Game Theory ā models ā Voluntary Lottery
confidence 94% Ā· To analyze this tension, we provide a game-theoretic thought experiment: a voluntary lottery in which authors accept a chance of random pre-review rejection...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The volume of scientific manuscripts is growing faster than the capacity to evaluate them, yet the institutions that govern peer review have remained largely unchanged. The result is a widening mismatch: reviewer scarcity, noisier assessments, and declining confidence in editorial decisions. Every scientist wants better reviews, but review quality depends on the total burden, which no single author can shift. To isolate this tension, we provide a game-theoretic thought experiment: a voluntary lottery in which authors accept a chance of random pre-review rejection, reducing reviewer burden and improving the quality of surviving evaluations. We show that a Nash equilibrium emerges in which authors voluntarily enter the lottery. Scientists who care about the literature they read, not just the papers they publish, will opt in, raising the quality of published science for all.
Tags
Links
- Source: https://arxiv.org/abs/2604.27900v1
- Canonical: https://arxiv.org/abs/2604.27900v1
Trouble viewing inline? Open PDF directly ā
Full Text
83,868 characters extracted from source content.
Expand or collapse full text
Can We Volunteer Out of the Peer Review Crisis? Theo Tang 1 , Toby Handfield 2 , and Julian Garcia 1,* 1 Department of Data Science and Artificial Intelligence, Monash University, Melbourne, Australia 2 SOPHIS, Monash University, Melbourne, Australia * Corresponding author: julian.garcia@monash.edu May 1, 2026 Abstract The volume of scientific manuscripts is growing faster than the capacity to evaluate them, yet the institutions that govern peer review have remained largely unchanged. The result is a widening mismatch: reviewer scarcity, noisier assessments, and declining confidence in editorial decisions. Every scientist wants better reviews, but review quality depends on the total burden, which no single author can shift. To analyze this tension, we provide a game- theoretic thought experiment: a voluntary lottery in which authors accept a chance of random pre-review rejection, reducing reviewer burden and improving the quality of surviving eval- uations. We show that a Nash equilibrium emerges in which authors voluntarily enter the lottery. Scientists who care about the literature they read, not just the papers they publish, will opt in, raising the quality of published science for all. 1 Introduction The peer review system is under mounting strain. The number of submitted manuscripts has been doubling roughly every decade across most fields [1] but the reviewer pool has not kept pace [2]. The consequences are predictable: reviewers face ever-larger piles of manuscripts, with less time to devote to each, and the quality of editorial decisions degrades [3]. The rise of large language models, which dramatically lower the cost of producing manuscripts, threatens to accelerate this trend further [4]. Yet the basic architecture of peer review has remained largely unchanged, cre- ating a growing mismatch between the volume of scientific output and the institutions designed to evaluate it. Empirical evidence confirms that this degradation is substantial. In a landmark consistency experiment at NeurIPS (a major machine-learning venue), two independent review panels dis- agreed on the fate of roughly one quarter of submissions [5, 6]. A replication at NeurIPS 2021, with five times as many submissions, found the same disagreement rate ā and showed that making a venue more selective would increase, not decrease, the arbitrariness of decisions [7]. The pattern is not confined to machine learning: in a study of 4,000 ESRC grant proposals, inter- reviewer correlations were as low as 0.2, and a single negative review roughly halved a proposalās 1 arXiv:2604.27900v1 [cs.GT] 30 Apr 2026 chances even when other reviews were strong [8]. Indeed, peer review has been called a game of chance [9]. These findings raise a basic question: what, if anything, can be done about the link between the scale of science and the rigorous enforcement of quality? Recent theoretical work has begun to formalise this problem. Noisier review encourages weaker papers to be submitted, which increases load and further degrades review quality [10, 11]. Game-theoretic and simulation models of peer review [12, 13, 14, 15] have explored how strategic behaviour by authors and reviewers interacts with review quality [16], but how much voluntary cooperation among authors could improve the system remains an open question. At root, this is a type of public goods problem. Scientists want good quality research to be published, but they also want their own research to be published. Individuals maximizing for the latter goal will tend to degrade our collective achievement of the former goal. We can then ask two questions: first, from a social plannerās point of view, could we improve the quality of published research by a kind of blanket reduction in participation: do ālessā science in order to get ābetterā science? Second, if scientists are sufficiently motivated by the collective quality of science, could some sort of voluntary scheme to reduce the volume of submissions be sustained, and how effective would such a scheme be compared to centralized institutions? To formalise this question, we study a voluntary lottery in which authors accept a chance of random pre-review rejection, reducing the reviewer burden and thereby reducing error in the review process. When review quality is sufficiently sensitive to burden, the random loss of a few papers is more than compensated by the improved accuracy of decisions on all survivors. The lottery is useful not as a policy prescription but because it isolates this tradeoff. It has a single free parameter (the rejection probability L), makes no quality judgments, and poses the cooperation problem sharply: each scientistās decision to participate affects the noise level faced by all. Unlike desk rejection quotas, it requires no editorial discretion and no information beyond each authorās willingness to enter. Lottery mechanisms even have some precedent in science funding, where random allocation among shortlisted proposals has been both proposed and piloted [17, 18]. Our analysis proceeds as follows. We first model the review process and show how the quality of accepted papers degrades as scale increases, then introduce a pre-review lottery as a burden- reduction mechanism and show that it improves the quality of published science (Section 2). We show that the optimal lottery takes the form of a quality threshold, with only low-quality papers entering the lottery (Section 3). We then analyse the strategic equilibrium in which self- interested authors choose their own lottery participation (Section 4), before discussing the impli- cations of AI-driven submission growth (Section 5). We show that even modest epistemic concern sustains voluntary participation, in turn, narrowing the gap between self-interested behaviour and the social optimum. 2 The cost of scale We model a scientific venue (journal or conference) that receives N submissions. Each submission i has true quality q i ā[0,1]. The editor observes a noisy signal of each paperās quality: the review 2 10 1 10 2 10 3 Number of submissionsN 0.65 0.70 0.75 0.80 Average accepted quality Ģ q A β= 0.06 β= 0.3 β= 0.5 02500500075001000012500 Number of submissionsN 0.05 0.10 0.15 0.20 0.25 Review noise Ļ B β= 0.06 β= 0.3 β= 0.5 ICLR data Figure 1: Scale and the quality of published science. (A) Average accepted quality Ģ q as a function of venue size N for three noise elasticities β, computed via Monte Carlo simulation (M= 10,000 replications; Ļ= 0.3, α= 10%). Higher β produces steeper quality loss as venues grow. (B) Empirical support for noise scaling: estimated review noise Ļ from ICLR submission data (2017ā2025; see SI Section S3). score for paper i is drawn from a truncated normal distribution, with mean q i and standard deviation Ļ, where Ļ captures the noise inherent in the review process [12]. The editor ranks submissions by score and accepts the top αN papers, yielding an acceptance rate α. We assume that review noise increases with the reviewer burden. More papers per reviewer mean less time per assessment and a thinner pool of qualified referees, both of which degrade the quality of evaluations. We model this by tracking the fraction Ģ w of the N submissions that receive full review. The effective noise follows a power law: Ļ eff =ĻĀ· Ģ w β (1) whereĻ is the baseline noise (when all submissions are reviewed) andβ > 0 is the noise elasticity ā the rate at which noise increases with load. When Ģ w= 1 (all papers reviewed), Ļ eff =Ļ; any mechanism that reduces the fraction reaching review ( Ģ w < 1) also reduces noise. The power- law form is consistent with empirical data and has some desirable modelling properties (see SI Section S3 for details). Figure 1 illustrates the central prediction: as the number of submissions grows, review noise increases and the quality of accepted papers degrades. This motivates the search for mecha- nisms that can break the link between scale and noise. Various mechanisms already target this link (submission windows, reviewer mandates, excess-paper fees) but remain ad hoc (see Dis- cussion). We study a simple mechanism: a lottery in which papers are randomly rejected before review with probability L, surviving with probability 1ā L. The logic is straightforward: fewer surviving submissions reduce reviewer burden, which translates into lower effective noise. If all papers enter the lottery, a fraction Ģ w=(1ā L) survives to review, and the effective noise becomes 3 Ļ eff = ĻĀ·(1ā L) β . The lottery can only reduce noise, and does so more strongly when β is large. The tradeoff is that some papers, including potentially good ones, are lost before review. Whether this tradeoff is worthwhile depends on how much noise the lottery eliminates relative to the fraction of good papers it sacrifices. Note that the lottery does not change how many papers are published because the acceptance fraction α remains fixed: journals fill a fixed number of publication slots [19], and acceptance rates at major venues are empirically stable even as submissions grow [20]. The lottery only changes how accurately these papers are selected. When the noise elasticity is large enough, even blanket adoption improves quality; more generally, the lotteryās value comes from its interaction with strategic behaviour. We next ask who should participate, for maximum benefit to science, and then who will participate when scientists act in their own interest. 3 The optimal lottery We derive these results using a continuous approximation that yields tractable expressions for acceptance probabilities, quality, and optimal participation. The venue accepts a fixed fraction α of submissions, not a fixed number, so both papers and acceptances scale with N; thus, the limit converts a discrete ranking problem into a continuous threshold problem. Each author chooses a lottery participation level p(q)ā[0,1], and a submission with partic- ipation p survives the lottery with probability w= 1ā L p. The aggregate survival fraction Ģ w= R 1 0 w(q) f(q) dq, where f(q) is the quality density, determines the effective noise: Ļ eff =ĻĀ· Ģ w β . The acceptance probability for a paper of quality q is determined by the(1āα) quantile of the marginal performance distribution (see SI Section S2 for the full derivation). The continuous approximation closely matches Monte Carlo simulations (lines vs. markers in Figures 3 and 4; systematic comparison in SI Section S1). We first map the parameter space where full adoption of the lottery improves review quality. Figure 2 shows how the gain depends on baseline noise Ļ and noise elasticity β. Below the dashed contour, the lottery removes good papers without reducing noise enough to compensate; above it, the noise reduction dominates and quality improves across a broad region. Venues with higher noise or lower acceptance rates fall squarely in the regime where lotteries help, though precise calibration remains difficult given the limited empirical data available (see SI Section S3). The lottery works best when participation is selective: low-quality papers enter, reducing reviewer burden, while high-quality papers do not enter, thus having maximal chance of being accepted. The socially optimal rule (maximising the average quality of published science; since α is fixed, this is equivalent to maximising total quality) takes the form of a threshold: all papers below a quality cutoff Ļ ā enter the lottery, while those above do not (Figure 2, panel B). Low- quality papers contribute to reviewer burden but have little chance of acceptance; removing them reduces noise at almost no cost. High-quality papers would likely be accepted, so removing them sacrifices real value (see SI Section S2 for the proof). 4 0.20.40.60.8 Baseline noiseĻ 0.1 0.2 0.5 1.0 2.0 4.0 8.0 Noise elasticity β noisier reviews noise growsfaster with load A 0.000.250.500.751.00 Paper qualityq 0.0 0.2 0.4 0.6 0.8 1.0 Lottery participation p ( q ) B Ļ= 0.4,β= 0.5 Ļ= 0.6,β= 2 Ļ= 0.7,β= 6 ā5 0 5 10 15 20 Quality gain (%) Figure 2: Under full adoption, lotteries help when noise is sufficiently high; the optimal rule adapts to the regime. (A) Quality gain (colour) when all scientists enter the lottery, as a function of baseline noise Ļ and noise elasticity β, with acceptance rate α= 10% and L= 0.20. The dashed contour marks the boundary where gain equals zero; above and to the right, the lottery improves quality. Markers indicate three representative parameter combinations shown in (B). (B) Socially optimal participation profiles p(q) at the three marked points: deeper into the region where lotteries help, the optimal threshold rises and more papers enter the lottery. See SI Section S3 for how the gain varies with acceptance rate. Figure 2B shows how the optimal threshold adapts: deeper into the region where lotteries help, more papers enter. The quality improvement grows with noise (Figure 2A), confirming that the mechanism is most effective when review noise is sufficiently high. 4 Voluntary participation We now ask whether scientists would voluntarily participate when acting in their own interest. Each scientist with paper quality q i chooses a lottery participation probability p i ā[0,1]. The journal, too, is a strategic player: by putting in more effort, it can lower the noise Ļ, but at an increasing marginal cost. The lottery rejection probability L is fixed by the venue exogenously (see SI Section S4 for alternative assumptions). Each scientistās utility combines a private publication benefit, an epistemic term, and a rejec- tion cost: U i = b(1ā L p i ) A(q i ) Ģ q+ s Ģ qā c q i 1ā(1ā L p i ) A(q i ) (2) where A(q i ) is the acceptance probability for a paper of quality q i , b is the private benefit of pub- lication, s captures epistemic motivation ā concern for the overall quality of published literature regardless of oneās own outcome ā and c is the cost of rejection, proportional to paper quality. To tractably analyse this finite-N game, we use a continuous approximation that treats the quality distribution as smooth but retains the finite impact of an individualās deviation on the aggregate reviewer burden (see SI Section S2 for the derivation). Since the payoff is linear in(b, s, c), equilibria are invariant under rescaling(b, s, c)ā(Ī»b,Ī»s,Ī»c) 5 for any Ī»> 0. Only two quantities affect strategic behaviour: the ratio r= b/(b+ s), which cap- tures the privateāepistemic balance, and the normalised cost c/(b+ s). We hold the normalised cost fixed throughout and vary r: when r= 1, scientists are purely self-interested; as r decreases, they internalise the collective benefit of better review quality, and the equilibrium shifts toward the social optimum. Following Zollman et al. [12], the journal chooses the review noise Ļ to maximise a payoff that trades off the quality of accepted papers against the cost of thorough review: Ī J = Ģ qā Ģ w(1+Ļ) āk (3) where k > 0 governs how steeply the cost of noise reduction increases as noise falls. Because the lottery removes papers before review, the total review cost scales with the surviving fraction Ģ w. The lottery thus operates through three channels: it shifts the quality composition of the reviewed pool, it reduces noise by lightening the reviewer burden, and it lowers the total cost of review. The Nash equilibrium is a joint fixed point: scientists choose p i given the journalās noise level, and the journal chooses Ļ given the scientistsā participation strategy. Figure 3 compares the Nash equilibrium to the social optimum. Self-interested scientists under-participate: the threshold below which, in equilibrium, scientists submit to the lottery is lower than the social optimum, and the quality gap between the two widens with noise (Fig- ure 3B). Yet even partial voluntary participation improves on the no-lottery baseline. How much of the optimal gain is realised depends on the balance between private and epistemic motivation. Figure 4 shows that rational scientists sustain non-trivial lottery participation in equilibrium, provided they are sufficiently epistemic in motivation. The equilibrium has a clean structure: each scientistās best response is all-or-nothing, and because the incentive to participate decreases with paper quality, the Nash equilibrium is a threshold rule ā all papers below a cutoff Ļ ā enter, all above opt out (see SI Section S2 for the proof). The threshold increases as the privateāepistemic ratio r decreases (i.e., as scientists place more weight on collective quality). The gap between the social optimum and the Nash equilibrium narrows as r decreases, illustrating that epistemic concern for the quality of science is a partial substitute for centralised coordination. Voluntary lotteries need no mandate, but they require scientists to see the quality of published science as partly their problem. Monte Carlo simulations confirm that the continuous approximation is conservative: finite populations produce higher participation thresholds than the analytical prediction, so more papers enter the lottery than the theory predicts (SI Table S1), because discrete stochastic effects systematically favour voluntary participation. 5 The AI scaling pressure Generative AI is lowering the cost of producing manuscripts faster than it is improving the ca- pacity to evaluate them [21]. As submission volume outpaces reviewer capacity, review noise Ļ rises. In our model, even a moderate increase in noise reduces accepted quality by 14%; a 6 0.00.20.40.60.81.0 Paper qualityq 0.0 0.2 0.4 0.6 0.8 1.0 Lottery participation p ( q ) A Nash Social optimum 0.20.40.60.8 Review noiseĻ 0.6 0.7 0.8 0.9 Average accepted quality Ģ q B Nash lottery Optimal lottery No lottery Figure 3: Self-interest limits voluntary participation relative to the social optimum. (A) Par- ticipation profiles p(q): the Nash equilibrium at r= 0.33 (solid) has a lower threshold than the social optimum (dashed); self-interested scientists under-participate. (B) Average accepted quality Ģ q as a function of review noise under three regimes: the scientistsā equilibrium lottery at r= 0.33 (solid red), the socially optimal lottery (dashed blue), and no lottery (dashed grey). Lines show the continuous approximation; markers show Monte Carlo simulation (N= 100). Each point shows the scientist-side equilibrium at a given noise level Ļ; varying Ļ reveals how the gap between equilibrium and optimal participation grows with review noise. Parameters: β= 8, L= 0.10, α= 10%. doubling of noise costs 22%. The lottery can recover much of this loss. At the noise elasticity used throughout our analysis (β= 8), the socially optimal lottery more than compensates for a doubling of noise, restoring quality to above its pre-increase level. The phase diagram (Figure 2) maps the full picture: the gain grows with both Ļ and β, meaning the worse the noise problem, the more the lottery helps. These percentages may understate the damage at highly selective venues. When a journal accepts 10% of submissions, most decisions fall near the noise margin; increased noise does not degrade all decisions equally but concentrates among borderline cases, swapping papers that should have been accepted for papers that should not. The evidence that submission growth is already outpacing review capacity is substantial. Large-language-model access cuts professional writing time by roughly 40% [21], more than one-fifth of computer-science preprints on arXiv show measurable evidence of large-language- model modification [4], and at least 13% of biomedical abstracts do the same [22]. Submis- sion volumes at major machine-learning venues have grown exponentially over the past decade, and this growth has coincided with widespread perceptions of arbitrary editorial decisions ā a pattern consistent with the quality degradation our model predicts. The NeurIPS consistency experiments [5, 7] show that review noise was already substantial before AI-driven accelera- tion. AI-driven submission growth is likely to push this further, and AI-assisted research may amplify exploitation of well-explored problems at the expense of genuinely novel inquiry [23], compounding the signal-to-noise problem that reviewers face. 7 0.00.20.40.60.81.0 Paper qualityq 0.0 0.2 0.4 0.6 0.8 1.0 Average participation p ( q ) A r= 0.67 r= 0.50 r= 0.33 0.20.40.60.8 Privateāepistemic ratior 0 20 40 60 80 100 Optimal quality gain captured (%) B Nash equilibrium Figure 4: Epistemic concern sustains voluntary participation and improves quality. (A) Equi- librium participation profiles p(q) for three values of the privateāepistemic ratio: r= 0.67 (two- thirds private), r= 0.50 (equal balance), and r= 0.33 (one-third private). Lines show the con- tinuous approximation; dots show Monte Carlo simulation averages (N= 100). Noise is fixed at Ļ= 0.3. Lower r (more epistemic) produces a higher participation threshold: more scientists enter the lottery when they weigh collective quality more heavily. (B) Fraction of the optimal quality gain captured by the Nash equilibrium, as a function of r. As epistemic concern increases (lower r), the Nash equilibrium closes the gap to the social optimum, rising from near zero at r= 0.9 to more than 80% at r= 0.10. Parameters: Ļ= 0.3, β= 8, L= 0.10, α= 10%, c= 0.1, b= 1. Higher noise pushes venues deeper into the region where the lottery helps, increasing the potential gain from collective action. These gains represent the ceiling on what any voluntary burden-reduction mechanism can achieve through coordinated participation. How much of that ceiling is realised depends on the epistemic concern r established in the previous section: the more scientists internalise the collective benefit, the closer the outcome approaches the optimum. Whether noise increases because of expanding research communities or because AI lowers pro- duction costs, the underlying problem is the same: reviewer overload, and the case for collective action strengthens as the gap between production and evaluation widens. 6 Discussion The voluntary lottery is deliberately simple ā its value lies less in the specific mechanism than in what it reveals about the potential for collective action in peer review. Our central finding is that if scientists could collectively reduce the review burden ā even through a mechanism as crude as random pre-review rejection ā the resulting improvement in review quality would more than compensate for the papers lost, provided the noise elasticity is sufficiently strong. Our model does not capture the self-screening feedback posited by Bergstrom and Gross [11], whereby noisier review attracts more marginal submissions; because this cycle is absent, the quality gains we report may be conservative. 8 In our model, the privateāepistemic ratio r determines how close voluntary participation comes to the social optimum: epistemic concern substitutes for centralised coordination [24, 25]. The lottery threshold shares a structural parallel with the self-screening mechanism of Zollman et al. [12], where submission costs drive bottom-up self-selection. Where Zollman et al. show that self-screening arises from imposed costs, we show that a similar quality stratification can be sustained voluntarily through epistemic concern ā without the welfare costs of imposed barriers. This voluntary-versus-imposed distinction matters in practice. Venues have already begun ex- perimenting with burden-reduction mechanisms: some philosophy journals restrict submissions to defined windows of the year [26], and computer science venues have introduced submis- sion fees [27]. While more structural reforms could address the scale problem directly, they face large coordination challenges and resistance. The lottery isolates this tradeoff and shows that imperfect adoption helps. A venue need not enforce optimal participation to benefit. Neither of the mechanismās requirements (sufficient noise elasticityβ and sufficient epistemic concern r) is a strong assumption. Review noise is a universal feature of selective venues [28, 29], and behavioural evidence consistently shows that people contribute to collective goods when they perceive the enterprise as worthwhile [24] ā a description that fits most scientistsā relationship to their field. The analysis requires only that some scientists care, not that all do: the SI shows the mechanism is robust to heterogeneous epistemic concern (SI Section S4). Because each scientistās influence on aggregate quality shrinks as the community grows, com- munity size matters. This is not a limitation of the mechanism but a reflection of a structural con- straint on collective action. Peer review is a commons [2]: every scientist benefits from rigorous evaluation whether or not they contribute to providing it. In small communities, reputation-based cooperation can sustain the commons, because each memberās behaviour is observable and de- fection carries reputational cost [30]. As the community grows, monitoring becomes infeasible, individual influence dilutes, and the incentive to free-ride dominates. Our model captures the dilution channel directly. At N= 100, voluntary participation yields a meaningful quality gain; at N= 500, this roughly halves. But even a modest increase in epistemic concern compensates: at N= 500, shifting from r= 0.5 to r= 0.3 roughly doubles the quality gain, recovering the level seen in smaller communities (SI Table S1). Without epistemic concern (r ā 1), the mechanism produces negligible gains regardless of community size. The breakdown of reputation-based co- operation at scale reinforces this prediction in practice. This pattern is characteristic of commons governance more broadly. Ostrom [31] identifies monitoring, graduated sanctions, and clearly defined boundaries as design principles for sustain- able collective management, conditions that hold naturally in specialised journals and workshops but break down at mega-conference scale, where reviewers are drawn from an anonymous global pool and no editor can track individual contributions. Ostromās eighth design principle (nested enterprises for larger-scale commons) suggests the structural response: federated review systems in which mega-venues decompose into community-scale tracks, each small enough to restore the conditions under which voluntary cooperation is self-sustaining. Some venues already approxi- 9 mate this through area-based review organisation; our results suggest that community structure, not just reviewer incentives, is key to sustaining review quality. The mechanismās requirements (sufficient noise elasticity and epistemic concern) are most likely met where the problem is most acute. Beygelzimer et al. [7] show that making a venue more selective increases the arbitrariness of its decisions, because the fraction of acceptances that fall within the review noise margin grows as the pool shrinks. Our lottery is most effective in exactly this regime: the venues where peer review is most arbitrary are those where voluntary collective action to reduce noise would yield the greatest gains. The peer review crisis, in other words, contains the seed of its own remedy: the worse the problem, the stronger the incentive for scientists who care about the quality of published science to act. Materials and Methods We model a scientific venue where scientists submit N papers of quality q i ā[0,1], an editor accepts the top fractionα based on noisy review scores, and review noise scales with the fraction of papers reaching review (Ļ eff = ĻĀ· Ģ w β ; see Section 2 for full details). Scientists may enter a voluntary pre-review lottery; the journal chooses review effort to maximise its own payoff. The Nash equilibrium is a joint fixed point of scientist and journal best responses. The relevant strategies are threshold strategies ā step functions where all scientists with q i ā¤ Ļ enter the lottery and those with q i > Ļ do not. In both the social optimum and the Nash equilibrium, the threshold structure arises because the cost of losing a paper grows with quality while the noise-reduction benefit does not (SI Section S2). Nash equilibria are identified by searching over all threshold strategies Ļā[0,1], computing the journalās best-response noise level, and verifying that no scientist can improve their payoff by unilateral deviation. Monte Carlo simulations use N= 100 scientists and M= 5,000 replications per parameter combination. The continuous approximation is evaluated on a 50-point quality grid via numerical integration. The complete formal model specification is given in SI Section S1; derivations of the acceptance probability, lottery mechanism, payoff functions, and Nash equilibrium algorithm follow in SI Section S2. Source code is available at https://github.com/juliangarcia/lottery-peer-review. References [1] MichĆØle Kovanis, RaphaĆ«l Porcher, Philippe Ravaud, and Ludovic Trinquart. The global bur- den of journal peer review in the biomedical literature: Strong imbalance in the collective enterprise. PLoS ONE, 11(11):e0166387, 2016. doi: 10.1371/journal.pone.0166387. [2] Michael E. Hochberg, Jonathan M. Chase, Nicholas J. Gotelli, Alan Hastings, and Shahid Naeem. The tragedy of the reviewer commons. Ecology Letters, 12(1):2ā4, 2009. doi: 10.1111/j.1461-0248.2008.01276.x. 10 [3] Chakkrit Tantithamthavorn, Nicole Novielli, Ayushi Rastogi, Olga Baysal, and Bram Adams. Blended PC peer review model: Process and reflection. In ACM SIGSOFT Software Engineer- ing Notes, 2025. doi: 10.1145/3735931.3735937. [4] Weixin Liang, Yaohui Zhang, Hancheng Cao, Binglu Wang, Daisy Yi Ding, Xinyu Yang, Kailas Vodrahalli, Siqi He, Daniel Scott Smith, Yian Yin, Daniel McFarland, and James Zou. Monitoring AI-modified content at scale: a case study on the impact of ChatGPT on AI conference peer reviews. Nature, 640:461ā469, 2025. doi: 10.1038/s41586-024-08520-w. [5] Eric Price. The NIPS experiment. Blog post, http://blog.mrtz.org/2014/12/15/ the-nips-experiment.html, 2014. Accessed 2026-03-30. [6] Corinna Cortes and Neil D. Lawrence. Inconsistency in conference peer review: Revisiting the 2014 NeurIPS experiment. arXiv preprint arXiv:2109.09774, 2021. [7] Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan. Has the machine learning review process become more arbitrary as the field has grown? The NeurIPS 2021 consistency experiment. arXiv preprint arXiv:2306.03262, 2023. [8] John Jerrim and Robert de Vries. Are peer-reviews of grant proposals reliable? An analysis of Economic and Social Research Council (ESRC) funding applications. The Social Science Journal, 60(1):91ā109, 2020. doi: 10.1080/03623319.2020.1728506. [9] Bryan D. Neff and Julian D. Olden. Is peer review a game of chance? BioScience, 56(4): 333ā340, 2006. doi: 10.1641/0006-3568(2006)56[333:IPRAGO]2.0.CO;2. [10] JĆ©rĆ“me Adda and Marco Ottaviani. Grantmaking, grading on a curve, and the paradox of relative evaluation in nonmarkets. Quarterly Journal of Economics, 139(2):1255ā1319, 2024. doi: 10.1093/qje/qjad056. [11] Carl T. Bergstrom and Kevin Gross. Screening, sorting, and the feedback cycles that imperil peer review. PLOS Biology, 24(2):e3003650, 2026. doi: 10.1371/journal.pbio.3003650. [12] Kevin J. S. Zollman, Julian Garcia, and Toby Handfield. Academic journals, incentives, and the quality of peer review: A model. Philosophy of Science, 91(1):186ā203, 2023. doi: 10.1017/psa.2023.132. [13] Leonid Tiokhin, Karthik Panchanathan, DaniĆ«l Lakens, Simine Vazire, Thomas Morgan, and Kevin Zollman. Honest signaling in academic publishing. PLoS ONE, 16(2):e0246675, 2021. doi: 10.1371/journal.pone.0246675. [14] Stefan Thurner and Rudolf Hanel. Peer-review in a world with rational scientists: Toward selection of the average. European Physical Journal B, 84(4):707ā711, 2011. doi: 10.1140/ epjb/e2011-20545-7. 11 [15] Yichi Zhang, Fang-Yi Yu, Grant Schoenebeck, and David Kempe. A system-level analysis of conference peer review. Proceedings of the 23rd ACM Conference on Economics and Compu- tation, pages 1041ā1080, 2022. doi: 10.1145/3490486.3538306. [16] Thomas Feliciani, Junwen Luo, Lai Ma, Pablo Lucas, Flaminio Squazzoni, Ana Marusic, et al. A scoping review of simulation models of peer review. Scientometrics, 121(1):555ā 594, 2019. doi: 10.1007/s11192-019-03205-w. [17] Ferric C. Fang and Arturo Casadevall. Research funding: The case for a modified lottery. mBio, 7(2):e00422ā16, 2016. doi: 10.1128/mBio.00422-16. [18] Kevin Gross and Carl T. Bergstrom. Contest models highlight inherent inefficiencies of sci- entific funding competitions. PLoS Biology, 17(1):e3000065, 2019. doi: 10.1371/journal. pbio.3000065. [19] David Card, Stefano DellaVigna, Patricia Funk, and Nagore Iriberri. What do editors max- imize? Evidence from four economics journals. Review of Economics and Statistics, 102(1): 195ā217, 2020. doi: 10.1162/rest_a_00839. [20] Emery D. Berger. Cs conference acceptance rates. https://github.com/emeryberger/ csconferences, 2025. Accessed: 2026-04-27. [21] Shakked Noy and Whitney Zhang. Experimental evidence on the productivity effects of generative artificial intelligence. Science, 381(6654):187ā192, 2023. doi: 10.1126/science. adh2586. [22] Dmitry Kobak, Rita GonzĆ”lez-MĆ”rquez, EmÅke-Ćgnes HorvĆ”t, and Jan Lause. Delving into ChatGPT usage in academic writing through excess vocabulary. PLoS ONE, 19(5): e0297826, 2024. doi: 10.1371/journal.pone.0297826. [23] Qianyue Hao, Fengli Xu, Yong Li, and James Evans. Artificial intelligence tools expand scientistsā impact but contract scienceās focus. Nature, 649(8099):1237ā1243, 2026. doi: 10.1038/s41586-025-09922-y. [24] Ernst Fehr and Urs Fischbacher. Why social preferences matter ā the impact of non-selfish motives on competition, cooperation and incentives. Economic Journal, 112(478):C1āC33, 2002. doi: 10.1111/1468-0297.00027. [25] Remco Heesen and Liam Kofi Bright. Is peer review a good idea? British Journal for the Philosophy of Science, 72(3):635ā663, 2021. doi: 10.1093/bjps/axz029. [26] Journals that close submissions part of the year ā The Philoso- phersā Cocoon, April 2026.URL https://web.archive.org/web/ 20260407075957/https://philosopherscocoon.com/2024/04/05/ journals-that-close-submissions-part-of-the-year/. 12 [27] IJCAI-ECAI 2026.Primary paper initiative. https://2026.ijcai.org/ primary-paper-initiative/, 2025. Announced November 2025. Accessed April 2026. [28] Rafael DāAndrea and James P. OāDwyer. Can editors save peer review from peer reviewers? PLoS ONE, 12(10):e0186111, 2017. doi: 10.1371/journal.pone.0186111. [29] Glenn Ellison. Evolving standards for academic publishing: A q-r theory. Journal of Political Economy, 110(5):994ā1034, 2002. doi: 10.1086/341871. [30] Michihiro Kandori. Social norms and community enforcement. Review of Economic Studies, 59(1):63ā80, 1992. doi: 10.2307/2297925. [31] Elinor Ostrom. Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press, 1990. doi: 10.1017/CBO9780511807763. 13 Supplementary Information: Can we volunteer out of the peer review crisis? Theo Tang, Toby Handfield, and Julian Garcia April 30, 2026 Contents S1 Complete model specification2 S1.1 Review process. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .2 S1.2 Lottery mechanism. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .2 S1.3 Payoffs and strategies. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .2 S1.4 Equilibria and socially optimal lotteries. . . . . . . . . . . . . . . . . . . . . . . . . . . . .3 S1.5 Monte Carlo simulation. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .3 S1.6 Parameters. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .3 S2 Analytical approximation4 S2.1 Acceptance probability. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .4 S2.2 Special cases. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .5 S2.3 Lottery mechanism derivations. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .5 S2.4 Payoff functions and Nash equilibrium. . . . . . . . . . . . . . . . . . . . . . . . . . . . .6 S2.5 Nash equilibria are threshold strategies. . . . . . . . . . . . . . . . . . . . . . . . . . . . .9 S2.6 Validation of the analytical approximation. . . . . . . . . . . . . . . . . . . . . . . . . . .11 S3 Empirical calibration of noise elasticity11 S3.1 Dataset and estimation ofĻ. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .12 S3.2 Functional form and calibratedβ. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .13 S4 Extensions and robustness15 S4.1 Sensitivity to the privateāepistemic ratio. . . . . . . . . . . . . . . . . . . . . . . . . . . .15 S4.2 Robustness to the quality distribution. . . . . . . . . . . . . . . . . . . . . . . . . . . . . .16 S4.3 Robustness to heterogeneous epistemic concern. . . . . . . . . . . . . . . . . . . . . . .16 S4.4 Strategic use of the lottery. . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . . .17 1 S1 Complete model specification This section provides a detailed description of the model summarised in the main text. We specify the review process and noise-scaling mechanism, define the lottery and its interaction with review quality, introduce the payoff functions for scientists and the journal, and describe the Monte Carlo simulation used to validate the analytical results. We then report convergence checks and theε-Nash characterisation of the equilibrium. S1.1 Review process A scientific venue receivesNsubmissions, one from each scientist. Each scientistihas a paper of true qualityq i drawn independently from a uniform distribution on[0,1]. The venue accepts a fixed fractionαof submissions. To decide which papers to accept, an editor observes a noisy signal of each paperās quality, X i ā¼TruncatedNormal(q i ,Ļ 2 eff ,[0,1]), ranks all papers by their signal, and accepts the topαN. The key feature of the review process is that noise increases with the reviewer burden. More papers per reviewer mean less time per assessment and a thinner pool of qualified referees. We model this as a power law: when a fraction Ģ wof submissions reach review, the effective noise is Ļ eff =ĻĀ· Ģ w β (S1) whereĻis the baseline noise (when all submissions are reviewed) andβ>0 is the noise elasticity ā the rate at which noise grows with the review burden. This functional form is empirically motivated by ICLR review data (see Section S3). When Ģ w=1 (all papers reviewed),Ļ eff =Ļ; any mechanism that reduces the fraction reaching review ( Ģ w<1) also reduces noise. S1.2 Lottery mechanism We now introduce a pre-review lottery that can reduce the number of papers reaching the editor. Each scientist chooses a participation probabilityp i ā[0,1]. A paper whose author choosesp i enters the lottery with probabilityp i ; if it enters, it is randomly rejected before review with probabilityLā(0,1), surviving with probability 1āL. A paper that does not enter the lottery (p i =0) proceeds directly to review. Each lottery entrant survives independently with probability 1āL, and each non-entrant survives with certainty. Because the lottery reduces the fraction of papers reaching review ( Ģ w<1), it also reducesĻ eff , improving review quality for the papers that remain. S1.3 Payoffs and strategies Each scientist with paper qualityq i choosesp i to maximise their individual payoff: U i =bĀ·1 published Ā· Ģ q+sĀ· Ģ qācĀ·q i Ā·(1ā1 published )(S2) where1 published indicates whether paperisurvives the lottery and is accepted by the editor, and Ģ qis the average quality of accepted papers. The three terms in the scientistsā utility capture distinct incentives: ā¢bā„0 weights the private benefit of publication, scaled by the quality of the venue; ā¢sā„0 capturesepistemic concernā concern for the overall quality of published science, enter- ing the payoff regardless of the scientistās own outcome; ā¢cā„0 is the rejection cost coefficient, scaled by paper qualityq i . 2 Whens=0, the scientist is purely self-interested; assgrows, the scientist internalises the collec- tive benefit of better review quality. The payoff is scale-invariant: only the privateāepistemic ratio r=b/(b+s)and the normalised costc/(b+s)affect equilibrium behaviour. Throughout the SI we work with the parameters ( b , s , c ) directly; the correspondence is r = 1 / ( 1 + s ) when b = 1. Each scientistās strategy is a choice ofp i ā[0,1]. While this is formally a continuous choice, the equilibrium best response is binary:p ā i ā0,1for all agents (Proposition 1). Furthermore, the set of agents who choose to enter (p i =1) always forms a lower intervali:q i ā¤Ļfor some threshold Ļ(Proposition2). Following Zollman et al.[1], the journal chooses the baseline review noiseĻto maximise: Ī J = Ģ qā Ģ w(1+Ļ) āk ,(S3) wherek>0 governs how steeply the cost of noise reduction increases as noise falls and Ģ wis the fraction of submissions surviving to review, so that the cost scales with the number of papers actually reviewed. S1.4 Equilibria and socially optimal lotteries A Nash equilibrium is a joint fixed point: scientists choosep i given the journalās noise level, and the journal choosesĻgiven the scientistsā participation strategy. Details of the equilibrium algorithm we use to identify equilibria are given in Section S2.4. The social optimum is defined by the thresholdĻthat maximises the average quality of accepted papers: Ļ ā =argmax Ļ Ģ q(Ļ),(S4) where Ģ q(Ļ)=E[ 1 T ā jāaccepted q j ]andTis the number of accepted papers, estimated by Monte Carlo simulation overMindependent replications. S1.5 Monte Carlo simulation For the finite stochastic game, we estimate payoffs via Monte Carlo simulation: for each strategy pro- file,Mindependent replications of the review process are drawn, and expected payoffs are computed as sample averages. Nash equilibria are identified by searching over threshold strategies and veri- fying via deviation checks ā for each candidate threshold, we check whether any agent gains from unilateral deviation. Best-response dynamics provide a complementary solution method: starting from an initial strategy profile, agents sequentially update to their best response until convergence. S1.6 Parameters The model has three groups of exogenous parameters: venue characteristics (α,L,β), payoff weights (b,s,c), and the journalās cost parameterk. The endogenous variables ā each scien- tistās participationp i , the journalās noise choiceĻ, and the resulting quality of accepted papers Ģ q ā are determined in equilibrium. Unless otherwise stated, the following default values are used throughout: 3 Parameter Symbol Default value Population sizeN100 Acceptance rateα10% Baseline noise (reference value)Ļ0.3 Noise elasticityβ8 Lottery rejection probabilityL0.10 Journal review cost parameterk8 Publication benefitb1 Rejection costc0.1 Epistemic concerns1 MC replicationsM5,000 We setβ=8 as an illustrative, high-elasticity parameter; while empirical data from a single confer- ence venue suggests lower baseline elasticities in practice (see SectionS3), using an amplified value clearly demonstrates the mechanismās structural dynamics and boundary conditions. S2 Analytical approximation We use a continuous approximation that replaces the discrete ranking problem with a continuous threshold problem with tractable closed-form expressions for acceptance probabilities, quality, and payoffs. This section derives these expressions and proves the threshold structure of Nash equi- libria. The continuous model provides structural insight (proving threshold existence, corner best responses, and monotonicity), while Monte Carlo simulation serves as quantitative ground truth. S2.1 Acceptance probability The simulation in SectionS1uses a finite population ofNscientists. In this section we derive an analytical expression for the acceptance probability under the continuous approximation. The key insight is that the acceptance rateαdoes not vanish asNgrows, because the venue accepts a fixed fractionof submissions: both the number of papers and the number of acceptances scale withN. Under the continuous approximation, the discrete ranking ofNpapers is replaced by a continu- ous threshold problem in which each paper competes against a population distribution of perceived qualities rather than a finite set of competitors. Figure S1confirms that this approximation is accu- rate even for moderateN. Each paper has true qualityqdrawn fromF(q)with densityf(q)on[0,1]. The editor observes a noisy signalX q ā¼TruncatedNormal(q,Ļ 2 ,[0,1])with density: g(x|q)= Ļ xāq Ļ Ļ Ā Ī¦ Ā 1āq Ļ Ā āΦ āq Ļ Ā (S5) whereĻandΦare the standard normal PDF and CDF, respectively. The perceived quality of a randomly drawn paper follows the mixture distribution with marginal density: h(y)= ā« 1 0 g(y|q)f(q)dq.(S6) The acceptance thresholdy ā is the(1āα)quantile of this distribution: ā« y ā 0 h(y)d y=1āα.(S7) 4 0.00.20.40.60.81.0 Paper qualityq 0.0 0.2 0.4 0.6 0.8 1.0 Acceptance probability SimulationN= 20 SimulationN= 50 SimulationN= 500 Continuous approx. Figure S1:Continuous approximation matches simulation.Acceptance probability as a function of paper quality: Monte Carlo simulation (markers,N=20,50,500) versus the continuous approx- imation (solid black line). Agreement improves asNincreases, confirming that the approximation is accurate for moderate populations. A paper of true qualityqis accepted whenever its noisy signal exceeds this threshold, giving the acceptance probability: P(accept|q)= ā« 1 y ā g(x|q)dx=1āG(y ā |q)(S8) whereG(Ā·|q)is the CDF of TruncatedNormal(q,Ļ 2 ,[0,1]). In practice, we computeh(y)by numerical integration overq, findy ā by standard numerical root-finding on the quantile equation, and then evaluateP(accept|q)for each quality level. S2.2 Special cases Two limiting cases build intuition for the acceptance probability. When noise vanishes (Ļā0) and qualities are uniformly distributed (f(q)=1), the editor observes true quality perfectly. Acceptance reduces to a deterministic cutoff: every paper above quality 1āαis accepted and every paper below is rejected: y ā =1āα,(S9) P(accept|q)= Ģ 1 ifq>1āα, 0 ifqā¤1āα. (S10) When noise is small but positive, this sharp step function softens into a smooth sigmoid. The threshold remains neary ā ā1āα, but papers close to the boundary now have intermediate accep- tance probabilities: P(accept|q)ā1āΦ Ā y ā āq Ģ Ļ Ā (S11) where the effective width Ģ Ļaccounts for truncation effects. Larger noise means a wider band of uncertainty around the cutoff. S2.3 Lottery mechanism derivations Each submission survives the lottery with probability: w(q)=1āL p(q)(S12) 5 wherep(q)ā[0,1]is the participation level andLā[0,1]is the lottery rejection probability. The average survivor fraction is: Ģ w= ā« 1 0 w(q)f(q)dq.(S13) The overall lottery participation rate is: Ģ p= ā« 1 0 p(q)f(q)dq.(S14) The effective noise decreases with the fraction of submissions removed by the lottery: Ļ eff =ĻĀ· Ģ w β (S15) whereβ>0 is the noise elasticity ā the rate at which noise increases with load. Since Ģ wā¤1, the lottery can only reduce noise. Acceptance among survivors.The lottery biases the reviewed pool. The acceptance rate among survivors is: α sur = α Ģ w .(S16) The performance density of a randomly selected survivor is: h sur (y)= ā« 1 0 g(y|q) w(q)f(q) Ģ w dq.(S17) The acceptance threshold among survivors satisfies: ā« y ā 0 h sur (y)d y=1āα sur .(S18) Average quality of accepted papers.The average quality of accepted papers is: Ģ q= ā« 1 0 q w(q)f(q)[1āG(y ā |q)]dq ā« 1 0 w(q)f(q)[1āG(y ā |q)]dq .(S19) S2.4 Payoff functions and Nash equilibrium Journal payoff and best response.The journalās payoff trades off the quality of accepted papers against the cost of thorough review: Ī J = Ģ qā Ģ w(1+Ļ) āk (S20) wherek>0 governs how steeply the cost of noise reduction increases as noise falls and the factor Ģ wreflects the per-paper nature of review costs. Higherkmakes thorough review more expensive, pushing the journal toward tolerating more noise at equilibrium. The Ģ wfactor in the cost term is not redundant with the noise-reduction channelĻ eff =ĻĀ· Ģ w β . The two operate on different parts of the payoff: noise reduction is a demand-side effect (fewer papers improve each review), while cost scaling is a supply-side effect (fewer papers require less total review effort). Under this interpretation, the lottery benefits the journal through both channels simultaneously. Given a scientist participation strategyp(q), the journalās best response is theĻthat maximises Ī J . 6 0.00.20.40.60.81.0 Review noiseĻ 0.0 0.2 0.4 0.6 0.8 Journal payoff Full lottery adoption (p= 1) No lottery (p= 0) Figure S2:The lottery raises the journalās peak payoff and shifts its optimum to higher noise. Journal payoff as a function of review noiseĻunder no lottery (p=0, dashed) and full lottery adoption (p=1, solid), withk=8,α=10%,L=0.10,β=8. Markers indicate the payoff- maximising noise levelĻ ā under each regime. Because the lottery absorbs part of the submission load before review, the journal can tolerate a higher noise level at its optimum (Ļ ā =0.43 vs. 0.29) while achieving a higher peak payoff (0.80 vs. 0.64). The shaded region shows the payoff gain from the lottery across the full noise range. Scientist payoff.Replacing the indicator function in Equation (S2), with the lottery survival and acceptance probability gives: U(q)=b(1āLp)A(q) Ģ q+s Ģ qāc q 1ā(1āLp)A(q) (S21) whereA(q)=1āG(y ā |q)is the acceptance probability conditional on surviving the lottery. The three terms are: ā¢b(1āLp)A(q) Ģ qā publication benefit; ā¢s Ģ qā epistemic concern; ā¢c q[1ā(1āLp)A(q)]ā rejection cost. Deviation payoff.To check whether a strategy profilep(q)constitutes a Nash equilibrium, we com- pute the deviation payoff: what happens when a single scientist of qualityqunilaterally changes their participation fromp(q)to some alternativep ā² , while all other scientists maintainp(Ā·). The de- viatorās change affects the aggregate survival fraction Ģ wand thereforeĻ eff , the acceptance threshold, and Ģ q. The strategyp(q)is a best response if no unilateral deviation improves the scientistās payoff. In a finite venue, opting out also creates a direct replacement effect by freeing an acceptance slot for a marginal paper. However, numerical analysis shows this direct effect is dominated by the indirect noise-reduction channel by a factor of 10 to 100, because it requires a paper to have both a high baseline acceptance probability and a quality far below the threshold; these conditions rarely overlap. For analytical clarity, our continuous approximation isolates the dominant noise-reduction channel. 7 Social optimum.The social optimum maximises Ģ q, the average quality of accepted papers. The optimal policy is a threshold rule: papers withqā¤Ļenter the lottery, while higher-quality papers do not. To see why, note that removing a paper of qualityqreduces noise (benefiting all survivors) but also removes its contribution to the accepted pool. The noise-reduction benefit is independent of which paper is removed, while the cost of removal increases inq, so the optimal rule removes papers in order of increasing quality, yielding a threshold. The boundary case atĻ=1 is illustrative. Here, full adoption is quality-blind, so noise falls to Ļ(1āL) β but the acceptance rate among survivors rises toα/(1āL). AtĻ=0 the second effect dominates: Ģ q full =1āα/[2(1āL)]<1āα/2 ā full lottery is worse than no lottery. Epistemic vs welfare social optimum.The social optimum maximises published quality, rather than the welfare of scientists. Under threshold strategies and uniform quality on[0,1], the two objectives are equivalent. To see this, note that aggregate welfare is W= ā« 1 0 U(q)dq=(bα+s+cα) Ģ qā c 2 , a strictly increasing affine function of Ģ q, so maximising welfare and maximising quality select the sameĻ ā regardless ofb,s,c,Ļ, orβ. The intuition is that the lottery reallocates quality across the accept/reject split, but cannot change either the total quality (fixed at ā« 1 0 q dq=1/2) or the accepted mass (fixed atαby venue capacity). Accepted-quality is thereforeα Ģ qand rejected-quality is 1/2āα Ģ q. Welfare combines three terms: a publication benefit on accepted quality (bα Ģ q), an epistemic gain in venue quality (s Ģ q), and savings on rejection cost (cα Ģ q). All three terms grow with Ģ q. Nash equilibrium algorithm.The Nash equilibrium is a joint fixed point of two best-response mappings: scientists choosep(q)given the journalās noise levelĻ, and the journal choosesĻgiven p ( q ) . We solve this by searching over a grid of candidate Ļ values and, for each, computing the scientistsā fixed point. Scientistsā fixed point.GivenĻ, the scientistsā equilibrium participation functionp ā (q)is com- puted by relaxation-based fixed-point iteration: 1.Initialisep (0) (q)on a discretised quality grid (Npoints on[0,1]). 2.Compute the best response for eachqby grid search overp ā² ā0,1/M,2/M,...,1: BR (n) (q)=argmax p ā² U(q;p ā² ,p (n) āq ). 3.Update via squared-loss relaxation: p (n+1) (q)=p (n) (q)āĪ·Ā·2 p (n) (q)āBR (n) (q) (S22) whereĪ·=0.3 is the learning rate, followed by clipping to[0,1]. By Proposition 1, the best response is a corner solution and the convex update preserves[0,1], so in exact arithmetic the grid search and clipping are redundant. We retain both as numerical safeguards. 4.Repeat until max q |p (n+1) (q)āp (n) (q)|<Ī“. Joint equilibrium.For each candidateĻon the grid, we compute the scientistsā fixed pointp ā Ļ (q), then check whetherĻis the journalās best response givenp ā Ļ (q). A candidate is accepted as an equilibrium if the journalās best-responseĻ ā satisfies|Ļ ā āĻ|<Ī“ Ļ for a toleranceĪ“ Ļ =10 ā6 . Among all accepted candidates, we select the one with the highest journal payoff. 8 Convergence.In all parameter configurations tested, the relaxation scheme converges within 50 iterations to a final error below 10 ā3 . The squared-loss convergence history is monotonically de- creasing in typical runs. S2.5 Nash equilibria are threshold strategies The socially optimal policy is provably a threshold rule (see āSocial optimumā in SectionS2). Here we show that Nash equilibria also take this form: all scientists with papers below some quality cutoffĻ ā enter the lottery, and all scientists aboveĻ ā do not. The argument has two steps. First, each scientistās best response is all-or-nothing: enter fully or not at all. Second, the incentive to enter the lottery decreases with paper quality, producing a clean threshold. Each scientistās decision is all-or-nothing LetU i (p i )denote scientistiās payoff as a function of their own participation, holding all other sci- entistsā strategies fixed. We decomposeU i (p i )=D(p i )+R(p i ), whereDcaptures thedirect effectof participation on scientistiās own payoff andRtheindirect effectmediated through the population aggregates. The direct effect is D(p i )=āL A(q i )(b Ģ q+c q i )p i +const., whereA(q i )is the acceptance probability conditional on passing the lottery stage and Ģ qis the average quality of accepted papers.Dis linear inp i , soD ā² is a constant andD ā² =0. The indirect effectR enters the payoff through the population average Ģ w= 1 N ā j (1āLp j ), soā Ģ w/āp i =āL/N. Proposition 1(Corner best response).Whenever(1+s)N, where s is the epistemic-concern pa- rameter, the direct effect dominates: U i is strictly monotone on[0,1]and the best response satisfies p ā i ā0,1. Sketch.A linear function on[0,1]has no interior maximumāit is maximised at one of the endpoints. We show thatU i (p i )is approximately linear for largeN, so the best response is always a corner. For any reference pointPā[0,1], expandU i (p i )aroundP: U i (p i )=U i (P)+U ā² i (P)(p i āP)+ 1 2 U ā² i (ξ)(p i āP) 2 , for someξbetweenPandp i . The first two terms are linear inp i , so the question reduces to whether the quadratic correction is negligible (i.e., whether|U ā² i |is small). SinceDis linear inp i ,D ā² =0 exactly, soU ā² i =R ā² . The indirect effectRdepends onp i only through the population average Ģ w, so we can writeR(p i )=R( Ģ w(p i )). Using the chain rule to find the second derivative we find: R ā² = ā 2 R ā Ģ w 2 Ā ā Ģ w āp i Ā 2 + āR ā Ģ w ā 2 Ģ w āp 2 i . The second term vanishes because Ģ wis linear inp i . This leavesR ā² =(ā 2 R/ā Ģ w 2 )(ā Ģ w/āp i ) 2 . We bound each factor in turn. The first,ā 2 R/ā Ģ w 2 , measures how fast the indirect effectās slope changes with Ģ w. The indirect channel enters the payoff through mean accepted quality Ģ q, which carries combined weightb+s=1+s, soā 2 R/ā Ģ w 2 is bounded by a constant times(1+s). The second factor is(ā Ģ w/āp i ) 2 =L 2 /N 2 āthe āone scientist amongNā scaling. Multiplying: |U ā² i |=|R ā² |ā¤C(1+s) L 2 N 2 for some constantC. This bound holds for allp i ā[0,1], so it holds atξregardless of whereξfalls. Since|p i āP|ā¤1, the quadratic term is at most 1 2 C(1+s)L 2 /N 2 , which is negligible whenever (1+s)N. The payoff is approximately linear, and the best response lies at a corner:p ā i ā0,1. 9 SinceU i is approximately linear, its slope is approximately constantāwe only need the sign of U ā² i (P)=D ā² +R ā² (P). WhenA(q i )>0,D ā² <0 dominates and the payoff is decreasing:p ā i =0 (do not enter). WhenA(q i )ā0,D ā² ā0 andR ā² >0 dominates:p ā i =1 (enter). At the threshold quality Ļ ā , the two payoffs are equal by continuity, and the threshold is determined by whereā(q)crosses zero (Proposition2). The incentive to enter decreases with paper quality By Proposition1, every scientist plays eitherp=1 orp=0. Define the net incentive to enter for a scientist with paper qualityqas ā(q)=U(q,p=1,p āq )āU(q,p=0,p āq ). Proposition 2(Unique threshold).Under the conditions of Proposition1,ā(q)is continuous and strictly decreasing in q, withā(0)>0andā(1)<0. Consequently there is a unique threshold Ļ ā ā(0,1)such that all scientists with q<Ļ ā enter the lottery and all with q>Ļ ā do not. Sketch.We show (i)āis strictly decreasing on[0,1]and (i)ā(0)>0,ā(1)<0; continuity then gives a unique zero crossing. (i) Strict decrease.The dominant term ofāis the direct personal cost of entering, ā direct (q)=āL A(q)(b Ģ q+c q), whose derivative is dā direct dq =āL A ā² (q)(b Ģ q+c q)+A(q)c <0, strictly negative and of order one (independent ofN). Two mechanisms drive the decrease: bet- ter papers are more likely to be accepted (A ā² (q)>0), so they lose more from random pre-review rejection; and the rejection costc qgrows with paper quality. By the same chain-rule argument used in Proposition1, aggregate shifts in( Ģ w,Ļ eff ,y ā , Ģ q)con- tribute corrections todā/dqof sizeO((1+s)/N), small relative to the order-one dominant term at the paperās parameters. Hencedā/dq<0 throughout[0,1]. (i) Boundary values.Atq=0, the acceptance probabilityA(0)is essentially zero under any realistic noise level, so the direct entry cost vanishes. The aggregate benefit remains positive: a lowest-quality paper that enters the lottery and is randomly rejected reduces the noise in the review pool and slightly improves mean accepted quality. Henceā(0)>0. Atq=1, the direct cost L A(1)(b Ģ q+c)is substantial ā a top-quality paper is near-certain to be accepted without the lottery ā and dominates the aggregate benefit of sizeO(s/N). Henceā(1)<0. āis continuous, strictly decreasing, positive atq=0 and negative atq=1, so it crosses zero exactly once at someĻ ā ā(0,1). Robustness to the cost specification.The threshold structure does not depend on the quality- proportional costc q. With a flat rejection costc(independent ofq), the direct term becomes ā direct (q)=āL A(q)(b Ģ q+c), whose derivative isāL A ā² (q)(b Ģ q+c)<0 ā still strictly decreas- ing. TheA(q)channel alone suffices for monotonicity; the quality-proportional cost strengthens the result but is not required. Threshold persistence across population sizes A natural concern about Propositions1and2is whether the equilibrium threshold survives at larger populations. AsNgrows, each scientistās individual influence on the aggregate is of order 1/N, which might suggest that the incentive to enter vanishes and the equilibrium collapses toĻ ā =0. 10 We test this directly by running the Nash equilibrium search atNā20,50,100,200,500. TableS1reports the Monte Carlo equilibrium threshold, the analytical prediction, and the quality gain over the no-lottery baseline. The threshold decreases withNbut remains substantial across the tested range, and the quality gain stays positive at every population size. NĻ ā MC Ļ ā cont ā Ģ q MC 20 0.63 0.53 0.045 50 0.51 0.41 0.040 100 0.46 0.32 0.030 200 0.33 0.24 0.024 500 0.31 0.10 0.016 Table S1:Threshold persistence across population sizes.Monte Carlo equilibrium thresholdĻ ā MC , continuous approximationĻ ā cont , and quality gainā Ģ q MC over the no-lottery baseline. Parameters: Ļ=0.3,β=8,L=0.05,α=10%,s=1,c=0.1,M=5,000. Two features of the data deserve comment. First, the MC threshold sits consistently above the continuous approximation, and the gap widens withN: atN=500,Ļ ā MC ā0.31 whileĻ ā cont ā0.10. The continuous approximation is therefore a conservative lower bound on cooperation: by evalu- ating payoffs at the expected aggregate composition rather than averaging over discrete stochastic realisations, it smooths away the variance that systematically boosts the epistemic payoff of vol- unteering (a Jensenās-inequality effect that favours participation in finite populations). Second, the quality gainā Ģ qalso shrinks withN, fromā¼0.045 atN=20 toā¼0.016 atN=500. The mechanism therefore retains its qualitative character at large populations but becomes quantitatively weaker as individual influence dilutes. We also verified threshold structure empirically across motivation regimes: unrestricted best- response dynamics from 36 random starting conditions (N=100, three privateāepistemic ratios) all converge to threshold-shaped profiles with zero ordering violations. S2.6 Validation of the analytical approximation The continuous approximation shows close agreement with finite-population Monte Carlo simula- tion across the full noise range, for both the Nash equilibrium and the social optimum (FigureS3). To further characterise the accuracy of the continuous approximation at finiteN, TableS2reports the optimal thresholdĻ ā and quality gainā Ģ qacross several population sizes. Both converge asN increases, and the analytical predictions are accurate even at moderateN. ε-Nash characterisation.The analytical thresholdĻ ā is not an exact Nash equilibrium of the finite stochastic game (the maximum deviation gain is positive under Monte Carlo estimation), but it is an ε-Nash equilibrium[2]withεā0.003. This means no agent can gain more than 0.3% in expected payoff by unilateral deviation. S3 Empirical calibration of noise elasticity The power-law noise modelĻ eff =ĻĀ· Ģ w β requires two empirical inputs: the baseline noiseĻand the noise elasticityβ. Calibrating either number against real venues is hard because peer-review scores are almost never released: most journals and conferences treat reviewer ratings as confidential, and the published experimental studies[ 3,4]cover only a single venue-year with no variation inN. Machine-learning conferences on OpenReview are a rare exception, and ICLR in particular has over a decade of fully anonymised scores spanning a wide range of conference sizes, making it the best available setting for calibrating a size-dependent noise model. The numbers reported below should 11 0.10.20.30.40.50.60.70.8 Review noiseĻ 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 Ģ q (Nash) A r= 0.67 r= 0.50 r= 0.33 0.10.20.30.40.50.60.70.8 Review noiseĻ 0.55 0.60 0.65 0.70 0.75 0.80 0.85 0.90 0.95 Ģ q (social optimum) B Socially optimal lottery No lottery Figure S3:Continuous approximation matches finite-population Monte Carlo simulation.Av- erage quality of accepted papers Ģ q as a function of review noise Ļ . Lines show the continuous approximation; markers show Monte Carlo simulation (N=100,M=5,000).(A)Nash equilib- rium for three privateāepistemic ratios (r=0.67, 0.50, 0.33, corresponding tos=0.5, 1, 2). Higher epistemic concern (lowerr) sustains larger quality gains across all noise levels; the gains narrow and disappear beyondĻā0.5ā0.7, where the equilibrium collapses to no participation.(B)Social optimum and the no-lottery baseline. Continuous approx.Monte Carlo NĻ ā Ģ qā Ģ qĻ ā Ģ qā Ģ q 20 0.947 0.8578 0.0747 1.000 0.8448 0.0601 30 0.966 0.8510 0.0746 1.000 0.8406 0.0659 50 0.980 0.8455 0.0744 0.959 0.8394 0.0698 75 0.986 0.8427 0.0743 1.000 0.8396 0.0691 100 0.990 0.8413 0.0742 0.980 0.8380 0.0715 Table S2:Grid convergence: optimal threshold and quality gain.Continuous approximation versus Monte Carlo simulation for several population sizesN. Parameters:Ļ=0.3,β=8,L=0.10, α=10%,M=5,000. therefore be read as illustrative of what the mechanism looks like at one unusually open venue, not as venue-agnostic constants. S3.1 Dataset and estimation ofĻ Data source.We use anonymised review scores from ICLR submissions (2017ā2025), obtained from the ICLR dataset of GonzĆ”lez-MĆ”rquez and Kobak[5], a complete scrape of the OpenReview platform. 1 ICLR uses an ordinal 1ā10 score and assigns at least three reviewers per paper; the mean number of reviews per paper rises from 3.0 in 2017 to 4.1 in 2025 (TableS3). After restricting to papers withnā„2 scores, the panel contains roughly 35,000 papers across nine conference years. Mapping reviewer disagreement onto the model.The model in SectionS1.1treatsĻas the standard deviation of the editorāsaggregatequality signalX i , not of any individual reviewer. We therefore estimateĻin three steps: 1 Data and processing pipeline athttps://github.com/berenslab/iclr-dataset. 12 YearN Ģ n Ė Ļ rev Ė Ļ95% CI 2017489 3.1 0.096 0.055[0.052, 0.058] 2018967 3.0 0.111 0.064[0.061, 0.067] 2019 1,554 3.0 0.104 0.060[0.058, 0.061] 2020 2,561 3.0 0.155 0.089[0.087, 0.092] 2021 2,974 3.9 0.112 0.057[0.056, 0.058] 2022 3,375 3.9 0.131 0.067[0.066, 0.068] 2023 4,919 3.8 0.136 0.070[0.069, 0.072] 2024 7,261 3.9 0.134 0.069[0.068, 0.070] 2025 11,513 4.1 0.132 0.066[0.065, 0.067] Table S3:Per-year estimates of review noise from ICLR data.Nis the number of papers with nā„2 scores; Ģ nis the mean number of reviews per paper; Ė Ļ rev is the average within-paper standard deviation on the rescaled[0,1]scale; Ė Ļ= Ė Ļ rev / p Ģ nis the implied noise of the aggregate review signal. The 2020 conference (greyed in figures) is excluded from subsequent fits as a COVID-era outlier: a sudden expansion of the reviewer pool and the rapid switch to remote review elevated within-paper disagreement well above trend. 1.Rescaleeach raw score to the modelās quality scale, Ģ s p,r =(s p,r ā1)/9ā[0,1], wheres p,r is reviewerrās score for paperp. 2.Within-paper noise.For each paper withn p ā„2 reviews, compute the sample standard de- viation Ė Ļ rev,p =sd( Ģ s p,1 ,..., Ģ s p,n p ). This estimates the noise of a single reviewerās score around the paperās true quality. 3.Aggregate-signal noise.Convert to the noise of the mean ofn p reviews, Ė Ļ p = Ė Ļ rev,p / p n p (the usual averaging argument: averagingnindependent noisy scores produces a mean with p ntimes less noise). This is the standard error the editor would face when ranking by the mean review score, and it is the quantity that maps onto the modelāsĻ. Concretely, for 2025 the year-level averages are Ė Ļ rev,y =0.132 and Ģ n y =4.1, giving Ė Ļ y =0.132/ p 4.1ā 0.065 (Table S3). Because ICLR reviews every submission ( Ģ w=1), the resulting estimate is of the baselinenoiseĻinĻ eff =Ļ Ģ w β , not of some post-interventionĻ eff . We then average the per-paper Ė Ļ p across papers within each year to obtain the annual estimate Ė Ļ y , reporting a bootstrap 95% con- fidence interval (2,000 resamples); averaging (rather than pooling all scores) keeps within-paper reviewer disagreement separated from between-paper quality variation. Averaging Ė Ļ y across years (excluding 2020) yields the headline calibrationĻā0.065. Year-level estimates are reported in Table S3and lie in the narrow range[0.055,0.070]once 2020 is excluded. Statistical modelling assumptions.The estimator rests on four standard assumptions: (i) review- ersā rescaled scores are conditionally independent noisy estimates of paper quality; (i) the ordinal 1ā10 scale can be treated as cardinal; (i) noise is homoscedastic across the quality range within a year; and (iv) reviewer-paper assignment is exogenous. (i)ā(i) are textbook measurement as- sumptions; the within-paper reviewer disagreement we measure is in line with prior reports for ML venues[ 3,4]. (iv) is mildly violated by ICLRās bidding system, which probably biases Ė Ļdownward (expert reviewers disagree less), making our calibration a lower bound on the noise a generic editor would face. S3.2 Functional form and calibratedβ To estimateβwe fit four candidate forms relating the conference sizeNto the year-level noiseĻ(N), wherea, Ė b,N 0 , andĻ max are fitted constants: 13 1.Power law:Ļ(N)=aĀ·N Ė b 2.Logarithmic:Ļ(N)=a+ Ė blogN 3.Saturating:Ļ(N)=Ļ max (1āe āN/N 0 ) 4.Linear:Ļ(N)=a+ Ė bN By AIC, the power-law and logarithmic forms are nearly tied and both substantially outperform the saturating and linear alternatives. The data alone therefore do not select between the two top- ranked forms; the choice has to be made on structural grounds. Why we adopt the power law.The power-law and logarithmic forms fit the data equally well (tied by AIC), so the choice has to be made on what the parameterβmeansacross venues. Only the power law makes the lotteryās relative effect a function of the reduction rate Ģ walone: Ļ(N Ģ w) Ļ(N) = a(N Ģ w) Ė b a N Ė b = Ģ w Ė b ,(S23) the same fractional drop at every venue, regardless of raw sizeNor baselineĻ. The logarithmic form givesĻ(N Ģ w)āĻ(N)= Ė blog Ģ w: an absolute reduction identical for a venue withĻ=0.50 and one withĻ=0.05, which is implausible. A calibration under the power law therefore transports across venues without recalibration; under the log form, Ė bwould be a property of the specific venue we fit on. Calibrated estimate.Taking logs ofĻ(N)=aN β gives logĻ=loga+βlogN, a straight line in logālog coordinates whose slope isβ. We therefore estimateβby ordinary least squares of log Ė Ļ y on logN y , which is consistent with the multiplicative scaling argued for above. The fit yields Ė Ī²ā0.06 (R 2 =0.50, AIC=ā62.7). Including 2020 leaves Ė Ī²essentially unchanged at 0.058 but collapsesR 2 to 0.16, reflecting the atypical within-paper disagreement in that COVID-era conference; the point estimate is thus robust to the exclusion. The fit implies a noise-reduction factor 0.90 0.06 ā0.99 for a 10% reduction in the reviewed pool. For each(Ļ,α)pair, the gain from a full-participation lottery crosses zero at a critical valueβ ā ; Table S4reportsβ ā across several combinations. At the ICLR calibration (Ļ=0.065,α=0.10),β ā ā0.99; the threshold falls sharply with baseline noise, reachingβ ā ā0.08 atĻ=0.5. Across theĻrange shown,β ā > Ė Ī², but the gap narrows from sixteenfold at ICLR to within a factor of two at high noise. A venueās position relative to the feasibility boundary therefore depends onĻas much as onβā the joint dependence that the phase diagram in the main text visualises directly. Interpretation. Ė Ī²is the slope of aggregate review noise against submission volumeacrossICLR years ā a descriptive quantity that tells us how noise changed as the venue grew. The lottery mechanism responds to a different quantity: the within-year causal effect of reducing load on a fixed reviewer pool. These differ because the mean number of reviews per paper rose from 3.0 in 2017 to 4.1 in 2025 (Table S3), which mechanically dampens the across-year slope through the 1/ p nfactor in Ė Ļ. Other changes at the conference over this period may push in the same direction, but we cannot document them from data alone. The headline Ė Ī²ā0.06 should therefore be read as a lower bound on the modelāsβ: the calibration confirms the direction of the load-sensitivity channel (review noise scales with reviewer burden) without pinning down its magnitude, and reflects a single conference venue rather than a universal calibration target. The mechanism requiresβabove approximately 1 to activate (Table S4); our illustrativeβ=8 is well above this threshold. Because the model functions as an illustrative thought experiment rather than a literal policy projection, our strategic simulations use this high-elasticity parameter to isolate the theoretical mechanics of the pre-review lottery, making the systemās phase transitions and collective action trade-offs structurally visible. 14 ĻAcceptance rateαβ ā (critical) 0.065 10%0.99 0.065 20%>20 0.065 32%>20 0.110%0.48 0.210%0.20 0.310%0.12 0.332%0.49 0.510%0.08 Table S4:Critical noise elasticity.Minimumβfor which the full-participation lottery (L=0.10) produces positive quality gains over the no-lottery baseline, across representative combinations of baseline noiseĻand acceptance rateα. The threshold falls rapidly with bothĻand selectivity; at moderate noise (Ļā„0.2) and low acceptance rates, smallβsuffice. Table S5:Lottery participation decreases as scientists become more privately oriented.Nash equilibrium as a function of the privateāepistemic ratior=b/(b+s), withc=0.1 held fixed. Other parameters:N=50,Ļ=0.3,β=8,L=0.10,α=10%. r=b/(b+s)Ļ ā n in Ģ q Nash 0.10.71 37 0.823 0.20.63 33 0.819 0.30.59 28 0.814 0.50.39 22 0.807 0.70.37 18 0.798 0.90.14 7 0.787 S4 Extensions and robustness We examine the robustness of our results along three dimensions: sensitivity to the privateāepistemic balance, heterogeneity in epistemic concern across scientists, and the moral hazard that arises when the journal controls the lottery intensity. S4.1 Sensitivity to the privateāepistemic ratio The payoff function (Equation 4 in the main text) is linear in(b,s,c), so equilibria are invariant under rescaling(b,s,c)ā(Ī»b,Ī»s,Ī»c)for anyĪ»>0. Only the ratior=b/(b+s)and the normalised cost c/(b+s)affect strategic behaviour. Settingb=s=1 (i.e.r=0.5) is a convenience; here we verify that the qualitative results hold across the full range of privateāepistemic balance. TableS5reports the Nash equilibrium thresholdĻ ā , the number of participating scientistsn in , and the equilibrium quality Ģ qas a function ofr, withc=0.1 held fixed. The qualitative result ā that voluntary lottery participation is sustained in equilibrium and improves quality ā holds for all r<1. Asrincreases (scientists weight private benefit more heavily), the threshold drops, fewer scientists enter, and the quality gain over the no-lottery baseline diminishes. The quantitative gap between the Nash equilibrium and the social optimum widens as scientists become more privately oriented. The normalised rejection costc/(b+s)has a modest, monotonic effect: higher cost lowers the equilibrium threshold (fewer papers enter the lottery) but does not qualitatively change the dependence onr. Across a 50-fold range ofc/(b+s)(from 0.01 to 0.50), the fraction of the plannerās gain captured by the Nash equilibrium varies by roughly 5ā10 percentage points within a givenr, compared to a 30 percentage-point spread acrossrvalues. 15 Table S6:Robustness to the quality distribution.Nash equilibrium threshold, participation, and quality under uniform and Beta(2,5) quality distributions. All equilibria are threshold-shaped with zero violations. Parameters:Ļ=0.3,β=8,L=0.10,α=10%,N=100. Distributions n in Ģ q Nash Threshold? Uniform0.5 27 0.790Ć Uniform1.0 36 0.797Ć Uniform2.0 46 0.805Ć Beta(2,5) 0.5 10 0.417Ć Beta(2,5) 1.0 29 0.431Ć Beta(2,5) 2.0 54 0.451Ć S4.2 Robustness to the quality distribution The main analysis assumes paper qualities are uniformly distributed on[0,1]. To test robustness, we repeat the social-optimum and Nash equilibrium analyses with a right-skewed Beta(2,5) distribution, which concentrates mass on lower-quality papers (meanā0.29) and has fewer high-quality papers ā arguably a more realistic shape. TableS6reports the results. The threshold equilibrium structure survives intact: all six Nash equilibria under Beta(2,5) are threshold-shaped with zero ordering violations. The qualitative pat- tern ā more epistemic concern leads to more participation and higher quality ā is identical across both distributions. The social optimum also yields a threshold under Beta(2,5) (Ļ ā =0.614, com- pared toĻ ā =1.000 under uniform), with a similar quality gain. S4.3 Robustness to heterogeneous epistemic concern The main analysis assumes that all scientists share the same epistemic concern parameters. In prac- tice, scientists differ in how much weight they place on the collective quality of published science. A natural worry is that heterogeneity could unravel the mechanism via a weakest-link effect: if a sub- stantial fraction of scientists are purely self-interested (s=0), they might free-ride on the epistemic concern of others. To test this, we allows i to vary across scientists while holding the population mean fixed at Ģ s=1. We consider two illustrative scenarios alongside the homogeneous baseline: 1.Independent random variation: eachs i is drawn from a Beta(2,2) distribution, rescaled to mean 1. This represents moderate, quality-independent heterogeneity. 2.Weakest link (50% selfish): half the scientists haves i =0 (purely self-interested), the other half haves i =2 (strongly epistemic), preserving Ģ s=1. We solve for Nash equilibria using unrestricted sequential best-response dynamics (N=100, M=5,000 Monte Carlo replications per payoff estimate) rather than restricting to threshold strate- gies, since heterogeneity could in principle break the threshold structure. Table S7summarises the results. Independent random variation has essentially no effect on the mechanism: the quality gain and participation count are indistinguishable from the homogeneous baseline. This is because the threshold structure is driven by paper quality, not by individual moti- vation ā scientists with low-quality papers enter regardless of theirs i , while high-quality scientists stay out regardless. The weakest-link scenario is more demanding: half the population has no epistemic concern at all. Even so, the mechanism retains 88% of the homogeneous gain. The self-interested scientists do not enter the lottery, but the epistemic scientists do, and their participation is sufficient to reduce 16 Table S7:The voluntary lottery mechanism is robust to heterogeneous epistemic concern. Quality gainā Ģ qover the no-lottery baseline under three heterogeneity scenarios, with mean epis- temic concern held fixed at Ģ s=1. Even when half the population is purely self-interested, the mechanism retains 88% of the homogeneous gain. Parameters:Ļ=0.3,β=8,L=0.10,α=10%, N=100. ScenarioParticipantsā Ģ q% of baseline Homogeneous (s=1)40 0.040100% Independent random40 0.040100% 50% selfish (s i =0)35 0.03588% noise and improve quality for everyone. The threshold structure survives intact, with participation determined by paper quality within each motivation type. We also tested more extreme configurations (75% selfish), where the gain drops to 69% of the baseline but never vanishes. Positive correlation between quality and epistemic concern (ānoblesse obligeā) is the most harmful configuration, reducing the gain to 77%, because it concentrates epis- temic motivation among scientists whose papers would not benefit from entering the lottery. These results suggest that the mechanism does not require uniform epistemic concern to func- tion. A sufficient condition is that enough scientists care about the quality of the published record to sustain a non-trivial participation threshold. S4.4 Strategic use of the lottery Throughout the main analysis, the lottery intensityLis an exogenous protocol parameter. A natural question is what happens when the journal controls both the review noiseĻand the lottery intensity L, choosing both to maximise its payoff. The intuition is straightforward. The effective noiseĻ eff =ĻĀ· Ģ w β depends on Ģ wthrough a power law with exponentβ. Even moderate reductions in the survival fraction Ģ wproduce large reductions in effective noise ā for instance, atβ=8 and Ģ w=0.55, the effective noise is less than 1% of the baseline. This makes the lottery an extraordinarily cost-effective substitute for review investment: the journal can achieve lower effective noise by increasingLthan by reducingĻdirectly. When allowed to choose both instruments, the journalās optimal strategy is to abandon review effort entirely (Ļā1) and rely on an aggressive lottery (Lā0.45). This is a classic moral hazard: the lottery is intended to complement peer review, but the journal uses it as a substitute. The journal free-rides on the lotteryās noise-reduction properties and stops investing in review quality. The surprising feature of this result is that average accepted quality still improves relative to the no-lottery baseline, because the lottery reduces reviewer burden so drastically that effective noise approaches zero, and the surviving pool is sorted with near-perfect accuracy. However, the mecha- nism is no longer functioning as intended ā it replaces rigorous evaluation with random screening, which may undermine trust in the review process even if aggregate quality metrics improve. This moral hazard motivates treatingLas a pre-committed protocol parameter rather than a journal-controlled lever. In practice, a venueās lottery policy would need to be announced credibly and set at a level acceptable to the research community ā much as acceptance rates are understood as venue characteristics rather than strategic instruments. Reputational costs, community norms, or institutional oversight would serve as natural constraints: a venue that setLtoo aggressively would face author flight and loss of prestige, limiting the scope for substitution. The analysis in the main text therefore fixesLexogenously, treating the lottery as a collective mechanism designed by the community rather than an instrument optimised by the journal. 17 Code and data availability The simulation engine, figure-generation scripts, and pre-computed result files used in this paper are available athttps://github.com/juliangarcia/lottery-peer-review. The ICLR review- score data analysed in SectionS3were obtained from the public release of GonzĆ”lez-MĆ”rquez and Kobak[5]. References [1]Kevin J. S. Zollman, Julian Garcia, and Toby Handfield. Academic journals, incentives, and the quality of peer review: A model.Philosophy of Science, 91(1):186ā203, 2023. doi:10.1017/ psa.2023.132. [2]Roy Radner. Collusive behavior in noncooperative epsilon-equilibria of oligopolies with long but finite lives.Journal of Economic Theory, 22(2):136ā154, 1980. doi:10.1016/0022-0531(80) 90037-X. [3]Corinna Cortes and Neil D. Lawrence. Inconsistency in conference peer review: Revisiting the 2014 NeurIPS experiment.arXiv preprint arXiv:2109.09774, 2021. [4]Alina Beygelzimer, Yann N. Dauphin, Percy Liang, and Jennifer Wortman Vaughan. Has the machine learning review process become more arbitrary as the field has grown? The NeurIPS 2021 consistency experiment.arXiv preprint arXiv:2306.03262, 2023. [5]Rita GonzĆ”lez-MĆ”rquez and Dmitry Kobak. Learning representations of learning represen- tations. InData-centric Machine Learning Research (DMLR) Workshop at ICLR 2024, 2024. URL https://arxiv.org/abs/2404.08403. Dataset:https://github.com/berenslab/ iclr-dataset. 18