Paper deep dive
How AI Aggregation Affects Knowledge
Daron Acemoglu, Tianyi Lin, Asuman Ozdaglar, James Siderius
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 4/10/2026, 2:58:52 AM
Summary
The paper extends the DeGroot model of social learning to analyze the impact of AI aggregators that train on population beliefs and feed synthesized signals back into the network. It identifies a 'learning gap' between long-run consensus and the efficient benchmark, demonstrating that rapid AI updating creates feedback loops that can lead to model collapse and fragile learning. The study compares global and local aggregation architectures, finding that while global aggregators can amplify social biases and worsen learning, local, topic-specific aggregators preserve informational diversity and robustly improve learning outcomes.
Entities (5)
Relation Signals (3)
AI Aggregator → creates → Feedback Loop
confidence 95% · This creates a feedback loop in which AI systems ingest beliefs that they have themselves helped generate.
Local Aggregator → improves → Learning
confidence 90% · Local aggregators trained on proximate or topic-specific data robustly improve learning in all environments.
Global Aggregator → worsens → Learning
confidence 90% · replacing specialized local aggregators with a single global aggregator worsens learning in at least one dimension of the state.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Artificial intelligence (AI) changes social learning when aggregated outputs become training data for future predictions. To study this, we extend the DeGroot model by introducing an AI aggregator that trains on population beliefs and feeds synthesized signals back to agents. We define the learning gap as the deviation of long-run beliefs from the efficient benchmark, allowing us to capture how AI aggregation affects learning. Our main result identifies a threshold in the speed of updating: when the aggregator updates too quickly, there is no positive-measure set of training weights that robustly improves learning across a broad class of environments, whereas such weights exist when updating is sufficiently slow. We then compare global and local architectures. Local aggregators trained on proximate or topic-specific data robustly improve learning in all environments. Consequently, replacing specialized local aggregators with a single global aggregator worsens learning in at least one dimension of the state.
Tags
Links
- Source: https://arxiv.org/abs/2604.04906v1
- Canonical: https://arxiv.org/abs/2604.04906v1
Trouble viewing inline? Open PDF directly →
Full Text
158,387 characters extracted from source content.
Expand or collapse full text
How AI Aggregation Affects Knowledge†thanks: We are grateful to numerous participants at the Applied and Computational Mathematics Seminar at Dartmouth College, the 2025 Annual Network Science in Economics Conference, the Tuck’s AI/ML Seminar Series, and the EC’25 Workshop on LLMs and Information Economics. Daron Acemoglu Tianyi Lin Asuman Ozdaglar James Siderius Massachusetts Institute of Technology, NBER, and CEPR, daron@mit.eduColumbia University, tl3335@columbia.eduMassachusetts Institute of Technology, asuman@mit.eduTuck School of Business at Dartmouth College, james.siderius@tuck.dartmouth.edu (April 6, 2026) Abstract Artificial intelligence (AI) changes social learning when aggregated outputs become training data for future predictions. To study this, we extend the DeGroot model by introducing an AI aggregator that trains on population beliefs and feeds synthesized signals back to agents. We define the learning gap as the deviation of long-run beliefs from the efficient benchmark, allowing us to capture how AI aggregation affects learning. Our main result identifies a threshold in the speed of updating: when the aggregator updates too quickly, there is no positive-measure set of training weights that robustly improves learning across a broad class of environments, whereas such weights exist when updating is sufficiently slow. We then compare global and local architectures. Local aggregators trained on proximate or topic-specific data robustly improve learning in all environments. Consequently, replacing specialized local aggregators with a single global aggregator worsens learning in at least one dimension of the state. Keywords: algorithmic bias, artificial intelligence, feedback loops, information aggregation, networks, social learning. JEL Classification: D80, D83, D85. 1 Introduction In recent years, generative artificial intelligence (GenAI) systems have become a leading interface through which individuals search for, synthesize, and interpret information (Cutler-2023-ChatGPT; Xu-2023-ChatGPT; Ayoub-2024-Head). Unlike traditional information intermediaries, these systems are trained directly on large-scale collections of human-generated content and generate (generally) unified responses to a wide range of queries. However, as GenAI tools have become more widely adopted, their outputs have started to shape the content later used for retraining (Wang-2023-Survey; Burtch-2024-Consequences; Burtch-2024-Generative). This creates a feedback loop in which AI systems ingest beliefs that they have themselves helped generate, blurring the distinction between original information and synthesized knowledge. A centralized aggregator can in principle improve decision-making by collecting and combining information from many dispersed sources. Yet when training data reflect endogenous belief formation in socially structured networks, aggregation can reshape not only collective learning outcomes but also the distribution of epistemic influence across groups. By combining and synthesizing population beliefs, aggregation architectures implicitly determine which signals receive greater weight in shaping AI output. If training data overrepresent certain groups or viewpoints, the resulting system may amplify those signals even in the absence of explicit discrimination. Thus, the central concern is not only predictive performance, but how aggregators, via their training data and responses, reallocate influence throughout human communities and interact with social segregation, feedback, and uncertainty about the underlying environment. To study these forces, we build on the DeGroot model of belief dynamics augmented with AI aggregation. The DeGroot model is characterized by a directed graph, where each edge represents the influence of one agent over the beliefs of another. This setting is attractive to study the learning implications of AI aggregation. First, it provides a tractable framework for the analysis of belief dynamics in the benchmark without AI aggregation. Second, it formalizes the influence of the training weights of AI models in a transparent manner — corresponding to the weights that an AI aggregator puts on the beliefs of different agents. Third, the influence of an AI aggregator on each agent can also be similarly incorporated into this setting, mapping directly to AI adoption. This formalization highlights that an AI aggregator feeds synthesized signals, based on its training weights, back into the network, creating feedback loops. We focus on long-run learning and compare outcomes with and without AI aggregation. When beliefs converge, we follow the literature and refer to the common limiting belief as the consensus. We evaluate this consensus against an efficient benchmark — the posterior mean that would arise under frictionless aggregation of all private signals. The difference between these objects, which we term the learning gap, measures mislearning induced by network structure and AI-mediated feedback. Because consensus is a weighted average of initial signals, the learning gap reflects not only aggregate efficiency loss but also distortions in the effective influence weights assigned to heterogeneous agents. Our first contribution is technical. We provide a closed-form characterization of the long-run consensus induced by AI-mediated learning. Building on perturbation methods in Schweitzer-1968-Perturbation, we show that introducing an AI aggregator into a DeGroot network yields a consensus that can be written explicitly as a function of the original network and a low-rank modification capturing AI training and feedback. This representation expresses the learning gap in closed form and makes transparent how aggregation reshapes influence weights. AI-mediated feedback effectively alters the social weighting structure through which initial information propagates. To sharpen intuition, we specialize our setting to a stylized two-group structure consisting of a majority island and a minority island. In practice, these islands can correspond to ideological-distinct communities, different geographies or demographic groups. Links are more likely within islands than across islands, capturing the common pattern of homophily or group-level segregation. For example, peers attending a common university are more likely to communicate and listen to others at that same university (Mcpherson-2001-Birds). This environment allows us to study how homophily and feedback jointly determine learning outcomes. When a global AI aggregator updates rapidly, its output closely tracks current population beliefs. Because those beliefs already reflect within-group reinforcement, especially within the majority group, the aggregator trains on endogenously distorted data. Feeding this output back into the population reinforces the same distortions, creating a recursive feedback loop between beliefs and training data. In this regime, the impact of an AI aggregator behaves less like information pooling and more like amplification of existing social structure. We formalize this fragility by assuming that the environment (the true network topology, the degree of segregation, and/or exact AI adoption patterns) are not known with precision, so an AI aggregator has to perform well across a range of “plausible” environments. We ask whether there exist training weights that improve information aggregation in the presence of an AI aggregator relative to the benchmark without AI aggregation across a range of environments. Our main result establishes that as updating becomes faster, such robust improvement becomes impossible. Here, robust improvement refers to improvement that holds across a class of networks and adoption patterns. When feedback is sufficiently strong, there is no positive-measure set of training weights that improves learning across admissible environments. Intuitively, rapid retraining repeatedly feeds AI-shaped beliefs back into the training data, reducing the effective diversity of independent information. The system ingests its own outputs. This mechanism parallels concerns described as model collapse: Even with abundant data, learning quality deteriorates when data increasingly reflect model-generated content rather than independent signals (Shumailov-2023-Curse; Gerstgrasser-2024-Model). Speed couples the impact of a global AI aggregator too tightly with current population beliefs, which were themselves shaped by the same AI aggregator. This feedback destroys robustness. This fragility has direct implications for fairness and aggregation of information in society. Because an AI aggregator reshapes effective influence weights, different training regimes implicitly redistribute epistemic power across groups. When environments differ in segregation or AI adoption, the same training design can amplify some group’s signals while attenuating others’. Thus robustness and fairness are structurally linked: The absence of a universally robust training weight implies that AI-based aggregation inevitably embeds distributional trade-offs. Unlike standard fairness notions based on predictive parity or classification error (Hardt-2016-Equality; Kleinberg-2017-Inherent), unfairness in our framework arises from endogenous reweighting of influence rather than disparate predictive error. Even when individual updating is symmetric and no explicit discrimination occurs, the presence of an AI aggregator systematically shifts whose information drives collective belief. The same feedback mechanism that generates aggregate fragility also produces distributional distortions in epistemic influence. We further characterize asymmetries between majority- and minority-weighted training. When training disproportionately reflects the majority island, data imbalance and social segregation reinforce one another: Majority beliefs already receive excess weight through within-group reinforcement, and majority-weighted training compounds this distortion. Learning deteriorates monotonically as homophily increases. By contrast, when training places greater weight on minority beliefs, AI can initially counteract baseline majority dominance, but its impact is non-monotone: with moderate segregation, minority bias protects minority information long enough to discipline the consensus, while with high segregation the same minority bias is amplified by AI-mediated feedback. Correcting underrepresentation is therefore not simply a matter of reweighting data; it interacts endogenously with network structure and feedback. Even well-intentioned interventions can fail when the social environment is imperfectly understood. Finally, we study an alternative architecture in which information aggregators are local and topic-specific. Rather than pooling beliefs into a single global system, the local aggregator model introduces multiple intermediaries (e.g., local newspapers or community-based websites) trained on restricted subsets of agents informative about specific topics. Each local aggregator exerts stronger influence within its constituency than across groups, and own effects dominate cross effects. This localization compartmentalizes feedback: Errors in one dimension do not automatically propagate to others, and informational diversity is preserved even under rapid updating. As a result, local aggregators robustly improve learning relative to the benchmark with no such aggregators. However, replacing specialized local aggregators with a global aggregator necessarily couples previously separate feedback loops and worsens learning along at least one dimension. The key design question is therefore not whether AI aggregates information, but how broadly it does so. An AI aggregator that pools beliefs across the entire population broadens the base of information, but also creates feedback loops, ultimately exacerbating the influence of some groups and rendering learning fragile. In contrast, architectures that restrict training to more localized or topic-relevant subsets preserve informational diversity and compartmentalize feedback, improving robustness even under rapid updating. Related Literature. Our model has built on the foundational literature on DeGroot learning and networked information aggregation (Degroot-1974-Reaching; Bala-1998-Learning; Demarzo-2003-Persuasion; Golub-2010-Naive; Acemoglu-2010-Spread; Acemoglu-2011-Opinion). These results demonstrate that decentralized social learning can aggregate dispersed information effectively under standard conditions. For example, Golub-2010-Naive show that in large networks, beliefs converge arbitrarily close to the truth so long as influence is sufficiently diffuse. Subsequent works extend these results to settings with sparse signals (Banerjee-2021-Naive) and richer belief updating rules (Jadbabaie-2012-Non). A complementary strand demonstrates that networked learning can systematically fail. Acemoglu-2010-Spread show that the presence of agents who remain anchored to initial beliefs can prevent efficient aggregation, leading to enduring belief distortions. More recently, Bohren-2021-Learning show that even without stubbornness, misspecified updating rules can generate systematic long-run errors. Our results align with this second strand, but identify a distinct mechanism: Mislearning arises not from individual stubbornness or incorrect inference, but from introducing an aggregator whose training data are endogenous and shaped by beliefs it previously influenced. Our work is related to, but distinct from, models of stubborn or influential agents (Acemoglu-2013-Opinion; Yildiz-2013-Binary; Ghaderi-2013-Opinion; Hunter-2022-Optimizing; Mostagir-2022-Society). In those models, mislearning typically arises because some agents do not fully update or engage in sustained persuasion, which often leads to persistent disagreement or polarization rather than full consensus. In contrast, in our framework beliefs converge to a unique consensus. However, that consensus can still be distorted, because an AI aggregator endogenously reshapes the effective weights placed on initial information via feedback loops. As a result, our paper highlights a specific form of inefficiency due to the reweighting of information induced by an AI aggregator itself — rather than those rooted in stubbornness or disagreement, emphasized in the previous literature. Another related literature studies how homophily and network structure shape opinion dynamics (Friedkin-1990-Social; Deffuant-2000-Mixing; Golub-2012-Homophily; Mostagir-2023-Social; Grabisch-2023-Design). These papers show that segregation can distort information aggregation even when agents update naïvely. Our contribution differs in two key aspects. First, we introduce an explicit aggregator node that collects and redistributes beliefs, altering the direction and intensity of information flows. Second, rather than studying segregation in isolation, we specify how segregation interacts with training imbalance, updating speed, and aggregation architecture, distinguishing settings where an AI aggregator mitigates network distortions from those where it amplifies them. Finally, our paper connects to emerging empirical and computational work on large language models and their interactions with humans and with one another (Argyle-2023-Out; Park-2022-Social; Park-2023-Generative; Fu-2023-Improving; Leng-2023-LLM; Xiong-2023-Examining; Chan-2024-Chateval; Du-2024-Improving; Filippas-2024-Large; Liang-2024-Encouraging; Papachristou-2025-Network; Chang-2025-LLM). While this literature documents emergent behaviors and network effects among LLMs, it is empirical and does not provide a theory of long-run learning under feedback. Our key contribution is to offer a theoretical framework that formalizes concerns often described informally as model/knowledge collapse (Shumailov-2024-AI; Dohmatob-2024-Tale; Peterson-2025-AI): When AI systems retrain rapidly on data they have themselves influenced, the effective diversity of information can shrink and learning can fail in large populations. By connecting this phenomenon to classical results in social learning, we clarify when and why centralized AI-based information aggregation improves or undermines collective knowledge. Paper Outline. Section 2 introduces the social learning model with a single global AI aggregator. Section 3 establishes the closed-form learning gap for general social networks. Section 4 specializes our model to a two-island setup and studies whether an AI aggregator can robustly improve learning. Section 5 analyzes how segregation and training imbalance interact. Section 6 introduces local, topic-specific aggregators, and compares their effects to those of a global aggregator. We conclude in Section 7. Proofs are presented in the appendix sections. 2 Model We study social learning in a population of n agents indexed by i∈1,…,ni∈\1,…,n\ who seek to learn an unknown scalar state θ∈ℝθ . Time is discrete and runs from t=0t=0 to infinity. Each agent i observes a single private signal si=θ+εis_i=θ+ _i, where εii=1n\ _i\_i=1^n are independent, zero-mean noise terms with finite variance, at time t=0t=0. There are no external signals thereafter. Agents update beliefs over time by observing others’ beliefs through a social network and, when present, by observing the output of an aggregator. Because private signals are unbiased and equally informative, we use the simple average of all private signals as the efficient benchmark: θ^≡1n∑i=1nsi=1n∑i=1npi(0). θ≡ 1n _i=1^ns_i= 1n _i=1^np_i(0). This benchmark corresponds to frictionless aggregation of all private information and serves as a reference point for evaluating learning outcomes. Baseline social learning. Let pi(t)p_i(t) denote agent i’s belief about θ at time t, and let p(t)=(p1(t),…,pn(t))⊤p(t)=(p_1(t),…,p_n(t)) . In the baseline, beliefs evolve according to the benchmark DeGroot learning rule, which takes the form p(t+1)=Tp(t),p(t+1)=Tp(t), where T∈ℝn×nT ^n× n is a row-stochastic matrix describing the network and accounts for an attention or trust matrix. The entry TijT_ij records how much weight agent i places on agent j’s current belief. For example, if agent i forms beliefs by listening to friends, coworkers, local media, or members of the same community, then the row TiT_i summarizes how these sources are weighted. We assume that T is strongly connected and aperiodic. Under these conditions, Golub-2010-Naive show that beliefs converge to a common limit: There exists a scalar p⋆p such that limt→∞pi(t)=p⋆,for all i. _t→∞p_i(t)=p , for all i. Throughout this paper, we refer to p⋆p as the consensus without aggregators (to contrast with the consensus with aggregators, described below). This consensus reflects the long-run belief generated by decentralized social learning alone. Social learning with a global AI aggregator. We introduce an AI aggregator, modeled as an information intermediary that produces a single observable signal based on current population beliefs and feeds this signal back into the network. At each time t, the aggregator forms a weighted average of agents’ beliefs: m(t)=∑i=1nαipi(t)m(t)= _i=1^n _ip_i(t), where α=(α1,…,αn)α=( _1,…, _n) is a 1×n1× n vector of non-negative weights satisfying ∑i=1nαi=1 _i=1^n _i=1. The training weights αi _i capture how strongly the beliefs of different agents or groups are represented in the data used to train or fine-tune the aggregator. Unequal weights may arise because some groups generate more content, are more visible online, receive more engagement, are more extensively digitized, or are deliberately reweighted by a platform. We initialize the aggregator with an uninformed seed, which is similar to how Banerjee-2021-Naive initialize uninformed agents in their model of naïve learning. This initialization implies that a(1)=m(0)a(1)=m(0) and p(1)=Tp(0)p(1)=Tp(0), so that the AI aggregator’s output is shaped by the beliefs of the agents in the population that it places positive training weight on. Thereafter, this output a(t)∈ℝa(t) evolves according to a(t+1)=ρa(t)+(1−ρ)m(t),for all t≥1,a(t+1)=ρ a(t)+(1-ρ)m(t), for all t≥ 1, where ρ∈(0,1)ρ∈(0,1) measures how quickly the aggregator refreshes in response to endogenously evolving population beliefs. A lower value of ρ places more weight on current population beliefs, while a higher value places more weight on the aggregator’s past output. Agents incorporate the output of the AI aggregator into their beliefs with varying weights. In particular, once the aggregator is available, population beliefs evolve according to pi(t+1)=(1−βi)∑j=1nTijpj(t)+βia(t),for all t≥1,p_i(t+1)=(1- _i) _j=1^nT_ijp_j(t)+ _ia(t), for all t≥ 1, where βi∈(0,1) _i∈(0,1) measures the extent to which agent i relies on the aggregator output for all i. Under similar regularity conditions to Golub-2010-Naive (see Proposition 1), beliefs again converge to a common limit: There exists a scalar p⋆p such that limt→∞pi(t)=p⋆,for all i. _t→∞p_i(t)=p , for all i. We refer to p⋆p as the consensus with a global AI aggregator. Learning performance and learning gap. We evaluate learning by comparing long-run consensus beliefs to the efficient benchmark θ θ defined above. Accordingly, we define the learning gaps without and with AI as Δ0≡|p⋆−θ^|,Δ1≡|p⋆−θ^|, _0≡|p - θ|, _1≡|p - θ|, where p⋆p and p⋆p denote the long-run consensuses without and with AI aggregation. The learning gap measures the extent of mislearning: it is zero if and only if decentralized learning fully aggregates private information, and it is positive whenever the consensus is away from the efficient benchmark. Throughout the paper, we say AI aggregation improves learning when Δ1<Δ0 _1< _0 and worsens learning when Δ1>Δ0 _1> _0. Remark — For expositional clarity, we focus in this section on a scalar state. The analysis extends to a multi-dimensional state, with learning occurring componentwise along each dimension. In Section 6, we develop this extension and allow different subsets of agents to be differentially informed about distinct topics. 3 General Network Models We first establish general results for arbitrary networks. In particular, we provide sufficient conditions under which beliefs will converge to a common limit when an aggregator is present. We then derive a closed-form characterization of the long-run consensus and the associated learning gap for any network structure. These results serve as the workhorse for the remainder of the analysis. 3.1 Convergence of Beliefs We begin by deriving conditions under which beliefs converge in the presence of a global AI aggregator. Recall that T denotes the matrix governing social learning among agents and let Γ denote the augmented transition matrix given by: Γ=(ρ(1−ρ)αβDiag(1−β)T), = ( array[]cρ&(1-ρ)α\\ β& Diag(1-β)T array ), where α∈ℝ1×nα ^1× n is the training weight vector and β∈ℝn×1β ^n× 1 is the AI adoption vector. Proposition 1. Suppose that T is strongly connected and aperiodic. Then, the augmented transition matrix Γ is strongly connected and aperiodic if: (i) ρ∈(0,1)ρ∈(0,1), (i) βi<1 _i<1 for all i, and (i) ∑i=1nβi>0 _i=1^n _i>0. Proposition 1 provides simple sufficient conditions for convergence. Indeed, Condition (i) ensures that the AI aggregator does not create an absorbing node disconnected from the population: With probability 1−ρ>01-ρ>0, the aggregator’s next output depends on current beliefs through α. Condition (i) guarantees that agents continue to place positive weight on social learning each period, so the strong connectivity of T is inherited by the agent-based subgraph in the augmented system. Condition (i) rules out the degenerate case in which no agent ever relies on AI, in which case the additional node is irrelevant for learning dynamics. Under these conditions, Γ is a row-stochastic matrix describing a finite-state Markov chain on n+1n+1 nodes that is strongly connected and aperiodic. By the Perron-Frobenius theorem for primitive stochastic matrices, Γ admits a unique stationary distribution π∈Δn+1π∈ ^n+1 on the augmented state space, and Γt→n+1π ^t 1_n+1π as t→∞t→∞. Here and throughout, k1_k is the k-dimensional column vector of ones. Consequently, for any initial condition p(0)p(0), beliefs converge to a common limit: There exists a scalar p⋆p such that a(t)→p⋆andpi(t)→p⋆for all i.a(t)→ p and p_i(t)→ p all i. where p⋆p is the consensus with the AI aggregator, as defined above. 3.2 Characterization of the Long-Run Consensus We next provide a closed-form characterization of the consensus with a global AI aggregator. Theorem 1. Suppose that ρ∈(0,1)ρ∈(0,1) and βi∈(0,1) _i∈(0,1) for all i. Then, the consensus with an AI aggregator satisfies p⋆=11+zn(α+zT)p(0).p = 11+z1_n(α+zT)p(0). (1) where z=(1−ρ)α(n−(n−Diag(β))T)−1z=(1-ρ)α(I_n-(I_n- Diag\,(β))T)^-1 and nI_n is a n×n× n identity matrix. Theorem 1 exploits the linear structure of the learning dynamics. In the absence of AI aggregation, DeGroot learning converges to a weighted average of initial beliefs determined by the stationary distribution of T. Introducing a global AI aggregator creates an endogenous feedback loop: current beliefs influence the aggregator’s output through the training weights α∈ℝ1×nα ^1× n, and this output in turn enters future belief updates with intensities β∈ℝn×1β ^n× 1. Rather than solving directly for the stationary distribution of the augmented system, the proof uses perturbation arguments for finite Markov chains (Schweitzer-1968-Perturbation). Mathematically, the aggregator induces a low-rank modification of baseline DeGroot dynamics, and the resulting closed-form consensus reveals how AI-mediated feedback reweights the influence of initial information. The expression shows that the final consensus can be interpreted as a weighted average of agents’ initial beliefs, where the weights reflect both direct persistence and AI-mediated aggregation through the network. The term α captures how much each agent’s own prior continues to matter, while the term zTzT captures how the AI aggregates information across the network and redistributes it back to agents. The scalar normalization ensures these weights sum to one. Economically, the AI aggregator reshapes influence: rather than beliefs diffusing purely through the network, the AI reweights and amplifies certain information paths, so that an agent’s impact on the final consensus depends both on their position in the network and on how the aggregator processes and feeds information back into the population. 4 How the Speed of AI Updating Affects Learning In this section, we specialize the analysis to the two-island model and ask whether there exist training weights that improve learning not just for one fixed environment, but across a range of admissible values of homophily and AI reliance. This is our notion of robust improvement. Specializing the analysis to the two-island model serves two purposes. First, it isolates in a minimal way how group-level asymmetries in representation and adoption interact with feedback to shape learning. Second, it provides a parsimonious environment in which heterogeneity is coarse but economically meaningful, allowing us to derive sharp fragility and mislearning results that would be obscured in fully general networks. Figure 1: Global aggregator architecture. Model. Agents are partitioned into two types, which we refer to as islands. Islands may correspond to ideological camps, geographic regions, demographic groups, or any salient dimension along which social interactions are more likely within than across groups. Agents of the same type are connected with probability ps∈(0,1)p_s∈(0,1), while agents of different types are connected with probability pd<psp_d<p_s. The ratio h=ps/pd>1h=p_s/p_d>1 captures the degree of homophily in the social network. Larger values of h correspond to more segregated communication structures, while h→1h→ 1 recovers a well-mixed population. There are n1n_1 agents on island 1 (“majority”) and n2=n−n1n_2=n-n_1 agents on island 2 (“minority”). We summarize relative group size by π=n1/n2∈(1,∞)π=n_1/n_2∈(1,∞). The two-island model is the simplest network that features within-group reinforcement, cross-group information flow, and systematic asymmetries in representation in training data. These features are central to the operation of the AI aggregator in practice, where training data often overrepresent some groups and adoption varies across the population. Various qualitative properties of richer networks — including echo chambers, amplification of majority views, and underrepresentation of minority signals — can be seen in this two-group structure. With two islands, the high-dimensional objects (T,α,β)(T,α,β) reduce to a small number of interpretable parameters, as illustrated in Figure 1. We also let α∈[0,1]α∈[0,1] denote the share of training weight placed on the majority island, with 1−α1-α placed on the minority island. We also let β1,β2∈(0,1) _1, _2∈(0,1) capture the reliance on the AI aggregator by the agents in the two islands, respectively. Then, the expected interaction matrix reduces to the 2×22× 2 matrix as follows, F=(hπhπ+11hπ+1πh+πh+π),F= pmatrix hπhπ+1& 1hπ+1\\ πh+π& hh+π pmatrix, where each entry gives the expected weight an agent places on opinions originating from each island. The matrix F encapsulates a simple form of within-group reinforcement in learning (each individual puts more weight onmembers of its own island) and abstracts from idiosyncratic network realizations (the structure of connections is symmetric within islands). This island setup makes explicit the three channels through which the AI aggregator affects learning: (i) data representation, captured by α; (i) adoption and reliance, captured by (β1,β2)( _1, _2); and (i) social amplification, governed by homophily h and relative group sizes π. Accordingly, the learning gaps without and with the AI aggregator are given by Δ0(h,π) _0(h,π) and Δ1(ρ,α,β1,β2,h,π) _1(ρ,α, _1, _2,h,π). Throughout this section, we define Δ⋆:=Δ1(ρ,α,β1,β2,h,π)−Δ0(h,π), := _1(ρ,α, _1, _2,h,π)- _0(h,π), (2) which measures how the AI aggregator changes the learning gap relative to decentralized learning alone. Thus, Δ⋆<0 <0 indicates that the aggregator improves learning, while Δ⋆>0 >0 indicates that it worsens learning. Fragility of AI aggregation. We now study how the speed of updating affects the robustness of information aggregation with AI. Throughout this subsection, we fix the relative size of the two groups π>1π>1 and consider variation along two dimensions. First, the degree of homophily h is assumed to vary over a compact interval [h¯,h¯][ h, h], where h¯,h¯ h, h are finite and satisfy certain conditions. Second, agents’ reliance on the AI aggregator is allowed to vary across groups, with (β1,β2)∈(0,1)2( _1, _2)∈(0,1)^2. We define Λρ:=α∈[0,1]∣Δ⋆(ρ,α,β1,β2,h,π)<0 for all h∈[h¯,h¯] and for all (β1,β2)∈(0,1)2. _ρ:=\α∈[0,1] (ρ,α, _1, _2,h,π)<0 for all h∈[ h, h] and for all ( _1, _2)∈(0,1)^2\. Thus, Λρ _ρ is the set of training weights that improve learning relative to the standard benchmark across this range of environments. We refer to Λρ _ρ as the robust improvement set. We focus on the robust improvement set because it is not reasonable to imagine that AI model parameters can be finely tuned exactly to the pattern of homophily and the precise usage patterns of different groups in society. With this focus, we require the AI aggregator to perform well across a range of environments. Theorem 2. Fix π>1π>1 and h¯,h¯ h, h such that h¯>2π h>2π, h¯>20π h>20π and h¯>h¯ h> h.111The conditions h¯>2π h>2π and h¯>20π h>20π are sufficient bounds that ensure the two-island structure exhibits meaningful segregation and majority amplification; they are not necessary and are imposed to simplify the analysis. Then, there exists a threshold ρ⋆:=ρ⋆(π,h¯,h¯)∈(0,1)ρ :=ρ (π, h, h)∈(0,1) such that 1. if ρ<ρ⋆ρ<ρ , then the robust improvement set is zero-measure: μ(Λρ)=0μ( _ρ)=0; 2. if ρ>ρ⋆ρ>ρ , then the robust improvement set is positive-measure: μ(Λρ)>0μ( _ρ)>0. Theorem 2 highlights that the scope for robust improvement depends on updating speed. When updating is sufficiently fast, the robust improvement set Λρ _ρ is zero-measure; when updating is sufficiently slow, Λρ _ρ is positive-measure. The intuition is that fast updating strengthens the feedback loop between current beliefs and future training data. Because the current beliefs already reflect homophily and within-group reinforcement, an aggregator that closely tracks them feeds the same distortions back into the population, and the resulting amplification depends sensitively on the realized network and AI-reliance profile. This leaves little room for training weights that improve learning robustly across admissible environments. By contrast, slow updating weakens this loop: the aggregator responds to a smoother history of beliefs rather than the current distorted cross-section, so bias is less tightly fed back into training data. In that regime, a nontrivial range of training weights can offset homophily across admissible environments, implying μ(Λρ)>0μ( _ρ)>0. Theorem 2 therefore identifies a tradeoff between speed and robustness: faster updating can make robust improvement harder. Remark — While Theorem 2 focuses on robust improvement, this criterion is motivated by the fact that, in practice, network structure and patterns of AI reliance are typically not known precisely and may vary across settings. By contrast, Appendix B studies learning in a fixed and fully specified environment, allowing for a more detailed characterization of how the aggregator shapes information aggregation when these features are known. 5 AI-Network Interaction on Learning In this section, we isolate how segregation and training imbalance interact, holding the pattern of AI reliance symmetric across groups. For this reason, we now impose β1=β2=β _1= _2=β and focus on the comparative statics of the learning gap with respect to network segregation. We also distinguish between two empirically and conceptually relevant training regimes: one in which the AI aggregator places substantial weight on the majority island, and one in which it places relatively greater weight on the minority island. 5.1 Strong Majority Bias We begin with a regime in which the AI aggregator places substantial weight on the majority island in its training data. This case captures environments in which data availability, visibility, or engagement are systematically skewed toward a dominant group. For example, platforms where majority users generate disproportionate volumes of content, or an AI aggregator is trained primarily on data from high-activity populations. In such environments, the AI aggregator does not merely reflect existing social biases; it risks amplifying them. Proposition 2. Suppose that α>π2π2+1α> π^2π^2+1. Then, we have Δ⋆>0 >0, and Δ1 _1 is monotonically increasing in the degree of homophily h. Proposition 2 shows that when α>π2/(π2+1)α>π^2/(π^2+1), majority-weighted training worsens learning relative to the standard social dynamics, and the learning gap increases monotonically with segregation. In this regime, the aggregator places too much weight on majority beliefs relative to the efficient benchmark. As segregation rises, majority opinions are reinforced more strongly within the dominant island before reaching the minority; feeding these beliefs into a majority-weighted aggregator then amplifies the same distortion. Thus, segregation and training imbalance reinforce one another: When training is sufficiently tilted toward the majority, greater segregation never improves learning. From a design perspective, Proposition 2 underscores that correcting data imbalance is not merely a fairness concern but a robustness requirement. When training data disproportionately reflect majority groups, greater segregation unambiguously worsens learning in the presence of a global aggregator. 5.2 Minority Bias Can biasing the AI aggregator’s training weights in favor of the minority group correct this bias? We next answer this question by considering the opposite regime, in which the global AI aggregator places greater weight on the minority island. This captures environments where AI models are deliberately designed to counteract majority dominance through reweighting schemes, fairness constraints, or targeted data collection. The effects of minority bias are more subtle than those of majority bias. Indeed, minority-weighted training can counteract the baseline tendency of segregated networks to overweight majority beliefs. However, doing so introduces a new tension: correcting one source of bias can lead to overcorrection once feedback and social learning are taken into account. As a result, the interaction between minority bias and network structure is inherently non-monotone. Proposition 3. There exists β⋆>0β >0 such that if α<12α< 12 and β<β⋆β<β , then the sign of Δ⋆ is ambiguous and its dependence on h is non-monotone. In particular, there exist 1<h¯<h¯<∞1< h< h<∞ such that: 1. Δ⋆>0 >0 and Δ1 _1 is decreasing in h over (1,h¯)(1, h); 2. Δ⋆<0 <0 and Δ1 _1 is non-monotone in h over (h¯,h¯)( h, h); 3. Δ⋆>0 >0 and Δ1 _1 is increasing in h over (h¯,∞)( h,∞). Proposition 3 shows that minority-weighted training improves learning only at intermediate levels of segregation. When segregation is low, placing extra weight on minority signals can over-correct and push the long-run consensus away from the efficient benchmark. When segregation is moderate, the same tilt offsets majority dominance and improves learning relative to the no-AI benchmark. When segregation is high, cross-group interaction becomes too weak to discipline the aggregator, so minority-weighted training again worsens learning. Thus, the effect of minority reweighting is non-monotone: it is beneficial when it counteracts majority bias, but detrimental when it either over-corrects or when limited cross-group interaction prevents information from being effectively aggregated. 6 Social Learning with Local Aggregators The analysis so far has focused on a single global aggregator that is trained on population-wide beliefs and feeds a unified signal back to all agents. This architecture captures large-scale systems, such as current large language models, that pool information broadly. In many environments, however, intermediated information aggregation can also be more localized and topic-specific. This can be because of pre-AI intermediaries such as newspapers, professional bodies and local associations, or because of domain-specific AI models that primarily train on information from local communities and are thus designed to be informative about particular issues relevant to these communities (even though their outputs may diffuse beyond those communities). This section studies how learning changes when aggregators are local rather than global. Figure 2: Local aggregator architecture. 6.1 Model with Local Aggregators Extended environment. We extend the baseline environment to a multidimensional state θ=(θ1,θ2)⊤∈ℝ2θ=( _1, _2) ^2, where θk _k represents the state of topic k. As before, agents are partitioned into two islands j∈1,2j∈\1,2\ with relative size π=n1/n2>1π=n_1/n_2>1 and homophily parameter h>1h>1 governing within- versus cross-island interaction. Let F denote the 2×22× 2 matrix from Section 4. Information is local (or topic-specific): island j is the population that is directly informative about topic θj _j. Indeed, each agent i on island j receives an unbiased private signal about θj _j: si,j=θj+εi,js_i,j= _j+ _i,j where εi,ji=1n\ _i,j\_i=1^n are independent, zero mean noise terms with finite variance, and receives no direct information about the other topic θj′ _j with j′≠j ≠ j. The assumption is not that only one island cares about a topic, but that first-hand signals and specialized expertise are concentrated locally (e.g., local health systems vs. local industries/labor markets), making initial information topic-specific. We normalize initial beliefs so that agents place zero belief on topics about which they are uninformed, i.e., pi,j′(0)=0p_i,j (0)=0 for j′≠j ≠ j. Let pk(t)∈ℝ2p_k(t) ^2 denote the vector of island-level beliefs about topic k at time t, with the no-aggregator dynamics in the following form of pk(t+1)=Fpk(t),for all k∈1,2.p_k(t+1)=Fp_k(t), for all k∈\1,2\. The efficient benchmark aggregates the informative signals topic by topic. Under the same diffuse prior and equal-variance signal structure as before, the benchmark is θ^=(θ^1,θ^2)=(1n1∑i∈Island 1si,1,1n2∑i∈Island 2si,2). θ=( θ_1, θ_2)= ( 1n_1 _i \,1s_i,1, 1n_2 _i \,2s_i,2 ). There are two local aggregators, indexed by k∈1,2k∈\1,2\, where local aggregator k is specialized to topic θk _k. Each local aggregator trains only on beliefs about its topic (see Figure 2). Formally, let A1=(1 0)A_1=(1\ \ 0) and A2=(0 1)A_2=(0\ \ 1) so that Akpk(t)A_kp_k(t) extracts beliefs about topic k. Each local aggregator produces an observable output ak(t)∈ℝa_k(t) that updates according to ak(t+1)=ρak(t)+(1−ρ)Akpk(t),for all k∈1,2,a_k(t+1)=ρ a_k(t)+(1-ρ)A_kp_k(t), for all k∈\1,2\, where ρ∈(0,1)ρ∈(0,1) governs the speed of updating. Here, the lower ρ corresponds to faster updating and stronger feedback. Local aggregators influence agents asymmetrically across islands. Let Bk=(βk1βk2)∈ℝ2, for all k∈1,2,B_k= pmatrix _k1\\ _k2 pmatrix ^2, for all k∈\1,2\, so that BkB_k collects island-by-island reliance on local aggregator k. In particular, B1=(β11β12),B2=(β21β22).B_1= pmatrix _11\\ _12 pmatrix, B_2= pmatrix _21\\ _22 pmatrix. Here, βkj∈[0,1) _kj∈[0,1) denotes the weight placed by island j on local aggregator k. Equivalently, the first index k labels the local aggregator (topic), and the second index j labels the island. A key feature of local aggregators is that each of them is primarily trusted by (and thus has stronger influence on) the population that is informative about its topic. Because Bk=(βk1,βk2)⊤B_k=( _k1, _k2) collects island-by-island reliance on local aggregator k, we impose the following asymmetry: β11>β12,β22>β21. _11> _12, _22> _21. (3) That is, island 11 relies more on the topic 11 aggregator than island 22 does, and island 22 relies more on the topic 22 aggregator than island 11 does. This assumption formalizes the idea that topic-relevant intermediaries have greater influence within their own communities than across communities, and rules out the degenerate case in which a local aggregator is relied upon more heavily by the island that is uninformed about its topic. Given local aggregator outputs, beliefs about each topic evolve as pk(t+1)=(2−Diag(Bk))Fpk(t)+Bkak(t),for all k∈1,2,p_k(t+1)=(I_2- Diag\,(B_k))Fp_k(t)+B_ka_k(t), all k∈\1,2\, where Diag(Bk) Diag\,(B_k) is the diagonal matrix with entries given by BkB_k. Under the same regularity conditions as in Section 3, the augmented system admits a unique consensus for each topic, yielding a limiting belief vector. By abuse of notation, we define p⋆:=(p1⋆,p2⋆),p :=(p_1 ,p_2 ), where pk⋆p_k denotes the consensus belief about topic k under local aggregators. Performance metric. We let the local-aggregation learning gap be the vector Δ2:=(|p1⋆−θ^1|,|p2⋆−θ^2|). _2:=(|p_1 - θ_1|,|p_2 - θ_2|). We compare Δ2 _2 to the no-aggregator benchmark vector Δ0 _0 (formed by applying the no-aggregator dynamics to each topic) and to the global-aggregator learning gap vector Δ1 _1 (formed by applying the global-aggregator dynamics to each topic). Each topic evolves under the global-aggregator rule applied to pk(t)p_k(t), with a shared training design across topics. Accordingly, Δ1 _1 is computed topic-wise by running the global-aggregator update on that topic’s beliefs. The key question is whether localization of training and influence improves learning and mitigates the feedback-driven fragility we identified in the presence of a global aggregator. We next compare learning under local aggregators to the no-aggregator benchmark and to learning under a single global aggregator. To avoid confusion with Sections 2-5, note that there Δ0 _0 and Δ1 _1 both denote the scalar gap to the efficient benchmark θ θ (which equals π+1 π+1 under our two-island normalization), whereas Δ0 _0, Δ1 _1 and Δ2 _2 here are all vectors of topicwise gaps to the topic truths under the unit normalization p1(0)=(1,0)⊤p_1(0)=(1,0) and p2(0)=(0,1)⊤p_2(0)=(0,1) (hence the efficient benchmark is given by (1,1)(1,1)). Throughout, we hold fixed the underlying primitives (i.e., signals, network structure, and agents’ updating rules) so that differences in outcomes arise solely from the architecture of aggregators. This allows us to isolate the economic forces introduced by scale and centralization, abstracting from differences in data quality or behavioral assumptions. 6.2 Local Aggregators versus the No-Aggregator Benchmark We first compare local aggregators to decentralized learning without any aggregators. Proposition 4. Learning is better across all topics under local aggregators than without any aggregators. That is, (Δ2)k<(Δ0)k( _2)_k<( _0)_k for each topic k∈1,2k∈\1,2\. Proposition 4 demonstrates that local aggregators improve learning relative to the no-aggregator benchmark. The reason is that each aggregator is topic-specific: aggregator k is trained only on beliefs about θk _k from the subgroup that is informative about that topic, so its input is more relevant and less noisy. Its influence is also disciplined, since each local aggregator is relied on more heavily by the island that is informative about its topic and less heavily by the other island. This allows topic-relevant information to spill across groups without generating the system-wide feedback distortions of a global aggregator. Unlike the global case in Theorem 2, where training reflects an endogenously distorted population-wide mixture of beliefs, local aggregators keep feedback in separate channels anchored to the informative subgroup, making learning more robust. Proposition 4 and Theorem 2 emphasize that the key design issue is not whether aggregator outputs cross groups — they do so under both architectures — but whether training data are globally pooled and endogenously contaminated or locally anchored to informative sources. A global aggregator magnifies feedback and this makes learning fragile, especially under uncertainty or fast updating, while local aggregators preserve informational discipline by tying each training process to the agents who observe the relevant state. 6.3 Limits of A Single Global Aggregator We proceed to compare learning under local aggregators and that under a single global aggregator in a multidimensional setting. By a single global aggregator, we do not mean a scalar intermediary that pools beliefs across topics and broadcasts one common numerical output. Rather, the model is parallel by topic: for each topic k, the aggregator produces a topic-specific signal/output and the within-topic belief-updating dynamics are run on that topic’s state. The sense in which the aggregator is single is that it is the same global architecture applied across topics (e.g., one common set of training weights α and the same adoption structure, when imposed) so that the induced map is identical across topics up to the topic’s inputs. Consequently, objects such as Δ1 _1 are defined and analyzed topicwise by applying the global-aggregator dynamics separately to each topic, and then comparing the resulting learning gaps across specifications. Theorem 3. Suppose a single global aggregator replaces the local aggregators. Then there exists at least one topic k⋆∈1,2k ∈\1,2\ for which learning is worse under a global aggregator than under local aggregators. That is, (Δ1)k⋆>(Δ2)k⋆( _1)_k >( _2)_k . Theorem 3 formalizes a basic limitation of global aggregation in multi-topic environments. Local aggregators are specialized: each topic is assigned an aggregator trained on beliefs from the subgroup that is informative about that topic, so training remains aligned with the relevant source of information even if outputs spill across islands. A single global aggregator, by contrast, applies one common training-and-feedback design across all topics. This shared design cannot simultaneously match different islands’ informational advantages: performing well on topic 1 requires placing weight on island 1, while performing well on topic 2 requires placing weight on island 2. These objectives conflict, so any global design that improves learning on one topic necessarily weakens it on another. Local aggregators avoid this problem by keeping training channels separate and topic-specific. Theorem 3 therefore complements the earlier results in two ways. First, it strengthens the message of Theorem 2: fragility is not only about updating speed or uncertainty over network structure, but also about the scope of AI-based aggregation. Second, it clarifies why Proposition 4 holds: local aggregators improve learning by preserving specialization and anchoring topic-specific aggregation to agents who are most informed about that topic. In short, global AI-based aggregation of information fails typically both because of feedback-driven amplification and because of intrinsic multi-topic coupling, whereas localized aggregation avoids both forces by construction. 7 Conclusion This paper studies how AI aggregation influences social learning. We extend the DeGroot model of belief dynamics by introducing an AI aggregator as an endogenous intermediary that both trains on and influences population beliefs. The DeGroot model provides a tractable framework in which this training can be formalized — as training weights attached to the beliefs of different agents. Our analysis highlights how the network structure (in particular, the degree of segregation and homophily) interacts with the training weights and the speed of updating of the global AI aggregator to shape belief dynamics. Our first set of results presents an important robustness tradeoff. When a single global aggregator updates rapidly, feedback between its outputs and its training data undermines robustness: small misspecifications in training weights or uncertainty about the social network are amplified rather than corrected. Beyond a threshold, no training design can robustly improve learning across plausible environments. This provides a formal account of feedback-driven failure, often described as model collapse, arising from endogenous redundancy rather than data scarcity. We explore the interaction between aggregators and group structure in greater detail: majority-weighted training interacts monotonically with segregation to worsen learning, as network reinforcement and data imbalance align. Minority-weighted training can initially improve learning by counteracting majority dominance, but its effects are non-monotone: increased segregation eventually weakens cross-group discipline and leads to overcorrection. Bias correction through centralized aggregation of information therefore depends critically on social structure and feedback. Finally, we compare global and local aggregators in a multidimensional setting. Local, topic-specific aggregators anchor training to populations that are informative about each dimension, compartmentalizing feedback and preserving informational diversity. This architecture avoids the system-wide coupling that drives fragility under the global aggregator. Moreover, no single global aggregator can replicate the performance of specialized local aggregators across all dimensions, revealing a fundamental limitation of centralized design. In summary, our results emphasize that a central design choice in AI is not whether information is aggregated, but how broad the information sources are for AI models, how quickly these updates take place, and how those updates are then fed back into the population. Scale and speed can be beneficial only insofar as feedback remains disciplined. Modular, localized architectures sacrifice breadth and scale, but preserve valuable specialization, yielding more reliable improvements in learning. There are many interesting areas for future research. First, the framework here can be extended so that there are multiple global aggregators with different training weights. Second, a more ambitious generalization would be to endogenize the reliance of different agents on different global and local AI aggregation (e.g., by making them more Bayesian in the weights they place on the various aggregators). Third, one could consider hybrid global-local architectures. Fourth, the overall network structure can be endogenized more generally, though this is typically challenging in the DeGroot setup. Finally, it would be interesting to experimentally investigate whether changing the training weights of AI aggregation along the lines of our analysis will modify the extent of effects in practice. References Appendix A Proofs We present all omitted proofs from the main body. A.1 Proofs from Section 3 Proof of Proposition 1. We show that Γ is strongly connected. First, consider any two agents i and j. Because T is strongly connected and βi<1 _i<1 for all i, agent i is reached from agent j and agent j is reached from agent i in the augmented graph Γ . Next, consider the aggregator and an arbitrary agent j. Because ∑i=1nαi=1 _i=1^n _i=1 and αi≥0 _i≥ 0 for all i, there exists some agent i∗i such that αi⋆>0 _i >0. Hence the aggregator is reached from agent i⋆i . Because T is strongly connected, agent i∗i is reached from agent j. Therefore, the aggregator is reached from agent j. Conversely, because ∑i=1nβi>0 _i=1^n _i>0, there exists some agent i⋆i such that βi⋆>0 _i >0. Hence agent i⋆i is reached from the aggregator. Because T is strongly connected, agent j is reached from agent i⋆i . Therefore, agent j is reached from the aggregator. Putting these pieces together yields that Γ is strongly connected. We next show that Γ is aperiodic. Because ρ∈(0,1)ρ∈(0,1), the aggregator has a self-loop. In addition, the subgraph induced by agents is aperiodic because T is aperiodic and βi<1 _i<1 for each i. Putting these pieces together yields the desired result. Proposition A.1. Let T∈ℝn×nT ^n× n be a regular Markov transition matrix with a unique stationary distribution s∈ℝ1×ns ^1× n. Let T∞T^∞ denote the rank-one matrix with s in every row, and define the fundamental matrix Y≡∑k=0∞(Tk−T∞)Y≡ _k=0^∞(T^k-T^∞). Let D∈ℝn×nD ^n× n be such that T^=T+D T=T+D is also regular, and let s^∈ℝ1×n s ^1× n denote the unique stationary distribution of T T. If n−DYI_n-DY is nonsingular, then s^−s=sDY(n−DY)−1 s-s=sDY(I_n-DY)^-1. Equivalently, s^=s(n−DY)−1 s=s(I_n-DY)^-1. Proof. This follows immediately from Schweitzer-1968-Perturbation. Proof of Theorem 1. Because T is strongly connected and aperiodic, there is a rank-1 matrix T∞T^∞ corresponding to the unique left-eigenvector s of eigenvalue one in every row. In this context, the fundamental matrix of T is defined by Y≡∑k=0∞(Tk−T∞)Y≡ _k=0^∞(T^k-T^∞). We claim the following form of the consensus as a function of ρ, α, β, s, and the fundamental matrix Y of network T. We define D∈ℝn×nD ^n× n, w^∈ℝ w and v v (a 1×n1× n vector) as follows, D=βα−Diag(β)T,D=βα- Diag(β)T, and v^=s(n−DY)−1,w^=11−ρv^β. v=s(I_n-DY)^-1, w= 11-ρ vβ. Then, the consensus is given by 11+w^(w^α+v^T)p(0). 11+ w( wα+ vT)p(0). To see this, we have (w^,v^)Γ=(w^,v^)( w, v) =( w, v) if and only if w^=ρw^+v^β,(1−ρ)w^α+v^(n−Diag(β))T=v^. w=ρ w+ vβ, (1-ρ) wα+ v(I_n- Diag\,(β))T= v. This implies that (w^,v^)Γ=(w^,v^)( w, v) =( w, v) if and only if w^=11−ρv^β w= 11-ρ vβ and v^(T+D)=v v(T+D)= v. Because D is a perturbation matrix such that T+DT+D is regular, Proposition A.1 implies v^−s=sDY(n−DY)−1 v-s=sDY(I_n-DY)^-1. Hence, v^=s+sDY(n−DY)−1=s(n−DY)−1 v=s+sDY(I_n-DY)^-1=s(I_n-DY)^-1. Because a(1)=αp(0)a(1)=α p(0) and p(1)=Tp(0)p(1)=Tp(0), we have p⋆=11+w^(w^a(1)+v^p(1))=11+w^(w^α+v^T)p(0).p = 11+ w( wa(1)+ vp(1))= 11+ w( wα+ vT)p(0). Finally, we define z=(1−ρ)α(n−(n−Diag(β))T)−1z=(1-ρ)α(I_n-(I_n- Diag(β))T)^-1. Then, we show the consensus is given by p⋆=11+zn(α+zT)p(0).p = 11+z1_n(α+zT)p(0). As a consequence of our previous argument, the consensus is given by 11+w^(w^α+v^T)p(0), 11+ w( wα+ vT)p(0), (4) where v^=s(n−DY)−1 v=s(I_n-DY)^-1 and w^=11−ρv^β w= 11-ρ vβ. Note that D=βα−Diag(β)TD=βα- Diag(β)T and Y=(n−T+ns)−1−nsY=(I_n-T+1_ns)^-1-1_ns. Because Dn=0D1_n=0, we have DY=(βα−Diag(β)T)(n−T+ns)−1DY=(βα- Diag(β)T)(I_n-T+1_ns)^-1. By applying the Woodbury identity and using the fact that sT=ssT=s, we have v v = = s(n−(Diag(β)T−βα)(n−T+ns+Diag(β)T−βα)−1) s(I_n-( Diag\,(β)T-βα)(I_n-T+1_ns+ Diag\,(β)T-βα)^-1) = = s(n−T+ns)(n−T+ns+Diag(β)T−βα)−1 s(I_n-T+1_ns)(I_n-T+1_ns+ Diag(β)T-βα)^-1 = = s(n−(n−Diag(β))T+ns−βα)−1. s(I_n-(I_n- Diag(β))T+1_ns-βα)^-1. For simplicity, we define Ω=n−(n−Diag(β))T =I_n-(I_n- Diag(β))T. This matrix is invertible because ‖(n−Diag(β))T‖∞=maxi(1−βi)<1\|(I_n- Diag(β))T\|_∞= _i(1- _i)<1. Then, we can rewrite v v in the following form of v^=s(Ω+(n−β)(sα))−1. v=s ( + pmatrix1_n&-β pmatrix pmatrixs\\ α pmatrix )^-1. Applying the Woodbury identity again yields v v = = s(Ω−1−Ω−1(n−β)(2+(sα)Ω−1(n−β))−1(sα)Ω−1) s ( ^-1- ^-1 pmatrix1_n&-β pmatrix (I_2+ pmatrixs\\ α pmatrix ^-1 pmatrix1_n&-β pmatrix )^-1 pmatrixs\\ α pmatrix ^-1 ) = = sΩ−1−(sΩ−1n−sΩ−1β)(1+sΩ−1n−sΩ−1βαΩ−1n1−αΩ−1β)−1(sΩ−1αΩ−1). s ^-1- pmatrixs ^-11_n&-s ^-1β pmatrix pmatrix1+s ^-11_n&-s ^-1β\\ α ^-11_n&1-α ^-1β pmatrix^-1 pmatrixs ^-1\\ α ^-1 pmatrix. Using the definition of Ω , we have Ω−1β=n ^-1β=1_n. Plugging this result into the above equality and using the fact that αn=sn=1 1_n=s1_n=1 yields v v = = sΩ−1−(sΩ−1n−1)(1+sΩ−1n−1αΩ−1n0)−1(sΩ−1αΩ−1) s ^-1- pmatrixs ^-11_n&-1 pmatrix pmatrix1+s ^-11_n&-1\\ α ^-11_n&0 pmatrix^-1 pmatrixs ^-1\\ α ^-1 pmatrix = = sΩ−1−1αΩ−1n(sΩ−1n−1)(01−αΩ−1n1+sΩ−1n)(sΩ−1αΩ−1) s ^-1- 1α ^-11_n pmatrixs ^-11_n&-1 pmatrix pmatrix0&1\\ -α ^-11_n&1+s ^-11_n pmatrix pmatrixs ^-1\\ α ^-1 pmatrix = = sΩ−1−1αΩ−1n(αΩ−1n−1)(sΩ−1αΩ−1) s ^-1- 1α ^-11_n pmatrixα ^-11_n&-1 pmatrix pmatrixs ^-1\\ α ^-1 pmatrix = = αΩ−1αΩ−1n. α ^-1α ^-11_n. Thus, we have w^=11−ρv^β=1(1−ρ)αΩ−1n w= 11-ρ vβ= 1(1-ρ)α ^-11_n. Plugging (w^,v^)( w, v) into Eq. (4) yields the consensus p⋆p . A.2 Closed-Form Learning Gaps (Corollary to Theorem 1) Using Theorem 1, we provide closed-form expressions for the learning gaps under a global AI aggregator and two local aggregators. For a global AI aggregator, we have scalar learning gaps Δ1 _1 (with AI aggregator) and Δ0 _0 (without an aggregator). For two local aggregators, we have the two-dimensional learning gaps Δ0 _0 (no aggregator), Δ1 _1 (global aggregator architecture), and Δ2 _2 (local aggregator architecture). The learning gap with a global aggregator. Suppose that h=ps/pd∈(1,∞)h=p_s/p_d∈(1,∞) and π=n1/n2∈(1,∞)π=n_1/n_2∈(1,∞). Then, we can rewrite α,β,Fα,β,F as follows, α=(α1−α),β=(β1β2),F=(hπhπ+11hπ+1πh+πh+π),p(0)=(10).α= pmatrixα&1-α pmatrix, β= pmatrix _1\\ _2 pmatrix, F= pmatrix hπhπ+1& 1hπ+1\\ πh+π& hh+π pmatrix, p(0)= pmatrix1\\ 0 pmatrix. and derive a closed-form characterization of the consensus p⋆p using Theorem 1 as follows, p⋆=11+z2(α+z(hπhπ+1πh+π)),p = 11+z1_2 (α+z pmatrix hπhπ+1\\ πh+π pmatrix ), (5) where z=(1−ρ)(α 1−α)(2−(2−Diag(β))F)−1.z=(1-ρ)(α\ \ \ 1-α)(I_2-(I_2- Diag\,(β))F)^-1. First, we claim that (2−(2−Diag(β))F)−1=(11−β1β1hπ+11)(hπ+1β1hπ+1(h+π)(β1hπ+1)(β2h+π)(β1hπ+1)−(1−β1)(1−β2)π)(1(1−β2)π(hπ+1)(h+π)(β1hπ+1)1).(I_2-(I_2- Diag(β))F)^-1= pmatrix1& 1- _1 _1hπ+1\\ &1 pmatrix pmatrix hπ+1 _1hπ+1&\\ & (h+π)( _1hπ+1)( _2h+π)( _1hπ+1)-(1- _1)(1- _2)π pmatrix pmatrix1&\\ (1- _2)π(hπ+1)(h+π)( _1hπ+1)&1 pmatrix. Indeed, we have 2−(2−Diag(β))F=(β1hπ+1hπ+1−1−β1hπ+1−(1−β2)πh+πβ2h+πh+π)≐(abcd),I_2-(I_2- Diag(β))F= pmatrix _1hπ+1hπ+1&- 1- _1hπ+1\\ - (1- _2)πh+π& _2h+πh+π pmatrix pmatrixa&b\\ c&d pmatrix, and obtain the desired result using the one-dimensional version of Schur complement as follows, (abcd)−1=(1−ba1)(1ad−bc)(1−ca1). pmatrixa&b\\ c&d pmatrix^-1= pmatrix1&- ba\\ &1 pmatrix pmatrix 1a&\\ & aad-bc pmatrix pmatrix1&\\ - ca&1 pmatrix. Then, we have z=(1−ρ)(α 1−α)(2−(2−Diag(β))F)−1=(1−ρ)(α 1−α)(11−β1β1hπ+11)(hπ+1β1hπ+1(h+π)(β1hπ+1)(β2h+π)(β1hπ+1)−(1−β1)(1−β2)π)(1(1−β2)π(hπ+1)(h+π)(β1hπ+1)1)=(1−ρ)(α(1−α)β1hπ+(1−αβ1)β1hπ+1)(hπ+1β1hπ+1(h+π)(β1hπ+1)(β2h+π)(β1hπ+1)−(1−β1)(1−β2)π)(1(1−β2)π(hπ+1)(h+π)(β1hπ+1)1)=(1−ρ)(α(hπ+1)β1hπ+1(h+π)((1−α)β1hπ+(1−αβ1))(β2h+π)(β1hπ+1)−(1−β1)(1−β2)π)(1(1−β2)π(hπ+1)(h+π)(β1hπ+1)1)=(1−ρ)((hπ+1)(αβ2h+(1−β2+αβ2)π)(β2h+π)(β1hπ+1)−(1−β1)(1−β2)π(h+π)((1−α)β1hπ+(1−αβ1))(β2h+π)(β1hπ+1)−(1−β1)(1−β2)π), array[]rclz&=&(1-ρ)(α\ \ \ 1-α)(I_2-(I_2- Diag(β))F)^-1\\ &=&(1-ρ)(α\ \ \ 1-α) pmatrix1& 1- _1 _1hπ+1\\ &1 pmatrix pmatrix hπ+1 _1hπ+1&\\ & (h+π)( _1hπ+1)( _2h+π)( _1hπ+1)-(1- _1)(1- _2)π pmatrix pmatrix1&\\ (1- _2)π(hπ+1)(h+π)( _1hπ+1)&1 pmatrix\\ &=&(1-ρ) pmatrixα& (1-α) _1hπ+(1-α _1) _1hπ+1 pmatrix pmatrix hπ+1 _1hπ+1&\\ & (h+π)( _1hπ+1)( _2h+π)( _1hπ+1)-(1- _1)(1- _2)π pmatrix pmatrix1&\\ (1- _2)π(hπ+1)(h+π)( _1hπ+1)&1 pmatrix\\ &=&(1-ρ) pmatrix α(hπ+1) _1hπ+1& (h+π)((1-α) _1hπ+(1-α _1))( _2h+π)( _1hπ+1)-(1- _1)(1- _2)π pmatrix pmatrix1&\\ (1- _2)π(hπ+1)(h+π)( _1hπ+1)&1 pmatrix\\ &=&(1-ρ) pmatrix (hπ+1)(α _2h+(1- _2+α _2)π)( _2h+π)( _1hπ+1)-(1- _1)(1- _2)π& (h+π)((1-α) _1hπ+(1-α _1))( _2h+π)( _1hπ+1)-(1- _1)(1- _2)π pmatrix, array which implies z2=(1−ρ)(((1−α)β1+αβ2)h2π+(1+(1−α)β1−(1−α)β2)hπ2+(1−αβ1+αβ2)h+(2−αβ1−(1−α)β2)π)β1β2h2π+β1hπ2+β2h+(β1+β2−β1β2)π,z1_2= (1-ρ)(((1-α) _1+α _2)h^2π+(1+(1-α) _1-(1-α) _2)hπ^2+(1-α _1+α _2)h+(2-α _1-(1-α) _2)π) _1 _2h^2π+ _1hπ^2+ _2h+( _1+ _2- _1 _2)π, (6) and z(hπhπ+1πh+π)=(1−ρ)(αβ2h2π+(1+β1−β2−αβ1+αβ2)hπ2+(1−αβ1)π)β1β2h2π+β1hπ2+β2h+(β1+β2−β1β2)π.z pmatrix hπhπ+1\\ πh+π pmatrix= (1-ρ)(α _2h^2π+(1+ _1- _2-α _1+α _2)hπ^2+(1-α _1)π) _1 _2h^2π+ _1hπ^2+ _2h+( _1+ _2- _1 _2)π. (7) Plugging Eq. (6) and Eq. (7) into Eq. (5) yields p⋆≡(αβ1β2+(1−ρ)αβ2)h2π+(αβ1+(1−ρ)(1+(1−α)(β1−β2)))hπ2+αβ2h+(α(β1+β2−β1β2)+(1−ρ)(1−αβ1))π(β1β2+(1−ρ)(β1−α(β1−β2))h2π+(β1+(1−ρ)(1+(1−α)(β1−β2)))hπ2+(β2+(1−ρ)(1−α(β1−β2)))h+(β1+β2−β1β2+(1−ρ)(2−β2−α(β1−β2)))π, array[]lp ≡\\ (α _1 _2+(1-ρ)α _2)h^2π+(α _1+(1-ρ)(1+(1-α)( _1- _2)))hπ^2+α _2h+(α( _1+ _2- _1 _2)+(1-ρ)(1-α _1))π( _1 _2+(1-ρ)( _1-α( _1- _2))h^2π+( _1+(1-ρ)(1+(1-α)( _1- _2)))hπ^2+( _2+(1-ρ)(1-α( _1- _2)))h+( _1+ _2- _1 _2+(1-ρ)(2- _2-α( _1- _2)))π, array which implies Δ1(ρ,α,β1,β2,h,π)≡|(αβ1β2+(1−ρ)αβ2)h2π+(αβ1+(1−ρ)(1+(1−α)(β1−β2)))hπ2+αβ2h+(α(β1+β2−β1β2)+(1−ρ)(1−αβ1))π(β1β2+(1−ρ)(β1−α(β1−β2))h2π+(β1+(1−ρ)(1+(1−α)(β1−β2)))hπ2+(β2+(1−ρ)(1−α(β1−β2)))h+(β1+β2−β1β2+(1−ρ)(2−β2−α(β1−β2)))π−π+1|. array[]l _1(ρ,α, _1, _2,h,π)≡\\ | (α _1 _2+(1-ρ)α _2)h^2π+(α _1+(1-ρ)(1+(1-α)( _1- _2)))hπ^2+α _2h+(α( _1+ _2- _1 _2)+(1-ρ)(1-α _1))π( _1 _2+(1-ρ)( _1-α( _1- _2))h^2π+( _1+(1-ρ)(1+(1-α)( _1- _2)))hπ^2+( _2+(1-ρ)(1-α( _1- _2)))h+( _1+ _2- _1 _2+(1-ρ)(2- _2-α( _1- _2)))π- π+1 |. array We also have by setting β1=β2=β _1= _2=β, Δ1(ρ,α,β,h,π)≡|αβ(β+1−ρ)h2π+(αβ+1−ρ)hπ2+αβh+(αβ(2−β)+(1−ρ)(1−αβ))πβ(β+1−ρ)h2π+(β+1−ρ)hπ2+(β+1−ρ)h+(2−β)(β+1−ρ)π−π+1|. _1(ρ,α,β,h,π)≡ | αβ(β+1-ρ)h^2π+(αβ+1-ρ)hπ^2+αβ h+(αβ(2-β)+(1-ρ)(1-αβ))πβ(β+1-ρ)h^2π+(β+1-ρ)hπ^2+(β+1-ρ)h+(2-β)(β+1-ρ)π- π+1 |. The learning gap with local aggregators. Suppose that h=ps/pd∈(1,∞)h=p_s/p_d∈(1,∞) and π=n1/n2∈(1,∞)π=n_1/n_2∈(1,∞). Then, we can rewrite A1,A2,B1,B2,FA_1,A_2,B_1,B_2,F as follows, A1=(10),A2=(01),B1=(β11β12),B2=(β21β22),F=(hπhπ+11hπ+1πh+πh+π),p1(0)=(10),p2(0)=(01).A_1= pmatrix1&0 pmatrix,\ A_2= pmatrix0&1 pmatrix,\ B_1= pmatrix _11\\ _12 pmatrix,\ B_2= pmatrix _21\\ _22 pmatrix,\ F= pmatrix hπhπ+1& 1hπ+1\\ πh+π& hh+π pmatrix,\ p_1(0)= pmatrix1\\ 0 pmatrix,\ p_2(0)= pmatrix0\\ 1 pmatrix. Because information is topic-specific, the initial belief profiles differ across topics. For topic 11, island 11 is the informed population, so we normalize the initial belief vector as p1(0)=(1,0)⊤p_1(0)=(1,0) (island 1 starts with a unit informational advantage and island 2 is uninformed). For topic 22, island 22 is the informed population, so the analogous normalization is p2(0)=(0,1)⊤p_2(0)=(0,1) . All subsequent expressions for pk⋆p_k and pk⋆p_k are linear in pk(0)p_k(0), and the learning-gap comparisons depend only on the induced influence weights; thus, without loss of generality, we work with these unit normalizations. Then, we derive a closed-form characterization of p1⋆p_1 and p2⋆p_2 using Theorem 1 as follows, p1⋆=11+z12⊤(1+z1(hπhπ+1πh+π)),p2⋆=11+z22⊤(0+z2(1hπ+1h+π)),p_1 = 11+z_11_2 (1+z_1 pmatrix hπhπ+1\\ πh+π pmatrix ), p_2 = 11+z_21_2 (0+z_2 pmatrix 1hπ+1\\ hh+π pmatrix ), where z1=(1−ρ)(1 0)(2−(2−Diag(B1))F)−1,z2=(1−ρ)(0 1)(2−(2−Diag(B2))F)−1.z_1=(1-ρ)(1\ \ \ 0)(I_2-(I_2- Diag\,(B_1))F)^-1, z_2=(1-ρ)(0\ \ \ 1)(I_2-(I_2- Diag\,(B_2))F)^-1. By using the same arguments, we have p1⋆(ρ,β11,β12,h,π)≡(1−ρ+β11)β12h2π+(1−ρ+β11)hπ2+β12h+(1−ρ+ρβ11+β12−β11β12)π(1−ρ+β11)β12h2π+(1−ρ+β11)hπ2+(β12+(1−ρ)(1−β11+β12))h+(2(1−ρ)+ρβ11+β12−β11β12)π,p2⋆(ρ,β21,β22,h,π)≡(1−ρ+β22)β21h2π+β21hπ2+(1−ρ+β22)h+((1−ρ)+β21+ρβ22−β21β22)π(1−ρ+β22)β21h2π+(β21+(1−ρ)(1+β21−β22))hπ2+(1−ρ+β22)h+(2(1−ρ)+β21+ρβ22−β21β22)π. array[]rclp_1 (ρ, _11, _12,h,π)&≡& (1-ρ+ _11) _12h^2π+(1-ρ+ _11)hπ^2+ _12h+(1-ρ+ρ _11+ _12- _11 _12)π(1-ρ+ _11) _12h^2π+(1-ρ+ _11)hπ^2+( _12+(1-ρ)(1- _11+ _12))h+(2(1-ρ)+ρ _11+ _12- _11 _12)π,\\ p_2 (ρ, _21, _22,h,π)&≡& (1-ρ+ _22) _21h^2π+ _21hπ^2+(1-ρ+ _22)h+((1-ρ)+ _21+ρ _22- _21 _22)π(1-ρ+ _22) _21h^2π+( _21+(1-ρ)(1+ _21- _22))hπ^2+(1-ρ+ _22)h+(2(1-ρ)+ _21+ρ _22- _21 _22)π. array As a consequence, we have Δ2(ρ,β11,β12,β21,β22,h,π)≡|(p1⋆(ρ,β11,β12,h,π),p2⋆(ρ,β21,β22,h,π))−(1,1)|. _2(ρ, _11, _12, _21, _22,h,π)≡|(p_1 (ρ, _11, _12,h,π),p_2 (ρ, _21, _22,h,π))-(1,1)|. Similarly, the efficient benchmark without any aggregator is (hπ2+πhπ2+h+2π,h+πhπ2+h+2π) ( hπ^2+πhπ^2+h+2π, h+πhπ^2+h+2π ). This leads to Δ0(h,π)≡|(hπ2+πhπ2+h+2π,h+πhπ2+h+2π)−(1,1)|. _0(h,π)≡ | ( hπ^2+πhπ^2+h+2π, h+πhπ^2+h+2π )-(1,1) |. By abuse of notation, we have Δ1(ρ,α,β1,β2,h,π)≡|(p1⋆(ρ,α,β1,β2,h,π),p2⋆(ρ,α,β1,β2,h,π))−(1,1)|. _1(ρ,α, _1, _2,h,π)≡ |(p_1 (ρ,α, _1, _2,h,π),p_2 (ρ,α, _1, _2,h,π))-(1,1) |. where pk⋆p_k denotes the topic-k consensus under the global-aggregator dynamics. Because p1(0)=(1,0)⊤p_1(0)=(1,0) and p2(0)=(0,1)⊤p_2(0)=(0,1) , we have p1⋆(ρ,α,β1,β2,h,π)+p2⋆(ρ,α,β1,β2,h,π)=1p_1 (ρ,α, _1, _2,h,π)+p_2 (ρ,α, _1, _2,h,π)=1. This leads to Δ1(ρ,α,β1,β2,h,π)≡|(p1⋆(ρ,α,β1,β2,h,π),1−p1⋆(ρ,α,β1,β2,h,π))−(1,1)|, _1(ρ,α, _1, _2,h,π)≡ |(p_1 (ρ,α, _1, _2,h,π),1-p_1 (ρ,α, _1, _2,h,π))-(1,1) |, where p1⋆(ρ,α,β1,β2,h,π)∈(0,1)p_1 (ρ,α, _1, _2,h,π)∈(0,1) is the topic-11 consensus. A.3 Proofs from Section 4 Proof of Theorem 2. We rewrite the learning gap with a global aggregator (ρ,α,β1,β2)(ρ,α, _1, _2) as Δ1(ρ,α,β1,β2,h,π)=|ϕ¯1(ρ,α,β1,β2,h,π)ϕ¯1(ρ,α,β1,β2,h,π)−π+1|, _1(ρ,α, _1, _2,h,π)= | φ_1(ρ,α, _1, _2,h,π) φ_1(ρ,α, _1, _2,h,π)- π+1 |, where ϕ¯1 φ_1 and ϕ¯1 φ_1 are defined by ϕ¯1(ρ,α,β1,β2,h,π) φ_1(ρ,α, _1, _2,h,π) = = (αβ1β2+(1−ρ)αβ2)h2π+(αβ1+(1−ρ)(1+(1−α)(β1−β2)))hπ2 (α _1 _2+(1-ρ)α _2)h^2π+(α _1+(1-ρ)(1+(1-α)( _1- _2)))hπ^2 +αβ2h+(α(β1+β2−β1β2)+(1−ρ)(1−αβ1))π, +α _2h+(α( _1+ _2- _1 _2)+(1-ρ)(1-α _1))π, ϕ¯1(ρ,α,β1,β2,h,π) φ_1(ρ,α, _1, _2,h,π) = = (β1β2+(1−ρ)(β1−α(β1−β2))h2π+(β1+(1−ρ)(1+(1−α)(β1−β2)))hπ2 ( _1 _2+(1-ρ)( _1-α( _1- _2))h^2π+( _1+(1-ρ)(1+(1-α)( _1- _2)))hπ^2 +(β2+(1−ρ)(1−α(β1−β2)))h+(β1+β2−β1β2+(1−ρ)(2−β2−α(β1−β2)))π. +( _2+(1-ρ)(1-α( _1- _2)))h+( _1+ _2- _1 _2+(1-ρ)(2- _2-α( _1- _2)))π. The learning gap without a global aggregator is Δ0(h,π)=|hπ2+πhπ2+h+2π−π+1|. _0(h,π)= | hπ^2+πhπ^2+h+2π- π+1 |. By definition, we have Λρ=α∈[0,1]∣Δ1(ρ,α,β1,β2,h,π)<Δ0(h,π),∀h∈[h¯,h¯],∀β1,β2∈(0,1) _ρ= \α∈[0,1] _1(ρ,α, _1, _2,h,π)< _0(h,π),\,∀ h∈[ h, h],∀ _1, _2∈(0,1) \ Fixing β1,β2∈(0,1) _1, _2∈(0,1) and h∈[h¯,h¯]h∈[ h, h], we have that Δ1(ρ,α,β1,β2,h,π)<Δ0(h,π) _1(ρ,α, _1, _2,h,π)< _0(h,π) if and only if 2π+1−hπ2+πhπ2+h+2π<ϕ¯1(ρ,α,β1,β2,h,π)ϕ¯1(ρ,α,β1,β2,h,π)<hπ2+πhπ2+h+2π. 2π+1- hπ^2+πhπ^2+h+2π< φ_1(ρ,α, _1, _2,h,π) φ_1(ρ,α, _1, _2,h,π)< hπ^2+πhπ^2+h+2π. Because ϕ¯1(ρ,α,β1,β2,h,π)>0 φ_1(ρ,α, _1, _2,h,π)>0, we have (2π+1−hπ2+πhπ2+h+2π)ϕ¯1(ρ,α,β1,β2,h,π)<ϕ¯1(ρ,α,β1,β2,h,π)<(hπ2+πhπ2+h+2π)ϕ¯1(ρ,α,β1,β2,h,π). ( 2π+1- hπ^2+πhπ^2+h+2π ) φ_1(ρ,α, _1, _2,h,π)< φ_1(ρ,α, _1, _2,h,π)< ( hπ^2+πhπ^2+h+2π ) φ_1(ρ,α, _1, _2,h,π). This yields two inequalities as follows, (2π+1−hπ2+πhπ2+h+2π)(β1(β2+1−ρ)h2π+(β1+(1−ρ)(1+β1−β2))hπ2+(β2+1−ρ)h+(β1+β2−β1β2+(1−ρ)(2−β2))π)−(1−ρ)((β1−β2+1)hπ2+π)<α((1−ρ)(β1−β2)(h2π+hπ2+h+π)(2π+1−hπ2+πhπ2+h+2π)+β2(β1+1−ρ)h2π+(β1−(1−ρ)(β1−β2))hπ2+β2h+(β1+β2−β1β2−(1−ρ)β1)π) array[]rcl&& ( 2π+1- hπ^2+πhπ^2+h+2π ) ( _1( _2+1-ρ)h^2π+( _1+(1-ρ)(1+ _1- _2))hπ^2+( _2+1-ρ)h .\\ && .+( _1+ _2- _1 _2+(1-ρ)(2- _2))π )-(1-ρ) (( _1- _2+1)hπ^2+π )\\ &<&α ((1-ρ)( _1- _2)(h^2π+hπ^2+h+π) ( 2π+1- hπ^2+πhπ^2+h+2π ) .\\ && .+ _2( _1+1-ρ)h^2π+( _1-(1-ρ)( _1- _2))hπ^2+ _2h+( _1+ _2- _1 _2-(1-ρ) _1)π ) array (8) and (hπ2+πhπ2+h+2π)(β1(β2+1−ρ)h2π+(β1+(1−ρ)(1+β1−β2))hπ2+(β2+1−ρ)h+(β1+β2−β1β2+(1−ρ)(2−β2))π)−(1−ρ)((β1−β2+1)hπ2+π)>α((1−ρ)(β1−β2)(h2π+hπ2+h+π)(hπ2+πhπ2+h+2π)+β2(β1+1−ρ)h2π+(β1−(1−ρ)(β1−β2))hπ2+β2h+(β1+β2−β1β2−(1−ρ)β1)π) array[]rcl&& ( hπ^2+πhπ^2+h+2π ) ( _1( _2+1-ρ)h^2π+( _1+(1-ρ)(1+ _1- _2))hπ^2+( _2+1-ρ)h .\\ && .+( _1+ _2- _1 _2+(1-ρ)(2- _2))π )-(1-ρ) (( _1- _2+1)hπ^2+π )\\ &>&α ((1-ρ)( _1- _2)(h^2π+hπ^2+h+π) ( hπ^2+πhπ^2+h+2π ) .\\ && .+ _2( _1+1-ρ)h^2π+( _1-(1-ρ)( _1- _2))hπ^2+ _2h+( _1+ _2- _1 _2-(1-ρ) _1)π ) array (9) The coefficients of α in Eq. (8) can be rewritten as β1π(hπ+1)+β2(h+π)+β1β2π(h2−1)+(1−ρ)π(h−1)(π+1)(hπ2+h+2π)(β1(hπ+1)E1+β2(h+π)E2), _1π(hπ+1)+ _2(h+π)+ _1 _2π(h^2-1)+ (1-ρ)π(h-1)(π+1)(hπ^2+h+2π) ( _1(hπ+1)E_1+ _2(h+π)E_2 ), where E1=hπ2−hπ+2h−π2+3π>0E_1=hπ^2-hπ+2h-π^2+3π>0 and E2=2hπ2−hπ+h+3π−1>0E_2=2hπ^2-hπ+h+3π-1>0. Similarly, the coefficient α in Eq. (9) can be rewritten as β1β2π(h2−1)+[β1π(hπ+1)+β2(h+π)][(1−ρ)h2π+hπ2+h+(1+ρ)π]hπ2+h+2π>0. _1 _2π(h^2-1)+ [ _1π(hπ+1)+ _2(h+π)][(1-ρ)h^2π+hπ^2+h+(1+ρ)π]hπ^2+h+2π>0. Both coefficients are strictly positive. Thus, we have α¯(ρ,β1,β2,h,π)<α<α¯(ρ,β1,β2,h,π), α(ρ, _1, _2,h,π)<α< α(ρ, _1, _2,h,π), where α¯(ρ,β1,β2,h,π) α(ρ, _1, _2,h,π) = = (hπ2+πhπ2+h+2π)(β1(β2+1−ρ)h2π+(β1+(1−ρ)(1+β1−β2))hπ2+(β2+1−ρ)h+(β1+β2−β1β2+(1−ρ)(2−β2))π)−(1−ρ)((1+β1−β2)hπ2+π)(1−ρ)(β1−β2)(h2π+hπ2+h+π)(hπ2+πhπ2+h+2π)+β2(β1+1−ρ)h2π+(β1−(1−ρ)(β1−β2))hπ2+β2h+(β1+β2−β1β2−(1−ρ)β1)π, ( hπ^2+πhπ^2+h+2π )( _1( _2+1-ρ)h^2π+( _1+(1-ρ)(1+ _1- _2))hπ^2+( _2+1-ρ)h+( _1+ _2- _1 _2+(1-ρ)(2- _2))π)-(1-ρ)((1+ _1- _2)hπ^2+π)(1-ρ)( _1- _2)(h^2π+hπ^2+h+π) ( hπ^2+πhπ^2+h+2π )+ _2( _1+1-ρ)h^2π+( _1-(1-ρ)( _1- _2))hπ^2+ _2h+( _1+ _2- _1 _2-(1-ρ) _1)π, and α¯(ρ,β1,β2,h,π) α(ρ, _1, _2,h,π) = = (2π+1−hπ2+πhπ2+h+2π)(β1(β2+1−ρ)h2π+(β1+(1−ρ)(1+β1−β2))hπ2+(β2+1−ρ)h+(β1+β2−β1β2+(1−ρ)(2−β2))π)−(1−ρ)((1+β1−β2)hπ2+π)(1−ρ)(β1−β2)(h2π+hπ2+h+π)(2π+1−hπ2+πhπ2+h+2π)+β2(β1+1−ρ)h2π+(β1−(1−ρ)(β1−β2))hπ2+β2h+(β1+β2−β1β2−(1−ρ)β1)π. ( 2π+1- hπ^2+πhπ^2+h+2π )( _1( _2+1-ρ)h^2π+( _1+(1-ρ)(1+ _1- _2))hπ^2+( _2+1-ρ)h+( _1+ _2- _1 _2+(1-ρ)(2- _2))π)-(1-ρ)((1+ _1- _2)hπ^2+π)(1-ρ)( _1- _2)(h^2π+hπ^2+h+π) ( 2π+1- hπ^2+πhπ^2+h+2π )+ _2( _1+1-ρ)h^2π+( _1-(1-ρ)( _1- _2))hπ^2+ _2h+( _1+ _2- _1 _2-(1-ρ) _1)π. We then show that α¯(ρ,β1,β2,h,π)<α¯(ρ,β1,β2,h,π), for all ρ,β1,β2∈(0,1) and h,π>1. α(ρ, _1, _2,h,π)< α(ρ, _1, _2,h,π), for all ρ, _1, _2∈(0,1) and h,π>1. (10) For simplicity, we define M1:=hπ2+πhπ2+h+2π,M2:=2π+1−M1,M_1:= hπ^2+πhπ^2+h+2π, M_2:= 2π+1-M_1, and D1:=β1(β2+1−ρ)h2π+(β1+(1−ρ)(1+β1−β2))hπ2+(β2+1−ρ)h+(β1+β2−β1β2+(1−ρ)(2−β2))π,D2:=(1−ρ)((β1−β2+1)hπ2+π),D3:=(1−ρ)(β1−β2)(h2π+hπ2+h+π),D4:=β2(β1+1−ρ)h2π+(ρβ1+(1−ρ)β2)hπ2+β2h+(β2(1−β1)+ρβ1)π. array[]rclD_1&:=& _1( _2+1-ρ)h^2π+( _1+(1-ρ)(1+ _1- _2))hπ^2+( _2+1-ρ)h\\ &&+( _1+ _2- _1 _2+(1-ρ)(2- _2))π,\\ D_2&:=&(1-ρ)(( _1- _2+1)hπ^2+π),\\ D_3&:=&(1-ρ)( _1- _2)(h^2π+hπ^2+h+π),\\ D_4&:=& _2( _1+1-ρ)h^2π+(ρ _1+(1-ρ) _2)hπ^2+ _2h+( _2(1- _1)+ρ _1)π. array Then, we have α¯(ρ,β1,β2,h,π)=M1D1−D2M1D3+D4,α¯(ρ,β1,β2,h,π)=M2D1−D2M2D3+D4. α(ρ, _1, _2,h,π)= M_1D_1-D_2M_1D_3+D_4, α(ρ, _1, _2,h,π)= M_2D_1-D_2M_2D_3+D_4. Because h,π>1h,π>1, we have 0<M1,M2<10<M_1,M_2<1. In addition, M1D3+D4>0M_1D_3+D_4>0 and M2D3+D4>0M_2D_3+D_4>0 because they are the coefficients of α in Eq. (8) and Eq. (9). A direct calculation yields α¯(ρ,β1,β2,h,π)−α¯(ρ,β1,β2,h,π)=(M1−M2)(D1D4+D2D3)(M1D3+D4)(M2D3+D4). α(ρ, _1, _2,h,π)- α(ρ, _1, _2,h,π)= (M_1-M_2)(D_1D_4+D_2D_3)(M_1D_3+D_4)(M_2D_3+D_4). We also have M1−M2=2π(h−1)(π−1)(π+1)(hπ2+h+2π)>0,M_1-M_2= 2π(h-1)(π-1)(π+1)(hπ^2+h+2π)>0, and D1D4+D2D3=(β1β2π(h2−1)+β1π(hπ+1)+β2(h+π))(β1β2π(h2−1) D_1D_4+D_2D_3=( _1 _2π(h^2-1)+ _1π(hπ+1)+ _2(h+π))( _1 _2π(h^2-1) +β1(h2π(1−ρ)+hπ2+πρ)+β2(h2π(1−ρ)+h+πρ)+(1−ρ)(h2π(1−ρ)+hπ2+h+π(1+ρ)))>0. + _1 (h^2π(1-ρ)+hπ^2+πρ)+ _2(h^2π(1-ρ)+h+πρ)+(1-ρ)(h^2π(1-ρ)+hπ^2+h+π(1+ρ)))>0. This yields Eq. (10). Putting these pieces together yields Λρ=[0,1]∩(supβ1,β2∈(0,1),h∈[h¯,h¯]α¯(ρ,β1,β2,h,π),infβ1,β2∈(0,1),h∈[h¯,h¯]α¯(ρ,β1,β2,h,π)). _ρ=[0,1]∩ ( _ _1, _2∈(0,1),h∈[ h, h] α(ρ, _1, _2,h,π), _ _1, _2∈(0,1),h∈[ h, h] α(ρ, _1, _2,h,π) ). (11) We introduce and prove two lemmas as follows, Lemma A.1. Suppose that π>1π>1 is fixed. Then, we have that α¯(ρ,β1,β2,h,π) α(ρ, _1, _2,h,π) and α¯(ρ,β1,β2,h,π) α(ρ, _1, _2,h,π) are continuous and strictly increasing in β1 _1 on the interval [0,1][0,1] for all ρ,β2∈(0,1)ρ, _2∈(0,1) and h>1h>1. Proof. We have ∂α¯∂β1(ρ,β1,β2,h,π)=P′(β1)Q(β1)−P(β1)Q′(β1)(Q(β1))2, ∂ α∂ _1(ρ, _1, _2,h,π)= P ( _1)Q( _1)-P( _1)Q ( _1)(Q( _1))^2, where P(β1)=(hπ2+πhπ2+h+2π)(β1(β2+1−ρ)h2π+(β1+(1−ρ)(1+β1−β2))hπ2+(β2+1−ρ)h+(β1+β2−β1β2+(1−ρ)(2−β2))π)−(1−ρ)((β1−β2+1)hπ2+π),Q(β1)=(1−ρ)(β1−β2)(h2π+hπ2+h+π)(hπ2+πhπ2+h+2π)+β2(β1+1−ρ)h2π+(β1−(1−ρ)(β1−β2))hπ2+β2h+(β1+β2−β1β2−(1−ρ)β1)π. array[]rclP( _1)&=& ( hπ^2+πhπ^2+h+2π ) ( _1( _2+1-ρ)h^2π+( _1+(1-ρ)(1+ _1- _2))hπ^2+( _2+1-ρ)h .\\ && .+( _1+ _2- _1 _2+(1-ρ)(2- _2))π )-(1-ρ) (( _1- _2+1)hπ^2+π ),\\ Q( _1)&=&(1-ρ)( _1- _2)(h^2π+hπ^2+h+π) ( hπ^2+πhπ^2+h+2π )\\ &&+ _2( _1+1-ρ)h^2π+( _1-(1-ρ)( _1- _2))hπ^2+ _2h+( _1+ _2- _1 _2-(1-ρ) _1)π. array It suffices to show that P′(β1)Q(β1)−P(β1)Q′(β1)>0P ( _1)Q( _1)-P( _1)Q ( _1)>0 for all ρ,β2∈(0,1)ρ, _2∈(0,1) and h,π>1h,π>1. Indeed, we have P′(β1)Q(β1)−P(β1)Q′(β1)=(β2π3(1−ρ)(h−1)2(h+1)2(hπ2+h+2π)2)L(ρ,β2,h,π),P ( _1)Q( _1)-P( _1)Q ( _1)= ( _2π^3(1-ρ)(h-1)^2(h+1)^2(hπ^2+h+2π)^2 )L(ρ, _2,h,π), where L(ρ,β2,h,π)=π(1+β2−ρ)h2+(π2+1)h+π(1−β2+ρ)L(ρ, _2,h,π)=π(1+ _2-ρ)h^2+(π^2+1)h+π(1- _2+ρ). For all ρ,β2∈(0,1)ρ, _2∈(0,1) and h>1h>1, we have L(ρ,β2,h,π)>0L(ρ, _2,h,π)>0. This yields the desired result. By abuse of notation, we also have ∂α¯∂β1(ρ,β1,β2,h,π)=P′(β1)Q(β1)−P(β1)Q′(β1)(Q(β1))2, ∂ α∂ _1(ρ, _1, _2,h,π)= P ( _1)Q( _1)-P( _1)Q ( _1)(Q( _1))^2, where P(β1)=(2π+1−hπ2+πhπ2+h+2π)(β1(β2+1−ρ)h2π+(β1+(1−ρ)(1+β1−β2))hπ2+(β2+1−ρ)h+(β1+β2−β1β2+(1−ρ)(2−β2))π)−(1−ρ)((β1−β2+1)hπ2+π),Q(β1)=(1−ρ)(β1−β2)(h2π+hπ2+h+π)(2π+1−hπ2+πhπ2+h+2π)+β2(β1+1−ρ)h2π+(β1−(1−ρ)(β1−β2))hπ2+β2h+(β1+β2−β1β2−(1−ρ)β1)π. array[]rclP( _1)&=& ( 2π+1- hπ^2+πhπ^2+h+2π ) ( _1( _2+1-ρ)h^2π+( _1+(1-ρ)(1+ _1- _2))hπ^2+( _2+1-ρ)h .\\ && .+( _1+ _2- _1 _2+(1-ρ)(2- _2))π )-(1-ρ) (( _1- _2+1)hπ^2+π ),\\ Q( _1)&=&(1-ρ)( _1- _2)(h^2π+hπ^2+h+π) ( 2π+1- hπ^2+πhπ^2+h+2π )\\ &&+ _2( _1+1-ρ)h^2π+( _1-(1-ρ)( _1- _2))hπ^2+ _2h+( _1+ _2- _1 _2-(1-ρ) _1)π. array It suffices to show that P′(β1)Q(β1)−P(β1)Q′(β1)>0P ( _1)Q( _1)-P( _1)Q ( _1)>0 for all ρ,β2∈(0,1)ρ, _2∈(0,1) and h,π>1h,π>1. Indeed, we have P′(β1)Q(β1)−P(β1)Q′(β1)=(π2(1−ρ)(h−1)(π+1)2(hπ2+h+2π)2)L(ρ,β2,h,π),P ( _1)Q( _1)-P( _1)Q ( _1)= ( π^2(1-ρ)(h-1)(π+1)^2(hπ^2+h+2π)^2 )L(ρ, _2,h,π), where L(ρ,β2,h,π)=L0(β2,h,π)((1−ρ)L1(β2,h,π)+ρL2(β2,h,π))L(ρ, _2,h,π)=L_0( _2,h,π)((1-ρ)L_1( _2,h,π)+ρ L_2( _2,h,π)) with L0(β2,h,π)=β2A(h,π)+B(h,π),L1(β2,h,π)=β2C(h,π)+D(h,π),L2(β2,h,π)=β2C(h,π)+E(h,π),L_0( _2,h,π)= _2A(h,π)+B(h,π), L_1( _2,h,π)= _2C(h,π)+D(h,π), L_2( _2,h,π)= _2C(h,π)+E(h,π), where A(h,π)=2h3π3−h3π2+h3π+3h2π2−h2π−2hπ3+hπ2−hπ−3π2+π,B(h,π)=2h2π4−2h2π3+2h2π2−2h2π+6hπ3−6hπ2+2hπ−2h+4π2−4π,C(h,π)=h2π2−h2π+2h2−2hπ2+4hπ−2h+π2−3π,D(h,π)=h2π2−h2π+2h2+hπ3−hπ2+5hπ−h+3π2−π,E(h,π)=hπ3+hπ2+hπ+h+2π2+2π. array[]rclA(h,π)&=&2h^3π^3-h^3π^2+h^3π+3h^2π^2-h^2π-2hπ^3+hπ^2-hπ-3π^2+π,\\ B(h,π)&=&2h^2π^4-2h^2π^3+2h^2π^2-2h^2π+6hπ^3-6hπ^2+2hπ-2h+4π^2-4π,\\ C(h,π)&=&h^2π^2-h^2π+2h^2-2hπ^2+4hπ-2h+π^2-3π,\\ D(h,π)&=&h^2π^2-h^2π+2h^2+hπ^3-hπ^2+5hπ-h+3π^2-π,\\ E(h,π)&=&hπ^3+hπ^2+hπ+h+2π^2+2π. array Because h,π>1h,π>1, we have A(h,π)=π(h−1)(h+1)(h(2π2−π+1)+(3π−1))> 0,B(h,π)=2(π−1)(h2π(π2+1)+h(3π2+1)+2π)> 0,D(h,π)=h2(π2−π+2)+h(π3−π2+5π−1)+π(3π−1)> 0, array[]rclA(h,π)&=&π(h-1)(h+1)(h(2π^2-π+1)+(3π-1))\ >\ 0,\\ B(h,π)&=&2(π-1)(h^2π(π^2+1)+h(3π^2+1)+2π)\ >\ 0,\\ D(h,π)&=&h^2(π^2-π+2)+h(π^3-π^2+5π-1)+π(3π-1)\ >\ 0, array and D(h,π)+C(h,π)=2h2(π2−π+2)+h(π3−3π2+9π−3)+4π(π−1)> 0,E(h,π)+C(h,π)=h2(π2−π+2)+h(π3−π2+5π−1)+π(3π−1)> 0. array[]rclD(h,π)+C(h,π)&=&2h^2(π^2-π+2)+h(π^3-3π^2+9π-3)+4π(π-1)\ >\ 0,\\ E(h,π)+C(h,π)&=&h^2(π^2-π+2)+h(π^3-π^2+5π-1)+π(3π-1)\ >\ 0. array In addition, we have that L0(β2,h,π)L_0( _2,h,π), L1(β2,h,π)L_1( _2,h,π) and L2(β2,h,π)L_2( _2,h,π) are all linear in β2 _2 and β2∈(0,1) _2∈(0,1). Putting these pieces together yields that L0(β2,h,π)>0L_0( _2,h,π)>0, L1(β2,h,π)>0L_1( _2,h,π)>0 and L2(β2,h,π)>0L_2( _2,h,π)>0 for all β2∈(0,1) _2∈(0,1) and h,π>1h,π>1. By definition of L(⋅)L(·), we have L(ρ,β2,h,π)>0L(ρ, _2,h,π)>0 for all ρ,β2∈(0,1)ρ, _2∈(0,1) and h,π>1h,π>1. This yields the desired result. Lemma A.2. Suppose that π>1π>1 is fixed and h¯>2π h>2π. Then, we have that α¯(ρ,1,β2,h,π) α(ρ,1, _2,h,π) is continuous and strictly decreasing in β2 _2 on the interval [0,1][0,1] for all ρ∈(0,1)ρ∈(0,1) and h≥h¯h≥ h. Proof. We have α¯(ρ,1,β2,h,π)=(2π+1−hπ2+πhπ2+h+2π)((β2+1−ρ)h2π+(1+(1−ρ)(2−β2))hπ2+(β2+1−ρ)h+(1+(1−ρ)(2−β2))π)−(1−ρ)((2−β2)hπ2+π)(1−ρ)(1−β2)(h2π+hπ2+h+π)(2π+1−hπ2+πhπ2+h+2π)+β2(2−ρ)h2π+(1−(1−ρ)(1−β2))hπ2+β2h+ρπ. α(ρ,1, _2,h,π)= ( 2π+1- hπ^2+πhπ^2+h+2π ) (( _2+1-ρ)h^2π+(1+(1-ρ)(2- _2))hπ^2+( _2+1-ρ)h+(1+(1-ρ)(2- _2))π )-(1-ρ) ((2- _2)hπ^2+π )(1-ρ)(1- _2)(h^2π+hπ^2+h+π) ( 2π+1- hπ^2+πhπ^2+h+2π )+ _2(2-ρ)h^2π+(1-(1-ρ)(1- _2))hπ^2+ _2h+ρπ. This implies ∂α¯(ρ,1,β2,h,π)∂β2=P′(β2)Q(β2)−P(β2)Q′(β2)(Q(β2))2, ∂ α(ρ,1, _2,h,π)∂ _2= P ( _2)Q( _2)-P( _2)Q ( _2)(Q( _2))^2, where P(β2)=(2π+1−hπ2+πhπ2+h+2π)((β2+1−ρ)h2π+(1+(1−ρ)(2−β2))hπ2+(β2+1−ρ)h+(1+(1−ρ)(2−β2))π)−(1−ρ)((2−β2)hπ2+π),Q(β2)=(1−ρ)(1−β2)(h2π+hπ2+h+π)(2π+1−hπ2+πhπ2+h+2π)+β2(2−ρ)h2π+(1−(1−ρ)(1−β2))hπ2+β2h+ρπ. array[]rclP( _2)&=& ( 2π+1- hπ^2+πhπ^2+h+2π ) (( _2+1-ρ)h^2π+(1+(1-ρ)(2- _2))hπ^2+( _2+1-ρ)h .\\ && .+(1+(1-ρ)(2- _2))π )-(1-ρ) ((2- _2)hπ^2+π ),\\ Q( _2)&=&(1-ρ)(1- _2)(h^2π+hπ^2+h+π) ( 2π+1- hπ^2+πhπ^2+h+2π )\\ &&+ _2(2-ρ)h^2π+(1-(1-ρ)(1- _2))hπ^2+ _2h+ρπ. array It suffices to show that P′(β2)Q(β2)−P(β2)Q′(β2)<0P ( _2)Q( _2)-P( _2)Q ( _2)<0 for all ρ∈(0,1)ρ∈(0,1), h≥h¯h≥ h and π>1π>1. Indeed, we have P′(β2)Q(β2)−P(β2)Q′(β2)=(π(1−ρ)(h−1)(π+1)2(hπ2+h+2π)2)L(ρ,h,π),P ( _2)Q( _2)-P( _2)Q ( _2)= ( π(1-ρ)(h-1)(π+1)^2(hπ^2+h+2π)^2 )L(ρ,h,π), where L(ρ,h,π)=L0(h,π)+(1−ρ)L1(h,π)L(ρ,h,π)=L_0(h,π)+(1-ρ)L_1(h,π) with L0(h,π)=−(hπ+1)(h(2π2−π+1)−π2+3π)R(h,π),L1(h,π)=−π(h−1)(h(2π2−π+1)+3π−1)R(h,π),R(h,π)=h3(π3−π2+2π)−h2(3π3−5π2+2π−2)−h(2π4−π3+5π2−4π)−(3π3−π2). array[]rclL_0(h,π)&=&-(hπ+1)(h(2π^2-π+1)-π^2+3π)R(h,π),\\ L_1(h,π)&=&-π(h-1)(h(2π^2-π+1)+3π-1)R(h,π),\\ R(h,π)&=&h^3(π^3-π^2+2π)-h^2(3π^3-5π^2+2π-2)-h(2π^4-π^3+5π^2-4π)-(3π^3-π^2). array In what follows, we prove that R(h,π)>0R(h,π)>0 for all h≥h¯h≥ h. Indeed, we have ∂3R∂h3(h,π)=6(π3−π2+2π)=6π(π2−π+2)>0. ∂^3R∂ h^3(h,π)=6(π^3-π^2+2π)=6π(π^2-π+2)>0. This implies that ∂2R∂h2(h,π) ∂^2R∂ h^2(h,π) is strictly increasing in h. Because h≥h¯>2πh≥ h>2π and π>1π>1, we have ∂2R∂h2(h,π)>∂2R∂h2(2π,π)=12π4−18π3+34π2−4π+4=12π2(π−1)2+4π(π2−1)+2π3+22π2+4>0. array[]rcl ∂^2R∂ h^2(h,π)&>& ∂^2R∂ h^2(2π,π)=12π^4-18π^3+34π^2-4π+4\\ &=&12π^2(π-1)^2+4π(π^2-1)+2π^3+22π^2+4>0. array This implies that ∂R∂h(h,π) ∂ R∂ h(h,π) is strictly increasing in h on the interval [h¯,+∞)[ h,+∞). Because h≥h¯>2πh≥ h>2π and π>1π>1, we have ∂R∂h(h,π)>∂R∂h(2π,π)=12π5−26π4+45π3−13π2+12π=12π2(π−1)3+π2(π2−1)+9π4+9π3+12π>0. array[]rcl ∂ R∂ h(h,π)&>& ∂ R∂ h(2π,π)=12π^5-26π^4+45π^3-13π^2+12π\\ &=&12π^2(π-1)^3+π^2(π^2-1)+9π^4+9π^3+12π>0. array This implies that R(h,π)R(h,π) is strictly increasing in h on the interval [h¯,+∞)[ h,+∞). Because h≥h¯>2πh≥ h>2π and π>1π>1, we have R(h,π)>R(2π,π)=8π6−24π5+38π4−21π3+17π2=8π3(π−1)3+13π3(π−1)+π4+17π2>0. array[]rclR(h,π)&>&R(2π,π)=8π^6-24π^5+38π^4-21π^3+17π^2\\ &=&8π^3(π-1)^3+13π^3(π-1)+π^4+17π^2>0. array Because h≥h¯>2πh≥ h>2π and π>1π>1, we have that h(2π2−π+1)+3π−1>0h(2π^2-π+1)+3π-1>0 and h(2π2−π+1)−π2+3π≥2π(2π2−π+1)−π2+3π=4π3−3π2+5π>0.h(2π^2-π+1)-π^2+3π≥ 2π(2π^2-π+1)-π^2+3π=4π^3-3π^2+5π>0. Putting these pieces together yields that L0(h,π),L1(h,π)<0L_0(h,π),L_1(h,π)<0 for all h≥h¯h≥ h and π>1π>1. By definition of L(⋅)L(·), we have L(ρ,h,π)<0L(ρ,h,π)<0 for all ρ∈(0,1)ρ∈(0,1), h≥h¯h≥ h and π>1π>1. This yields the desired result. Back to the original proof of Theorem 2, we see from the definition of α¯(⋅) α(·) that α¯(ρ,0,β2,h,π)=π(ρ(h2π−π)−(2h2π+hπ2+h))(h+π)(ρ(h2π−π)−(h2π+hπ2+h+π)),for all β2∈(0,1). α(ρ,0, _2,h,π)= π(ρ(h^2π-π)-(2h^2π+hπ^2+h))(h+π)(ρ(h^2π-π)-(h^2π+hπ^2+h+π)), for all _2∈(0,1). Using Lemma A.1 and Lemma A.2, we have infβ1,β2∈(0,1),h∈[h¯,h¯]α¯(ρ,β1,β2,h,π)=infβ2∈(0,1),h∈[h¯,h¯]α¯(ρ,0,β2,h,π)=infh∈[h¯,h¯]g¯(ρ,h,π),supβ1,β2∈(0,1),h∈[h¯,h¯]α¯(ρ,β1,β2,h,π)=suph∈[h¯,h¯]α¯(ρ,1,0,h,π)=suph∈[h¯,h¯]g¯(ρ,h,π), array[]rcl _ _1, _2∈(0,1),h∈[ h, h] α(ρ, _1, _2,h,π)&=& _ _2∈(0,1),h∈[ h, h] α(ρ,0, _2,h,π)= _h∈[ h, h] g(ρ,h,π),\\ _ _1, _2∈(0,1),h∈[ h, h] α(ρ, _1, _2,h,π)&=& _h∈[ h, h] α(ρ,1,0,h,π)= _h∈[ h, h] g(ρ,h,π), array (12) where g¯ g and g¯ g are given by g¯(ρ,h,π)=π(ρ(h2π−π)−(2h2π+hπ2+h))(h+π)(ρ(h2π−π)−(h2π+hπ2+h+π)),g¯(ρ,h,π)=(2π+1−hπ2+πhπ2+h+2π)((1−ρ)h2π+(3−2ρ)hπ2+(1−ρ)h+(3−2ρ)π)−(1−ρ)(2hπ2+π)(1−ρ)(h2π+hπ2+h+π)(2π+1−hπ2+πhπ2+h+2π)+ρhπ2+ρπ. array[]rcl g(ρ,h,π)&=& π(ρ(h^2π-π)-(2h^2π+hπ^2+h))(h+π)(ρ(h^2π-π)-(h^2π+hπ^2+h+π)),\\ g(ρ,h,π)&=& ( 2π+1- hπ^2+πhπ^2+h+2π )((1-ρ)h^2π+(3-2ρ)hπ^2+(1-ρ)h+(3-2ρ)π)-(1-ρ)(2hπ^2+π)(1-ρ)(h^2π+hπ^2+h+π) ( 2π+1- hπ^2+πhπ^2+h+2π )+ρ hπ^2+ρπ. array Monotonicity results. We prove the monotonicity results as follows, • infβ1,β2∈(0,1),h∈[h¯,h¯]α¯(ρ,β1,β2,h,π) _ _1, _2∈(0,1),h∈[ h, h] α(ρ, _1, _2,h,π) is increasing in ρ on the interval (0,1)(0,1). • supβ1,β2∈(0,1),h∈[h¯,h¯]α¯(ρ,β1,β2,h,π) _ _1, _2∈(0,1),h∈[ h, h] α(ρ, _1, _2,h,π) is decreasing in ρ on the interval (0,1)(0,1). Based on Eq. (12), it suffices to show that • g¯(ρ,h,π) g(ρ,h,π) is increasing in ρ on the interval (0,1)(0,1) for any h∈[h¯,h¯]h∈[ h, h]. • g¯(ρ,h,π) g(ρ,h,π) is decreasing in ρ on the interval (0,1)(0,1) for any h∈[h¯,h¯]h∈[ h, h]. Indeed, we have ∂g¯∂ρ(ρ,h,π)=π3(h−1)2(h+1)2(h+π)(ρ(h2π−π)−(h2π+hπ2+h+π))2. ∂ g∂ρ(ρ,h,π)= π^3(h-1)^2(h+1)^2(h+π)(ρ(h^2π-π)-(h^2π+hπ^2+h+π))^2. Because h≥h¯>1h≥ h>1, we have (h−1)2>0(h-1)^2>0. In addition, ρ∈(0,1)ρ∈(0,1). Thus, we have ρ(h2π−π)−(h2π+hπ2+h+π)<−(hπ2+h+2π)<0.ρ(h^2π-π)-(h^2π+hπ^2+h+π)<-(hπ^2+h+2π)<0. Putting these pieces together yields that ∂g¯∂ρ(ρ,h,π)>0 ∂ g∂ρ(ρ,h,π)>0 for any ρ∈(0,1)ρ∈(0,1) and any h∈[h¯,h¯]h∈[ h, h]. This implies that g¯(ρ,h,π) g(ρ,h,π) is increasing in ρ on the interval (0,1)(0,1) for any h∈[h¯,h¯]h∈[ h, h]. Proceeding a further step, we have ∂g¯∂ρ(ρ,h,π)=−(h−1)R(h,π)(hπ+1)(ρA(h,π)−B(h,π))2, ∂ g∂ρ(ρ,h,π)=- (h-1)R(h,π)(hπ+1)(ρ A(h,π)-B(h,π))^2, where R(h,π)=h3(2π5−3π4+6π3−3π2+2π)−h2(2π6+4π5−11π4+14π3−15π2+4π−2)−h(6π5+13π4−18π3+13π2−10π)−(5π4+10π3−11π2),A(h,π)=(h−1)(hπ2−hπ+2h−π2+3π),B(h,π)=(h+π)(hπ2−hπ+2h+3π−1). array[]rclR(h,π)&=&h^3(2π^5-3π^4+6π^3-3π^2+2π)-h^2(2π^6+4π^5-11π^4+14π^3-15π^2+4π-2)\\ &&-h(6π^5+13π^4-18π^3+13π^2-10π)-(5π^4+10π^3-11π^2),\\ A(h,π)&=&(h-1)(hπ^2-hπ+2h-π^2+3π),\\ B(h,π)&=&(h+π)(hπ^2-hπ+2h+3π-1). array Because ρ∈(0,1)ρ∈(0,1), we have ρA(h,π)−B(h,π)<A(h,π)−B(h,π)=−(π+1)(hπ2+h+2π)<0.ρ A(h,π)-B(h,π)<A(h,π)-B(h,π)=-(π+1)(hπ^2+h+2π)<0. Putting these pieces together yields that the sign of ∂g¯∂ρ(ρ,h,π) ∂ g∂ρ(ρ,h,π) is the same as −R(h,π)-R(h,π). In what follows, we prove that R(h,π)>0R(h,π)>0 for any h∈[h¯,h¯]h∈[ h, h]. Indeed, we have ∂3R∂h3(h,π)=6(2π5−3π4+6π3−3π2+2π)=6π(π2−π+2)(2π2−π+1)>0. ∂^3R∂ h^3(h,π)=6(2π^5-3π^4+6π^3-3π^2+2π)=6π(π^2-π+2)(2π^2-π+1)>0. This implies that ∂2R∂h2(h,π) ∂^2R∂ h^2(h,π) is strictly increasing in h. Because h≥h¯>2πh≥ h>2π and π>1π>1, we have ∂2R∂h2(h,π)>∂2R∂h2(2π,π)=20π6−44π5+94π4−64π3+54π2−8π+4=20π3(π−1)3+16π3(π2−1)+34π3(π−1)+8π(π−1)+6π3+46π2+4>0. array[]rcl ∂^2R∂ h^2(h,π)&>& ∂^2R∂ h^2(2π,π)=20π^6-44π^5+94π^4-64π^3+54π^2-8π+4\\ &=&20π^3(π-1)^3+16π^3(π^2-1)+34π^3(π-1)+8π(π-1)+6π^3+46π^2+4>0. array This implies that ∂R∂h(h,π) ∂ R∂ h(h,π) is strictly increasing in h on the interval [h¯,+∞)[ h,+∞). Because h≥h¯>2πh≥ h>2π and π>1π>1, we have ∂R∂h(h,π)>∂R∂h(2π,π)=16π7−52π6+110π5−105π4+102π3−29π2+18π,=16π3(π−1)4+12π2(π2−1)2+14π3(π−1)2+41π2(π−1)+11π4+31π3+18π>0. array[]rcl ∂ R∂ h(h,π)&>& ∂ R∂ h(2π,π)=16π^7-52π^6+110π^5-105π^4+102π^3-29π^2+18π,\\ &=&16π^3(π-1)^4+12π^2(π^2-1)^2+14π^3(π-1)^2+41π^2(π-1)+11π^4+31π^3+18π>0. array This implies that R(h,π)R(h,π) is strictly increasing in h on the interval [h¯,+∞)[ h,+∞). Because h≥h¯>2πh≥ h>2π, we have R(h,π)>R(2π,π)=π2(8π6−40π5+80π4−106π3+107π2−52π+39).R(h,π)>R(2π,π)=π^2(8π^6-40π^5+80π^4-106π^3+107π^2-52π+39). Because π>1π>1, we let t=π−1>0t=π-1>0 for simplicity. Then, we have 8π6−40π5+80π4−106π3+107π2−52π+39=(8t6+36−26t3)+t(8t4+12−11t).8π^6-40π^5+80π^4-106π^3+107π^2-52π+39=(8t^6+36-26t^3)+t(8t^4+12-11t). For the first term, we have 8t6+36−26t3≥242t3−26t3>0.8t^6+36-26t^3≥ 24 2t^3-26t^3>0. For the second term, we have 8t4+12−11t≥4(8⋅4⋅4⋅4)1/4t−11t=(1624−11)t>0.8t^4+12-11t≥ 4(8· 4· 4· 4)^1/4t-11t=(16 [4]2-11)t>0. Putting these pieces together yields that R(h,π)>0R(h,π)>0 for any h∈[h¯,h¯]h∈[ h, h]. Thus, we have that ∂g¯∂ρ(ρ,h,π)<0 ∂ g∂ρ(ρ,h,π)<0 for any ρ∈(0,1)ρ∈(0,1) and any h∈[h¯,h¯]h∈[ h, h]. This implies that g¯(ρ,h,π) g(ρ,h,π) is decreasing in ρ on the interval (0,1)(0,1) for any h∈[h¯,h¯]h∈[ h, h]. Boundary results. We prove the boundary results as follows, • The following statement holds true, supβ1,β2∈(0,1),h∈[h¯,h¯]α¯(ρ,β1,β2,h,π)≥infβ1,β2∈(0,1),h∈[h¯,h¯]α¯(ρ,β1,β2,h,π),for all ρ∈(0,12]. _ _1, _2∈(0,1),h∈[ h, h] α(ρ, _1, _2,h,π)≥ _ _1, _2∈(0,1),h∈[ h, h] α(ρ, _1, _2,h,π), for all ρ∈(0, 12]. • Suppose that ϵ∈(0,12(h¯π2+πh¯π2+h¯+2π−π+1))ε∈ (0, 12 ( hπ^2+π hπ^2+ h+2π- π+1 ) ) is fixed. Then, there exists δ∈(0,12)δ∈(0, 12) such that, for all ρ∈(1−δ,1)ρ∈(1-δ,1), the following statement holds true, infβ1,β2∈(0,1),h∈[h¯,h¯]α¯(ρ,β1,β2,h,π)≥h¯π2+πh¯π2+h¯+2π−ϵ,supβ1,β2∈(0,1),h∈[h¯,h¯]α¯(ρ,β1,β2,h,π)≤(2π+1−h¯π2+πh¯π2+h¯+2π)+ϵ. array[]rcl _ _1, _2∈(0,1),h∈[ h, h] α(ρ, _1, _2,h,π)&≥& hπ^2+π hπ^2+ h+2π-ε,\\ _ _1, _2∈(0,1),h∈[ h, h] α(ρ, _1, _2,h,π)&≤& ( 2π+1- hπ^2+π hπ^2+ h+2π )+ε. array Based on Eq. (12), it suffices to show that • The following statement holds true, suph∈[h¯,h¯]g¯(ρ,h,π)≥infh∈[h¯,h¯]g¯(ρ,h,π),for all ρ∈(0,12]. _h∈[ h, h] g(ρ,h,π)≥ _h∈[ h, h] g(ρ,h,π), for all ρ∈(0, 12]. • Suppose that ϵ∈(0,12(h¯π2+πh¯π2+h¯+2π−π+1))ε∈ (0, 12 ( hπ^2+π hπ^2+ h+2π- π+1 ) ) is fixed. Then, there exists δ∈(0,12)δ∈(0, 12) such that, for all ρ∈(1−δ,1)ρ∈(1-δ,1), the following statement holds true, infh∈[h¯,h¯]g¯(ρ,h,π)≥h¯π2+πh¯π2+h¯+2π−ϵ,suph∈[h¯,h¯]g¯(ρ,h,π)≤(2π+1−h¯π2+πh¯π2+h¯+2π)+ϵ. _h∈[ h, h] g(ρ,h,π)≥ hπ^2+π hπ^2+ h+2π-ε, _h∈[ h, h] g(ρ,h,π)≤ ( 2π+1- hπ^2+π hπ^2+ h+2π )+ε. Indeed, we have g¯(ρ,h,π)=π(ρ(h2π−π)−(2h2π+hπ2+h))(h+π)(ρ(h2π−π)−(h2π+hπ2+h+π)),g¯(ρ,h,π)=(2π+1−hπ2+πhπ2+h+2π)((1−ρ)h2π+(3−2ρ)hπ2+(1−ρ)h+(3−2ρ)π)−(1−ρ)(2hπ2+π)(1−ρ)(h2π+hπ2+h+π)(2π+1−hπ2+πhπ2+h+2π)+ρhπ2+ρπ. array[]rcl g(ρ,h,π)&=& π(ρ(h^2π-π)-(2h^2π+hπ^2+h))(h+π)(ρ(h^2π-π)-(h^2π+hπ^2+h+π)),\\ g(ρ,h,π)&=& ( 2π+1- hπ^2+πhπ^2+h+2π )((1-ρ)h^2π+(3-2ρ)hπ^2+(1-ρ)h+(3-2ρ)π)-(1-ρ)(2hπ^2+π)(1-ρ)(h^2π+hπ^2+h+π) ( 2π+1- hπ^2+πhπ^2+h+2π )+ρ hπ^2+ρπ. array For the first boundary result, it suffices to show g¯(ρ,h¯,π)>g¯(ρ,h¯,π),for all ρ∈(0,12]. g(ρ, h,π)> g(ρ, h,π), for all ρ∈(0, 12]. (13) Because π>1π>1, ρ∈(0,12]ρ∈(0, 12] and h¯>20π>1 h>20π>1, we have ρ(h¯2π−π)−(2h¯2π+h¯π2+h¯)<0,ρ(h¯2π−π)−(h¯2π+h¯π2+h¯+π)<0.ρ( h^2π-π)-(2 h^2π+ hπ^2+ h)<0, ρ( h^2π-π)-( h^2π+ hπ^2+ h+π)<0. Thus, we rewrite g¯(ρ,h¯,π)=π((2h¯2π+h¯π2+h¯)−ρ(h¯2π−π))(h¯+π)((h¯2π+h¯π2+h¯+π)−ρ(h¯2π−π))=π((2−ρ)h¯2π+(π2+1)h¯+ρπ)(h¯+π)((1−ρ)h¯2π+(π2+1)h¯+(1+ρ)π). g(ρ, h,π)= π((2 h^2π+ hπ^2+ h)-ρ( h^2π-π))( h+π)(( h^2π+ hπ^2+ h+π)-ρ( h^2π-π))= π((2-ρ) h^2π+(π^2+1) h+ρπ)( h+π)((1-ρ) h^2π+(π^2+1) h+(1+ρ)π). Note that (1−ρ)h¯2π+(π2+1)h¯+(1+ρ)π>(1−ρ)h¯2π>0(1-ρ) h^2π+(π^2+1) h+(1+ρ)π>(1-ρ) h^2π>0 and h¯+π>h¯>0 h+π> h>0. Thus, we have g¯(ρ,h¯,π)<(2−ρ)h¯2π+(π2+1)h¯+ρπ(1−ρ)h¯3=(2−ρ)π(1−ρ)h¯+π2+1(1−ρ)h¯2+ρπ(1−ρ)h¯3. g(ρ, h,π)< (2-ρ) h^2π+(π^2+1) h+ρπ(1-ρ) h^3= (2-ρ)π(1-ρ) h+ π^2+1(1-ρ) h^2+ ρπ(1-ρ) h^3. Because ρ∈(0,12]ρ∈(0, 12], we have 2−ρ1−ρ=1+11−ρ≤3,11−ρ≤2,ρ1−ρ≤1. 2-ρ1-ρ=1+ 11-ρ≤ 3, 11-ρ≤ 2, ρ1-ρ≤ 1. Putting these pieces together yields g¯(ρ,h¯,π)≤3πh¯+2(π2+1)h¯2+πh¯3,for all ρ∈(0,12]. g(ρ, h,π)≤ 3π h+ 2(π^2+1) h^2+ π h^3, for all ρ∈(0, 12]. Because h¯>20π h>20π, we have 3πh¯<320,2(π2+1)h¯2<2(π2+1)400π2<4π2400π2=1100,πh¯3<π8000π3=18000π2<18000. 3π h< 320, 2(π^2+1) h^2< 2(π^2+1)400π^2< 4π^2400π^2= 1100, π h^3< π8000π^3= 18000π^2< 18000. Putting these pieces together yields g¯(ρ,h¯,π)<320+1100+18000<0.17,for all ρ∈(0,12]. g(ρ, h,π)< 320+ 1100+ 18000<0.17, for all ρ∈(0, 12]. (14) Because π>1π>1 and h¯>2π h>2π, we have 12≤2π+1−h¯π2+πh¯π2+h¯+2π<1. 12≤ 2π+1- hπ^2+π hπ^2+ h+2π<1. Then, we have (2π+1−h¯π2+πh¯π2+h¯+2π)((1−ρ)h¯2π+(3−2ρ)h¯π2+(1−ρ)h¯+(3−2ρ)π)−(1−ρ)(2h¯π2+π)>12(1−ρ)h¯2π−(1−ρ)(2h¯π2+π)=(1−ρ)(12h¯2π−2h¯π2−π) array[]rcl&& ( 2π+1- hπ^2+π hπ^2+ h+2π )((1-ρ) h^2π+(3-2ρ) hπ^2+(1-ρ) h+(3-2ρ)π)-(1-ρ)(2 hπ^2+π)\\ &>& 12(1-ρ) h^2π-(1-ρ)(2 hπ^2+π)\ =\ (1-ρ)( 12 h^2π-2 hπ^2-π) array Because ρ∈(0,12]ρ∈(0, 12], we have (2π+1−h¯π2+πh¯π2+h¯+2π)((1−ρ)h¯2π+(3−2ρ)h¯π2+(1−ρ)h¯+(3−2ρ)π)−(1−ρ)(2h¯π2+π)>14h¯2π−h¯π2−12π. ( 2π+1- hπ^2+π hπ^2+ h+2π )((1-ρ) h^2π+(3-2ρ) hπ^2+(1-ρ) h+(3-2ρ)π)-(1-ρ)(2 hπ^2+π)> 14 h^2π- hπ^2- 12π. We also have (1−ρ)(h¯2π+h¯π2+h¯+π)(2π+1−h¯π2+πh¯π2+h¯+2π)+ρh¯π2+ρπ<(1−ρ)(h¯2π+h¯π2+h¯+π)+12(h¯π2+π)<(h¯2π+h¯π2+h¯+π)+12(h¯π2+π)=h¯2π+32h¯π2+h¯+32π. array[]rcl&&(1-ρ)( h^2π+ hπ^2+ h+π) ( 2π+1- hπ^2+π hπ^2+ h+2π )+ρ hπ^2+ρπ\\ &<&(1-ρ)( h^2π+ hπ^2+ h+π)+ 12( hπ^2+π)\ <\ ( h^2π+ hπ^2+ h+π)+ 12( hπ^2+π)\\ &=& h^2π+ 32 hπ^2+ h+ 32π. array Putting these pieces together yields g¯(ρ,h¯,π)>14h¯2π−h¯π2−12πh¯2π+32h¯π2+h¯+32π=14−πh¯−12h¯21+3π2h¯+1h¯π+32h¯2. g(ρ, h,π)> 14 h^2π- hπ^2- 12π h^2π+ 32 hπ^2+ h+ 32π= 14- π h- 12 h^21+ 3π2 h+ 1 hπ+ 32 h^2. Because h¯>20π h>20π and π>1π>1, we have 1h¯<120π<120,12h¯2<1800π2<1800,3π2h¯<340,1h¯π<120π2<120,32h¯2≤3800π2<3800. 1 h< 120π< 120, 12 h^2< 1800π^2< 1800, 3π2 h< 340, 1 hπ< 120π^2< 120, 32 h^2≤ 3800π^2< 3800. This implies 14−πh¯−12h¯2>14−120−1800=0.19875, 14- π h- 12 h^2> 14- 120- 1800=0.19875, and 1+3π2h¯+1h¯π+32h¯2<1+340+120+3800=1.12875.1+ 3π2 h+ 1 hπ+ 32 h^2<1+ 340+ 120+ 3800=1.12875. Putting these pieces together yields g¯(ρ,h¯,π)>0.198751.12875>0.17,for all ρ∈(0,12]. g(ρ, h,π)> 0.198751.12875>0.17, for all ρ∈(0, 12]. (15) Combining Eq. (14) and Eq. (15) yields the desired result in Eq. (13). For the second boundary result, we have infh∈[h¯,h¯]hπ2+πhπ2+h+2π=h¯π2+πh¯π2+h¯+2π,suph∈[h¯,h¯]2π+1−hπ2+πhπ2+h+2π=2π+1−h¯π2+πh¯π2+h¯+2π. _h∈[ h, h] \ hπ^2+πhπ^2+h+2π \= hπ^2+π hπ^2+ h+2π, _h∈[ h, h] \ 2π+1- hπ^2+πhπ^2+h+2π \= 2π+1- hπ^2+π hπ^2+ h+2π. (16) We rewrite g¯(ρ,h,π)=hπ2+πhπ2+h+2π−(1−ρ)(π3(h2−1)2(h+π)(hπ2+h+2π)(hπ2+h+2π+(1−ρ)π(h2−1))). g(ρ,h,π)= hπ^2+πhπ^2+h+2π-(1-ρ) ( π^3(h^2-1)^2(h+π)(hπ^2+h+2π)(hπ^2+h+2π+(1-ρ)π(h^2-1)) ). Because ρ∈(0,1)ρ∈(0,1) and h,π>1h,π>1, we have hπ2+h+2π+(1−ρ)π(h2−1)>hπ2+h+2π>0hπ^2+h+2π+(1-ρ)π(h^2-1)>hπ^2+h+2π>0. This implies g¯(ρ,h,π)>hπ2+πhπ2+h+2π−(1−ρ)(π3(h2−1)2(h+π)(hπ2+h+2π)2). g(ρ,h,π)> hπ^2+πhπ^2+h+2π-(1-ρ) ( π^3(h^2-1)^2(h+π)(hπ^2+h+2π)^2 ). Because h¯,h¯ h, h are finite, we have M1:=maxh∈[h¯,h¯]π3(h2−1)2(h+π)(hπ2+h+2π)2M_1:= _h∈[ h, h] \ π^3(h^2-1)^2(h+π)(hπ^2+h+2π)^2 \ is finite. This implies g¯(ρ,h,π)>hπ2+πhπ2+h+2π−(1−ρ)M1,for all h∈[h¯,h¯]. g(ρ,h,π)> hπ^2+πhπ^2+h+2π-(1-ρ)M_1, for all h∈[ h, h]. Choosing δ1:=min12,ϵM1∈(0,12) _1:= \ 12, εM_1 \∈(0, 12). If ρ∈(1−δ1,1)ρ∈(1- _1,1), we have g¯(ρ,h,π)>hπ2+πhπ2+h+2π−ϵ,for all h∈[h¯,h¯]. g(ρ,h,π)> hπ^2+πhπ^2+h+2π-ε, for all h∈[ h, h]. Taking the infimum over the interval [h¯,h¯][ h, h] and using Eq. (16) yields infh∈[h¯,h¯]g¯(ρ,h,π)≥h¯π2+πh¯π2+h¯+2π−ϵ≥0. _h∈[ h, h] g(ρ,h,π)≥ hπ^2+π hπ^2+ h+2π-ε≥ 0. (17) By abuse of notation, we rewrite g¯(ρ,h,π)=(2π+1−hπ2+πhπ2+h+2π)+(1−ρ)(P(h)−(2π+1−hπ2+πhπ2+h+2π)Q(h)hπ2+π+(1−ρ)Q(h)). g(ρ,h,π)= ( 2π+1- hπ^2+πhπ^2+h+2π )+(1-ρ) ( P(h)- ( 2π+1- hπ^2+πhπ^2+h+2π )Q(h)hπ^2+π+(1-ρ)Q(h) ). where P(h)=(2π+1−hπ2+πhπ2+h+2π)(h2π+2hπ2+h+2π)−(2hπ2+π),Q(h)=(2π+1−hπ2+πhπ2+h+2π)(h2π+hπ2+h+π)−(hπ2+π). array[]rclP(h)&=& ( 2π+1- hπ^2+πhπ^2+h+2π )(h^2π+2hπ^2+h+2π)-(2hπ^2+π),\\ Q(h)&=& ( 2π+1- hπ^2+πhπ^2+h+2π )(h^2π+hπ^2+h+π)-(hπ^2+π). array Because π>1π>1 and h¯>2π h>2π, we have 12≤2π+1−hπ2+πhπ2+h+2π<1,for all h∈[h¯,h¯]. 12≤ 2π+1- hπ^2+πhπ^2+h+2π<1, for all h∈[ h, h]. This implies Q(h)≥12(h2π−hπ2+h−π)=12(h−π)(hπ+1)>0,for all h∈[h¯,h¯].Q(h)≥ 12(h^2π-hπ^2+h-π)= 12(h-π)(hπ+1)>0, for all h∈[ h, h]. Because ρ∈(0,1)ρ∈(0,1), we have hπ2+π+(1−ρ)Q(h)>hπ2+π>π>1hπ^2+π+(1-ρ)Q(h)>hπ^2+π>π>1 for all h∈[h¯,h¯]h∈[ h, h]. Thus, we have g¯(ρ,h,π)−(2π+1−hπ2+πhπ2+h+2π)≤(1−ρ)|P(h)−(2π+1−hπ2+πhπ2+h+2π)Q(h)hπ2+π+(1−ρ)Q(h)|<(1−ρ)|P(h)−(2π+1−hπ2+πhπ2+h+2π)Q(h)|. g(ρ,h,π)- ( 2π+1- hπ^2+πhπ^2+h+2π )≤(1-ρ) | P(h)- ( 2π+1- hπ^2+πhπ^2+h+2π )Q(h)hπ^2+π+(1-ρ)Q(h) |<(1-ρ) |P(h)- ( 2π+1- hπ^2+πhπ^2+h+2π )Q(h) |. Because h¯,h¯ h, h are finite, we have M2:=maxh∈[h¯,h¯]|P(h)−(2π+1−hπ2+πhπ2+h+2π)Q(h)|M_2:= _h∈[ h, h] \ |P(h)- ( 2π+1- hπ^2+πhπ^2+h+2π )Q(h) | \ is finite. This implies g¯(ρ,h,π)<(2π+1−hπ2+πhπ2+h+2π)+(1−ρ)M2,for all h∈[h¯,h¯]. g(ρ,h,π)< ( 2π+1- hπ^2+πhπ^2+h+2π )+(1-ρ)M_2, for all h∈[ h, h]. Choosing δ2:=min12,ϵM2∈(0,12) _2:= \ 12, εM_2 \∈(0, 12). If ρ∈(1−δ2,1)ρ∈(1- _2,1), we have g¯(ρ,h,π)<(2π+1−hπ2+πhπ2+h+2π)+ϵ,for all h∈[h¯,h¯]. g(ρ,h,π)< ( 2π+1- hπ^2+πhπ^2+h+2π )+ε, all h∈[ h, h]. Taking the supremum over the interval [h¯,h¯][ h, h] and using Eq. (16) yields suph∈[h¯,h¯]g¯(ρ,h,π)≤(2π+1−h¯π2+πh¯π2+h¯+2π)+ϵ≤1. _h∈[ h, h] g(ρ,h,π)≤ ( 2π+1- hπ^2+π hπ^2+ h+2π )+ε≤ 1. (18) Combining Eq. (17) and Eq. (18) and choosing δ:=minδ1,δ2∈(0,12)δ:= \ _1, _2\∈(0, 12) yields the desired result. By definition of Λρ _ρ (see Eq. (11)), we have Λρ=[0,1]∩(supβ1,β2∈(0,1),h∈[h¯,h¯]α¯(ρ,β1,β2,h,π),infβ1,β2∈(0,1),h∈[h¯,h¯]α¯(ρ,β1,β2,h,π)). _ρ=[0,1]∩ ( _ _1, _2∈(0,1),h∈[ h, h] α(ρ, _1, _2,h,π), _ _1, _2∈(0,1),h∈[ h, h] α(ρ, _1, _2,h,π) ). The monotonicity results guarantee that Λρ1⊆Λρ2 _ _1 _ _2 if 0<ρ1≤ρ2<10< _1≤ _2<1. The boundary results guarantee that Λρ=∅ _ρ= for all ρ∈(0,12]ρ∈(0, 12] and there exists δ∈(0,12)δ∈(0, 12) such that Λρ≠∅ _ρ≠ for all ρ∈(1−δ,1)ρ∈(1-δ,1). Putting these pieces together yields that μ(Λρ)μ( _ρ) as a function of ρ on the interval (0,1)(0,1) is nondecreasing and satisfies that μ(Λρ)=0μ( _ρ)=0 for all ρ∈(0,12]ρ∈(0, 12] and μ(Λρ)>0μ( _ρ)>0 for all ρ∈(1−δ,1)ρ∈(1-δ,1). We define ρ⋆=supρ∈(0,1):μ(Λρ)=0ρ = \ρ∈(0,1):μ( _ρ)=0\. Then, the previous results guarantee that 12≤ρ⋆≤1−δ 12≤ρ ≤ 1-δ and the following statement holds true, 1. If ρ<ρ⋆ρ<ρ , then μ(Λρ)=0μ( _ρ)=0. 2. If ρ>ρ⋆ρ>ρ , then μ(Λρ)>0μ( _ρ)>0. This completes the proof. A.4 Proofs from Section 5 Proof of Proposition 2. We have Δ⋆≡Δ1(ρ,α,β,β,h,π)−Δ0(h,π)=|Δ¯1(ρ,α,β,h,π)|−Δ0(h,π), ≡ _1(ρ,α,β,β,h,π)- _0(h,π)=| _1(ρ,α,β,h,π)|- _0(h,π), where Δ¯1(ρ,α,β,h,π)=(1−ρ)(hπ2+π)+α(β(1+β−ρ)h2π+βhπ2+βh+β(1−β+ρ)π)(1−ρ)(hπ2+2π+h)+(β(1+β−ρ)h2π+βhπ2+βh+β(1−β+ρ)π)−π+1. _1(ρ,α,β,h,π)= (1-ρ)(hπ^2+π)+α(β(1+β-ρ)h^2π+β hπ^2+β h+β(1-β+ρ)π)(1-ρ)(hπ^2+2π+h)+(β(1+β-ρ)h^2π+β hπ^2+β h+β(1-β+ρ)π)- π+1. Fixing ρ∈(0,1)ρ∈(0,1) and π>1π>1, we have ∂Δ¯1∂h(ρ,α,β,h,π)=π(1−ρ)(β(α(π2+1)−π2)h2+2βπ(2α−1)h+β(α(π2+1)−π2)+π2−1)(1+β−ρ)(βh2π+hπ2+h+(2−β)π)2, ∂ _1∂ h(ρ,α,β,h,π)= π(1-ρ)(β(α(π^2+1)-π^2)h^2+2βπ(2α-1)h+β(α(π^2+1)-π^2)+π^2-1)(1+β-ρ)(β h^2π+hπ^2+h+(2-β)π)^2, (19) and ∂Δ¯1∂β(ρ,α,β,h,π)=−(1−ρ)(hπ2+π−α(hπ2+2π+h))((1+2β−ρ)h2π+hπ2+h+(1−2β+ρ)π)(1+β−ρ)2(βh2π+hπ2+h+(2−β)π)2, ∂ _1∂β(ρ,α,β,h,π)=- (1-ρ)(hπ^2+π-α(hπ^2+2π+h))((1+2β-ρ)h^2π+hπ^2+h+(1-2β+ρ)π)(1+β-ρ)^2(β h^2π+hπ^2+h+(2-β)π)^2, (20) and ∂Δ¯1∂α(ρ,α,β,h,π)=β(1+β−ρ)h2π+βhπ2+βh+β(1−β+ρ)π(1+β−ρ)(βh2π+hπ2+h+(2−β)π). ∂ _1∂α(ρ,α,β,h,π)= β(1+β-ρ)h^2π+β hπ^2+β h+β(1-β+ρ)π(1+β-ρ)(β h^2π+hπ^2+h+(2-β)π). (21) As a consequence, we have that ∂Δ¯1∂β(ρ,α,β,h,π)<0 ∂ _1∂β(ρ,α,β,h,π)<0 for any α<hπ2+πhπ2+2π+hα< hπ^2+πhπ^2+2π+h and ∂Δ¯1∂α(ρ,α,β,h,π)>0 ∂ _1∂α(ρ,α,β,h,π)>0. Notice that if α>π2π2+1α> π^2π^2+1, then Δ⋆>0 >0 and Δ1 _1 is monotonically increasing in the homophily, h. Indeed, we have α>π2π2+1>hπ2+πhπ2+2π+hα> π^2π^2+1> hπ^2+πhπ^2+2π+h for all h>1h>1. This implies Δ¯1(ρ,α,β,h,π)>minα,hπ2+πhπ2+2π+h−π+1=hπ2+πhπ2+2π+h−π+1=Δ0(h,π)>0. _1(ρ,α,β,h,π)> \α, hπ^2+πhπ^2+2π+h \- π+1= hπ^2+πhπ^2+2π+h- π+1= _0(h,π)>0. Thus, we have that Δ1(ρ,α,β,h,π)=|Δ¯1(ρ,α,β,h,π)|=Δ¯1(ρ,α,β,h,π) _1(ρ,α,β,h,π)=| _1(ρ,α,β,h,π)|= _1(ρ,α,β,h,π) and Δ⋆=|Δ¯1(ρ,α,β,h,π)|−Δ0(h,π)>0 =| _1(ρ,α,β,h,π)|- _0(h,π)>0. In addition, we have α(π2+1)−π2>0α(π^2+1)-π^2>0. Using Eq. (19), we have ∂Δ1∂h(ρ,α,β,h,π)=∂Δ¯1∂h(ρ,α,β,h,π)>0. ∂ _1∂ h(ρ,α,β,h,π)= ∂ _1∂ h(ρ,α,β,h,π)>0. This implies that Δ1 _1 is monotonically increasing in the homophily, h. Proof of Proposition 3. We show that, if α<12α< 12 and β<β⋆β<β , then sign(Δ⋆) sign( ) is ambiguous and Δ1 _1 is non-monotone in h. In particular, there exist 1<h¯<h¯<∞1< h< h<∞ such that 1. Δ⋆>0 >0 and Δ1 _1 is decreasing over h∈(1,h¯)h∈(1, h); 2. Δ⋆<0 <0 and Δ1 _1 is non-monotone over h∈(h¯,h¯)h∈( h, h); 3. Δ⋆>0 >0 and Δ1 _1 is increasing over h∈(h¯,∞)h∈( h,∞). We introduce and prove two lemmas as follows, Lemma A.3. Fixing ρ,β∈(0,1)ρ,β∈(0,1), α∈(0,12)α∈(0, 12) and π>1π>1. For each h>1h>1, let β1⋆(h)∈(0,1) _1 (h)∈(0,1) denote the (unique) threshold such that −Δ¯1(ρ,α,β,h,π)−Δ0(h,π)≥0- _1(ρ,α,β,h,π)- _0(h,π)≥ 0 if and only if β∈[β1⋆(h),1]β∈[ _1 (h),1]. Then, we define β1⋆:=suph>1β1⋆(h)∈(0,1]. _1 := _h>1 _1 (h)∈(0,1]. For any β<β1⋆β< _1 , there exists 1<h¯1<h¯1<∞1< h_1< h_1<∞ such that −Δ¯1(ρ,α,β,h,π)−Δ0(h,π)≥0,if 1≤h≤h¯1 or h≥h¯1,<0,otherwise.- _1(ρ,α,β,h,π)- _0(h,π) \ array[]l≥ 0,& if 1≤ h≤ h_1 \ or h≥ h_1,\\ <0,& otherwise. array . Proof. We have −Δ¯1(ρ,α,β,h,π)−Δ0(h,π)=P(h)(1+β−ρ)(π+1)(hπ2+2π+h)(βh2π+hπ2+h+(2−β)π),- _1(ρ,α,β,h,π)- _0(h,π)= P(h)(1+β-ρ)(π+1)(hπ^2+2π+h)(β h^2π+hπ^2+h+(2-β)π), where P(h)=A3(ρ,α,β,π)h3+A2(ρ,α,β,π)h2+A1(ρ,α,β,π)h+A0(ρ,α,β,π)P(h)=A_3(ρ,α,β,π)h^3+A_2(ρ,α,β,π)h^2+A_1(ρ,α,β,π)h+A_0(ρ,α,β,π) and A3(ρ,α,β,π)=−βπ(1+β−ρ)(α(π3+π2+π+1)−(π3−π2+2π)),A2(ρ,α,β,π)=β(1+β−ρ)(3π3−π2)+β(π3+π)(π2−π+2)−2(1−ρ)(π3+π)(π−1)−αβ(π+1)(π4+(4−2ρ)π2+2βπ2+1),A1(ρ,α,β,π)=αβ2(π4+π3+π2+π)−β2(π4−π3+2π)+2(1−ρ)(π−1)3+β(π3+π)(π(ρ+4)−(ρ+2)−α(π+1)(ρ+3))+β(ρ+1)(π2+π),A0(ρ,α,β,π)=π2(β2((2α−3)π+(2α+1))+(1+ρ)β((3−2α)π−(2α+1))+4(1−ρ)(π−1)). array[]rclA_3(ρ,α,β,π)&=&-βπ(1+β-ρ)(α(π^3+π^2+π+1)-(π^3-π^2+2π)),\\ A_2(ρ,α,β,π)&=&β(1+β-ρ)(3π^3-π^2)+β(π^3+π)(π^2-π+2)-2(1-ρ)(π^3+π)(π-1)\\ &&-αβ(π+1)(π^4+(4-2ρ)π^2+2βπ^2+1),\\ A_1(ρ,α,β,π)&=&αβ^2(π^4+π^3+π^2+π)-β^2(π^4-π^3+2π)+2(1-ρ)(π-1)^3\\ &&+β(π^3+π)(π(ρ+4)-(ρ+2)-α(π+1)(ρ+3))+β(ρ+1)(π^2+π),\\ A_0(ρ,α,β,π)&=&π^2(β^2((2α-3)π+(2α+1))+(1+ρ)β((3-2α)π-(2α+1))+4(1-ρ)(π-1)). array Clearly, the sign of −Δ¯1(ρ,α,β,h,π)−Δ0(h,π)- _1(ρ,α,β,h,π)- _0(h,π) is the same as the sign of P(h)P(h). Because α<12α< 12, we have A3(ρ,α,β,π)>0A_3(ρ,α,β,π)>0. Indeed, we have α(π3+π2+π+1)−(π3−π2+2π)<0α(π^3+π^2+π+1)-(π^3-π^2+2π)<0. This implies that limh→+∞P(h)=+∞ _h→+∞P(h)=+∞ and hence P(h)>0P(h)>0 for all sufficiently large h. In addition, because α<π+1α< π+1, we have −Δ¯1(ρ,α,β,1,π)−Δ0(1,π)=π+1−(1−ρ)(π2+π)+α(β(1+β−ρ)π+βπ2+β+β(1−β+ρ)π)(1−ρ)(π2+2π+1)+(β(1+β−ρ)π+βπ2+β+β(1−β+ρ)π)>π+1−maxα,π+1=0,- _1(ρ,α,β,1,π)- _0(1,π)= π+1- (1-ρ)(π^2+π)+α(β(1+β-ρ)π+βπ^2+β+β(1-β+ρ)π)(1-ρ)(π^2+2π+1)+(β(1+β-ρ)π+βπ^2+β+β(1-β+ρ)π)> π+1- \α, π+1 \=0, which implies P(1)>0P(1)>0. By definition, the function P(⋅)P(·) is a cubic polynomial with a strictly positive leading coefficient. Suppose that P(h)<0P(h)<0 for some h>1h>1. Then, the continuity of P(⋅)P(·) guarantees that there exists 1<h¯1<h¯1<∞1< h_1< h_1<∞ such that P(h)≥0,if 1≤h≤h¯1 or h≥h¯1,<0,otherwise.P(h) \ array[]l≥ 0,& if 1≤ h≤ h_1 \ or h≥ h_1,\\ <0,& otherwise. array . This together with the fact that the sign of −Δ¯1(ρ,α,β,h,π)−Δ0(h,π)- _1(ρ,α,β,h,π)- _0(h,π) is the same as the sign of P(h)P(h) yields the desired result. In what follows, we show that P(h)<0P(h)<0 for some h>1h>1 whenever α<12α< 12 and β<β1⋆β< _1 . Indeed, we show why β1⋆ _1 exists and is unique. Because α<hπ2+πhπ2+2π+hα< hπ^2+πhπ^2+2π+h, we have ∂Δ¯1∂β(ρ,α,β,h,π)<0 ∂ _1∂β(ρ,α,β,h,π)<0, implying that −Δ¯1(ρ,α,β,h,π)−Δ0(h,π)- _1(ρ,α,β,h,π)- _0(h,π) as a function of β is increasing over [0,1][0,1]. Fixing any h>1h>1, we have −Δ¯1(ρ,α,0,h,π)−Δ0(h,π)=2π+1−2(hπ2+π)hπ2+2π+h<0.- _1(ρ,α,0,h,π)- _0(h,π)= 2π+1- 2(hπ^2+π)hπ^2+2π+h<0. In addition, we have −Δ¯1(ρ,α,1,h,π)−Δ0(h,π)=2π+1−hπ2+πhπ2+2π+h−(1−ρ)(hπ2+π)+α((2−ρ)h2π+hπ2+h+ρπ)(1−ρ)(hπ2+2π+h)+((2−ρ)h2π+hπ2+h+ρπ).- _1(ρ,α,1,h,π)- _0(h,π)= 2π+1- hπ^2+πhπ^2+2π+h- (1-ρ)(hπ^2+π)+α((2-ρ)h^2π+hπ^2+h+ρπ)(1-ρ)(hπ^2+2π+h)+((2-ρ)h^2π+hπ^2+h+ρπ). This implies that −Δ¯1(ρ,α,1,h,π)−Δ0(h,π)- _1(ρ,α,1,h,π)- _0(h,π) as a function of α is strictly decreasing over [0,1][0,1]. Because α<12α< 12, we have −Δ¯1(ρ,α,1,h,π)−Δ0(h,π)>−Δ¯1(ρ,12,1,h,π)−Δ0(h,π)=(π−1)R(ρ)2(2−ρ)(π+1)(hπ2+2π+h)(h2π+hπ2+h+π)- _1(ρ,α,1,h,π)- _0(h,π)>- _1(ρ, 12,1,h,π)- _0(h,π)= (π-1)R(ρ)2(2-ρ)(π+1)(hπ^2+2π+h)(h^2π+hπ^2+h+π) where R(ρ)=R0(h,π)+R1(h,π)ρR(ρ)=R_0(h,π)+R_1(h,π)ρ and R1(h,π)=−h3π3+2h3π2−h3π+4h2π3−4h2π2+4h2π−3hπ3+6hπ2−3hπ−4π2,R0(h,π)=2h3π3−4h3π2+2h3π+h2π4−6h2π3+10h2π2−6h2π+h2+8hπ3−8hπ2+8hπ+8π2. array[]rclR_1(h,π)&=&-h^3π^3+2h^3π^2-h^3π+4h^2π^3-4h^2π^2+4h^2π-3hπ^3+6hπ^2-3hπ-4π^2,\\ R_0(h,π)&=&2h^3π^3-4h^3π^2+2h^3π+h^2π^4-6h^2π^3+10h^2π^2-6h^2π+h^2+8hπ^3-8hπ^2+8hπ+8π^2. array Then, we have R(1)=h3π3−2h3π2+h3π+h2π4−2h2π3+6h2π2−2h2π+h2+5hπ3−2hπ2+5hπ+4π2=(h+π)(hπ+1)(h(π−1)2+4π)> 0. array[]rclR(1)&=&h^3π^3-2h^3π^2+h^3π+h^2π^4-2h^2π^3+6h^2π^2-2h^2π+h^2+5hπ^3-2hπ^2+5hπ+4π^2\\ &=&(h+π)(hπ+1)(h(π-1)^2+4π)\ >\ 0. array Proceeding to R(0)=R0(h,π)R(0)=R_0(h,π). For simplicity, we let x=π−1>0x=π-1>0 and y=h−1>0y=h-1>0. Then, we have R0(h,π)=x4y2+2x3y3+2x4y+4x3y2+2x2y3+x4+10x3y+4x2y2+8x3+18x2y+24x2+16xy+32x+8y+16>0.R_0(h,π)=x^4y^2+2x^3y^3+2x^4y+4x^3y^2+2x^2y^3+x^4+10x^3y+4x^2y^2+8x^3+18x^2y+24x^2+16xy+32x+8y+16>0. Because R(ρ)R(ρ) is linear in ρ and R(0),R(1)>0R(0),R(1)>0, we have that R(ρ)>0R(ρ)>0 for ∀ρ∈(0,1)∀ρ∈(0,1). This implies that −Δ¯1(ρ,α,1,h,π)−Δ0(h,π)>0- _1(ρ,α,1,h,π)- _0(h,π)>0. Putting these pieces together yields that there exists β1⋆(h)∈(0,1) _1 (h)∈(0,1) such that −Δ¯1(ρ,α,β,h,π)−Δ0(h,π)≥0- _1(ρ,α,β,h,π)- _0(h,π)≥ 0 if and only if β∈[β1⋆(h),1]β∈[ _1 (h),1]. For each fixed h>1h>1, the preceding argument implies that there exists a unique β1⋆(h)∈(0,1) _1 (h)∈(0,1) such that −Δ¯1(ρ,α,β,h,π)−Δ0(h,π)≥0- _1(ρ,α,β,h,π)- _0(h,π)≥ 0 if and only if β∈[β1⋆(h),1]β∈[ _1 (h),1]. We define β1⋆:=suph>1β1⋆(h)∈(0,1] _1 := _h>1 _1 (h)∈(0,1]. In what follows, we show that P(h)<0P(h)<0 for some h>1h>1 whenever α<12α< 12 and β<β1⋆β< _1 . Indeed, if β<β1⋆β< _1 , then by the definition of supremum there exists some hβ>1h_β>1 such that β<β1⋆(hβ)β< _1 (h_β). This implies −Δ¯1(ρ,α,β,hβ,π)−Δ0(hβ,π)<0- _1(ρ,α,β,h_β,π)- _0(h_β,π)<0. Thus, we have P(hβ)<0P(h_β)<0, as desired. Lemma A.4. Fixing ρ,β∈(0,1)ρ,β∈(0,1), α∈(0,12)α∈(0, 12) and π>1π>1. For each h>1h>1, let β2⋆(h)∈(0,1) _2 (h)∈(0,1) denote the (unique) threshold such that −Δ¯1(ρ,α,β,h,π)≤0- _1(ρ,α,β,h,π)≤ 0 if and only if β∈[β2⋆(h),1]β∈[ _2 (h),1]. Then, we define β2⋆:=suph>1β2⋆(h)∈(0,1]. _2 := _h>1 _2 (h)∈(0,1]. For any β<minβ2⋆,π−12π−2α(π+1)β< \ _2 , π-12π-2α(π+1)\, there exists h0>1h_0>1 and 1<h¯2<h¯2<∞1< h_2< h_2<∞ such that ∂Δ¯1∂h(ρ,α,β,h,π)≥0,if 1≤h≤h0,<0,otherwise. ∂ _1∂ h(ρ,α,β,h,π) \ array[]l≥ 0,& if 1≤ h≤ h_0,\\ <0,& otherwise. array . and Δ¯1(ρ,α,β,h,π)≤0,if 1≤h≤h¯2 or h≥h¯2,>0,otherwise. _1(ρ,α,β,h,π) \ array[]l≤ 0,& if 1≤ h≤ h_2 \ or h≥ h_2,\\ >0,& otherwise. array . Proof. We have ∂Δ¯1∂h(ρ,α,β,h,π)=π(1−ρ)Q(h)(1+β−ρ)(βh2π+hπ2+h+(2−β)π)2, ∂ _1∂ h(ρ,α,β,h,π)= π(1-ρ)Q(h)(1+β-ρ)(β h^2π+hπ^2+h+(2-β)π)^2, where Q(h)=β(α(π2+1)−π2)h2+2βπ(2α−1)h+β(α(π2+1)−π2)+π2−1Q(h)=β(α(π^2+1)-π^2)h^2+2βπ(2α-1)h+β(α(π^2+1)-π^2)+π^2-1. Because α<12α< 12, we have α(1+π2)−π2<0α(1+π^2)-π^2<0. This implies that limh→+∞Q(h)=−∞ _h→+∞Q(h)=-∞. Because β<π−12π−2α(π+1)β< π-12π-2α(π+1), we have Q(1)=(π+1)(2β(α(π+1)−π)+π−1)>0.Q(1)=(π+1) (2β (α(π+1)-π )+π-1 )>0. Note that the function Q(⋅)Q(·) is a quadratic polynomial with a strictly negative leading coefficient. Thus, we have that there exists h0>1h_0>1 such that Q(h)≥0,if 1≤h≤h0,<0,otherwise.Q(h) \ array[]l≥ 0,& if 1≤ h≤ h_0,\\ <0,& otherwise. array . This together with the fact that the sign of ∂Δ¯1∂h(ρ,α,β,h,π) ∂ _1∂ h(ρ,α,β,h,π) is the same as the sign of Q(h)Q(h) yields the desired result. Because α<12α< 12, we have Δ¯1(α,β,1,π)<0 _1(α,β,1,π)<0. As proved before, there exists h0>1h_0>1 such that ∂Δ¯1∂h(ρ,α,β,h,π)≥0,if 1≤h≤h0,<0,otherwise. ∂ _1∂ h(ρ,α,β,h,π) \ array[]l≥ 0,& if 1≤ h≤ h_0,\\ <0,& otherwise. array . It suffices to show that Δ¯1(ρ,α,β,h,π)>0 _1(ρ,α,β,h,π)>0 for some h>1h>1 whenever α<12α< 12 and β<β2⋆β< _2 . Indeed, we show why β2⋆ _2 exists and is unique. Because α<hπ2+πhπ2+2π+hα< hπ^2+πhπ^2+2π+h, we have ∂Δ¯1∂β(ρ,α,β,h,π)<0 ∂ _1∂β(ρ,α,β,h,π)<0, implying that Δ¯1(ρ,α,β,h,π) _1(ρ,α,β,h,π) as a function of β is decreasing over [0,1][0,1]. Fixing any h>1h>1, we have Δ¯1(ρ,α,0,h,π)=hπ2+πhπ2+2π+h−π+1>0. _1(ρ,α,0,h,π)= hπ^2+πhπ^2+2π+h- π+1>0. In addition, we have Δ¯1(ρ,α,1,h,π)=(1−ρ)(hπ2+π)+α((2−ρ)h2π+hπ2+h+ρπ)(1−ρ)(hπ2+2π+h)+((2−ρ)h2π+hπ2+h+ρπ)−π+1. _1(ρ,α,1,h,π)= (1-ρ)(hπ^2+π)+α((2-ρ)h^2π+hπ^2+h+ρπ)(1-ρ)(hπ^2+2π+h)+((2-ρ)h^2π+hπ^2+h+ρπ)- π+1. This implies that Δ¯1(ρ,α,1,h,π) _1(ρ,α,1,h,π) as a function of α is strictly increasing over [0,1][0,1]. Because α<12α< 12, we have Δ¯1(ρ,α,1,h,π)<Δ¯1(ρ,12,1,h,π)=(π−1)V(ρ)2(2−ρ)(π+1)(h+π)(hπ+1), _1(ρ,α,1,h,π)< _1(ρ, 12,1,h,π)= (π-1)V(ρ)2(2-ρ)(π+1)(h+π)(hπ+1), where V(ρ)=V0(h,π)+V1(h,π)ρV(ρ)=V_0(h,π)+V_1(h,π)ρ and V1(h,π)=π(h−1)2,V0(h,π)=−(hπ2+h+2π+2hπ(h−1)).V_1(h,π)=π(h-1)^2, V_0(h,π)=-(hπ^2+h+2π+2hπ(h-1)). Then, we have V(1)=−(hπ2+h+π+h2π)<0,V(0)=−(hπ2+h+2π+2hπ(h−1))<0.V(1)=-(hπ^2+h+π+h^2π)<0, V(0)=-(hπ^2+h+2π+2hπ(h-1))<0. Because V(ρ)V(ρ) is linear in ρ and V(0),V(1)<0V(0),V(1)<0, we have that V(ρ)<0V(ρ)<0 for all ρ∈(0,1)ρ∈(0,1). This implies that Δ¯1(ρ,α,1,h,π)<0 _1(ρ,α,1,h,π)<0. Putting these pieces together yields that there exists β2⋆(h)∈(0,1) _2 (h)∈(0,1) such that Δ¯1(ρ,α,β,h,π)≤0 _1(ρ,α,β,h,π)≤ 0 if and only if β∈[β2⋆(h),1]β∈[ _2 (h),1]. For each fixed h>1h>1, the preceding argument implies that there exists a unique β2⋆(h)∈(0,1) _2 (h)∈(0,1) such that Δ¯1(ρ,α,β,h,π)≤0 _1(ρ,α,β,h,π)≤ 0 if and only if β∈[β2⋆(h),1]β∈[ _2 (h),1]. We define β2⋆:=suph>1β2⋆(h)∈(0,1] _2 := _h>1 _2 (h)∈(0,1]. In what follows, we show that Δ¯1(ρ,α,β,h,π)>0 _1(ρ,α,β,h,π)>0 for some h>1h>1 whenever α<12α< 12 and β<β2⋆β< _2 . Indeed, if β<β2⋆β< _2 , then by the definition of supremum there exists some hβ>1h_β>1 such that β<β2⋆(hβ)β< _2 (h_β). This implies Δ¯1(ρ,α,β,hβ,π)>0 _1(ρ,α,β,h_β,π)>0, as desired. Back to the original claim of Proposition 3. We set β⋆=minβ1⋆,β2⋆,π−12π−2α(π+1)∈(0,1)β = \ _1 , _2 , π-12π-2α(π+1)\∈(0,1). By Lemma A.3, we have that there exists 1<h¯1<h¯1<∞1< h_1< h_1<∞ such that −Δ¯1(ρ,α,β,h,π)−Δ0(h,π)≥0,if 1≤h≤h¯1 or h≥h¯1,<0,otherwise.- _1(ρ,α,β,h,π)- _0(h,π) \ array[]l≥ 0,& if 1≤ h≤ h_1 \ or h≥ h_1,\\ <0,& otherwise. array . If 1≤h≤h¯11≤ h≤ h_1 or h≥h¯1h≥ h_1, we have that Δ1(ρ,α,β,h,π)=−Δ¯1(ρ,α,β,h,π)≥Δ0(h,π) _1(ρ,α,β,h,π)=- _1(ρ,α,β,h,π)≥ _0(h,π) because Δ0(h,π)≥0 _0(h,π)≥ 0. Otherwise, we consider: Δ¯1(ρ,α,β,h,π)≥0 _1(ρ,α,β,h,π)≥ 0 or Δ¯1(ρ,α,β,h,π)<0 _1(ρ,α,β,h,π)<0. For the former case, we have Δ1(ρ,α,β,h,π)−Δ0(h,π)=Δ¯1(ρ,α,β,h,π)−Δ0(h,π)<maxα,hπ2+πhπ2+2π+h−hπ2+πhπ2+2π+h=0. _1(ρ,α,β,h,π)- _0(h,π)= _1(ρ,α,β,h,π)- _0(h,π)< \α, hπ^2+πhπ^2+2π+h \- hπ^2+πhπ^2+2π+h=0. For the latter case, we have Δ1(ρ,α,β,h,π)−Δ0(h,π)=−Δ¯1(ρ,α,β,h,π)−Δ0(h,π)<0. _1(ρ,α,β,h,π)- _0(h,π)=- _1(ρ,α,β,h,π)- _0(h,π)<0. Putting these pieces together yields Δ⋆≡Δ1(ρ,α,β,h,π)−Δ0(h,π)>0,if 1<h<h¯1 or h>h¯1,<0,if h¯1<h<h¯1. ≡ _1(ρ,α,β,h,π)- _0(h,π) \ array[]l>0,& if 1<h< h_1 \ or h> h_1,\\ <0,& if h_1<h< h_1. array . (22) By Lemma A.4, we have that there exists h0>1h_0>1 and 1<h¯2<h¯2<∞1< h_2< h_2<∞ such that ∂Δ¯1∂h(ρ,α,β,h,π)>0,if 1<h<h0,<0,if h>h0. ∂ _1∂ h(ρ,α,β,h,π) \ array[]l>0,& if 1<h<h_0,\\ <0,& if h>h_0. array . and Δ¯1(ρ,α,β,h,π)<0,if 1<h<h¯2 or h>h¯2,>0,if h¯2<h<h¯2. _1(ρ,α,β,h,π) \ array[]l<0,& if 1<h< h_2 \ or h> h_2,\\ >0,& if h_2<h< h_2. array . Because Δ1(ρ,α,β,h,π)=|Δ¯1(ρ,α,β,h,π)| _1(ρ,α,β,h,π)=| _1(ρ,α,β,h,π)|, we have 1. Δ1 _1 is decreasing if 1<h<minh0,h¯21<h< \h_0, h_2\; 2. Δ1 _1 is non-monotone if minh0,h¯2<h<maxh0,h¯2 \h_0, h_2\<h< \h_0, h_2\; 3. Δ1 _1 is increasing if h>maxh0,h¯2h> \h_0, h_2\. In addition, we have Δ¯1(ρ,α,β,h,π)<−Δ0(h,π)<0 _1(ρ,α,β,h,π)<- _0(h,π)<0 if 1≤h≤h¯11≤ h≤ h_1 or h≥h¯1h≥ h_1. This implies that h¯1≤h¯2 h_1≤ h_2 and h¯1≥h¯2 h_1≥ h_2. Putting these pieces together with Eq. (22) yields the desired result with h¯=minh0,h¯1 h= \h_0, h_1\ and h¯=maxh0,h¯1 h= \h_0, h_1\. This completes the proof. A.5 Proofs from Section 6 Proof of Proposition 4. It suffices to show that |p1⋆(ρ,β11,β12,h,π)−1| |p_1 (ρ, _11, _12,h,π)-1| < < |hπ2+πhπ2+h+2π−1|=h+πhπ2+h+2π, | hπ^2+πhπ^2+h+2π-1|\ =\ h+πhπ^2+h+2π, |p2⋆(ρ,β21,β22,h,π)−1| |p_2 (ρ, _21, _22,h,π)-1| < < |h+πhπ2+h+2π−1|=hπ2+πhπ2+h+2π. | h+πhπ^2+h+2π-1|\ =\ hπ^2+πhπ^2+h+2π. Using the definition of p1⋆(ρ,β11,β12,h,π)p_1 (ρ, _11, _12,h,π), we have |p1⋆(ρ,β11,β12,h,π)−1|=(1−ρ)(1−β11+β12)h+(1−ρ)π(1−ρ+β11)β12h2π+(1−ρ+β11)hπ2+(β12+(1−ρ)(1−β11+β12))h+(2(1−ρ)+ρβ11+β12−β11β12)π.|p_1 (ρ, _11, _12,h,π)-1|= (1-ρ)(1- _11+ _12)h+(1-ρ)π(1-ρ+ _11) _12h^2π+(1-ρ+ _11)hπ^2+( _12+(1-ρ)(1- _11+ _12))h+(2(1-ρ)+ρ _11+ _12- _11 _12)π. Because (1−ρ+β11)β12h2π≥0,1−ρ+β11≥1−ρ,β12≥0,ρβ11+β12−β11β12≥0,(1-ρ+ _11) _12h^2π≥ 0, 1-ρ+ _11≥ 1-ρ, _12≥ 0, ρ _11+ _12- _11 _12≥ 0, we have |p1⋆(ρ,β11,β12,h,π)−1|≤(1−ρ)(1−β11+β12)h+(1−ρ)π(1−ρ)hπ2+(1−ρ)(1−β11+β12)h+2(1−ρ)π.|p_1 (ρ, _11, _12,h,π)-1|≤ (1-ρ)(1- _11+ _12)h+(1-ρ)π(1-ρ)hπ^2+(1-ρ)(1- _11+ _12)h+2(1-ρ)π. Because 1−ρ>01-ρ>0, we have |p1⋆(ρ,β11,β12,h,π)−1|≤(1−β11+β12)h+πhπ2+(1−β11+β12)h+2π.|p_1 (ρ, _11, _12,h,π)-1|≤ (1- _11+ _12)h+πhπ^2+(1- _11+ _12)h+2π. Then, we have (1−β11+β12)h+πhπ2+(1−β11+β12)h+2π−h+πhπ2+h+2π=−(β11−β12)hπ(hπ+1)(hπ2+(1−β11+β12)h+2π)(hπ2+h+2π)<β11>β120. (1- _11+ _12)h+πhπ^2+(1- _11+ _12)h+2π- h+πhπ^2+h+2π=- ( _11- _12)hπ(hπ+1)(hπ^2+(1- _11+ _12)h+2π)(hπ^2+h+2π) _11> _12<0. Putting these pieces together yields |p1⋆(ρ,β11,β12,h,π)−1|<h+πhπ2+h+2π.|p_1 (ρ, _11, _12,h,π)-1|< h+πhπ^2+h+2π. (23) Using the definition of p2⋆(ρ,β21,β22,h,π)p_2 (ρ, _21, _22,h,π), we have |p2⋆(ρ,β21,β22,h,π)−1|=(1−ρ)(1+β21−β22)hπ2+(1−ρ)π(1−ρ+β22)β21h2π+(β21+(1−ρ)(1+β21−β22))hπ2+(1−ρ+β22)h+(2(1−ρ)+β21+ρβ22−β21β22)π.|p_2 (ρ, _21, _22,h,π)-1|= (1-ρ)(1+ _21- _22)hπ^2+(1-ρ)π(1-ρ+ _22) _21h^2π+( _21+(1-ρ)(1+ _21- _22))hπ^2+(1-ρ+ _22)h+(2(1-ρ)+ _21+ρ _22- _21 _22)π. Because (1−ρ+β22)β21h2π≥0,β21≥0,1−ρ+β22≥1−ρ,β21+ρβ22−β21β22≥0,(1-ρ+ _22) _21h^2π≥ 0, _21≥ 0, 1-ρ+ _22≥ 1-ρ, _21+ρ _22- _21 _22≥ 0, we have |p2⋆(ρ,β21,β22,h,π)−1|≤(1−ρ)(1+β21−β22)hπ2+(1−ρ)π(1−ρ)(1+β21−β22)hπ2+(1−ρ)h+2(1−ρ)π.|p_2 (ρ, _21, _22,h,π)-1|≤ (1-ρ)(1+ _21- _22)hπ^2+(1-ρ)π(1-ρ)(1+ _21- _22)hπ^2+(1-ρ)h+2(1-ρ)π. Because 1−ρ>01-ρ>0, we have |p2⋆(ρ,β21,β22,h,π)−1|≤(1+β21−β22)hπ2+π(1+β21−β22)hπ2+h+2π.|p_2 (ρ, _21, _22,h,π)-1|≤ (1+ _21- _22)hπ^2+π(1+ _21- _22)hπ^2+h+2π. Then, we have (1+β21−β22)hπ2+π(1+β21−β22)hπ2+h+2π−hπ2+πhπ2+h+2π=(β21−β22)hπ2(h+π)((1+β21−β22)hπ2+h+2π)(hπ2+h+2π)<β21<β220. (1+ _21- _22)hπ^2+π(1+ _21- _22)hπ^2+h+2π- hπ^2+πhπ^2+h+2π= ( _21- _22)hπ^2(h+π)((1+ _21- _22)hπ^2+h+2π)(hπ^2+h+2π) _21< _22<0. Putting these pieces together yields |p2⋆(ρ,β21,β22,h,π)−1|<hπ2+πhπ2+h+2π.|p_2 (ρ, _21, _22,h,π)-1|< hπ^2+πhπ^2+h+2π. (24) This completes the proof. Proof of Theorem 3. From Proposition 4, we have |p1⋆(ρ,β11,β12,h,π)−1|<h+πhπ2+h+2π,|p2⋆(ρ,β21,β22,h,π)−1|<hπ2+πhπ2+h+2π.|p_1 (ρ, _11, _12,h,π)-1|< h+πhπ^2+h+2π, |p_2 (ρ, _21, _22,h,π)-1|< hπ^2+πhπ^2+h+2π. Thus, we have |p1⋆(ρ,β11,β12,h,π)−1|+|p2⋆(ρ,β21,β22,h,π)−1|<1|p_1 (ρ, _11, _12,h,π)-1|+|p_2 (ρ, _21, _22,h,π)-1|<1. Because p1⋆(ρ,α,β1,β2,h,π)∈(0,1)p_1 (ρ,α, _1, _2,h,π)∈(0,1), we have |p1⋆(ρ,α,β1,β2,h,π)−1|+|p1⋆(ρ,α,β1,β2,h,π)|=1.|p_1 (ρ,α, _1, _2,h,π)-1|+|p_1 (ρ,α, _1, _2,h,π)|=1. Suppose, toward a contradiction, that (Δ1)k≤(Δ2)k( _1)_k≤( _2)_k for all k∈1,2k∈\1,2\. Then, we have (Δ1)1+(Δ1)2≤(Δ2)1+(Δ2)2=|p1⋆(ρ,β11,β12,h,π)−1|+|p2⋆(ρ,β21,β22,h,π)−1|<1.( _1)_1+( _1)_2≤( _2)_1+( _2)_2=|p_1 (ρ, _11, _12,h,π)-1|+|p_2 (ρ, _21, _22,h,π)-1|<1. However, we have (Δ1)1+(Δ1)2=|p1⋆(ρ,α,β1,β2,h,π)−1|+|p1⋆(ρ,α,β1,β2,h,π)|=1.( _1)_1+( _1)_2=|p_1 (ρ,α, _1, _2,h,π)-1|+|p_1 (ρ,α, _1, _2,h,π)|=1. This yields the contradiction. Thus, there exists at least one topic k⋆∈1,2k ∈\1,2\ such that (Δ1)k⋆>(Δ2)k⋆( _1)_k >( _2)_k . This completes the proof. Appendix B Fixed Two-Island Environment In this appendix subsection, we study a different question from that in Theorem 2. We remain within the stylized two-island environment analyzed in the main text but treat its parameters (h,π,β1,β2)(h,π, _1, _2) as fixed and known. The training weights can therefore be calibrated to this particular environment. The objective is to characterize when a global aggregator improves learning pointwise in this fixed two-island environment, rather than whether a single training design is robustly beneficial across a range of admissible environments. Proposition B.1. Fix a two-island environment with parameters (h,π,β1,β2)(h,π, _1, _2). Then there exist α¯(ρ)<α¯(ρ)∈(0,1) α(ρ)< α(ρ)∈(0,1) such that: Δ⋆(ρ,α,β1,β2,h,π)≤0if α∈[max0,α¯(ρ),α¯(ρ)],>0if α∈[0,max0,α¯(ρ))∪(α¯(ρ),1]. (ρ,α, _1, _2,h,π) cases≤ 0& if α∈ [ \0, α(ρ)\, α(ρ) ],\\ >0& if α∈[0, \0, α(ρ)\)∪( α(ρ),1]. cases Proof of Proposition B.1. As in the proof of Theorem 2, we have Δ1(ρ,α,β1,β2,h,π)−Δ0(h,π)≤0, _1(ρ,α, _1, _2,h,π)- _0(h,π)≤ 0, if and only if α¯(ρ,β1,β2,h,π)≤α≤α¯(ρ,β1,β2,h,π), α(ρ, _1, _2,h,π)≤α≤ α(ρ, _1, _2,h,π), where α¯(ρ,β1,β2,h,π) α(ρ, _1, _2,h,π) = = (hπ2+πhπ2+h+2π)(β1(β2+1−ρ)h2π+(β1+(1−ρ)(1+β1−β2))hπ2+(β2+1−ρ)h+(β1+β2−β1β2+(1−ρ)(2−β2))π)−(1−ρ)((β1−β2+1)hπ2+π)(1−ρ)(β1−β2)(h2π+hπ2+h+π)(hπ2+πhπ2+h+2π)+β2(β1+1−ρ)h2π+(β1−(1−ρ)(β1−β2))hπ2+β2h+(β1+β2−β1β2−(1−ρ)β1)π, ( hπ^2+πhπ^2+h+2π )( _1( _2+1-ρ)h^2π+( _1+(1-ρ)(1+ _1- _2))hπ^2+( _2+1-ρ)h+( _1+ _2- _1 _2+(1-ρ)(2- _2))π)-(1-ρ)(( _1- _2+1)hπ^2+π)(1-ρ)( _1- _2)(h^2π+hπ^2+h+π) ( hπ^2+πhπ^2+h+2π )+ _2( _1+1-ρ)h^2π+( _1-(1-ρ)( _1- _2))hπ^2+ _2h+( _1+ _2- _1 _2-(1-ρ) _1)π, and α¯(ρ,β1,β2,h,π) α(ρ, _1, _2,h,π) = = (2π+1−hπ2+πhπ2+h+2π)(β1(β2+1−ρ)h2π+(β1+(1−ρ)(1+β1−β2))hπ2+(β2+1−ρ)h+(β1+β2−β1β2+(1−ρ)(2−β2))π)−(1−ρ)((β1−β2+1)hπ2+π)(1−ρ)(β1−β2)(h2π+hπ2+h+π)(2π+1−hπ2+πhπ2+h+2π)+β2(β1+1−ρ)h2π+(β1−(1−ρ)(β1−β2))hπ2+β2h+(β1+β2−β1β2−(1−ρ)β1)π. ( 2π+1- hπ^2+πhπ^2+h+2π )( _1( _2+1-ρ)h^2π+( _1+(1-ρ)(1+ _1- _2))hπ^2+( _2+1-ρ)h+( _1+ _2- _1 _2+(1-ρ)(2- _2))π)-(1-ρ)(( _1- _2+1)hπ^2+π)(1-ρ)( _1- _2)(h^2π+hπ^2+h+π) ( 2π+1- hπ^2+πhπ^2+h+2π )+ _2( _1+1-ρ)h^2π+( _1-(1-ρ)( _1- _2))hπ^2+ _2h+( _1+ _2- _1 _2-(1-ρ) _1)π. In what follows, we show that α¯(ρ,β1,β2,h,π)∈(0,1),for all ρ,β1,β2∈(0,1) and h,π>1. α(ρ, _1, _2,h,π)∈(0,1), for all ρ, _1, _2∈(0,1) and h,π>1. (25) Indeed, we let Nα¯N_ α and Dα¯D_ α denote the numerator and denominator of α¯(ρ,β1,β2,h,π) α(ρ, _1, _2,h,π), respectively. A direct rearrangement yields Nα¯=π(β1π((hπ+1)2+(1−ρ)hπ(h2−1))+β2((h+π)(hπ+1)+(1−ρ)π(h2−1))+β1β2π(h2−1)(hπ+1))hπ2+h+2π,Dα¯=β1β2π(h2−1)+(β1π(hπ+1)+β2(h+π))(h2π(1−ρ)+hπ2+h+π(1+ρ))hπ2+h+2π. array[]rclN_ α&=& π( _1π((hπ+1)^2+(1-ρ)hπ(h^2-1))+ _2((h+π)(hπ+1)+(1-ρ)π(h^2-1))+ _1 _2π(h^2-1)(hπ+1))hπ^2+h+2π,\\ D_ α&=& _1 _2π(h^2-1)+ ( _1π(hπ+1)+ _2(h+π))(h^2π(1-ρ)+hπ^2+h+π(1+ρ))hπ^2+h+2π. array Because ρ,β1,β2∈(0,1)ρ, _1, _2∈(0,1) and h,π>1h,π>1, we have 1−ρ>0,hπ2+h+2π>0,h2−1>0,hπ+1>0,h+π>0,1-ρ>0, hπ^2+h+2π>0, h^2-1>0, hπ+1>0, h+π>0, This implies that Nα¯>0N_ α>0 and Dα¯>0D_ α>0. Thus, we have α¯(ρ,β1,β2,h,π)>0 α(ρ, _1, _2,h,π)>0. We also have Dα¯−Nα¯=β1π((1−ρ)π(h2−1)+h2π+hπ2+h+π)+β1β2π(h2−1)(h+π)+β2(hπ(1−ρ)(h2−1)+(h+π)2)hπ2+h+2π.D_ α-N_ α= _1π((1-ρ)π(h^2-1)+h^2π+hπ^2+h+π)+ _1 _2π(h^2-1)(h+π)+ _2(hπ(1-ρ)(h^2-1)+(h+π)^2)hπ^2+h+2π. Because ρ,β1,β2∈(0,1)ρ, _1, _2∈(0,1) and h,π>1h,π>1, we have (1−ρ)π(h2−1)+h2π+hπ2+h+π>0,π(h2−1)(h+π)>0,hπ(1−ρ)(h2−1)+(h+π)2>0.(1-ρ)π(h^2-1)+h^2π+hπ^2+h+π>0, π(h^2-1)(h+π)>0, hπ(1-ρ)(h^2-1)+(h+π)^2>0. This implies that Dα¯−Nα¯>0D_ α-N_ α>0. Because Nα¯>0N_ α>0 and Dα¯>0D_ α>0, we have α¯(ρ,β1,β2,h,π)<1 α(ρ, _1, _2,h,π)<1. Putting these pieces together yields Eq. (25). Because α¯(ρ,β1,β2,h,π)<α¯(ρ,β1,β2,h,π)∈(0,1) α(ρ, _1, _2,h,π)< α(ρ, _1, _2,h,π)∈(0,1) for all ρ,β1,β2∈(0,1)ρ, _1, _2∈(0,1) and h,π>1h,π>1 (see Eq. (10)), the interval [max0,α¯(ρ,β1,β2,h,π),α¯(ρ,β1,β2,h,π)][ \0, α(ρ, _1, _2,h,π)\, α(ρ, _1, _2,h,π)] is nonempty. Proposition B.1 shows that improvement requires correction, not simply more weight on minority signals. In the two-island environment, the no-AI benchmark overweights majority information because beliefs circulate disproportionately within the larger group. Lowering α helps only if it offsets this distortion by the right amount: if α is too high, the aggregator reinforces majority dominance, while if α is too low, it over-corrects toward the minority island. The beneficial set is therefore an interior interval rather than a monotone region. This is a pointwise result for a fixed, known environment (h,π,β1,β2)(h,π, _1, _2); it does not imply that the same training weights improve learning robustly across nearby environments.