Paper deep dive
When Is Emergent Consensus Real? A Measured Coupling Gain and a Validity Diagnostic for LLM Agent Societies
Dongxu Yang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 96%
Last extracted: 7/9/2026, 6:39:16 AM
Summary
The paper introduces a measured 'coupling gain' (γ) to quantify LLM agent susceptibility to neighbours' opinions, demonstrating it is stable, model-distinguishing, and context-dependent. It establishes that classical dynamics (Friedkin-Johnsen, signed-Laplacian) organize consensus and polarization regimes using measured coefficients, shows LLMs do not spontaneously backfire, and provides a validity diagnostic to separate genuine averaging from model-prior artifacts, notably re-analyzing prior work to reveal conflation of mechanisms.
Entities (12)
Relation Signals (10)
Deepseek → hascouplinggain → 0.43
confidence 99% · DeepSeek 0.43[0.42,0.44]
Claude → hascouplinggain → 0.15
confidence 99% · Claude 0.15[0.15,0.15]
Coupling Gain (γ) → measures → Agent Susceptibility
confidence 98% · γ is the sensitivity of agent i’s updated opinion to its neighbours’ stated opinions
Authenticity Diagnostic → separates → Genuine Averaging
confidence 97% · separates genuine averaging from model-prior artifacts
Chuang et al. (2023) → conflates → Genuine Averaging and Model-Prior Artifact
confidence 96% · conflates two mechanisms—genuine averaging on debatable claims with a model-prior artifact on settled facts
Signed-Laplacian → organizes → Polarization Regime
confidence 95% · signed-Laplacian / structural balance criterion for polarization
Friedkin-Johnsen Dynamics → organizes → Consensus/Pluralism Regime
confidence 95% · Friedkin–Johnsen for the consensus/pluralism axis
Group Coupling → predicts → Society Convergence
confidence 95% · modality-matched group coupling does... predicts a society's convergence
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:LLM "agent societies" are studied via demonstrations of emergent consensus or polarization -- with no measurable control parameter, no theory of when each regime appears, and no test of whether an outcome is a genuine social dynamic or a model artifact. We introduce the coupling gain gamma, measured per-agent by counterfactually perturbing a neighbour's stated opinion. (i) gamma is stable and model-distinguishing -- across five frontier models it spans 0.15-0.43 (n=20, 95% CIs <= 0.025), paraphrase-invariant; social-neighbour gamma roughly equals numeric-anchor gamma, so gamma is evidence-coupling, not uniquely social. (ii) Classical dynamics with measured (not assumed) coefficients organise the regime: Friedkin-Johnsen for consensus/pluralism, signed-Laplacian/structural-balance for polarization. (iii) Frontier LLMs do not spontaneously backfire (beta <= 0), so default societies do not self-polarize -- polarization is always induced; the beta>0 branch arises only in the FJ surrogate, never in the agents. (iv) A randomized-initial-condition diagnostic -- the (slope, bias) of final vs. initial opinion -- separates genuine averaging from model-prior artifacts (boundary-censoring ruled out by construction via interior-valued facts); applied to a published "emergent consensus" result (Chuang et al. 2023) it reveals a model-specific conflation: averaging on debatable claims, prior-artifact on settled facts. (v) Coupling is context-dependent: pairwise gamma does not predict multi-neighbour outcomes -- it can order them backwards -- whereas a modality-matched group coupling does (sixteen closed+open models, Pearson r=-0.70, permutation p=0.008). The regime laws take this matched coupling, not the single-neighbour gamma: emergent consensus must be read from coupling in the target interaction. We contribute a measurement protocol and a validity instrument, not new theory.
Tags
Links
- Source: https://arxiv.org/abs/2606.22203v1
- Canonical: https://arxiv.org/abs/2606.22203v1
Trouble viewing inline? Open PDF directly →
Full Text
38,438 characters extracted from source content.
Expand or collapse full text
When Is Emergent Consensus Real? A Measured Coupling Gain and a Validity Diagnostic for LLM Agent Societies Dongxu Yang DeepLethe Correspondence: wayland0916@gmail.com (Preprint — ) Abstract LLM “agent societies” are largely studied through demonstrations that report emergent consensus, cooperation, or polarization—with no measurable control parameter, no theory of when each regime appears, and no test of whether an emergent outcome is a genuine social dynamic or an artifact of the underlying model. We introduce the coupling gain γ, a per-agent quantity measured directly from an LLM by counterfactually perturbing a neighbour’s stated opinion. We show: (i) γ is stable and model-distinguishing—across five frontier models it ranges 0.150.15–0.430.43 (n=20n=20 reps, bootstrap 95% CIs ≤0.025≤ 0.025 wide), and is invariant to prompt paraphrase; a control finds γ for a social neighbour ≈γ≈γ for an impersonal numeric anchor, so γ is an evidence-coupling coefficient rather than a uniquely social one. (i) The macro regime is organised by classical dynamics with measured (not assumed) coefficients: Friedkin–Johnsen for the consensus/pluralism axis, and the signed-Laplacian / structural balance criterion for polarization. (i) On the opinion-update tasks tested, frontier LLMs do not spontaneously backfire (β≤0β≤ 0): default societies do not spontaneously polarize—polarization is always induced (the active-polarization β>0β>0 branch appears only in the FJ surrogate, not in the LLM agents). (iv) A randomized-initial-condition validity diagnostic—the (slope, bias) of final vs. initial opinion—separates genuine averaging from model-prior artifacts; we rule out a boundary-censoring confound by construction using interior-valued numeric facts, and apply the diagnostic to a published “emergent consensus” result (Chuang et al. 2023), showing it conflates two mechanisms—genuine averaging on debatable claims with a model-prior artifact on settled facts—and that the effect is model-specific. (v) Probing transfer, coupling is context-dependent: the pairwise γ does not predict multi-neighbour society outcomes—it can order them backwards—whereas a modality-matched group coupling does (across sixteen closed-and-open models it predicts a society’s convergence with Pearson r=−0.70r=-0.70, permutation p=0.008p=0.008). The coefficient that enters the regime laws of (i) is therefore this matched coupling, not the single-neighbour γ of (i); emergent consensus must be read from coupling measured in the target interaction, not a nominal influenceability. Our contribution is a measurement protocol and a validity instrument, not new theory. 1 Introduction Building a society of LLM agents and observing “emergent” behaviour has become a popular methodology, but the dominant works are systems/demos: they report phenomena (a rumour spreads, agents form cliques, opinions converge) without (a) a measurable parameter that controls the outcome, (b) a theory predicting which regime appears, or (c) a test of whether the phenomenon is a genuine social dynamic or an artifact of model sycophancy/homogenization. We supply the missing rigour. Our unit of analysis is a single measured quantity, the coupling gain γ: how strongly an agent moves its stated opinion toward a neighbour’s. From a measured coupling and the influence network we obtain a falsifiable regime prediction—though, as we show, the coupling that governs a multi-neighbour society must itself be measured in the matched interaction—and from a perturbation test, a lie-detector for emergent consensus. Contributions (measurement-first, with a falsification test for aggregation). 1. A measured coupling gain γ via counterfactual perturbation—the first per-agent susceptibility coefficient measured directly on LLMs; stable, model-distinguishing, paraphrase-invariant, and (by control) an evidence-coupling rather than uniquely-social quantity. 2. A (slope, bias) authenticity diagnostic separating genuine averaging from model-prior artifacts, with closed form κ=η/(1−γ)κ=η/(1-γ), applied to re-analyse a published result (Chuang et al. 2023), showing it conflates two mechanisms. 3. A negative result: no frontier LLM spontaneously backfires (β≤0β≤ 0), so default societies cannot spontaneously polarize. 4. A transfer finding with a falsification test: the measurement transfers to natural-language and vector-rating tasks, but coupling’s value does not—pairwise and group susceptibility are negatively associated (Spearman ρ=−0.48ρ=-0.48 over 15 models, sign-robust to leave-one-model-out; suggestive at p=0.07p=0.07), and only a modality-matched group coupling predicts multi-neighbour outcomes. We prove (Prop. 4) that any additive aggregation forbids such a negative association, so the data are directionally inconsistent with the DeGroot/FJ/bounded-confidence family (consensus-conditional aggregation is one consistent explanation). All code and per-run logs—including the contaminated-vs-clean convergence logs and the audit script—are released.111https://github.com/deeplethe/llm-coupling-gain The regime predictions (Props. 1–3) are applications of Friedkin–Johnsen [9] and the signed-Laplacian / structural-balance criterion [10], with γ,βγ,β measured on the LLM rather than assumed. 2 Related Work LLM social simulation (Generative Agents [1], Concordia [2], AgentSociety [3]) are systems/platforms with no control/order parameter, no regime theorem, no spectral link; we add the quantitative layer. Emergent conventions (Ashery, Aiello & Baronchelli [4]) show a genuine disorder→ transition but with random pairwise mixing—no network, no spectrum; they defer network embedding. Validity critiques (PIMMUR [5]; Barrie & Törnberg [6]) argue much reported emergence is artifact via qualitative checklists / data-leakage arguments; we provide the quantitative instrument. Classical opinion dynamics (DeGroot [8]; Friedkin–Johnsen [9]; bounded confidence) establish that the spectrum of the influence matrix governs consensus vs. disagreement; our novelty is that γ is measured on the LLM, and the regime is a falsifiable prediction. (We verified that prior LLM opinion-dynamics papers use LLM-generated updates with no measured gain and no spectral analysis.) Single-agent self-conditioning [11] is the N=1N=1 analogue. Concurrent LLM opinion-dynamics work studies related effects without our measured-gain + validity combination: imposed self-vs-social weighting under varying topology [12] (a design parameter, not a counterfactually-measured γ); drivers of LLM social conformity [13]; LLMs achieving structural social balance given explicit signed interactions [14] (we instead ask whether opinion-update agents spontaneously form repulsive coupling, and find they do not); validity via persona temporal stability [15] and operational replication [16]; and foundational principles of LLM opinion dynamics [17]. None measures a per-agent coupling gain by counterfactual perturbation, nor supplies an init-invariance authenticity test. 3 The Coupling Gain Definition 1 (Society as a coupled map). N agents on a row-stochastic influence matrix W; agent i holds opinion xi∈[0,100]x_i∈[0,100]. One round: xit+1=Fi(xit,xjt:Wij>0)x_i^t+1=F_i(x_i^t,\x_j^t:W_ij>0\), FiF_i the LLM (black box; no update rule assumed). Definition 2 (Coupling gain). γi _i is the sensitivity of agent i’s updated opinion to its neighbours’ stated opinions: present neighbour opinion v across a grid, regress the updated opinion u(v)u(v) on v; γi _i is the slope. γ=0γ=0: ignores neighbours; γ=1γ=1: fully adopts. Linearising F about a fixed point gives the Friedkin–Johnsen (FJ) form xt+1=γWxt+(1−γ)x0x^t+1=γ Wx^t+(1-γ)x^0; γ is the measured FJ susceptibility. 4 Theory (applications of classical dynamics) Proposition 1 (Consensus vs. pluralism). For FJ dynamics with scalar γ∈[0,1)γ∈[0,1) and connected aperiodic row-stochastic W, the stationary profile is x∗=(1−γ)(I−γW)−1x0x^*=(1-γ)(I-γ W)^-1x^0, a convex average of x0x^0, reached at rate γ. As γ→0γ→ 0, x∗→x0x^*→ x^0 (full diversity); as γ→1γ→ 1, x∗→x^*→ consensus at the Perron-weighted average. Scope of Prop. 1 (anchoring matters). Prop. 1 assumes persistent FJ anchoring to the initial opinion x0x^0. If agents instead anchor to their evolving opinion (pure DeGroot), every connected graph reaches consensus for any γ>0γ>0; the “γ sustains pluralism” claim then holds only on sparse/disconnected networks or when the prompt induces self-anchoring (which our stubborn-agent runs supply, Sec. 5). We measure γ under the FJ protocol and report this dependence rather than assuming it away (App. D). Proposition 2 (Authenticity diagnostic). With an exogenous prior attractor p of pull η (doubly-stochastic W), the consensus mean is m∗=(1−κ)m0+κpm^*=(1-κ)m_0+κ p with κ=η/(1−γ)κ=η/(1-γ). Hence regressing final mean on initial mean gives slope=1−κslope=1-κ and bias=κ(p−m0)bias=κ(p-m_0): κ≈0κ\!≈\!0 (slope ≈1≈ 1, bias ≈0≈ 0) is genuine averaging (REAL); κ→1κ\!→\!1 (slope →0→ 0) is the model’s prior imposed regardless of the group (ARTIFACT). Proposition 3 (Polarization threshold). Write signed-coupling dynamics xt+1=(I−γLs)x^t+1=(I-γ L_s)x with signed Laplacian LsL_s. The society polarizes iff LsL_s has a negative eigenvalue (the destabilizing eigenvector is the community/Fiedler vector). In a two-block mean field the community-difference mode grows by g=1+2γβ(aout/d)g=1+2γβ(a_out/d): β<0β<0 consensus, β=0β=0 frozen (bounded confidence), β>0β>0 active polarization (this branch is a surrogate prediction). Empirically β≤0β≤ 0 for all measured models, so the β>0β>0 branch is never realized on agents—polarization is induced-only. Proofs (Banach/Neumann for Prop. 1; mean-field + signed-Laplacian for Prop. 3, cf. [10]) are in the appendix. Prop. 2’s κ-identity passes a parameter-free check (Sec. 5). Which coupling enters these laws? The γ above is a single scalar; Sec. 5 shows the single-neighbour γ is not interchangeable with the multi-neighbour coupling that actually drives a society, so the regime laws must be fed the coupling measured in the matched interaction. We state this carefully. A linearised multi-neighbour update with contraction rate ρ<1ρ<1 still has a unique stable consensus by Banach (reached at rate ρ); our data identify the ordering of ρ across models with the group pull pftp_ft—yielders contract, resisters do not—and the held-out forecast (Sec. 5) confirms exactly this ordering. We do not claim the convex-average / Perron-value refinements of Prop. 1 for the general multi-neighbour map: those require the row-stochastic averaging structure and need not survive a possibly anti-weighting update. The honest statement is ordinal: the coupling that controls whether a society contracts is the matched group coupling, tracked by pftp_ft, and the single-neighbour γ mis-estimates it (indeed orders societies backwards). The next proposition explains why. Proposition 4 (Additivity test: pairwise vs. group coupling). Call a one-round update additively separable if it depends on the neighbours only through S=∑jwjκ(xj)S= _jw_j\,κ(x_j) with weights wj≥0w_j≥ 0, κ nondecreasing, and G strictly increasing in S (i.e. x+=G(x,S)x^+=G(x,S); DeGroot, Friedkin–Johnsen, and bounded-confidence are all of this form). Then the susceptibility to a coherent group of m neighbours at value v and to a single neighbour at v (own stance fixed) satisfy pgroup=(∑j≤mwj/w1)γ1≥γ1p_group= ( _j≤ mw_j/w_1 )\, _1≥ _1 under sum-pooling, or pgroup=γ1p_group= _1 under degree-normalised mean-pooling. Hence pgroupp_group and γ1 _1 are sign-aligned with pgroup≥γ1p_group≥ _1: no additively separable model can have γ1>0 _1>0 yet pgroup≈0p_group≈ 0, nor a negative γ1 _1–pgroupp_group correlation across agents. The data lean against additivity for free-text deliberation (suggestively). With numeric neighbours the group coupling amplifies the pairwise one (γgrp≈3γ _grp≈3γ, Fig. 6A)—additive-consistent. With free-text neighbours, however, the group pull pftp_ft is negatively associated with γ across models (Spearman ρ=−0.48ρ=-0.48 over 1515 models, −0.70-0.70 on the 55 closed; robust in sign to leave-one-model-out—negative dropping any single model, including DeepSeek, whose γ=0.43γ=0.43 yet pft≈0.01p_ft≈0.01). Under additivity Prop. 4 forbids any negative association, so the data are directionally inconsistent with additive aggregation—though at this panel size the cross-model effect is suggestive (permutation p=0.07p=0.07), not a significant falsification, and a larger panel is needed. Read together with the no-backfire result (which excludes a repulsive kernel), this points to non-additive, consensus-conditional aggregation—why the matched group coupling pftp_ft, not the pairwise γ, tracks the society’s outcome and why γ orders societies backwards. 5 Experiments and Results Setup: OpenRouter API; agents = LLMs; opinions 0–100100; controlled tasks. Models: deepseek-v4-pro, gpt-5.5, claude-opus-4.8, gemini-3.5-flash, qwen3.7-max. 5.1 The coupling gain: stable and model-distinguishing Measured directly, with tight CIs (Fig. 1). γ with bootstrap 95% CI (n=20n=20): DeepSeek 0.43[0.42,0.44]0.43\,[0.42,0.44], Qwen 0.28[0.27,0.29]0.28\,[0.27,0.29], Gemini 0.25[0.25,0.26]0.25\,[0.25,0.26], GPT-5.5 0.18[0.17,0.19]0.18\,[0.17,0.19], Claude 0.15[0.15,0.15]0.15\,[0.15,0.15]; CIs ≪ inter-model gaps. γ is invariant to prompt paraphrase, and γ(social neighbour) ≈ γ(numeric anchor) (DeepSeek 0.470.47 vs 0.500.50; Claude 0.110.11 vs 0.090.09; Fig. 2): an evidence-coupling, not uniquely social, quantity. Figure 1: Coupling gain γ per model (n=20n=20 reps, bootstrap 95% CI). Figure 2: Sycophancy control: γ is paraphrase-invariant, and a social neighbour gives nearly the same γ as an impersonal numeric anchor—so γ is an evidence-coupling, not a uniquely social, quantity. 5.2 Regimes: pluralism, consensus, and induced polarization No spontaneous backfire (negative result). With a strongly-opinionated agent facing a hostile neighbour, all five models move toward the neighbour or are inert (β≤0β≤ 0) on opinion-update tasks: default societies do not spontaneously form repulsive coupling, so they cannot spontaneously polarize (cf. [14], which studies balance given explicit signed interactions—a different setting). This is not a restatement of sycophancy training: sycophancy concerns agreement with a single interlocutor and is silent about cross-community dynamics—a model that mirrors whichever neighbour it faces could still amplify a community split. That none does (β≤0β≤ 0) is an empirical fact about multi-agent structure, not a corollary of the training objective. Regime vs. γ×γ×density (P1). Real societies (final spread ± , K=3K=3): on the sparse ring, stubborn Claude (γ=0.15γ=0.15) preserves pluralism (6.9±1.96.9± 1.9) while DeepSeek (γ=0.43γ=0.43) consensuses (4.9±0.94.9± 0.9); on the complete graph both consensus (Claude 0.20.2, DeepSeek 3.63.6). Matches Prop. 1. Induced polarization. On a two-community SBM initialised low/high, DEFAULT agents converge while confirmation-bias agents freeze, across all five models: gap-reduction ratio DEFAULT 0.380.38–0.550.55 vs. CONFIRM 1.021.02–1.091.09 (disjoint bands). Formal test: paired across models, CONFIRM−DEFAULT=0.58±0.06CONFIRM-DEFAULT=0.58± 0.06, t(4)=23.5t(4)=23.5, p<10−4p<10^-4. Figure 3: Default agents converge across communities; confirmation-bias agents freeze (gap-ratio →1→ 1). Polarization is induced, not spontaneous. 5.3 A validity diagnostic and a re-analysis of Chuang et al. The diagnostic matrix (Table 1, Fig. 4). We run the diagnostic over 6 issues × 4 models at K=5K=5 with bias 95% CIs. All four models average faithfully on debatable claims (slope ≈1≈ 1, bias ≈0≈ 0); the consensus becomes prior-dominated only on settled-fact claims, and the pattern is model-dependent: Claude and GPT-5.5 produce artifacts on both vaccine and flat-earth (bias −27-27 to −46-46, CIs excluding 0), DeepSeek only on flat-earth (it averages vaccine), and Gemini averages everything (the diagnostic correctly reports no artifact). This corrects Chuang et al.’s [7] blanket “inherent accuracy bias → consensus”: it conflates genuine averaging (debatable claims) with a model-prior artifact (settled facts), and the effect is model-specific, not a universal property. Figure 4: Authenticity diagnostic, 4 models × 6 issues, K=5K=5 (bias 95% CI bars). Debatable claims cluster at REAL (slope≈ 1, bias≈ 0); settled-fact claims are prior-dominated ARTIFACTs for Claude/GPT, flat-earth only for DeepSeek, never for Gemini. Table 1: Diagnostic battery: prior-pull bias (mean ± 95% CI, K=5K=5) for 4 models. REAL == bias ≈0≈ 0 (genuine averaging); bold == ARTIFACT (CI excludes 0, prior overrides the group). Artifacts concentrate on settled-fact claims and are model-specific. issue DeepSeek Claude GPT-5.5 Gemini remote work +0.8±0.6+0.8±0.6 −0.9±0.3-0.9±0.3 +0.6±0.4+0.6±0.4 −0.3±0.3-0.3±0.3 social media −0.2±0.8-0.2±0.8 −1.0±0.6-1.0±0.6 +0.3±0.2+0.3±0.2 −0.0±0.2-0.0±0.2 nuclear power +0.7±0.7+0.7±0.7 −0.7±0.1-0.7±0.1 +0.4±0.3+0.4±0.3 −0.2±0.2-0.2±0.2 GMO safety +1.0±0.6+1.0±0.6 +1.5±0.4+1.5±0.4 +0.6±0.3+0.6±0.3 −0.1±0.2-0.1±0.2 vaccines harmful −6.5±1.8-6.5±1.8 −28.3±2.2-28.3 2.2 −26.6±4.8-26.6 4.8 −0.4±0.2-0.4±0.2 Earth is flat −18.4±5.5-18.4 5.5 −46.2±1.6-46.2 1.6 −46.1±3.3-46.1 3.3 −0.2±0.3-0.2±0.3 Censoring ruled out by construction (interior facts, Table 2). On opinion claims, the prior-attractor and [0,100][0,100] censoring are confounded (the artifact sits at the boundary, p≈0p≈ 0). We break the confound with numeric-fact claims whose answer is interior (≈71/60/21≈ 71/60/21): a society started at init=15=15 is pulled up to the interior value and one at init=90=90 down to it—convergence from both sides, impossible under floor-censoring. The per-cell slope of final-on-init (slope ≈0≈ 0: init-invariant attractor; ≈1≈ 1: averaging) is model-dependent: Qwen is a clean attractor on all three facts, GPT-5.5/Claude on two, DeepSeek and Gemini average. The diagnostic detects genuine prior-domination when present and correctly reports its absence otherwise. Table 2: Interior-fact init-slope (slope of final-mean on init-mean, mean ± std, K=3K=3; p = independently solo-elicited prior). Bold == slope≈ 0 with final→ interior p: a genuine attractor (censoring-immune, since the pull is toward an interior value from both sides). Model-dependent. model earth-water (p=71p=71) body-water (p=60p=60) air-oxygen (p=21p=21) Qwen 0.00±0.000.00 0.00 0.01±0.020.01 0.02 0.00±0.000.00 0.00 GPT-5.5 0.10±0.180.10 0.18 0.72±0.360.72±0.36 0.00±0.000.00 0.00 Claude 0.07±0.030.07 0.03 0.72±0.220.72±0.22 0.00±0.000.00 0.00 DeepSeek 1.01±0.031.01±0.03 0.95±0.040.95±0.04 0.34±0.310.34±0.31 Gemini 1.01±0.031.01±0.03 0.92±0.160.92±0.16 0.89±0.010.89±0.01 Figure 5: Interior-fact convergence (Earth-water, p=71p=71). A flat line (slope≈ 0, Qwen) is init-invariant convergence to the interior value—an upward pull from init=15=15 that floor-censoring cannot produce; the diagonal (DeepSeek) is averaging. 5.4 Does the coupling gain transfer? The measurement transfers; the value does not (Figs. 6–7). We test whether the pairwise γ predicts behaviour beyond the scalar single-neighbour task. (i) Modality. Measuring γ with a natural-language neighbour argument instead of a number shifts its value: GPT-5.5 rises 0.18→0.350.18→ 0.35 while DeepSeek falls 0.43→0.270.43→ 0.27 (Fig. 6A)—consistent with γ being evidence-coupling (a persuasive paragraph carries more evidence than a digit). (i) Aggregation. Facing a group of five neighbours rather than one amplifies coupling ∼3× 3× for the stubborn models (γgrp _grp: Claude 0.15→0.420.15→ 0.42, GPT 0.18→0.670.18→ 0.67; n=30n=30, R2≥0.92R^2≥ 0.92). On a four-axis vector-rating task the measurement also transfers but adds little: most models arithmetic-average each axis (γd≈0.5 _d≈ 0.5 at R2≈1R^2≈ 1; Claude the exception, 0.120.12–0.440.44). (i) The macro outcome is predicted by group, not pairwise, coupling—a held-out test. We measure pftp_ft (the susceptibility of an agent’s stance to a free-text group centred away from it) once, then run six-agent free-text town-halls (K=5K=5) on all five models as a held-out forecast. pftp_ft predicts the convergence outcome 5/5 on the complete graph (upgraded below to a powered n=16n=16 correlation across closed++open models, r=−0.70r=-0.70, p=0.008p=0.008): the high-pftp_ft yielders (Claude 0.220.22, GPT 0.210.21) converge (mean final spread 3.03.0, 4.64.6) while the low-pftp_ft resisters (DeepSeek, Gemini, Qwen; ≤0.01≤ 0.01) stay plural (mean spread 5959–7676; Figs. 6B, 7). The pairwise γ orders the societies backwards—DeepSeek has the highest pairwise γ yet holds, while the two lowest-γ models (Claude, GPT) converge—and pairwise γ and pftp_ft are negatively associated (Spearman ρ=−0.70ρ=-0.70 on these five; −0.48-0.48 across all 15 models with both measured, sign-robust to leave-one-model-out but only suggestive, p=0.07p=0.07; cf. Prop. 4). External validity (five conditions). We re-run the forecast across a connectivity sweep (ring deg-22, ring deg-44, complete), a larger society (N=10N=10), and an asymmetric initialisation. The predicted 2-vs-3 separation holds in every condition: the two high-pftp_ft models end below all three low-pftp_ft models (ring deg-22: 35,3635,36 vs 6464–7272; ring deg-44: 8.6,7.68.6,7.6 vs 5757–7575; complete/asym: 2.0,7.42.0,7.4 vs 4141–6868; N=10N=10: 4.4,114.4,11 vs 6666–7676). Absolute convergence is connectivity-dependent: yielders’ mean final spread falls with degree (ring ∼35 35, ring2 ∼8 8, complete ≤5≤5) while resisters stay ∼57 57–8080, so the sparse ring’s partial convergence is a mixing-rate effect, not a predictor failure. The ordering is robust to topology, density, size, and initialisation. Data hygiene: a failed API call is detected (retried, then the whole run discarded), never silently recorded as a held opinion; all reported numbers are over complete runs only (the same safeguard is applied in the surrogate-society code). Takeaway: the counterfactual measurement transfers to every setting, but coupling is context-dependent—pairwise and group susceptibility can even anti-correlate—so society-level behaviour must be read from coupling measured in the matched interaction, not from a nominal one-on-one influenceability. This sharpens the validity message rather than weakening it. Figure 6: Coupling is context-dependent and only group coupling predicts the society. (A) A group of five neighbours amplifies coupling ∼3× 3× over a single neighbour, and a natural-language neighbour shifts it again (GPT up, DeepSeek down). (B) The free-text group pull pftp_ft splits yielders (Claude/GPT) from resisters (DeepSeek/Gemini/Qwen); the pairwise γ (diamonds) orders them backwards—DeepSeek has the highest pairwise γ yet the lowest group pull. Figure 7: Free-text six-agent societies (K=5K=5, mean opinion spread per round). The two high-group-pull models (Claude, GPT) converge; the three low-group-pull models (DeepSeek, Gemini, Qwen) stay split—a held-out 5/55/5 match for pftp_ft. DeepSeek has the highest pairwise γ yet holds: the macro outcome tracks group, not pairwise, coupling. Reproducibility on open-weight models, and an independent prior check. Because closed frontier APIs drift over time, we replicate the protocol on eleven open-weight models (public weights, no drift: Llama-3.1-8B/70B, Llama-3.3-70B, Llama-4-Maverick, Qwen-2.5-7B/72B, Mistral-Large/Small, Mixtral-8x22B, Gemma-2-27B, DeepSeek-Chat). Pairwise γ is again stable and model-distinguishing (0.250.25–0.340.34; bootstrap 95% CIs ≤0.025≤ 0.025). The open set spans the coupling range—yielders (Qwen-2.5-7B pft=0.51p_ft=0.51, Mistral-Large 0.460.46) and resisters (Llama-3.1-70B −0.10-0.10)—so the full yielder/resister separation reproduces on open weights, and the pft→p_ft\!→\!convergence relation becomes a quantitative, powered law rather than a five-model binary: across all sixteen models (closed ++ open), pftp_ft predicts the society’s final spread with Pearson r=−0.70r=-0.70 (95%95\% bootstrap CI [−0.87,−0.48][-0.87,-0.48]; r=−0.75r=-0.75 dropping the lone outlier), Spearman ρ=−0.66ρ=-0.66 (permutation p=0.008p=0.008). The relation is noisier on the diverse open set—one outlier (Mixtral-8x22B, high pftp_ft yet holds) shows pftp_ft is a significant but imperfect predictor across architectures. Separately, to check that Prop. 2’s parameter-free test is not circular, we re-elicit the prior p through an independent neutral factual-rating channel (no discussion framing): it matches the diagnostic’s p (flat-earth ≈0≈0, vaccines ≈0≈0–22, and the ground-truth control “71%71\% water” ≈100≈100 for the true statement), so p is identified independently of the dynamics it is used to explain. 6 Limitations Small societies (N=6N=6–1010), few rounds; the free-text convergence (Fig. 7) is shown on five models at K=5K=5; the predictor pftp_ft (measured on all five at n=30n=30) forecasts the convergence split held-out 5/55/5 on the complete graph and preserves the yielder/resister separation across the external-validity conditions (ring deg-22/deg-44/complete, N=10N=10, asymmetric init), with absolute convergence connectivity-dependent (slower on the sparse ring); still-larger societies, learned networks, and non-opinion tasks remain future work; the propositions use a scalar-γ linearisation whose limits we now demonstrate empirically (Fig. 6); γ is evidence-coupling, not uniquely social; the active-polarization regime (β>0β>0) is never observed on real agents (only on the FJ surrogate); η is identified phenomenologically and the opinion prior p is solo-elicited (a mild circularity in Prop. 2’s parameter-free check that an independent p calibration would close—the interior-fact p is ground truth and so immune). We are explicit about statistical scale: the closed-model held-out forecast is 5/55/5 on five models, but the pft→p_ft\!→\!convergence relation it rests on is now a powered correlation over sixteen closed++open models (permutation p=0.008p=0.008); the induced-polarization paired test, however, still has only four degrees of freedom (t(4)=23.5t(4)=23.5), and societies remain small (N=6N=6–1010), so several claims rest on large, non-overlapping effect sizes rather than large n. All numbers, per-rep logs, and the contaminated-run audit are released. 7 Conclusion On 0–100100 opinion-update tasks across five frontier models, a single measured quantity—the coupling gain γ—plus a backfire coefficient β organises the macro outcome into consensus, pluralism, and (induced) polarization and predicts which appears; with the (slope, bias) diagnostic (censoring-immune via interior facts) it separates genuine social dynamics from model artifacts. We claim a reusable measurement protocol, not a universal law of agent societies: the active-polarization branch (β>0β>0) is realized only in the FJ surrogate (never measured on the agents), and pairwise γ does not transfer to multi-neighbour, natural-language interaction—there the macro outcome is governed by the modality-matched group coupling pftp_ft (Prop. 4), not the single-neighbour γ. The aim is to replace demonstration with measurement. Appendix A Proof of Proposition 1 (consensus vs. pluralism) The map T(x)=γWx+(1−γ)x0T(x)=γ Wx+(1-γ)x^0 is affine with linear part γWγ W. In the ∞-norm ‖γW‖∞=γmaxi∑jWij=γ<1\|γ W\|_∞=γ _i _jW_ij=γ<1 (row-stochastic W), so T is a contraction; by Banach’s theorem it has a unique fixed point with ‖xt−x∗‖∞≤γt‖x0−x∗‖∞\|x^t-x^*\|_∞≤γ^t\|x^0-x^*\|_∞ (rate γ). Solving x∗=γWx∗+(1−γ)x0x^*=γ Wx^*+(1-γ)x^0 gives x∗=(1−γ)(I−γW)−1x0x^*=(1-γ)(I-γ W)^-1x^0 (I−γWI-γ W invertible as ρ(γW)=γ<1ρ(γ W)=γ<1). Expanding (I−γW)−1=∑k≥0(γW)k(I-γ W)^-1= _k≥ 0(γ W)^k, all terms are entrywise ≥0≥ 0 and the row sums of (1−γ)∑kγkWk(1-γ) _kγ^kW^k equal (1−γ)∑kγk=1(1-γ) _kγ^k=1, so each xi∗x^*_i is a convex average of x0x^0. As γ→0γ→ 0, x∗→x0x^*→ x^0; as γ→1γ→ 1, (1−γ)(I−γW)−1→π⊤(1-γ)(I-γ W)^-1 1π (π the left Perron vector), i.e. consensus at π⊤x0π x^0. For reversible W the stationary spread is monotone decreasing in γ. □ Appendix B Proof of Proposition 2 (authenticity slope and bias) With an exogenous attractor p of pull η, xt+1=γWxt+(1−γ−η)x0+ηpx^t+1=γ Wx^t+(1-γ-η)x^0+η p1 (0≤γ+η<10≤γ+η<1, W doubly stochastic). The fixed point is x∗=(I−γW)−1[(1−γ−η)x0+ηp]x^*=(I-γ W)^-1[(1-γ-η)x^0+η p1]. Averaging and using that doubly-stochastic W preserves the mean, avg((I−γW)−1y)=avg(y)/(1−γ)avg((I-γ W)^-1y)=avg(y)/(1-γ). With m0=avg(x0)m_0=avg(x^0) and κ=η/(1−γ)κ=η/(1-γ), m∗=(1−γ−η)m0+ηp1−γ=(1−κ)m0+κpm^*= (1-γ-η)m_0+η p1-γ=(1-κ)m_0+κ p, so slope=dm∗/dm0=1−κslope=dm^*/dm_0=1-κ and bias=m∗−m0=κ(p−m0)bias=m^*-m_0=κ(p-m_0). □ Parameter-free check. Claude on “the Earth is flat”: slope −0.07⇒κ≈1-0.07 κ≈ 1, predicting bias ≈p−m0≈ p-m_0; with solo-elicited p≈2p≈ 2 and grid mean m0≈48m_0≈ 48, p−m0≈−46p-m_0≈-46, vs. measured −46.2±1.6-46.2± 1.6. Appendix C Proof sketch of Proposition 3 (polarization threshold) Write xt+1=(I−γLs)x^t+1=(I-γ L_s)x with signed random-walk Laplacian Ls=R−ML_s=R-M, Mij=Aijσij/diM_ij=A_ij _ij/d_i (σ=+1σ=+1 within, −β-β across community), R=diag(∑jMij)R=diag( _jM_ij). An eigenpair (μ,v)(μ,v) of LsL_s yields T=I−γLsT=I-γ L_s eigenvalue 1−γμ1-γμ, which grows iff μ<0μ<0: the society polarizes iff LsL_s has a negative eigenvalue, along its (community/Fiedler) eigenvector. In a balanced two-block mean field the difference δ=mA−mBδ=m_A-m_B obeys δt+1=(1+2γβaout/d)δ^t+1=(1+2γβ\,a_out/d)\,δ: β<0β<0 consensus, β=0β=0 frozen, β>0β>0 active polarization at rate ∝βλmax β _ of the cross-community block. Empirically β≤0β≤ 0 for all five models. □ Appendix D Proof of Proposition 4 (additivity test) Write x+=G(x,S)x^+=G(x,S) with S=∑jwjκ(xj)S= _jw_jκ(x_j), wj≥0w_j≥ 0, G strictly increasing in S (GS>0G_S>0 on the relevant range: more neighbour agreement does not repel) and κ nondecreasing (κ′≥0κ ≥ 0). [The one-directional falsification used below needs only GS≥0G_S≥ 0; the biconditional uses the strict GS>0G_S>0.] Fix own stance x=ox=o. A single neighbour at v gives S1=w1κ(v)S_1=w_1κ(v) and γ1=GS(o,S1)w1κ′(v) _1=G_S(o,S_1)\,w_1κ (v); a coherent group of m at v gives Sm=(∑j≤mwj)κ(v)S_m= ( _j≤ mw_j )κ(v) and pgroup=GS(o,Sm)(∑j≤mwj)κ′(v)p_group=G_S(o,S_m)\, ( _j≤ mw_j )κ (v). Each is a product of nonnegative factors sharing GSκ′G_Sκ , so γ1,pgroup≥0 _1,p_group≥ 0 are sign-aligned and pgroup=0⇔γ1=0p_group=0 _1=0; hence γ1>0⇒pgroup>0 _1>0 p_group>0. Under degree-normalised mean-pooling the two signals coincide (S1=Sm=κ(v)S_1=S_m=κ(v)), giving exactly pgroup=γ1p_group= _1; under sum-pooling pgroup/γ1=(∑j≤mwj/w1)GS(o,Sm)/GS(o,S1)≥1p_group/ _1= ( _j≤ mw_j/w_1 )\,G_S(o,S_m)/G_S(o,S_1)≥ 1 unless G is strictly concave in S. In every case the cross-agent map γ1↦pgroup _1 p_group is nonnegative; a negative correlation, or γ1>0 _1>0 with pgroup≈0p_group≈ 0, is impossible. Contrapositive: such an observation falsifies additive separability. □ Appendix E Provenance and the boundary-censoring control Props. 1–3 apply classical results: Friedkin–Johnsen [9] / DeGroot [8] (consensus axis) and the signed-Laplacian / structural-balance criterion [10] (polarization); the novel atoms are the counterfactual measurement of γ,βγ,β on LLMs and the κ-identity. FJ vs. DeGroot (measured for the movers). Regressing each agent’s new stance on its initial stance x0x^0, its previous stance xt−1x^t-1, and the neighbour mean: for the models that actually move (the yielders Claude/GPT) the weight on x0x^0 is near zero (|wx0|≤0.08|w_x^0|≤ 0.08) and that on xt−1x^t-1 dominates (R2≥0.94R^2≥ 0.94)—they anchor to their evolving opinion (DeGroot), not the initial one (FJ). The near-stationary resisters move too little to separate x0x^0 from xt−1x^t-1 (the two regressors are collinear there), so their anchoring is not identified; but only the movers’ anchoring bears on the pluralism caveat. Under DeGroot a connected graph consensuses for any γ>0γ>0, so we restrict the “γ sustains pluralism” claim to sparse networks or prompt-induced self-anchoring—consistent with the ring/dense split in Sec. 5. Censoring: on opinion claims the η-attractor and the [0,100][0,100] bound coincide at p≈0p≈ 0; interior-valued numeric facts break the confound, since an upward pull from init=15init=15 to an interior p cannot be floor-censoring. Appendix F Experimental details Agents are LLMs via OpenRouter (deepseek-v4-pro, gpt-5.5, claude-opus-4.8, gemini-3.5-flash, qwen3.7-max); opinions on 0–100100; networks ring / ER / complete / two-community SBM. Replication: pairwise γ at n=20n=20 (CIs are nonparametric bootstrap, B=20,000B=20,000 resamples; script released); P1 and induced at K=3K=3 (robustness N=16N=16, K=5K=5); diagnostic battery K=5K=5; flat-earth K=10K=10; interior facts K=3K=3. Transfer: free-text and group coupling at n=30n=30; free-text societies (N=6N=6, also N=10N=10; T=4T=4) at K=5K=5 on complete, ring (deg-22), and ring-2 (deg-44) graphs and an asymmetric initialisation. A failed API call is retried, then the run is discarded (never recorded as a held opinion); reported numbers are over complete runs only. All per-run logs (and the contaminated-run audit) are released. References [1] J. S. Park et al. Generative Agents: Interactive Simulacra of Human Behavior. UIST 2023. arXiv:2304.03442. [2] A. S. Vezhnevets et al. Generative agent-based modeling with actions grounded in physical, social, or digital space using Concordia. arXiv:2312.03664. [3] J. Piao et al. AgentSociety: Large-Scale Simulation of LLM-Driven Generative Agents Advances Understanding of Human Behaviors and Society. arXiv:2502.08691. [4] A. Ashery, L. M. Aiello, A. Baronchelli. Emergent social conventions and collective bias in LLM populations. Science Advances, 2025. arXiv:2410.08948. [5] J. Zhou et al. The PIMMUR Principles: Ensuring Validity in Collective Behavior of LLM Societies. arXiv:2509.18052. [6] C. Barrie, P. Törnberg. Emergent LLM behaviors are observationally equivalent to data leakage. arXiv:2505.23796. [7] Y.-S. Chuang et al. Simulating Opinion Dynamics with Networks of LLM-based Agents. arXiv:2311.09618. [8] M. H. DeGroot. Reaching a Consensus. JASA, 1974. [9] N. E. Friedkin, E. C. Johnsen. Social influence and opinions. J. Math. Sociology, 1990. [10] C. Altafini. Consensus problems on networks with antagonistic interactions. IEEE TAC, 2013. [11] A. Sinha et al. The Illusion of Diminishing Returns: Measuring Long Horizon Execution in LLMs. arXiv:2509.09677. [12] C. Han et al. Conformity Dynamics in LLM Multi-Agent Systems: The Roles of Topology and Self-Social Weighting. arXiv:2601.05606. [13] H. Zhong et al. Disentangling the Drivers of LLM Social Conformity: An Uncertainty-Moderated Dual-Process Mechanism. arXiv:2508.14918. [14] P. Cisneros-Velarde. Large Language Models can Achieve Social Balance. arXiv:2410.04054. [15] J. Gonnermann-Müller et al. Stable Personas: Dual-Assessment of Temporal Stability in LLM-Based Human Simulation. arXiv:2601.22812. [16] A. Tomašević et al. Towards Operational Validation of LLM-Agent Social Simulations: A Replicated Study of a Reddit-like Technology Forum. arXiv:2508.21740. [17] P. Cisneros-Velarde et al. On the Principles behind Opinion Dynamics in Multi-Agent Systems of Large Language Models. arXiv:2406.15492.