Paper deep dive
Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study
Lizhuo Zhang, Mengmeng Tang, Chenfeng Long, Xiaoyong Tang, Xiang Luo
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/20/2026, 5:08:31 AM
Summary
This paper analyzes the failure of self-consistency (majority voting) in Large Language Models (LLMs), specifically GPT-4.1, by decomposing the 'wrong-consensus agreement' index (Gamma). The study introduces a counterfactual decomposition to distinguish between mechanical agreement driven by per-case answer preferences and a residual driven by run-level heterogeneity or shared bias. Using GPQA-Diamond and AIME benchmarks, the authors find that on multiple-choice tasks, per-case preferences explain most agreement (81-93%), whereas on open-domain tasks, a significant unexplained residual remains (1.56-2.80 Gamma units), indicating that agreement is graded evidence rather than certification of correctness.
Entities (10)
Relation Signals (9)
Chenfeng Long → affiliatedwith → Hunan Agricultural University
confidence 95% · Chenfeng Long... Affiliation: College of Information and Intelligence, Hunan Agricultural University
Lizhuo Zhang → affiliatedwith → Hunan Agricultural University
confidence 95% · Lizhuo Zhang Affiliation: College of Information and Intelligence, Hunan Agricultural University
GPT-4.1 → evaluatedon → GPQA-Diamond
confidence 95% · On GPT-4.1 the decomposition shows benchmark-associated direction... On multiple-choice GPQA-Diamond
GPT-4.1 → evaluatedon → AIME
confidence 95% · On open-domain AIME, the mechanical preference explains only 59-78%
Self-Consistency → suffersfrom → Backfire
confidence 92% · A self-consistency backfire on hard questions is reproduced
Pluralistic Agreement Index (Gamma) → decomposedinto → Preference-Unexplained Residual
confidence 90% · Gamma... is decomposed into a mechanical component... and a preference-unexplained residual
Pluralistic Agreement Index (Gamma) → decomposedinto → Mechanical Component
confidence 90% · Gamma... is decomposed into a mechanical component... and a preference-unexplained residual
Preference-Unexplained Residual → dominatesin →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Majority voting over multiple LLM samples is widely used to raise answer accuracy, yet its gain varies erratically: on hard questions it can even backfire. This paper gives a quantitative account of this failure. A pluralistic agreement index Gamma is defined as the expected fraction of the samples of a wrong run that agree with the consensus, normalized by a reference scale d=(1-p)/(C-1), and is decomposed into a mechanical component (what a vote delivers given only a per-case answer preference) and a preference-unexplained residual. The mechanical null is difficulty-matched and leak-free: each case is resimulated at its own accuracy and option preference, estimated from the case's other runs, so no run predicts its own agreement. On GPT-4.1 the decomposition shows benchmark-associated direction (an observational ordering over n=4 cells per benchmark, not a significance claim). On multiple-choice GPQA-Diamond, the per-case answer preference explains 81-93% of the held-out test-run agreement index: the shared-bias-dominates account over-claims here, because a wrong but attractive option the whole cohort latches onto is captured by the per-case preference channel (whether that preference is induced by shared training bias is not identified). On open-domain AIME, the mechanical preference explains only 59-78% (21-29% if shrunk to pure noise), and a preference-unexplained residual of 1.56-2.80 Gamma units survives, which a run-level preference-heterogeneity reference more than absorbs (1.4-2.1). A self-consistency backfire on hard questions is reproduced (binned voting gap down to -0.09, coupled CI [-0.12,-0.07]), and the highest-agreement bin reaches an accuracy of only 0.42-0.83, a 1.2-3.6x lift over base rate: agreement is graded evidence, not certification. No new voting method is proposed; code and evidence are committed and reproducible.
Tags
Links
- Source: https://arxiv.org/abs/2608.18795v1
- Canonical: https://arxiv.org/abs/2608.18795v1
Trouble viewing inline? Open PDF directly →
Full Text
81,527 characters extracted from source content.
Expand or collapse full text
Decomposing Wrong-Consensus Agreement in LLM Self-Consistency: A GPT-4.1 Case Study Lizhuo Zhang Affiliation: College of Information and Intelligence, Hunan Agricultural University, Changsha 410128, China Affiliation: Yuelushan Laboratory, Changsha 410128, China Affiliation: College of Information and Intelligence, Hunan Agricultural University, Changsha 410128, China Affiliation: College of Information and Intelligence, Hunan Agricultural University, Changsha 410128, China Affiliation: Yuelushan Laboratory, Changsha 410128, Chinaelong@hunau.edu.cn Affiliation: Yuelushan Laboratory, Changsha 410128, China Mengmeng Tang Affiliation: College of Information and Intelligence, Hunan Agricultural University, Changsha 410128, China Affiliation: College of Information and Intelligence, Hunan Agricultural University, Changsha 410128, China Affiliation: College of Information and Intelligence, Hunan Agricultural University, Changsha 410128, China Chenfeng Long (corresponding author) Affiliation: College of Information and Intelligence, Hunan Agricultural University, Changsha 410128, China Affiliation: Yuelushan Laboratory, Changsha 410128, China Affiliation: College of Information and Intelligence, Hunan Agricultural University, Changsha 410128, China Affiliation: College of Information and Intelligence, Hunan Agricultural University, Changsha 410128, China Affiliation: Yuelushan Laboratory, Changsha 410128, Chinaelong@hunau.edu.cn Affiliation: Yuelushan Laboratory, Changsha 410128, China Xiaoyong Tang Affiliation: Yuelushan Laboratory, Changsha 410128, China Affiliation: Yuelushan Laboratory, Changsha 410128, Chinaelong@hunau.edu.cn Affiliation: School of Computer Science and Technology, Changsha University of Science and Technology, Changsha 410114, China Affiliation: Yuelushan Laboratory, Changsha 410128, China Xiang Luo Affiliation: Information Department, China Guangdong Tobacco Meizhou Ltd, Meizhou, 514000, China August 19, 2026 Abstract Majority voting over multiple LLM samples is widely used to raise answer accuracy, yet its gain varies erratically: on hard questions it can even backfire. This paper gives a quantitative, descriptive account of this failure. A pluralistic agreement index Γ is defined as the expected fraction of the samples of a wrong run that agree with the consensus, normalized by a reference scale d=(1−p)/(C−1)d=(1-p)/(C-1), and this index is decomposed into a mechanical component (what a vote would deliver given only a per-case answer preference) and a preference-unexplained residual (agreement left once that preference is held fixed). The mechanical null is difficulty-matched and leak-free: each case is resimulated at its own accuracy and option preference, estimated from the case’s other runs, so no run predicts its own agreement. On GPT-4.1 the decomposition shows benchmark-associated direction (an observational ordering over n=4n=4 cells per benchmark, not a significance claim). On multiple-choice GPQA-Diamond, the mechanical per-case answer preference explains ≈81≈ 81–93%93\% of the held-out test-run agreement index Γemp(t) _emp^(t): the “shared-bias dominates” account over-claims on this data, because a wrong but attractive option the whole cohort latches onto is captured by the per-case preference channel. (Whether that preference itself is induced by shared training bias is not identified by this decomposition.) On open-domain AIME, by contrast, the mechanical preference explains only 5959–78%78\% at λ=1λ=1 (2121–29%29\% if the preference estimate is shrunk to pure noise; the residual is largest there, so the direction is preference-robust), and a preference-unexplained residual of 1.561.56–2.802.80 Γ units survives, a residual that a run-level preference-heterogeneity reference more than absorbs (ϕdm=1.4 _dm=1.4–2.12.1), consistent with run-to-run preference variation as its channel. A self-consistency backfire on hard questions is reproduced (binned voting gap down to −0.09-0.09, coupled CI [−0.12,−0.07][-0.12,-0.07] on the most negative bin), and the highest-agreement bin is shown to reach an empirical high-agreement accuracy of only 0.420.42–0.830.83, a 1.21.2–3.6×3.6× lift over base rate: agreement is graded evidence, not certification. The decomposition is deliberately model-agnostic and defined entirely in the Method; no new voting method is proposed. Code and evidence are committed and reproducible. Keywords majority voting; self-consistency; plurality ceiling; shared bias; uncertainty estimation 1 Introduction Sampling-based reasoning (drawing multiple answers and taking a majority vote) is the workhorse of self-consistency [29]. Practitioners rely on the heuristic that “more samples and more agreement mean a more reliable answer.” A large body of evidence, however, shows the picture is more subtle: agreement can be high while the answer is still wrong [6], and on hard questions majority voting can even reduce accuracy relative to a single sample [2]. Existing work documents these phenomena (Section 2) but no LLM self-consistency study has applied a counterfactual decomposition to wrong-consensus agreement: how much of the observed agreement is a generic plurality effect versus model-specific correlated error? The answer is built as a hierarchy of counterfactual nulls for an agreement index. Contribution. This paper defines a pluralistic agreement index Γ (Section 3) and measures wrong-consensus agreement on real per-run LLM data against a hierarchy of progressively richer counterfactual nulls: uniform i.i.d.→per-case preference, i.i.d.→per-case preference with run-level heterogeneity,uniform i.i.d.\;→\;per-case preference, i.i.d.\;→\;per-case preference with run-level heterogeneity, each null granting the data strictly more mechanical structure than the last. The headline quantity is the mechanical coverage ϕ=Γrival/Γemp(t)φ= _rival/ _emp^(t), where Γrival _rival is a leak-free reference that grants each case its per-case option preference (preference and accuracy estimated from the case’s other runs only), and Γemp(t) _emp^(t) is the empirical index on the same held-out test runs; ϕφ is the fraction of observed wrong-consensus agreement that a fixed-preference i.i.d. counterfactual reproduces, and 1−ϕ1-φ the preference-unexplained residual. On GPT-4.1 the coverage shows benchmark-associated direction (an observational ordering at n=4n=4 cells per benchmark, not a significance claim; Section 5.1): on multiple-choice GPQA-Diamond, ϕ∈[0.806,0.927]φ∈[0.806,0.927] and the preference-unexplained residual is small (a popular-but-wrong option is captured by the per-case preference channel), whereas on open-domain AIME ϕ∈[0.586,0.781]φ∈[0.586,0.781] and a preference-unexplained residual of 1.561.56–2.802.80 Γ units survives, a residual that the run-level-heterogeneity null more than absorbs (Appendix D). This is a descriptive counterfactual analysis, not a direct identification of the error mechanism, and it does not identify whether the per-case preference itself is induced by shared training bias. The uniform null ρ=Γiid/Γemp∈[0.340,0.546]ρ= _iid/ _emp∈[0.340,0.546] is a strictly more conservative coverage and is reported alongside. Which level of the null hierarchy (uniform, fixed per-case preference, or run-heterogeneous preference) reproduces the data is answered cell by cell in Section 5.1. This is a descriptive, understanding contribution. It proposes no new voting method and claims no new backfire phenomenon (that is [2]); its claim is the quantitative counterfactual hierarchy for Γ . 2 Related work Self-consistency and its limits. Wang et al. [29] introduced sampling-based self-consistency and showed substantial gains over greedy decoding. A growing literature documents that these gains are not uniform. Ding [6] provides a large-scale audit of self-consistency and publicly releases per-run data; the method consumes their data but adds a quantitative decomposition they do not provide. Bahuguna [2] shows self-consistency can backfire on hard questions, with a binned analysis but without an agreement-index decomposition; its v2 revision states that the mechanism of the plurality-agreement gate’s failure remains an open problem. The counterfactual hierarchy is a descriptive decomposition of that wrong-consensus agreement, not a mechanism identification. The regime where agreement-based uncertainty collapses (a model overconfident on the same wrong answer across samples) is exactly the wrong-consensus regime quantified in [11]. A complementary line argues that majority voting has a ceiling that perturbation diversity does not raise, because error correlations are identical [7]; that ceiling is given a quantitative, wrong-run-conditional form. At the cross-model level, a capability-controlled audit finds shared co-failure, not diversity, is the stable correlate of majority-vote gain [13]; the decomposition here is the within-model analogue, built on a hierarchy of per-case nulls rather than an ensemble-level audit. Closest at the mechanism level, Chen et al. [3] show that majority-vote accuracy is non-monotonic in the number of calls and attribute it to a mixture of easy and hard queries within a task; the wrong-consensus decomposition is a per-cell, wrong-run-conditional complement to that aggregate-level account. The wrong-consensus regime decomposed here is the “self-consistent error” regime of Tan et al. [26], who formally define it, show its frequency does not decrease with model scale, and find all four detection-method families struggle on it; the contribution is not detection but a quantitative counterfactual hierarchy for the agreement itself. The anytime-valid statistical certification of self-consistency in [23] analyzes the same sampling counts used here but certifies mode uniqueness rather than decomposing agreement. The “shared bias” account of confident-yet-wrong behavior has academic antecedents: the probability-driven systematic error patterns of [21] are a shared-bias signature in the same spirit (McCoy et al. measure error patterns shaped by the shared pretraining task, not calibration curves). The preference-unexplained residual is an empirical pattern compatible with this account, quantified in a model-agnostic way and, critically, only on the open-domain cells where a substantial preference-unexplained component survives (Section 5.1). Agreement as confidence. Using consistency as a confidence signal is its own literature. Calibration of model confidences [10, 5, 12]; consistency-based confidence for generative tasks [20, 17, 8]; the “consistency hypothesis” formalized and tested across tasks [30]; confidence–consistency unification through minimum Bayes risk [28]; semantic-entropy successors [22]. Measurement-protocol discipline for agreement statistics [24] governs the reporting conventions followed here (Section 3.4). The AUROC analysis (Section 5.4) sits inside this literature: agreement is a graded but noisy confidence signal (AUROC 0.610.61–0.850.85), and the ceiling analysis (lift 1.21.2–3.6×3.6×, never near 1.01.0) quantifies exactly how graded. Two conventions from this literature bear on the design: (i) semantic clustering of near-duplicate answers is standard for open-ended tasks [17, 8], whereas the AIME C is a string-level mean distinct count; the string-level convention is kept because ϕφ is invariant to C (Section 3.1) and because the per-case preference null operates on observed labels, but a semantic-clustering sensitivity remains future work; (i) temperature/decoding settings are first-order determinants of sampling behavior and are not recorded in Ding’s per-run files, so the results are strictly about that collection pipeline. Ensemble decompositions. The “mechanical vs. correlated” distinction predates LLMs: bias-variance decompositions of zero-one loss [14], the bias-variance-covariance decomposition of [27], and ensemble ambiguity [16] all separate what vote marginals explain from what error covariance explains; on the LLM side, the sharpened answer marginal is the object that per-case preference estimation targets [1]. The quantities here map onto that vocabulary: Γemp _emp is a wrong-run-conditional agreement index, Γrival _rival its value under per-case marginals with independence, and δ the covariance-like remainder. What is new here is not the algebra but its LLM instantiation: the wrong-run conditioning, the leave-one-out per-case preference and accuracy, and the difficulty matching, plus the empirical finding that the marginals carry ≈81≈ 81–93%93\% on constrained multiple-choice and only 5959–78%78\% on open-domain tasks. A literature search (August 2026, arXiv API) found no prior work decomposing wrong-consensus agreement of a single model’s samples against counterfactual nulls, nor applying chance-corrected agreement coefficients to agreement among samples themselves; the closest works are the audits cited above. Novelty boundary. The paper is explicit about what it does not claim: it does not discover that self-consistency backfires (that is [2]), and it proposes no new voting method. The contribution is the Γ decomposition itself and its leak-free per-case-preference mechanical coverage ϕφ, a quantitative agreement-index quantity, and a benchmark-associated one (mechanical on multiple-choice, substantial preference-unexplained residual on open-domain), absent from prior work. Note on prior series. In vision, the companion diagnostic [19] applies a related consensus analysis (with a cross-model consensus control) to annotation difficulty; the present paper’s agreement decomposition is self-contained and does not depend on it. 3 Method 3.1 Setup and notation The data consist of a set of question–case instances. For each case, a model is sampled K times under a fixed prompt. This yields per-run counts: let C be the number of answer options considered, p the single-sample accuracy (the fraction of all N×KN× K single answers that are correct), and majmaj a consensus label obtained by plurality over the K answers. The consensus accuracy is the fraction of cases where the plurality label is correct. • C is chosen by benchmark: GPQA-Diamond has C=4C=4; for AIME and other open-ended tasks the mean number of distinct answer strings per run is used (Section 3.4). This scalarization affects only the absolute scale of Γ : the numerator [α∣wrong]E[α ] is computed over actual answer labels and does not depend on C, while both Γrival _rival and Γemp _emp are divided by the same baseline d=(1−p)/(C−1)d=(1-p)/(C-1), so the ratio ϕ=Γrival/Γempφ= _rival/ _emp cancels C−1C-1 exactly. This is verified empirically: rerunning the AIME cells with C fixed to 99 or to 2020 leaves ϕφ bit-identical across the two fixed-C runs, and at equal settings (nsim=2×103n_sim=2× 10^3) the fixed-C runs and the mean-distinct run agree to within one digit in the third decimal (results/kappa_rival_csens_C9/C20.json vs. kappa_rival_tie_argmin.json). • p is the single-sample accuracy, not the consensus accuracy. Both are reported. • Ties in the plurality are broken by lowest class id (argminargmin). Note that K=50K=50 is fixed by the data and is even, so ties are possible and resolved by this rule. Empirically the plurality is tied in 7.5%7.5\% of AIME runs and 0.8%0.8\% of GPQA runs (results/data_audit.json); on AIME the deterministic rule can therefore favor numerically small answers. This channel is quantified directly: recomputing every run’s majority with a uniform random tie-break among the tied labels (three independent seeds, applied to both the empirical majority and the simulated runs) shifts ϕφ by at most 0.0090.009 in any cell (results/kappa_rival_tie_rnd1/2/3.json vs. kappa_rival_tie_argmin.json), so the tie-break rule does not materially affect the decomposition. 3.2 The agreement index Γ The central object is an index of how much the samples of a wrong-consensus run coalesce around the (possibly wrong) plurality label. Define, for each case, α=1K∑i=1Kansweri=maj,α\;=\; 1K _i=1^K1\answer_i=maj\, the self-consistency of the K samples (this is exactly the quantity Ding [6] tabulates; it is denoted α to keep C reserved for the option count). A run is one realization of the K samples; a run is wrong when its plurality label differs from the ground truth (maj≠gtmaj ). All Γ quantities in this paper are conditioned on the run being wrong, and in every simulated counterfactual the run’s plurality label is recomputed from the simulated draws and the same conditioning is applied; the procedure is given in the pseudocode of Section 3.3. The reference scale is d=1−pC−1,d\;=\; 1-pC-1, the per-option error mass under independence given single-sample accuracy p. The quantity d is a scale normalization, not a chance-correction claim: the observed wrong-vote distribution is not uniform (the rival null itself shows this), so d is a reference value, not an assumed data model. The empirical agreement index is defined as Γemp=[α∣run wrong]d. _emp\;=\; E[α wrong]d. A wrong run whose samples cluster tightly on the consensus (an “attractive but wrong” state) yields α well above the reference scale d, hence large Γemp _emp. This quantifies the ceiling: it measures how “sticky” the samples of wrong runs are to a wrong consensus. Relation to chance-corrected agreement coefficients. Γ shares the chance-correction idea of Cohen’s κ, Scott’s π, and Fleiss’ κ [4, 25, 9, 15], but it is not Cohen’s κ: it conditions on wrong runs, normalizes by the per-option error mass d, and is decomposed against simulated counterfactuals rather than a chance estimator. To avoid persistent confusion with the classic coefficients, the index is denoted Γ throughout (an earlier draft used κ). Ensemble-diversity measures (e.g. [18]) are conceptually adjacent (they quantify error correlation among classifiers) but operate on classifier outputs, not on sampled runs of a single model. 3.3 Mechanical counterfactuals A mechanical counterfactual asks: how much of the observed agreement would survive if the model’s within-case error were not correlated, leaving only the mechanical averaging of independent votes? The answer depends on what is held fixed about each case. Two counterfactuals are built, both i.i.d. multinomial simulations (10510^5 draws per cell) that differ only in how the incorrect votes are distributed over options. Uniform wrong-vote null Γiid _iid. The most neutral reference places the samples of a wrong run uniformly over the wrong options (the correct option carries none of the incorrect votes). This is the classic independent-vote ceiling: agreement arises purely from the mechanical concentration of ≥⌈K/2⌉≥ K/2 votes on one answer, with no per-case attraction. Because pooling over difficulty destroys the wrong-plurality phenomenon (a pooled-p variant predicts ≤0.6%≤ 0.6\% wrong-consensus on three of four GPQA cells (8.8%8.8\% on the fourth) where the observed share is 4242–60%60\%), each case is always held at its observed difficulty pip_i; the appendix shows the pooled-p control failing by two orders of magnitude (Appendix A), which is why difficulty-matching is the informative baseline. Note that the normalization d=(1−p)/(C−1)d=(1-p)/(C-1) uses the cell-level p for the empirical index and the rival null, while the uniform per-question null normalizes by the mean accuracy of its own simulated population (case-weighted p¯i p_i, typically within a few percent of the cell-level p, up to about 5%5\% across cells), so the three Γ values sit on nearly one scale; difficulty matching affects only where the simulated wrong votes fall, not the normalization. Leak-free per-case preference null Γrival _rival. The uniform null ties the model’s incorrect votes to no option at all. But an LLM rarely errs uniformly: on a multiple-choice item a single plausible-but-wrong distractor can attract the whole cohort, and on an open-domain item certain answers are systematically preferred. Such a per-case answer preference is a mechanical property of the item under this decomposition (whether the preference itself originates in shared training bias is not identified). To separate it, each case’s option preference is estimated from its other runs, and incorrect votes are resimulated according to that hold-out preference. Formally, for each case i, as in a leave-one-out scheme, the per-case option counts are permuted/aggregated into a preference distribution q^i q_i, and the K votes of each simulated run are drawn from a multinomial equipped with q^i q_i and the case’s held-out accuracy pip_i. (Implementation detail: q^i q_i and the case accuracy pip_i are both estimated on the other runs of case i only (leave-one-out on both the preference and the accuracy) so no run predicts its own agreement; CLI flags --rival-mode case and --min-wrong 1. A case is eligible as a test when it has at least one wrong run and at least two runs in total; every wrong run of an eligible case is used once as a held-out test.) The resulting Γrival _rival is the agreement the same per-case preferences would generate if voters were otherwise independent, i.e., the maximal agreement attributable to per-case mechanical attraction. Because the eligible population is a subset of all wrong runs, Γemp(t) _emp^(t), the empirical index restricted to exactly the held-out test runs, is also computed, and the mechanical coverage on the same population is defined: ϕ=Γrival/Γemp(t)φ= _rival/ _emp^(t); the full-cell Γemp _emp is reported alongside and differs from Γemp(t) _emp^(t) by at most 4%4\% in any cell. A cell-level variant is also recorded: Γrival,pool _rival,pool uses a single average preference per cell instead of per-case preferences. It is far below Γrival _rival in every cell (e.g. Γrival,pool=1.01 _rival,pool=1.01 vs. Γrival=4.26 _rival=4.26 on gpt-4.1 AIME), so per-case structure, not any global preference, carries the mechanical agreement. Note that Γrival,pool _rival,pool is not directly comparable to Γiid _iid: its simulation support is the cell-wide set of distinct answers, so on open-domain cells the plurality is diluted far beyond the C-option uniform null, while on multiple-choice it inherits a global letter preference and sits slightly above the uniform null. It is reported here, inline, as a negative control. Pseudocode. The conditioning event is identical in the empirical index and in every mechanical reference: the run’s plurality label, recomputed from its own draws, differs from the ground truth. For the empirical index: (1) for each run, set maj=argmaxccount(a1,…,aK)maj=argmax_c\,count(a_1,…,a_K) (ties to the lowest class id); (2) keep runs with maj≠gtmaj , average their α, and set Γemp=α¯wrong/d _emp= α_wrong/d. For a reference: (3) draw M simulated runs; each draws K votes with the correct option chosen with probability pip_i, and wrong votes uniform over the C−1C-1 wrong options (Γiid _iid) or proportional to the held-out per-case preference q^i q_i (Γrival _rival); (4) recompute maj(m)maj^(m) from the simulated draws, keep only simulated runs with maj(m)≠gtmaj^(m) , average their α(m)α^(m), and set Γref=α(m)¯sim-wrong/d _ref= α^(m)_sim-wrong/d with the same cell-level d as in step 2. Decomposition and mechanical coverage. The headline quantity is the mechanical coverage ϕ=ΓrivalΓemp(t),φ\;=\; _rival _emp^(t), defined on the held-out test-run population of Section 3.3, a ratio of nonnegative quantities that, up to simulation noise, lies in [0,1][0,1]: it is the fraction of the observed agreement index Γemp(t) _emp^(t) that a leak-free, per-case-preference i.i.d. reference reproduces, and a value at or slightly above 11 (within bootstrap noise) means the mechanical reference already over-reproduces the observed agreement, leaving no positive residual. The complement δ= 1−ϕ=Γemp(t)−ΓrivalΓemp(t)δ\;=\;1-φ\;=\; _emp^(t)- _rival _emp^(t) is the agreement that survives even after every per-case option preference is granted to the mechanical reference: the preference-unexplained residual. Two readings follow: • ϕ≈1φ≈ 1 (δ≈0δ≈ 0): the agreement index is fully captured by a per-case answer preference (an attractive but wrong option the whole cohort latches onto). This does not identify whether the preference itself is induced by shared training bias. • ϕφ substantially below 11 (δ>0δ>0): a residual the per-case preference cannot explain remains; the samples of a wrong run cluster on the consensus even after their per-case option preference is removed. This is consistent with shared, correlated error (the pretraining-bias account [21]); the mechanism reading is interpretive, not identified. The assumption-free name of δ is the preference-unexplained residual; “shared-bias residual” is used as a mnemonic, and the paper does not claim to identify its mechanism; see Section 6. Empirically, Γrival≥Γiid _rival≥ _iid in all eight cells. Because the two references use different support conventions on open-domain tasks, this ordering is treated as an observed sanity check rather than a mathematical guarantee. (Within a fixed support, concentrating wrong-vote mass onto the observed per-case preference could not reduce agreement relative to the uniform reference; the observed ordering is therefore reported as a sanity check.) So ϕφ is a larger mechanical ceiling than the uniform null (ρ=Γiid/Γempρ= _iid/ _emp), which is reported alongside as a strictly more conservative lower bound. (Γrival,pool _rival,pool is not ordered against Γiid _iid, for the support reasons just stated.) Because LLM draws are not strictly i.i.d., the reference does not incorporate within-case sampling correlation. One direction is nevertheless identified for the exchangeable (positively correlated) draws of temperature sampling: within-case correlation concentrates a run’s votes, so it can raise Γemp(t) _emp^(t) above the correlation-free reference but never lower it. (For arbitrary, e.g. deliberately de-correlated, sampling this need not hold; the argument below is confined to the exchangeable case.) Hence the observed coverage ϕφ understates the true mechanical coverage, the observed residual δ is an upper bound on the correlation-free preference-unexplained share, and a high ϕφ is identified in direction: a null without correlation already reproduces ≥81%≥ 81\% of the GPQA index, so an account attributing the amplification primarily to shared bias over-claims regardless of how much sampling correlation exists. The quantity Γrival _rival is therefore treated as a benchmark reference rather than a certified bound in the other direction. The honest conclusion is that even granting every per-case preference to a benchmark that reproduces Γemp _emp on GPQA, a residual δ>0δ>0 remains on AIME: what a per-case-preference-only reference cannot explain. 3.4 Conventions All conventions are fixed here so Γ is comparable across cells (and across future work): 1. p = single-sample accuracy, never consensus accuracy. 2. C=4C=4 for GPQA-Diamond; for AIME, C=C= the mean number of distinct answer strings per run. 3. Ties broken by argminargmin class id. 4. Both i.i.d. counterfactuals are difficulty-matched (per-case pip_i, leave-one-out for the rival null); the headline null is the leak-free per-case preference Γrival _rival (preference estimated on other runs only, --rival-mode case); the uniform null and pooled-p control are reported as conservative/negative references. 5. Bootstrap CIs (B=104B=10^4) on every headline Γ and on ϕφ; differences (consensus −- single-sample) use coupled bootstrap to respect the within-case pairing. Both a run-level bootstrap and a case-clustered bootstrap are reported (cases resampled with replacement, all units of a sampled case moving together); ϕφ’s clustered CI resamples the shared test-case population jointly on both sides of the ratio, while the run-level ϕφ CI resamples numerator and denominator independently and is therefore conservative. The clustered CI is the primary one (reported in the decomposition table of Section 5.1), and both are conditional on the estimated preferences q^i q_i (whose own noise is probed by the shrinkage analysis of Appendix C). 6. The robustness appendices (shrinkage, answer-space, dispersion) are descriptive sensitivity analyses, not a family of hypothesis tests; no multiple-comparison correction is implied, and the p-values reported in Section 5.1 are the only inferential statements, used descriptively. The fixed notation is collected in Table 1. Table 1: Notation. Symbol Meaning Γemp _emp empirical agreement index (=[α∣wrong run]/d=E[α run]/d) α self-consistency of a run (share of votes for its plurality) Γiid _iid uniform wrong-vote i.i.d. null Γrival _rival leak-free per-case-preference i.i.d. null Γemp(t) _emp^(t) empirical index on the held-out test-run population ϕφ mechanical coverage =Γrival/Γemp(t)= _rival/ _emp^(t) δ=1−ϕδ=1-φ preference-unexplained residual ρ uniform coverage =Γiid/Γemp= _iid/ _emp ϕdm _dm coverage under the run-heterogeneity (Dirichlet-multinomial) null d reference scale (1−p)/(C−1)(1-p)/(C-1) p single-sample accuracy; pip_i per-case C option count (GPQA: 4; AIME: mean per-run distinct answers) K samples per run (50; Qwen arm 16/32) q^i q_i per-case wrong-answer preference (leave-one-out) λ shrinkage of q^i q_i toward uniform (Appendix C) α dispersion of the run-heterogeneity null (Appendix D) 4 Data The analysis uses the public per-run self-consistency data of Ding [6], which tabulates, for each case and model, the K=50K=50 per-run answers, per-run correctness, self-consistency α, single-sample accuracy p, and the majority label. Eight model–benchmark–prompt cells are analyzed: • gpt-4.1, gpt-4.1-mini, gpt-4.1-nano; • GPQA-Diamond and Ding’s 196-case AIME subset (problems spanning 1983–2025; the subset is inherited from the release, not chosen by the authors); • zero-shot and chain-of-thought prompting (where reported). This is a secondary re-analysis: no model is run; only per-run rows are re-analyzed, so every number below is reproducible from a committed parquet file. Terminology is fixed once. A case is a (question, model, prompt) instance. A run is one independent K=50K=50 sampling pass; a case carries one or more runs (1–16 in these cells; results/data_audit.json), and every analysis except the paired one uses all runs. Ding’s data additionally marks each run with one of two sampling conditions (a/b). Only cases carrying at least one run under each condition enter the paired champion-fragility analysis of Section 5.3: for gpt-4.1-mini these are 178 of the 198 GPQA cases and 175 of the 196 AIME cases (353 in total), the remaining 41 cases carrying a single condition. Counts in each table are therefore labeled in the unit the analysis actually uses: runs for the Γ and ceiling tables, pairs for the fragility table. Sampling settings (temperature, top-p) are inherited from Ding’s collection pipeline and are not recorded in the per-run files used here; they are neither controlled nor independently verified. Axis and condition labels. Ding’s files label every run with an axis (A/B/C) and a condition (a/b). The release’s majority labels are used as-is: for 30 of the 5,300 runs (results/data_audit.json) the release’s majority_answer is not the raw argmax of that run’s own answer counts (untied), and these labels are inherited rather than recomputed, so the empirical wrong-run membership follows the release’s own convention (the simulations use argmin tie-breaking, stated in the Conventions). Model and prompt identity come from the model column of the release’s own case-level table, joined to the per-run table through the shared (axis,condition)(axis,condition) keys; the resolution is therefore internal to the release, not an external reconstruction. For the cells used here the resolution is: A/a→A/a\!→\!nano-ZS, A/b→A/b\!→\!mini-ZS, B/a→B/a\!→\!mini-ZS, B/b→B/b\!→\!mini-CoT, C/a→C/a\!→\!mini-ZS, C/b→C/b\!→\!4.1-ZS. Two consequences follow. First, the paired champion-fragility analysis is possible only for mini-ZS: 4.1 appears only in C/bC/b and nano only in A/aA/a, so neither has two conditions to pair. Second, the mini-ZS pairs span three axes (A/bA/b, B/aB/a, C/aC/a), and the axes carry slightly different case subsets and consensus accuracy (e.g. AIME 0.320.32/0.360.36/0.240.24; GPQA 0.510.51/0.520.52/0.520.52; results/data_audit.json). The files do not record whether a/b or the axes differ in any sampling parameter; the GPQA axes are statistically close, while on AIME the C/aC/a arm is lower-accuracy, so the AIME flip rate should be read as instability under nominally identical (model, prompt, case) settings that may differ in batch or condition, not as a certified same-parameter comparison. It is reported as an estimate consistent with an upper bound, with this caveat. A mislabeled model would misattribute rows to cells but would not change the decomposition structure (the Γ construction is model-agnostic). Table 2 gives the per-cell sample sizes. The counts differ across cells because Ding’s collection design varies per arm (e.g. mini-ZS runs appear under three axes, 4.1 under one); the data are reported as collected. Table 2: Per-cell sample sizes. Cases = distinct (question, model, prompt) instances; runs = K=50K=50 sampling passes; wrong runs = runs whose plurality label is wrong (the effective sample for Γemp _emp); test runs = held-out wrong runs entering the rival null and Γemp(t) _emp^(t) (Section 3.3). Model Prompt cases runs wrong runs test runs 4.1 AIME-ZS 180 450 375 335 4.1 GPQA-ZS 182 450 234 205 4.1-mini AIME-CoT 172 425 237 212 4.1-mini AIME-ZS 196 1325 919 917 4.1-mini GPQA-CoT 180 425 178 152 4.1-mini GPQA-ZS 198 1325 641 641 4.1-nano AIME-ZS 178 450 347 313 4.1-nano GPQA-ZS 178 450 269 244 5 Results 5.1 Mechanical coverage across a hierarchy of nulls: multiple-choice vs. open-domain Table 3 reports, for all eight cells, the empirical agreement index Γemp _emp, the uniform null Γiid _iid, the leak-free per-case preference null Γrival _rival, and the mechanical coverage ϕ=Γrival/Γemp(t)φ= _rival/ _emp^(t) with case-clustered 95%95\% CIs (B=104B=10^4). Two structural facts hold everywhere, before any benchmark split. First, in every cell the empirical index is several times its reference scale (Γemp∈[3.56,7.09] _emp∈[3.56,7.09]); because Γ ’s absolute scale depends on C, cross-benchmark statements are made only on the C-invariant ϕφ, never on Γ magnitudes. This confirms that the samples of wrong runs are strongly attracted to the consensus. Second, the per-case preference null is far more explanatory than the uniform null (Γrival≫Γiid _rival _iid); much of the naive “shared bias” is in fact the mechanical attraction of incorrect voters to a popular per-case answer. Table 3: Agreement-index decomposition. Γrival _rival is the leak-free per-case-preference i.i.d. null (preference and accuracy from other runs); ϕ=Γrival/Γemp(t)φ= _rival/ _emp^(t) is the mechanical coverage on the held-out test-run population, δ=1−ϕδ=1-φ the preference-unexplained residual; ρ=Γiid/Γempρ= _iid/ _emp is the strictly more conservative uniform null. CI column: 95% case-clustered percentile bootstrap (B=104B=10^4), untruncated; a CI that reaches above 11 (or below 00) is reported as-is. Simulation seed 0, nsim=105n_sim=10^5; Monte Carlo SE ≤4×10−3≤ 4× 10^-3 Γ units in every cell. Per-cell sample sizes are in Table 2. Model Prompt p Γemp _emp Γrival _rival ϕφ ϕφ CI Γiid _iid 4.1 AIME-ZS 0.123 6.26 4.26 0.687 [0.640, 0.733] 2.13 4.1 GPQA-ZS 0.474 4.94 4.51 0.915 [0.879, 0.948] 2.14 4.1-mini AIME-CoT 0.375 5.69 4.15 0.726 [0.703, 0.813] 2.53 4.1-mini AIME-ZS 0.260 6.76 3.96 0.586 [0.554, 0.605] 2.49 4.1-mini GPQA-CoT 0.550 4.65 4.28 0.927 [0.874, 0.974] 2.54 4.1-mini GPQA-ZS 0.492 4.54 3.66 0.806 [0.771, 0.834] 2.17 4.1-nano AIME-ZS 0.163 7.09 5.55 0.781 [0.733, 0.811] 2.70 4.1-nano GPQA-ZS 0.385 3.56 3.21 0.902 [0.864, 0.938] 1.82 The central finding is a benchmark-associated direction in ϕφ (an observational ordering over n=4n=4 cells per benchmark), though the two benchmark ranges are not cleanly separated: • Multiple-choice GPQA-Diamond (C=4C=4). The mechanical per-case preference reproduces most of the agreement index: ϕ∈[0.806,0.927]φ∈[0.806,0.927], with preference-unexplained residuals δ of 0.070.07–0.190.19 (0.340.34–0.880.88 Γ units). Here a wrong but popular option is largely captured by the per-case preference channel, so the “shared-bias dominates” account over-claims on multiple-choice, with the caveat that the origin of the preference itself (training bias or not) is not identified by this decomposition. • Open-domain AIME (C≈9C≈ 9–2020). The mechanical preference reproduces only 5959–78%78\% (ϕ∈[0.586,0.781]φ∈[0.586,0.781]), leaving a preference-unexplained residual of 1.561.56–2.802.80 Γ units (δ∈[0.22,0.41]δ∈[0.22,0.41]). Whereas on multiple-choice the set of contenders is fixed, on open-domain a wrong run’s samples cluster on the consensus even after conditioning the counterfactual on the estimated per-case option preference. Appendix D calibrates an explanatory null for this residual: a reference that grants each case run-level preference heterogeneity (Dirichlet-multinomial, dispersion matched to the observed within-case cross-run spread of the plurality share) more than reproduces the empirical index (ϕdm=1.4 _dm=1.4–2.12.1 on AIME vs. 0.850.85–1.011.01 on GPQA). The dispersion parameter is fitted on a random held-out half of the eligible cases and evaluated on the other half, so the parameter never touches the data it explains: the residual is more than absorbed by (hence consistent with) run-to-run preference variation on open-domain tasks, and should be read as a signature consistent with an upper bound, not an identified mechanism. This residual is the estimate most exposed to design choices: its sensitivity to the preference estimate (shrinkage) and to the option-support convention is probed in Appendix C, which shows the residual is conservative under shrinkage. The GPQA/AIME split is real but modest: the four GPQA cells are the four largest ϕφ values, and an exact permutation enumeration over the (84)=70 84=70 benchmark labelings gives p=2/70≈0.029p=2/70≈ 0.029; it is reported as an ordering statement, not as evidence of significance (the cells are not exchangeable units: three nested model sizes, one pipeline). The mean difference is 0.190.19 (95% coupled bootstrap CI [0.11,0.28][0.11,0.28]), and the case-clustered CIs of the two boundary cells (4.1-mini GPQA-ZS [0.77,0.83][0.77,0.83] vs. 4.1-nano AIME-ZS [0.73,0.81][0.73,0.81]) overlap. The finding is therefore described as direction-consistent and partially overlapping, not as a clean separation. Figure 1 shows all eight cells with their case-clustered CIs. The two benchmarks also differ in difficulty (GPQA p∈[0.39,0.55]p∈[0.39,0.55] vs. AIME p∈[0.12,0.38]p∈[0.12,0.38]); within-benchmark, ϕφ is not monotone in p (on AIME the minimum ϕ=0.586φ=0.586 occurs at mid-difficulty p=0.26p=0.26, while the lowest- and highest-p cells both exceed 0.680.68; on GPQA the lowest-p cell has ϕ=0.90φ=0.90): the observed within-benchmark pattern is not monotonic in single-sample accuracy, so there is no simple evidence for a monotonic difficulty-only explanation. Difficulty and answer-space openness remain confounded. As a continuous complement to the binary benchmark comparison, across the eight cells ϕφ is positively associated with single-sample accuracy (Spearman rs=0.81r_s=0.81, p=0.015p=0.015) and negatively with answer-space size (rs=−0.76r_s=-0.76 vs. C, p=0.028p=0.028) and with the open/closed indicator (rs=−0.87r_s=-0.87, p=0.005p=0.005). The three candidate drivers are themselves strongly confounded with each other (n=8n=8), so the data cannot separate difficulty from answer-space openness: the honest reading is that ϕφ is lower on hard, open-domain, high-cardinality cells, an associational pattern, not an identified benchmark effect. (The Spearman p-values assume exchangeable cells and are descriptive; no multiplicity adjustment is implied.) Figure 1: Mechanical coverage ϕφ by cell, with 95% case-clustered bootstrap CIs (B=104B=10^4). Blue: GPQA-Diamond; orange: AIME (Wong colorblind-safe palette). The benchmark difference is direction-consistent (mean difference 0.190.19, permutation p≈0.03p≈ 0.03, n=4n=4 cells per benchmark) but the boundary cells overlap (mini GPQA-ZS [0.77,0.83][0.77,0.83] vs. nano AIME-ZS [0.73,0.81][0.73,0.81]). Taken together, “shared bias” is not a single channel: on constrained multiple-choice it is largely absorbed by the per-case preference, and on open-domain a residual beyond the preference survives. The uniform null is strictly more conservative (ρ=Γiid/Γemp∈[0.340,0.546]ρ= _iid/ _emp∈[0.340,0.546] in every cell), so the observed index is never less mechanical than the uniform ceiling implies; the leak-free null merely shows how much of the gap is attributable to per-case preference rather than correlated error. 5.2 Backfire on hard questions and a low empirical high-agreement ceiling The backfire phenomenon of [2] is reproduced in the difficulty-binned view: cases are split into five bins at the 20/40/60/80 percentiles of per-case accuracy (ties fall into the upper bin, so bins are not exactly equal-frequency when many cases share an accuracy), and the voting gap (consensus accuracy minus single-sample accuracy) is computed per bin. Majority voting helps on easy questions but is neutral or harmful on hard ones. The most negative voting gaps occur on the hardest bins; three representative cells: • gpt-4.1-nano GPQA-ZS, cases with per-case p≈0.13p≈ 0.13: gap =−0.09=-0.09, coupled case-level bootstrap CI [−0.12,−0.07][-0.12,-0.07]. • gpt-4.1-mini GPQA-CoT, hardest bin (p≈0.05p≈ 0.05): gap =−0.05=-0.05, coupled case-level bootstrap CI [−0.07,−0.03][-0.07,-0.03]. • gpt-4.1-mini GPQA-ZS (p≈0.24p≈ 0.24): gap =−0.01=-0.01, coupled CI [−0.04,0.02][-0.04,0.02] (crosses zero). These intervals come from a coupled case-level bootstrap (B=104B=10^4; consensus and single-sample accuracy of a sampled case move together), the appropriate paired unit for the gap. The bins are fixed once on the full data before resampling, so each bootstrap draw resamples cases within the same fixed bin; because a bin is narrow in per-case single-sample accuracy, the paired gap varies across cases mainly through the coupling of the two accuracies rather than through residual within-bin dispersion in pip_i. This is what makes the intervals tight (the hardest-bin CI spans only 0.050.05) despite the modest per-bin sample sizes. The two hardest-bin gaps exclude zero, while the third crosses zero: the difficulty-binned backfire has its CI excluding zero on the two hardest bins and is directionally consistent with [2], reported as a replication of the pattern, not as an independent precision claim. The global gap is positive because easy cases dominate the sample; the difficulty-binned (not pooled) gap is reported so the backfire is visible and honestly scoped. Figure 2 shows the full binned picture for all eight cells. Figure 2: Voting gap (consensus −- single-sample accuracy) against per-case difficulty (bins at the 20/40/60/80 percentiles of p), with coupled case-level bootstrap intervals on each bin. Panels (a)–(h): the eight model–benchmark–prompt cells of Table 2. Negative values on the hardest bins are the backfire signature. A second, complementary ceiling signature appears in the agreement-binned view. “Ceiling” is used in the empirical sense only (the observed accuracy of the highest-agreement bin), not as a theoretical upper bound on achievable accuracy. Runs are split into five equal-quantile (quintile) bins on the run’s self-consistency α (the same equal-quantile rule as the difficulty bins above); the “highest-agreement bin” is the top quintile. Table 4 reports the consensus accuracy within the highest-agreement bin of each cell, alongside the cell base rate p. Two facts hold. First, the ceiling is a substantial lift over the base rate: from 1.2×1.2× (4.1 GPQA-ZS, 0.590.59 vs. p=0.47p=0.47) to 3.6×3.6× (4.1-nano AIME-ZS, 0.590.59 vs. p=0.16p=0.16), so high agreement does carry predictive signal, as the AUROC analysis of Section 5.4 also shows. Second, the ceiling is nevertheless far below 1.01.0: it reaches only 0.420.42–0.830.83, and for the weaker models (gpt-4.1-nano, gpt-4.1 on AIME) it is as low as 0.420.42–0.590.59. High self-consistency is therefore informative but far from reliable, consistent with the decomposition of Section 5.1 (where open-domain cells retain a substantial preference-unexplained residual). Table 4: Consensus accuracy in the highest-agreement bin (top quintile of run self-consistency α), with the cell base rate p (single-sample accuracy) and 95% Wilson CIs. Because many runs sit at α=1.0α=1.0, the top quintile can contain more than 20% of runs and for three GPQA cells (4.1 GPQA-ZS, mini GPQA-CoT, mini GPQA-ZS) consists of α=1.0α=1.0 runs alone. “Ceiling” is how far high agreement reaches toward a correct consensus: it is a clear lift over p but stays far below 1.01.0. n is the number of runs falling in that bin (a run is one K=50K=50 sampling pass; Section 4). Model Prompt ceiling CI p n (runs) 4.1 AIME-ZS 0.42 [0.32, 0.52] 0.12 93 4.1 GPQA-ZS 0.59 [0.52, 0.65] 0.47 227 4.1-mini AIME-CoT 0.83 [0.74, 0.89] 0.38 95 4.1-mini AIME-ZS 0.74 [0.69, 0.79] 0.26 290 4.1-mini GPQA-CoT 0.83 [0.74, 0.89] 0.55 92 4.1-mini GPQA-ZS 0.74 [0.69, 0.78] 0.49 445 4.1-nano AIME-ZS 0.59 [0.48, 0.68] 0.16 92 4.1-nano GPQA-ZS 0.55 [0.45, 0.65] 0.39 93 5.3 Champion fragility: the winner itself is not stable The mechanical coverage ϕφ and the low ceiling both concern the strength of a single consensus answer. A complementary question is stability: given two K=50K=50 sampling runs of the same case and prompt under nominally identical settings, does the plurality winner agree? Ding’s per-run data marks each run with one of two sampling conditions (a/b); for gpt-4.1-mini, 353 of the 394 cases carry at least one run under each condition (178 GPQA-ZS and 175 AIME-ZS; Section 4), and each such case’s two conditions are paired, taking the first run of each condition when a case has several. One structural fact shapes the interpretation: every mini-ZS pair is cross-axis (condition b exists only under axis A and condition a only under axes B/C), so this section measures cross-batch winner stability, not same-parameter replication; the AIME flip rate in particular confounds axis (batch) differences with sampling randomness. As a sensitivity, pairing only A/bA/b against B/aB/a (excluding the lower-accuracy C/aC/a arm; results/champion_axis_sensitivity.json) changes the AIME flip rate from 0.820.82 to 0.800.80 (n=155n=155) and leaves the GPQA rate at 0.410.41 (n=160n=160): the qualitative finding is unchanged. The first-run choice is immaterial: recomputing with the last run or a random run changes the overall flip rate by at most 0.060.06 (results/champion_tierun_sensitivity.json). The champion flip rate is defined as the fraction of pairs on which the two runs select different plurality winners; this requires no new sampling and no assumptions beyond the pairing. Table 5 reports the flip rate overall and binned by the self-consistency α of the run from condition a. Two findings stand out. First, the flip rate is substantial: 0.400.40 on GPQA and 0.820.82 on AIME, where open-ended answers give many distinct candidates that a K=50K=50 sample cannot reliably separate. The two benchmark-level rates are therefore not directly comparable as a measure of fragility per se: the AIME rate conflates winner instability with the larger candidate set. Second, high agreement only weakly stabilizes the winner: on GPQA the flip rate falls from 0.640.64 in the lowest bin to 0.170.17 at α≥0.98α≥ 0.98, but it does not vanish; even nearly unanimous runs pick a different champion 17%17\% of the time. Champion fragility is thus a fourth, complementary signature that consensus-driven confidence is mechanically fragile even before shared error is considered. Table 5: Champion flip rate across the two sampling conditions of the same case and prompt (gpt-4.1-mini). “Flip rate” is the fraction of pairs where the plurality winner differs between the case’s two sampling conditions; 95% bootstrap CIs in parentheses; α bins use the self-consistency of the run from condition a, split into three equal-quantile bins (hence the reported boundaries). Benchmark Bin flip rate n (pairs) GPQA-ZS overall 0.40 (0.33, 0.47) 178 GPQA-ZS α∈[0.28,0.72)α∈[0.28,0.72) 0.64 (0.52, 0.76) 58 GPQA-ZS α∈[0.72,0.98)α∈[0.72,0.98) 0.44 (0.30, 0.58) 50 GPQA-ZS α≥0.98α≥ 0.98 0.17 (0.09, 0.26) 70 AIME-ZS overall 0.82 (0.77, 0.87) 175 AIME-ZS α∈[0.06,0.28)α∈[0.06,0.28) 0.91 (0.81, 0.98) 54 AIME-ZS α∈[0.28,0.84)α∈[0.28,0.84) 0.79 (0.69, 0.89) 61 AIME-ZS α≥0.84α≥ 0.84 0.78 (0.68, 0.88) 60 5.4 Jensen gap: soft agreement adds nothing over hard A standard response to winner fragility is to replace hard majority voting with soft, probability-weighted aggregation. In this setting only the aggregate answer counts of each run are available (no per-sample logits), so the faithful analogue is a comparison of two scoring rules on the same hard winner: hard confidence H=maxcpcH= _cp_c (the self-consistency α) versus soft confidence S=∑cpc2S= _cp_c^2, the probability that two independent draws agree. The Jensen gap J=S−H2≥0J=S-H^2≥ 0 isolates the agreement mass that lies beyond the plurality answer. Table 6 reports AUROC against plurality correctness and the mean gap. Table 6: Soft vs. hard agreement as confidence signals. H=maxcpcH= _cp_c is self-consistency; S=∑cpc2S= _cp_c^2; J=S−H2J=S-H^2. AUROC predicts whether the plurality answer is correct. Cells aggregate all runs of a model on a benchmark (across prompts); pooled row is global. J split by plurality correctness is shown in the last two columns. Relative to the first version of this manuscript, the cell membership changed because model identity is now resolved through the (axis,condition)(axis,condition) map of Section 4 (the previous axis-only map mislabeled some runs); the pooled row is unchanged. Model Benchmark AUROC(H) AUROC(S) mean J J, wrong J, correct 4.1 AIME 0.739 0.738 0.052 0.055 0.036 4.1 GPQA 0.607 0.608 0.027 0.034 0.020 4.1-mini AIME 0.847 0.837 0.041 0.054 0.016 4.1-mini GPQA 0.678 0.677 0.047 0.062 0.033 4.1-nano AIME 0.820 0.806 0.041 0.046 0.025 4.1-nano GPQA 0.621 0.630 0.055 0.062 0.045 Pooled 0.770 0.769 0.044 0.055 0.027 Two findings matter. First, no appreciable AUROC improvement is observed from the distributional score S over the plurality concentration H (0.7690.769 vs. 0.7700.770 pooled, within 0.020.02 on every cell; per-cell 0.610.61–0.850.85), so agreement carries real but bounded predictive signal, consistent with the ceiling analysis above. This is a deliberately honest null result: soft weighting does not beat hard majority, so it is reported as an understanding contribution, not a new voting method. Second, the Jensen gap is diagnostic: it is roughly twice as large on runs whose plurality answer is wrong (mean 0.0550.055) than on correct ones (0.0270.027). When the winner is wrong, the empirical distribution is more spread across candidates. This unifies the previous signatures: wrong runs show higher distributional dispersion than correct ones: the mass that a distributional score would redistribute is exactly the mass that carries the error, and the available scores do not improve on plurality concentration. 6 Limitations and conclusion The principal limitation is the i.i.d. assumption (Section 3.3): temperature draws are correlated. Under exchangeable (positively correlated) draws, the direction is one-sided: correlation concentrates a run’s votes, so δ is an upper bound on the correlation-free residual and a high ϕφ is identified in direction (Section 3.3); for arbitrary correlation structures the direction is not identified, so both counterfactuals are treated as benchmark references rather than certified bounds; the residual δ is a descriptive signature, not a point-identified mechanism. External validity is limited to the GPT-4.1 family in the main text; a second family (Qwen3.5-9B, Appendix B) shows a direction-consistent uniform share, but that arm is small (n=70n=70 on MMLU-Pro, no CIs), uses K=16/32K=16/32 rather than K=50K=50, lacks a rival null, and conflates difficulty with answer-space openness; a full cross-family rival decomposition is left to future work. The difficulty-binned backfire is a small-sample analysis for some cells; bootstrap CIs are reported and underpowered subsets are not over-interpreted. No fix is proposed; the honest boundary is that this paper is an understanding contribution. Nonetheless, the descriptive pattern is consistent and practically relevant across the robustness checks. The decomposition is benchmark-associated in direction: on multiple-choice GPQA-Diamond the mechanical per-case answer preference accounts for ≈81≈ 81–93%93\% of the agreement index; on open-domain AIME the mechanical preference accounts for only 5959–78%78\% and a preference-unexplained residual of 1.561.56–2.802.80 Γ units survives (the four GPQA cells are the four largest ϕφ values; exact permutation p=2/70p=2/70, an ordering statement at n=4n=4 per benchmark; Section 5.1). This reframes self-consistency confidence: on constrained multiple-choice, a popular-but-wrong answer is captured by the per-case preference channel (the whole cohort is attracted to it), though the origin of that preference is not identified; on open-domain, high agreement carries a substantial preference-unexplained component, which Appendix D shows is more than absorbed by a calibrated run-level preference-heterogeneity null, consistent with the pretraining-bias account of systematic, pretraining-shaped error patterns [21]. Either way, high agreement is graded evidence of correctness, not certification: the samples of wrong runs are themselves attracted to the consensus, so agreement saturates well below certainty (ceiling 0.420.42–0.830.83; Section 5.2). What this design does identify, in one place. (i) A high mechanical coverage is identified in direction even under sampling correlation: the correlation-free reference reproduces ≥81%≥ 81\% of the GPQA index, so shared-bias-dominant accounts over-claim there (Section 3.3). (i) The wrong-run marginal is benchmark-associated in direction, an observational ordering at n=4n=4 cells per benchmark (Section 5.1). (i) The within-case cross-run dispersion of the plurality share is 22–3×3× larger on AIME than GPQA, a dispersion signature that does not depend on the dispersion null’s fit (Appendix D). (iv) High agreement is graded evidence of correctness, never certification: the empirical high-agreement accuracy is 0.420.42–0.830.83 with 1.21.2–3.6×3.6× lift (Section 5.2), and the champion flip rate stays substantial even at high agreement (0.170.17 GPQA at α≥0.98α≥ 0.98; 0.780.78 AIME at α≥0.84α≥ 0.84) (Section 5.3). Falsifiability. The headline ϕφ (or, more conservatively, the uniform ρ) is refutable in principle: if the per-case preference reference were sufficient, then ϕφ would concentrate near 11 and the shared residual Γemp(t)−Γrival _emp^(t)- _rival, and hence δ=1−ϕδ=1-φ, would not survive. The observed ϕ∈[0.586,0.927]φ∈[0.586,0.927] is heterogeneous in direction: on GPQA it approaches 11 (the preference reference largely suffices, so a correlated-error reading of multiple-choice is unsupported), while on AIME ϕ∈[0.586,0.781]φ∈[0.586,0.781] leaves a residual that survives; this falsifies a purely per-case preference account of open-domain saturation. The paper is explicit about what this does not rule out: positive within-case sampling correlation is a third channel that the i.i.d. reference does not incorporate, and the design cannot separate it from δ. The claim is accordingly a channel decomposition of the wrong-run marginal (per-case preference vs. residual), not a decomposition of error sources. Thus the claim that GPT-4.1 self-consistency saturation is largely mechanical on multiple-choice but carries a robust residual on open-domain is a falsifiable empirical statement about the wrong-run marginal, not a tautology about correlated draws. Conclusion. Wrong-consensus agreement in LLM self-consistency is decomposed in this paper into a mechanical plurality effect and a preference-unexplained residual, measured by the mechanical coverage ϕφ. On constrained multiple-choice GPQA-Diamond, a leak-free per-case preference reference reproduces ≥81%≥ 81\% of the agreement index, a benchmark-mechanical saturation; on open-domain AIME a preference-unexplained residual survives (ϕφ as low as 0.5860.586), which a calibrated run-heterogeneity null more than absorbs (Appendix D), consistent with a pretraining-bias account of systematic error. The result reframes self-consistency confidence as graded evidence of correctness, never certification. These conclusions are bounded by the i.i.d. assumption; future work should extend the rival null across model families (the current Qwen arm is too small, lacks a rival null, and uses different K), separate within-case sampling correlation from δ, and turn the channel decomposition into explicit uncertainty estimates. Statements. Data and code availability: the per-run data are a public release of Ding [6]; all analysis scripts and evidence files (JSON) that produced every number in this paper are committed in full (scripts and data tracked in the project’s version control; an anonymized copy will be made available through the venue’s anonymous-repository upload, e.g. the review system’s anonymized-upload field, so reviewers can access it directly rather than on request), and can be regenerated from the raw parquet files by the commands in each script’s header. Funding: supported in part by the Yuelushan Laboratory Breeding Program (Grant YLS-2026-ZY01002), in part by the National Natural Science Foundation of China (Grant 62372064), and in part by the Meizhou Tobacco Science Research Project (Grant 202404). The funders had no role in study design, data collection and analysis, decision to publish, or preparation of the manuscript. Competing interests: none declared. Ethics: this work re-analyzes existing model outputs and involves no human subjects; no new model inference was performed. Acknowledgments This work was supported in part by the Yuelushan Laboratory Breeding Program under Grant YLS-2026-ZY01002, in part by the National Natural Science Foundation of China under Grant 62372064, and in part by the Meizhou Tobacco Science Research Project under Grant 202404. References [1] A. Arzhantsev, O. Sakhi, and N. Chopin (2026) Self-consistency via marginal sharpening. External Links: 2605.28142 Cited by: §2. [2] U. Bahuguna (2026) When self-consistency backfires: majority vote hurts the majority of hard science problems for small LLMs. Note: v2 of 2026-08-15 External Links: 2608.11403 Cited by: §1, §1, §2, §2, §5.2, §5.2. [3] L. Chen, J. Q. Davis, B. Hanin, P. Bailis, I. Stoica, M. Zaharia, and J. Zou (2024) Are more LLM calls all you need? Towards the Scaling Properties of Compound AI Systems. In Advances in Neural Information Processing Systems (NeurIPS), Vol. 37, p. 45767–45790. External Links: 2403.02419, Document Cited by: §2. [4] J. Cohen (1960) A coefficient of agreement for nominal scales. Educational and Psychological Measurement 20 (1), p. 37–46. External Links: Document Cited by: §3.2. [5] S. Desai and G. Durrett (2020) Calibration of pre-trained transformers. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 295–302. External Links: Document Cited by: §2. [6] K. Ding (2026) When LLMs agree, are they right? Auditing Self-Consistency and Cross-Model Agreement as Confidence Signals. External Links: 2607.08065 Cited by: §1, §2, §3.2, §4, §6. [7] S. Fadnavis, P. Kanakaraj, and F. Wyss (2026) Beyond consensus: trace-level synthesis in mixture of agents. External Links: 2605.29116 Cited by: §2. [8] S. Farquhar, J. Kossen, L. Kuhn, and Y. Gal (2024) Detecting hallucinations in large language models using semantic entropy. Nature 630, p. 625–630. External Links: Document Cited by: §2. [9] J. L. Fleiss (1971) Measuring nominal scale agreement among many raters. Psychological Bulletin 76 (5), p. 378–382. External Links: Document Cited by: §3.2. [10] C. Guo, G. Pleiss, Y. Sun, and K. Q. Weinberger (2017) On calibration of modern neural networks. In International Conference on Machine Learning (ICML), Cited by: §2. [11] K. Hamidieh, V. Thost, W. Gerych, M. Yurochkin, and M. Ghassemi (2026) Complementing self-consistency with cross-model disagreement for uncertainty quantification. External Links: 2604.17112 Cited by: §2. [12] S. Kadavath, T. Conerly, A. Askell, T. Henighan, et al. (2022) Language models (mostly) know what they know. arXiv preprint arXiv:2207.05221. Cited by: §2. [13] D. Kim (2026) Are diversity metrics measuring diversity? A Capability-Controlled Audit of Majority-Vote Gain in LLM Ensembles. External Links: 2607.20768 Cited by: §2. [14] R. Kohavi and D. H. Wolpert (1996) Bias plus variance decomposition for zero-one loss functions. In International Conference on Machine Learning (ICML), p. 275–283. Cited by: §2. [15] K. Krippendorff (1970) Bivariate agreement coefficients for reliability of data. Sociological Methodology 2, p. 139–150. External Links: Document Cited by: §3.2. [16] A. Krogh and J. Vedelsby (1995) Neural network ensembles, cross validation, and active learning. In Advances in Neural Information Processing Systems (NIPS), Cited by: §2. [17] L. Kuhn, Y. Gal, and S. Farquhar (2023) Semantic uncertainty: linguistic invariances for uncertainty estimation in natural language generation. In International Conference on Learning Representations (ICLR), Cited by: §2. [18] L. I. Kuncheva and C. J. Whitaker (2003) Measures of diversity in classifier ensembles and their relationship with the ensemble accuracy. Machine Learning 51 (2), p. 181–207. External Links: Document Cited by: §3.2. [19] Y. Ma and L. Zhang (2026) AnchorScore: a CLIP-based diagnostic of MLLM annotation difficulty. External Links: 2608.16690 Cited by: §2. [20] P. Manakul, A. Liusie, and M. J. F. Gales (2023) SelfCheckGPT: zero-resource black-box hallucination detection for generative large language models. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing (EMNLP), p. 9004–9017. External Links: Document Cited by: §2. [21] R. T. McCoy, S. Yao, D. Friedman, M. D. Hardy, and T. L. Griffiths (2024) Embers of autoregression show how large language models are shaped by the problem they are trained to solve. Proceedings of the National Academy of Sciences 121 (41), p. e2322420121. External Links: Document Cited by: §2, 2nd item, §6. [22] D. Nguyen, A. Payani, and B. Mirzasoleiman (2025) Beyond semantic entropy: boosting LLM uncertainty quantification with pairwise semantic similarity. External Links: 2506.00245 Cited by: §2. [23] H. Ota, N. Iwase, Y. Ichihara, J. Komiyama, and M. Imaizumi (2026) CITE: anytime-valid statistical inference in LLM self-consistency. External Links: 2605.05873 Cited by: §2. [24] D. Rao and C. Callison-Burch (2026) Agreement metrics for LLM-as-judge evaluation: what to report and why. External Links: 2606.00093 Cited by: §2. [25] W. A. Scott (1955) Reliability of content analysis: the case of nominal scale coding. Public Opinion Quarterly 19 (3), p. 321–325. External Links: Document Cited by: §3.2. [26] H. Tan, F. Sun, S. Liu, D. Su, Q. Cao, X. Chen, J. Wang, X. Cai, Y. Wang, H. Shen, and X. Cheng (2025) Too consistent to detect: a study of self-consistent errors in LLMs. In Empirical Methods in Natural Language Processing (EMNLP), p. 4755–4765. External Links: Document, 2505.17656 Cited by: §2. [27] N. Ueda and R. Nakano (1996) Generalization error of ensemble estimators. In IEEE International Conference on Neural Networks (ICNN), p. 90–95. External Links: Document Cited by: §2. [28] R. Vashurin, M. Goloburda, A. Ilina, A. Rubashevskii, P. Nakov, A. Shelmanov, and M. Panov (2025) Uncertainty quantification for LLMs through minimum Bayes risk: bridging confidence and consistency. External Links: 2502.04964 Cited by: §2. [29] X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou (2023) Self-consistency improves chain of thought reasoning in language models. In International Conference on Learning Representations (ICLR), Note: arXiv:2203.11171 Cited by: §1, §2. [30] Q. Xiao, D. Bhattacharjya, B. Ganesan, R. Marinescu, K. Mirylenka, N. H. Pham, M. Glass, and J. Lee (2025) The consistency hypothesis in uncertainty quantification for large language models. External Links: 2506.21849 Cited by: §2. Appendix A The pooled-p control fails This appendix deliberately reports a counterfactual that fails, because its failure is what justifies difficulty-matching in both i.i.d. nulls of Section 3.3. Table 7 gives, for each cell, the empirical share of runs whose plurality label is wrong (wcempwc_emp), the share predicted by the difficulty-matched reference (wcperqwc_perq), and the share predicted by a pooled-p reference (wcpooledwc_pooled). Table 7: Pooled-p is a failed control. wcwc is the share of runs whose plurality label is wrong: empirical (wcempwc_emp), predicted by the difficulty-matched i.i.d. reference (wcperqwc_perq), and predicted by a pooled-p reference (wcpooledwc_pooled). A difficulty-matched i.i.d. reference reproduces the empirical consensus–wrong share; a pooled-p reference collapses it (by 0.80.8 to 33 orders of magnitude on GPQA). Model Prompt wcempwc_emp wcperqwc_perq wcpooledwc_pooled 4.1 AIME-ZS 0.833 0.710 0.568 4.1 GPQA-ZS 0.520 0.462 0.006 4.1-mini AIME-CoT 0.558 0.398 0.001 4.1-mini AIME-ZS 0.694 0.363 0.020 4.1-mini GPQA-CoT 0.419 0.292 0.000 4.1-mini GPQA-ZS 0.484 0.344 0.003 4.1-nano AIME-ZS 0.771 0.640 0.130 4.1-nano GPQA-ZS 0.598 0.464 0.088 On three of four GPQA cells, pooling over difficulty pushes the predicted wrong-consensus share to ≤0.6%≤ 0.6\% (and to 8.8%8.8\% on 4.1-nano GPQA-ZS), whereas the observed share is 4242–60%60\% and the difficulty-matched reference recovers 2929–46%46\%. The pooled-p reference is thus a poor counterfactual for LLM data: it attributes essentially all correctness to a single global accuracy, treating every case as if it were of average difficulty, and thereby understates by 0.80.8 to 33 orders of magnitude across cells how often a wrong plurality forms. Note also that even the difficulty-matched reference falls short of the empirical wrong-consensus share by 0.060.06–0.330.33 in every cell (largest on AIME, e.g. 0.3630.363 vs. 0.6940.694 for 4.1-mini AIME-ZS): this shortfall is exactly the cross-run correlation beyond per-case difficulty that the Γ decomposition quantifies as δ>0δ>0 on open-domain cells, so Appendix A and Table 3 corroborate each other. Any agreement mechanism estimated against a pooled-p reference would be structurally misled; this is why a difficulty-matched reference is the informative baseline for the Γ decomposition on LLMs. Appendix B Second model family (Qwen3.5-9B) To check whether the mechanical coverage is an artifact of the GPT-4.1 family, earlier sampling of Qwen3.5-9B in this project is re-used (ctx4k) on two multiple-choice benchmarks with a different protocol: MMLU (C=4C=4, K=16K=16, n=500n=500 questions) and MMLU-Pro (C=10C=10, K=32K=32, n=70n=70). This arm predates the leak-free rival null, so it reports the uniform per-question null ρ=Γiid,perq/Γempρ= _iid,perq/ _emp; on GPT-4.1 Γrival≥Γiid _rival≥ _iid holds in all eight cells, but the analogous ordering is unverified for Qwen3.5-9B (no rival null was computed there: the Qwen collection stores one run per question, so a leave-one-out per-case preference is undefined on it), so ρ is read as a same-metric point of contact, not as a certified bound in the other direction. Two facts replicate. First, the mechanical coverage is of the same order on a different family: ρ=0.556ρ=0.556 (MMLU) and ρ=0.293ρ=0.293 (MMLU-Pro), bracketing the GPT-4.1 uniform range [0.340,0.546][0.340,0.546]. Second, the coverage declines in the same direction as difficulty and/or answer-space openness increase: within GPT-4.1 the rival share falls from ≥0.806≥ 0.806 (GPQA) to ≤0.781≤ 0.781 (AIME), a change in which answer-space openness and difficulty are confounded, and within Qwen3.5-9B the uniform coverage falls from 0.5560.556 (MMLU, p=0.74p=0.74) to 0.2930.293 (MMLU-Pro, p=0.56p=0.56, also larger C and K). The two arms are not merged into one table: K, benchmark, prompt, and sampling pipeline all differ, so only this coarse, direction-consistent comparison is warranted. The Qwen arm is committed native evidence (analysis/llm_selfconsistency.py, results/anchoring_llm_selfconsistency_report.json). Appendix C Preference-estimate shrinkage The per-case preference q^i q_i is estimated from the other runs of a case, so on open-domain cases it is a noisy estimate. To bound the influence of this estimation noise, the rival null is rerun with the estimated preference shrunk toward uniform over its support, qi(λ)=λq^i+(1−λ)⋅uniformq_i(λ)=λ q_i+(1-λ)·uniform, for λ∈1,0.5,0λ∈\1,0.5,0\, where the support at λ=0λ=0 is the set of answer labels that appear in the case’s other runs after subtracting the test run’s counts and dropping the ground truth and _UNPARSEABLE_. (Settings: nsim=2×104n_sim=2× 10^4, seed 0, --min-wrong 1; Monte Carlo SEs are reported per cell in the evidence files results/kappa_rival_shrink_lam*.json.) Why λ=0λ=0 is not directly comparable to Γiid _iid. The uniform null draws over C−1C-1 options with C the per-run mean distinct-answer count, whereas the rival null at λ=0λ=0 draws over the per-case wrong-label support. These option spaces differ materially on open-domain cells: across the AIME cells the mean rival support is 19.519.5–69.469.4 labels (median 1717–5757, max 250250) against C=9C=9–2020 (results/kappa_support_audit.json). A uniform distribution over a larger option set dilutes the plurality, so Γrival(λ=0)<Γiid _rival(λ=0)< _iid on AIME is the expected direction, not an inconsistency: the two references are uniform on different option vocabularies, and Jensen’s inequality argument applies only when the two supports coincide (as they roughly do on GPQA, where the mean rival support is 1.71.7–2.62.6 against C=4C=4). The comparison of ϕ(λ=0)φ(λ=0) against ρ in Table 8 is therefore directional evidence about the preference estimate, not a like-for-like benchmark comparison; the only like-for-like comparison is across λ values within a column, which uses the identical support at every λ. Table 8: Mechanical coverage ϕφ under preference-estimate shrinkage. λ (shrinkage) 11 0.50.5 00 GPQA ϕφ range [0.808,0.927][0.808,0.927] [0.681,0.840][0.681,0.840] [0.593,0.781][0.593,0.781] AIME ϕφ range [0.584,0.777][0.584,0.777] [0.386,0.479][0.386,0.479] [0.206,0.292][0.206,0.292] All three columns share the same λ-invariant denominator Γemp(t) _emp^(t) (verified bit-identical across the three runs). Two mechanical notes. First, the rival numerator averages over test runs with a defined simulated expectation (a test run with no simulated wrong plurality contributes no draw): at nsim=2×104n_sim=2× 10^4 this is ≥76%≥ 76\% of test runs per cell (≥87%≥ 87\% in six of eight cells), and at the main-table setting (10510^5) the draw coverage is 7878–96%96\% across cells (the shortfall reflects high-accuracy cases whose simulation never produces a wrong plurality). Second, the λ=1λ=1 column reproduces Table 3 within Monte Carlo error (nsim=2×104n_sim=2× 10^4 here vs. 10510^5 there; differences ≤0.009≤ 0.009). Two conclusions. First, on GPQA the mechanical coverage stays above 0.50.5 even when the estimated preference is treated as pure noise (λ=0λ=0): the multiple-choice result does not ride on the preference estimate. Second, on AIME ϕφ declines monotonically with λ within a fixed support, so the residual 1−ϕ1-φ is largest at λ=0λ=0: the headline AIME residual (at λ=1λ=1) is, in this direction, conservative rather than inflated by estimation noise. The benchmark split thus survives shrinkage in direction at every level, with only its magnitude attenuated. Monte Carlo SEs of these Γ values are ≤4×10−3≤ 4× 10^-3 Γ units (reported per cell in the evidence files). A draw-coverage convention is disclosed: test runs whose simulation never produced a wrong plurality contribute to the denominator Γemp(t) _emp^(t) but not to the numerator. Counting them as zero in the numerator (the most conservative alternative) lowers ϕφ by 0.030.03–0.170.17 per cell (largest, 0.170.17, on the lowest-coverage cell at 76%76\% draw coverage; results/kappa_shrink_conservative.json), and the benchmark direction survives: GPQA ϕφ stays 0.750.75–0.840.84 against AIME 0.510.51–0.710.71. Appendix D Calibrated explanatory null: run-level preference heterogeneity The rival null fixes each case’s wrong-label preference q^i q_i and treats every run of the case as an i.i.d. draw from it. A stronger reference also grants the case run-level preference heterogeneity: before drawing its K votes, each simulated run draws its own preference q∗∼Dirichlet(αq^i)q^* (α\, q_i) and then draws its wrong votes from q∗q^*. α→∞α→∞ recovers the multinomial rival; small α means the runs of a case disagree strongly about which wrong answers are attractive. One α per cell is fitted by moment-matching the simulated cross-run variance of the plurality share to the observed within-case cross-run variance of the plurality share, fitted on a random held-out half of the eligible cases and evaluated on the other half (cases with ≥3≥ 3 wrong runs; nsim=2×103n_sim=2× 10^3, seed 0; results/kappa_rival_dispersion.json). The fit and evaluation halves contain 88–8282 eligible cases per cell (results/kappa_rival_dispersion.json). The denominator of ϕdm _dm is the empirical index over the dispersion-eligible population (cases with ≥3≥ 3 wrong runs), not the test-subset index Γemp(t) _emp^(t). Table 9 reports the fitted α per cell with the observed and simulated cross-run variance and the resulting coverage ϕdm _dm. Table 9: Run-level preference heterogeneity: fitted Dirichlet α and coverage of the calibrated null ϕdm _dm. Model Prompt α fit varobsvar_obs varDMvar_DM ϕdm _dm 4.1 AIME-ZS 1.5 0.026 0.026 1.38 4.1 GPQA-ZS 1.5 0.012 0.003 1.01 4.1-mini AIME-CoT 0.5 0.038 0.010 1.68 4.1-mini AIME-ZS 0.5 0.055 0.021 1.74 4.1-mini GPQA-CoT 25 0.016 0.004 0.89 4.1-mini GPQA-ZS 50 0.021 0.003 0.85 4.1-nano AIME-ZS 1.0 0.048 0.027 2.10 4.1-nano GPQA-ZS 1.0 0.022 0.008 1.01 Three results. First, the observed run-level dispersion of the plurality share is 22–3×3× larger on AIME than on GPQA, i.e., open-domain runs of the same case show substantially greater cross-run dispersion in which wrong answers are attractive; this dispersion itself is the empirical signature of cross-run structure beyond a fixed per-case marginal. Second, on GPQA the fitted α spans 1.01.0–5050 (moderate; two cells sit at 1.01.0–1.51.5) and ϕdm _dm sits near the multinomial rival (0.850.85–1.011.01): the multiple-choice conclusion is unchanged. Third, on AIME the fitted α is 0.50.5–1.51.5 and ϕdm=1.38 _dm=1.38–2.102.10: a reference that grants the observed run-level heterogeneity more than reproduces the empirical agreement index. Two honest caveats apply. First, the dispersion parameter is fitted on a random held-out half of the eligible cases and evaluated on the held-out half, so the parameter never touches the data it explains, an independent evaluation, not merely a calibrated fit. Second, the Dirichlet form under-disperses the observed spread in every cell, worse on GPQA (2.72.7–6.1×6.1×) than on AIME (1.01.0–3.7×3.7×), so the fitted α is a lower envelope and the ϕdm _dm values under-grant heterogeneity. (The ratio ranges are computed from the unrounded variances (the table above shows them to three decimals), so they do not always reproduce from the displayed entries; each ratio is the observed variance divided by the fitted-α variance of the same cell. The fit is a moment match on the grid 0.2,0.5,1.0,1.5,3,6,12,25,50,∞\0.2,0.5,1.0,1.5,3,6,12,25,50,∞\; no fitted value sits at the grid floor. Across three independent fit/eval split seeds, the ranges are stable: GPQA ϕdm∈[0.76,1.05] _dm∈[0.76,1.05], AIME ∈[1.36,2.10]∈[1.36,2.10].) The per-seed spread of each simulated varDMvar_DM is small (relative SE ≤7%≤ 7\% across the three seeds at nsim=2×103n_sim=2× 10^3), so the dispersion null is a robust lower-envelope diagnostic rather than a precision target. The preference-unexplained residual on AIME is therefore more than absorbed by run-level preference heterogeneity: on open-domain tasks the runs of a case do not share one attractive wrong answer; they differ in which wrong answers attract them, a correlated-error structure at the level of run-to-run preference variation. The residual δ should accordingly be read as a signature consistent with this channel, not as an identified mechanism. Author contributions Conceptualization and methodology: L. Zhang; software, formal analysis, investigation, data curation, and visualization: M. Tang and L. Zhang; writing (original draft preparation): L. Zhang and M. Tang; writing (review and editing): C. Long, X. Tang, X. Luo, and L. Zhang; supervision: L. Zhang, C. Long, and X. Tang; resources and funding acquisition: C. Long, X. Tang, and X. Luo.