Paper deep dive
The Consistency Illusion: How Multi-Agent Debate Hides Reasoning Misalignment
Xiaoyang Wang, Christopher C. Yang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/9/2026, 3:50:56 AM
Summary
This paper identifies a 'consistency illusion' in multi-agent LLM debate systems for medical question answering, where answer-level consensus masks divergent reasoning. It introduces CARA (Cross-Agent Reasoning Alignment) to measure reasoning alignment and proposes the Grounded Debate Protocol (GDP) to fix misalignment by enforcing structured claims, grounds, and stances. Experiments on MedQA-USMLE and MedThink-Bench show GDP significantly improves alignment without architectural changes.
Entities (10)
Relation Signals (10)
Standard Debate → causes → Consistency Illusion
confidence 95% · debate reduces detectable contradictions between agents while simultaneously decreasing the semantic similarity of their reasoning chains
Grounded Debate Protocol → improves → Cross-Agent Reasoning Alignment
confidence 95% · gdp produces large, consistent alignment improvements across all three conditions tested
CARA → measures → Cross-Agent Reasoning Alignment
confidence 95% · CARA measures whether agents that agree on an answer also agree on the reasoning behind it
M3-GDP → outperforms → M4
confidence 95% · gdp delivers a large alignment gain on both datasets ... rising to d=+2.26 and d=+2.22 against independent voting
M3-GDP → outperforms → M3
confidence 95% · M3-GDP r1 is the highest cell (0.952/0.953), and M3 r1 is the lowest (0.892/0.881)
Consistency Illusion → characterizedby → Reduced Contradictions & Decreased Semantic Similarity
confidence 90% · Debate produces surface harmony while reasoning diverges further: the empirical signature of the consistency illusion
MedQA-USMLE → evaluatedby → CARA
confidence 90% · Applying CARA to a standard debate system on two medical QA benchmarks, MedQA-USMLE and MedThink-Bench
MedThink-Bench → →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Multi-agent LLM systems for medical question answering often treat consensus as a reliability signal: if multiple agents agree on an answer, it is presumed trustworthy. However, answer-level consensus does not entail reasoning-level alignment. We introduce CARA (Cross-Agent Reasoning Alignment), a family of automated metrics that measure whether agents who agree on an answer also agree on the reasoning. Applying CARA to a standard debate system on two medical QA benchmarks, MedQA-USMLE and MedThink-Bench, we identify the consistency illusion: a failure mode where debate reduces detectable contradictions between agents while simultaneously decreasing the semantic similarity of their reasoning chains; agents appear to agree more but reason less consistently. To improve this misalignment, we propose the Grounded Debate Protocol (GDP), a prompt-level intervention that requires agents to commit to named medical facts and take explicit stances on other agents' claims. GDP produces large, consistent alignment improvements, with Cohen's d ranging from +1.43 to +1.99, across two datasets and two backbone models, without adding LLM calls or modifying system architecture. Our results motivate cross-agent reasoning alignment as a quantity to audit alongside accuracy in safety-critical domains.
Tags
Links
- Source: https://arxiv.org/abs/2606.08457v1
- Canonical: https://arxiv.org/abs/2606.08457v1
Trouble viewing inline? Open PDF directly →
Full Text
91,405 characters extracted from source content.
Expand or collapse full text
The Consistency Illusion: How Multi-Agent Debate Hides Reasoning Misalignment Xiaoyang Wang1, Christopher C. Yang1 1Drexel University xw388@drexel.edu, chris.yang@drexel.edu Abstract Multi-agent LLM systems for medical question answering treat consensus as a reliability signal: if multiple agents agree on an answer, it is presumed trustworthy. However, answer-level consensus does not entail reasoning-level alignment. We introduce CARA (Cross-Agent Reasoning Alignment), a family of automated metrics that measure whether agents who agree on an answer also agree on the reasoning. Applying cara to a standard debate system on two medical QA benchmarks (MedQA-USMLE and MedThink-Bench), we identify the consistency illusion: a failure mode where debate reduces detectable contradictions between agents while simultaneously decreasing the semantic similarity of their reasoning chains; agents appear to agree more but reason less consistently. To improve this misalignment, we propose the Grounded Debate Protocol (gdp), a prompt-level intervention that requires agents to commit to named medical facts and take explicit stances on other agents’ claims. gdp produces large, consistent alignment improvements (Cohen’s d=+1.43d=+1.43 to +1.99+1.99) across two datasets and two backbone models, without adding LLM calls or modifying system architecture. Our results motivate cross-agent reasoning alignment as a quantity to audit alongside accuracy in safety-critical domains. The Consistency Illusion: How Multi-Agent Debate Hides Reasoning Misalignment Xiaoyang Wang1, Christopher C. Yang1 1Drexel University xw388@drexel.edu, chris.yang@drexel.edu 1 Introduction Multi-agent LLM systems are increasingly used for medical question answering. Frameworks such as MedAgents (Tang et al., 2024), MDAgents (Kim et al., 2024), and MCC (Sun et al., 2026) deploy several LLM instances that debate, consult, and vote to answer clinical questions, and report accuracy gains over single-agent baselines on benchmarks such as MedQA. Because the dominant evaluation pipeline primarily reports macroscopic answer accuracy under majority vote, the underlying trust assumption is straightforward: if multiple agents independently converge on the same answer, that answer is reliable. We argue this trust assumption is incomplete: answer-level consensus does not entail reasoning-level alignment. Figure 1 illustrates the gap on a high-stakes case: three agents independently agree that atropine is the first-line treatment for symptomatic bradycardia, yet their underlying rationales invoke three mutually exclusive pharmacological targets (β1 _1-adrenergic agonism, M2M_2-muscarinic blockade, and acetylcholinesterase inhibition). The agents reach the correct answer through medically incompatible reasoning; we term this pattern the consistency illusion. This paper therefore measures cross-agent reasoning alignment: whether agents that converge on an answer also align in the reasoning that produced it. This question is orthogonal to both macroscopic answer accuracy and individual, single-agent reasoning faithfulness. Figure 1: The consistency illusion: three agents independently agree on the clinically correct answer (atropine for symptomatic bradycardia), yet their rationales invoke three mutually exclusive pharmacological targets. Empirical 2D visualization on D2 (MedThink-Bench) appears in Appendix N. This phenomenon has gone undetected because existing evaluations measure something else. Medical multi-agent systems report accuracy and answer agreement, while reasoning-faithfulness methods were built to audit a single agent’s trace rather than alignment across agents; to our knowledge, no prior method measures whether agents that agree on an answer also agree on the reasoning behind it. Recent theoretical critiques of multi-agent deliberation (Choi et al., 2025; Zhu et al., 2025) characterize debate dynamics as a voting-driven martingale rather than genuine synthesis. They leave open a complementary question central to safety-critical settings: when agents do agree, does the reasoning behind that agreement hold up? To answer this question, we introduce cara (Cross-Agent Reasoning Alignment), a family of automated metrics that decompose each agent’s response into reasoning steps and score pairwise alignment among agents that share an answer, using a hybrid of NLI-based contradiction detection and embedding similarity. cara introduces no additional LLM calls and operates post hoc on traces produced by ordinary system inference. Applying cara to a standard three-agent debate system on MedQA-USMLE (N=500N=500) and MedThink-Bench (N=499N=499), we find that debate reduces explicit contradictions between agents (CR↓\, ) while simultaneously reducing the semantic similarity of their reasoning chains (SIM↓\, ). Debate produces surface harmony while reasoning diverges further: the empirical signature of the consistency illusion. The effect is pronounced in open-ended reasoning, where the unconstrained reasoning space gives divergence room to surface: on MedThink-Bench it is large and replicates across both backbones (Cohen’s d=−0.30d=-0.30 for Qwen, −1.32-1.32 for Llama-3), and it is muted on closed-form multiple-choice (MedQA), consistent with the narrower reasoning surface. To test whether this failure is fundamental to multi-agent debate or merely a protocol-level deficiency, we propose the Grounded Debate Protocol (gdp), a prompt-level intervention requiring each reasoning step to pair a Claim with a named Ground and forcing explicit Stance statements toward other agents’ claims. gdp produces large, consistent alignment gains across all three conditions tested (Cohen’s d=+1.43d=+1.43 on D1 Qwen, +1.62+1.62 on D2 Qwen, and +1.99+1.99 on D2 Llama-3). Together, our results identify a previously unmeasured failure mode in multi-agent medical QA, name and quantify it, and demonstrate that lightweight prompt-level grounding substantially improves cross-agent reasoning alignment. We release the code and computed cara scores publicly111Code and data: https://anonymous.4open.science/r/consistency-illusion-code-3629/. 2 Related Work Multi-agent systems for medical QA. A growing family of medical MAS—MedAgents (Tang et al., 2024), MDAgents (Kim et al., 2024), MCC (Sun et al., 2026), TeamMedAgents (Mishra et al., 2025), the foundational debate framework of Du et al. (2024), and clinical-simulation systems like Agent Hospital (Li et al., 2025) and AgentClinic (Schmidgall et al., 2026)—deploy multiple LLM instances that debate, consult, or vote on medical questions, reporting accuracy gains over single-agent baselines on benchmarks such as MedQA. A common thread unites these architectures: they treat multi-agent answer-level consensus as a proxy for reliable clinical reasoning. None, however, audits whether agents that converge on the same option actually share consistent underlying rationales. Critiques of multi-agent deliberation. A growing body of work questions whether multi-agent debate improves reasoning beyond what voting or repeated sampling already provides. Choi et al. (2025) prove formally that debate dynamics form a martingale—gains are attributable to majority voting, not deliberation—a conclusion reinforced empirically by Zhu et al. (2025), Li et al. (2024), Wang et al. (2024), Tran and Kiela (2026), and Zhang et al. (2025b). Behavioral studies document specific failure modes: agents may reach consensus by copying explanations (Pitre et al., 2025), switch answers under sycophantic pressure (Yao et al., 2025), or degrade in multi-turn dialogue (Laban et al., 2026). Most closely related, the problem drift literature (Becker et al., 2026) measures how debate performance degrades over turns; cara asks a complementary question: when agreement does emerge, is the underlying reasoning aligned? Prior critiques attack the debate process; our work attacks an assumption about the debate outcome—that surface consensus implies reasoning alignment. Evaluation of multi-agent reasoning. A distinct evaluation lineage measures whether reported rationales support or reflect a single agent’s answer. Lanham et al. (2023) introduced CoT truncation and mistake-injection probes; Matton et al. (2025) extended these to medical QA, and Afolabi et al. (2025) show that reasoning steps often do not causally drive predictions. Automated metrics include ROSCOE (Golovneva et al., 2023), NLI-based faithfulness checking (Falke et al., 2019), and C-SHAP (Parcalabescu and Frank, 2024), but these are designed for single-trace verification of individual model outputs. Even the methods that do apply to multi-agent settings verify each trace in isolation; to our knowledge, none measures cross-agent reasoning alignment among agents that share an answer. How multi-agent aggregation affects collective reasoning coherence has therefore remained uncharacterized. 3 CARA: Measuring Reasoning Alignment cara measures whether agents that agree on an answer also agree on the reasoning behind it—a quantity that answer accuracy and single-agent faithfulness leave unmeasured. It is orthogonal to correctness: rather than checking whether a rationale is medically right, it asks whether agents who reach the same answer reach it for compatible reasons. 3.1 Agreement Set We run N agents on a clinical question q. Each agent AiA_i produces an answer aia_i and a reasoning chain Ri=(ri1,…,riKi)R_i=(r_i^1,…,r_i^K_i) of KiK_i steps. The system answer a a is the majority vote over ai\a_i\. We define the agreement set as the agents who voted for a a: (q)=i:ai=a^.S(q)=\i:a_i= a\. (1) We compute cara only when |(q)|≥2|S(q)|≥ 2. This restriction is the point: it isolates the safety-critical case where surface agreement may hide divergent reasoning. 3.2 Step-Level Alignment We split each rationale RiR_i into steps using a deterministic text-segmentation pipeline. For each ordered pair (Ai,Aj)(A_i,A_j) in the agreement set with i≠ji≠ j, we build an alignment matrix ij∈ℝKi×KjM_ij ^K_i× K_j with entries ij[k,l]=align(rik,rjl)M_ij[k,l]=align(r_i^k,r_j^l). Our metric is cara-hyb; cara-nli and cara-sim are its two components, reported separately in Table 9 to expose what drives alignment. cara-nli returns an NLI label in −1,0,+1\-1,0,+1\ for contradiction, neutral, or entailment. cara-sim returns the cosine similarity of step embeddings. cara-hyb combines both, letting contradictions override similarity: alignhyb(rik,rjl)=−1if PNLI>τ,cos(ik,jl)otherwise,align_hyb(r_i^k,r_j^l)= cases-1&if P_NLI>τ,\\ (e_i^k,e_j^l)&otherwise, cases (2) where PNLIP_NLI is the NLI contradiction probability for (rik,rjl)(r_i^k,r_j^l), ike_i^k is the L2-normalized embedding of rikr_i^k, and τ is the contradiction threshold. This hybrid addresses the known insensitivity of embeddings to negation (Marelli et al., 2014). For validation, cara-llm uses GPT-4o as an LLM judge to calibrate cara-hyb against holistic human-style judgments, reported in Section 5.3. 3.3 Question-Level Aggregation and Diagnostics Best-match alignment. Because agents produce different numbers of steps (Ki≠KjK_i≠ K_j), alignment is asymmetric. We define best-match alignment by matching each step in AiA_i to its best match in AjA_j: matchj(rik)=maxl∈1,…,Kjalign(rik,rjl).match_j(r_i^k)= _l∈\1,…,K_j\align(r_i^k,\,r_j^l). (3) Symmetrized pairwise score. The pairwise cara score symmetrizes by averaging both directions: CARA(i,j)=12[1Ki∑k=1Kimatchj(rik)+1Kj∑l=1Kjmatchi(rjl)].CARA(i,j)= 12 [ 1K_i\! _k=1^K_i\,match_j(r_i^k)\\ + 1K_j\! _l=1^K_j\,match_i(r_j^l) ]. (4) Question-level aggregation. The question-level score averages over all pairs in the agreement set: CARA(q)=2|(q)|(|(q)|−1)×∑i<ji,j∈(q)CARA(i,j).CARA(q)= 2|S(q)|\,(|S(q)|-1)\\ × _ subarrayci<j\\ i,j (q) subarrayCARA(i,j). (5) Variants with negative range (cara-nli, cara-hyb) are rescaled to [0,1][0,1] via (CARA+1)/2(CARA+1)/2. Corpus-level scores are the mean of CARA(q)CARA(q) over questions with |(q)|≥2|S(q)|≥ 2. Contradiction Rate (CR). We also report the Contradiction Rate: the fraction of best-matched step pairs in the agreement set flagged as contradictions by the NLI filter. CR(q)=|(i,j,k):i≠j,matchj(rik)=−1|∑i∈(q)(|(q)|−1)Ki.CR(q)= | \(i,j,k):i≠ j,\ match_j(r_i^k)=-1 \ | _i (q)(|S(q)|-1)K_i. (6) True alignment shows up as semantic similarity rising (cara-sim ↑ ) with contradictions falling (CR↓ ). When the two diverge—contradictions fall but similarity also falls—we see the consistency illusion. 4 Grounded Debate Protocol: A Protocol-Level Intervention 4.1 Motivation and Protocol Design cara shows that free-form debate can produce answer agreement without aligned reasoning, a result we establish in Section 6. gdp tests whether this failure is an unresolvable feature of language-model interaction or merely a protocol-level deficiency. Standard debate prompts allow three failure modes: vague clinical claims; sycophantic answer adoption without engaging with peer evidence (Yao et al., 2025); and explanation copying with topic drift (Pitre et al., 2025; Becker et al., 2026). The common root is that nothing in standard prompts requires agents to commit to specific reasoning steps or substantively engage with other agents’ claims. To address this, gdp enforces a structured, prompt-level communication protocol. It requires each agent to structure its output into up to three machine-parseable fields: • Claim: A singular, atomic, and falsifiable clinical assertion. • Ground: a named medical fact, mechanism, or guideline that supports the claim (e.g., disease, drug target, receptor pathway). • Stance (debate rounds only): one of Agree, Disagree, Extend toward a specific claim from another agent, with a one-sentence justification; a Disagree requires a counter-Ground. Agents emit Claim+Ground pairs in the independent round (r0r_0) and add Stance in the debate round (r1r_1). An anti-sycophancy rule instructs each agent to switch answers between r0r_0 and r1r_1 only after seeing a more compelling Ground from another agent or finding a factual error in its own reasoning. A worked clinical example contrasting an unstructured debate log with a synchronized gdp trace is provided in Appendix K. 4.2 Integration Properties gdp is a prompt-level change: the Claim/Ground/Stance format is added through the system message, with agent routing and the number of LLM calls unchanged. The only overhead is output length: tokens grow by ∼9% 9\% from the explicit field markers. Applied to the symmetric debate of Du et al. (2024), this gives M3-GDP, which we evaluate in Section 5.1. Traces are parsed post-hoc by regex on the field markers; this parsing supports logging and analysis only, as cara operates on the raw response text regardless of parse outcome. 5 Experimental Setup 5.1 Study Design We evaluate whether debate-induced answer consensus is accompanied by reasoning alignment among agents that converge on the same answer. Table 10 summarizes our experimental matrix of three systems, each with three agents under majority voting: M3 is symmetric one-round debate (Du et al., 2024), producing pre-debate (r0r_0) and post-debate (r1r_1) traces; M4 is an independent-vote control with no cross-agent communication, specifically isolating the sampling-and-voting contribution from debate (Choi et al., 2025); M3-GDP keeps the M3 architecture and the same LLM-call budget but rewrites all prompts to require the gdp Claim+Ground+Stance format (Section 4). Full configurations and prompts are in Appendix J. 5.2 Datasets and Backbones D1: MedQA-USMLE (Jin et al., 2021) provides closed-form four-option USMLE-style multiple-choice questions; we sample N=500N=500 stratified by exam stage with a fixed seed. D2: MedThink-Bench (Zhou et al., 2025) spans ten medical domains with 33–1010 answer options and expert-annotated reasoning trajectories (“Scoring Points”); we use N=499N=499 of 500 available items (one excluded for incompatible option format). D2’s wider option space gives reasoning more room to diverge than D1. Primary trials use Qwen 2.5 72B Instruct (Qwen et al., 2025); we replicate on D2 (N=100N=100) with Llama 3.3 70B Instruct (Grattafiori et al., 2024) to test the effect-direction consistency rather than cross-backbone statistical significance. 5.3 CARA Computation and Validation cara-hyb implementation. cara-hyb decomposes each agent response into reasoning steps and computes pairwise alignment via two complementary signals: an NLI hard-filter on contradictions (DeBERTa-v3 (Laurer et al., 2024)) and a sentence-embedding cosine on non-contradictory step pairs (Stella (Zhang et al., 2025a)). The pipeline runs post hoc with no additional LLM calls; model versions, threshold, decomposition rules, and runtime are in Appendix E. Agent–expert validation (cara-GT). For D2 only, cara-GT applies the same pipeline with one agent replaced by the expert Scoring-Points reference, measuring agent-to-expert alignment and expert coverage, detailed in Appendix C. LLM-judge calibration (cara-llm). cara-llm validates cara-hyb against GPT-4o-as-judge (OpenAI, 2024) on a sample of agent pairs: the scores show strict tercile monotonicity (and a moderate linear correlation), capturing complementary aspects of alignment—step-level semantic overlap (cara-hyb) vs. holistic logical coherence (GPT-4o judge). Sample design, correlation values, and tercile statistics are in Appendix F. Human-expert validation. Because reliable annotation of medical reasoning requires medically-trained experts, we validate on a deliberately focused set of 5050 agreement-set pairs rated by three independent annotators, with the full setup in Appendix H. Inter-annotator agreement is high (Spearman ρ∈[0.93,0.99]ρ∈[0.93,0.99]; 100%100\% within-1 agreement), and aggregate ratings correlate positively with cara-hyb (Spearman ρ=0.26ρ=0.26) and cara-sim (ρ=0.31ρ=0.31), with tercile means rising monotonically (4.00/4.65/4.714.00/4.65/4.71 for low/mid/high cara-hyb strata). The bounded magnitude reflects range restriction within the agreement set (cara-hyb∈[0.76,0.98] cara-hyb∈[0.76,0.98]), paralleling the cara-llm pattern. 5.4 Statistical Analysis We report Cohen’s d with 95%95\% percentile bootstrap CIs (B=10,000B=10,000, seed 4242); |d|>0.8|d|>0.8 is Tier A, 0.5<|d|≤0.80.5<|d|≤0.8 Tier B, and 0.2<|d|≤0.50.2<|d|≤0.5 Tier C. Accuracy comparisons use continuity-corrected McNemar tests. Questions with fewer than two agents choosing the same answer are undefined; we report this undefined rate per system and check survivorship bias by imputing the undefined cases. Per-cell counts are in Appendix M, and the full imputation in Appendix B and Section 2. 6 Results We report cara-hyb, cara-sim, and CR for all three systems on D1 and D2 with Qwen 2.5 72B, plus a cross-backbone replication on D2 (N=100N=100) with Llama 3.3 70B. All comparisons are within-dataset and within-backbone; we do not compare absolute cara values across datasets. D1 (MedQA, N=500N=500) D2 (MedThink, N=499N=499) System Round HYB SIM CR HYB SIM CR M3-GDP r1 (debate) 0.952 0.912 0.098 0.953 0.914 0.127 M3-GDP r0 (indep.) 0.912 0.835 0.117 0.912 0.836 0.137 M4 r0 (indep.) 0.895 0.801 0.108 0.895 0.801 0.161 M3 r0 (indep.) 0.895 0.801 0.104 0.896 0.801 0.144 M3 r1 (debate) 0.892 0.787 0.035† 0.881 0.769 0.054† Table 1: cara-hyb, cara-sim, and CR per cell on D1 and D2. Higher HYB and SIM, and lower CR, indicate better alignment. Green row = M3-GDP r1 (highest alignment); Red row = M3 r1 (lowest, consistency illusion). Per-cell HYB SDs are 0.0210.021–0.0590.059. N ranges 424424–500500 (D1) and 428428–499499 (D2) due to undefined questions. †M3 r1 CR is deflated by survivorship bias (Section 2). 6.1 Main Results and Dual-Layer Decomposition The rank ordering is identical on D1 and D2: M3-GDP r1 is the highest cell (0.9520.952/0.9530.953), and M3 r1 is the lowest (0.8920.892/0.8810.881)—standard debate alignment falls below independent voting. For M3 from r0 to r1, both CR and SIM decrease (CR 0.104→0.0350.104→0.035 on D1, 0.144→0.0540.144→0.054 on D2; SIM 0.801→0.7870.801→0.787 on D1, 0.801→0.7690.801→0.769 on D2); fewer contradictions with less semantic alignment is the consistency illusion (mechanism in Section 7). In contrast, gdp debate produces the genuine-alignment signature.From r0 to r1, SIM rises (0.835→0.9120.835→0.912 on D1; 0.836→0.9140.836→0.914 on D2) while CR falls (0.117→0.0980.117→0.098 on D1; 0.137→0.1270.137→0.127 on D2). Both protocols reduce contradictions, but only gdp simultaneously increases reasoning similarity—CR↓ with SIM↑ , the true-alignment pattern defined in Section 3 and the mirror image of the illusion. A D1 bootstrap test underscores that M3’s low CR is anomalous: gdp r1 CR exceeds M3 r1 (Δ=+0.063 =+0.063, CI [+0.042,+0.087][+0.042,+0.087]), and M3 r1 CR falls even below independent voting (Δ=−0.072 =-0.072 vs. M4, CI [−0.089,−0.056][-0.089,-0.056])—standard debate suppresses visible contradictions without aligning reasoning. 6.2 GDP Effect and the Consistency Illusion Figure 2 shows the three key pairwise comparisons as a forest plot of Cohen’s d across all three dataset–backbone combinations. Figure 2: Cohen’s d for three key pairwise comparisons, one panel per dataset–backbone. Error bars: 95%95\% bootstrap CIs (B=10,000B=10,000); filled markers exclude zero; shading marks Tier A (|d|>0.8|d|>0.8). gdp comparisons (top two rows) reach Tier A across all conditions; the consistency-illusion comparison (bottom) is significantly negative on D2 (both backbones). gdp delivers a large alignment gain on both datasets: d=+1.43d=+1.43 on D1 and d=+1.62d=+1.62 on D2 against M3r1_r_1, rising to d=+2.26d=+2.26 and d=+2.22d=+2.22 against independent voting (all four CIs above zero). Standard debate, in contrast, is alignment-neutral on D1 (d=−0.08d=-0.08, CI contains zero) and alignment-degrading on D2 (d=−0.30d=-0.30, CI [−0.43,−0.17][-0.43,-0.17]). The D1–D2 gap tracks both question complexity and undefined-question selection bias. The gdp effect splits in two: Claim+Ground format alone (GDPr0_r_0 vs. M4r0_r_0) yields d=+0.62d=+0.62 on D1; adding Stance-mediated debate (GDPr1_r_1 vs. GDPr0_r_0) yields d=+1.51d=+1.51, 2.4×2.4× the format component but only meaningful when there are structured claims to reference. On survivorship-bias sensitivity, M3 r1 excludes 15.2%15.2\% (D1) and 14.2%14.2\% (D2) of questions where no majority emerged. These exclusions are non-random—harder questions survive less often—so the sample skews toward easier cases. Under worst-case imputation (cara-hyb =0.50=0.50, CR=1.0=1.0 for every undefined question, both far beyond any observed value), M3 r1’s adjusted cara-hyb stays below M4 r0 (0.8320.832 vs. 0.8950.895 on D1; 0.8270.827 vs. 0.8950.895 on D2): the consistency-illusion conclusion is robust. The CR finding, however, reverses (0.182>0.1080.182>0.108 on D1; 0.188>0.1610.188>0.161 on D2), so we interpret CR alongside the undefined rate. 6.3 Cross-Backbone Validation We replicate all three systems on D2 (N=100N=100) with Llama 3.3 70B Instruct. All five key comparisons agree in sign across both backbones (Table 13): gdp is Tier A positive (Llama d=+1.99d=+1.99, Qwen d=+1.62d=+1.62), the illusion is negative (Llama d=−1.32d=-1.32, Tier A; Qwen d=−0.30d=-0.30, Tier C), standard debate degrades alignment r0→ 1 while gdp debate improves it, and the format effect is positive. The illusion is stronger on Llama partly because Llama’s M3 r1 undefined rate (2%2\%) is far lower than Qwen’s (14%14\%), so more hard cases survive into the measured sample. 6.4 Agent–Expert Alignment (cara-GT) cara-GT replaces one agent with the expert Scoring-Points reference (D2 only; Section 5.3). It mirrors the inter-agent finding: gdp moves agents toward expert reasoning (d=+0.32d=+0.32, CI [+0.20,+0.45][+0.20,+0.45]), and standard debate moves them away (d=−0.21d=-0.21, CI [−0.34,−0.08][-0.34,-0.08]). Expert-coverage tells a sharper story (Figure 3): standard debate loses 8.68.6p of expert scoring-point coverage; gdp recovers 5.45.4p. Figure 3: Expert scoring-point coverage on D2: standard debate (M3 r1r_1) drops 8.68.6p from the independent-voting baseline; gdp (M3-GDP r1r_1) recovers 5.45.4p. 6.5 Accuracy cara measures reasoning alignment, not answer correctness; majority-vote accuracy is reported separately. Standard debate does not change accuracy on either dataset (D1 Δ=+0.4 =+0.4p, D2 Δ=0.0 =0.0p; both p>0.85p>0.85), consistent with the martingale characterization (Choi et al., 2025). gdp shifts accuracy by −2.4-2.4p on Qwen (both D1 and D2) and −6.0-6.0p on Llama D2; none reaches significance at α=0.05α=0.05. D2’s lower absolute accuracy (∼41% 41\% baseline) reflects its ∼7.5 7.5 average answer options versus D1’s 44, giving an ≈3×≈3× random-performance ratio. M3 r1 GDP r1 Severity Mode Definition D1 D2 D1 D2 Severe FM1 Complementary reasoning: different, non-contradictory reasoning paths. 23 18 0 0 FM4 Sycophantic convergence: agent adopts the majority answer with zero reasoning steps. 11 15 0 0 Moderate FM2 Granularity mismatch: step-count ratio >2>2 (depth difference, not factual clash). 12 15 5 2 FM3 Contradictory premises: agents cite conflicting facts but converge on the correct answer. 2 1 12 15 FM5 Terminology divergence: different medical terms used for the same concept. 2 1 0 0 Mild FM6 Partial overlap: residual baseline misalignment without extreme signals. 0 0 31 32 Severe modes eliminated by gdp (FM1 + FM4): 34 33 0 0 Table 2: Failure mode taxonomy on the 50 lowest-cara-hyb correct-answer cases per system, with definitions and counts on both datasets. gdp eliminates both severe modes (FM1 + FM4: 34/5034/50 and 33/5033/50 under M3 → 0/500/50 under gdp on D1 and D2 respectively), shifting residual misalignment toward the mild baseline FM6. FM3’s rise under gdp is consistent with the CR analysis in §6.1: the structured format surfaces contradictions that vague free text would hide. 7 Analysis We answer four research questions about the consistency illusion, the mechanism by which gdp improves it, and the generality of the findings. RQ1: What mechanism produces the consistency illusion, and how does GDP improve it? Answer: Standard debate creates the illusion via two failures—contradiction smoothing without reasoning convergence, and sycophantic convergence—and gdp reverses both via a format effect (d=+0.62d=+0.62) amplified by a Stance-mediated debate interaction effect (d=+1.51d=+1.51, 2.4×2.4× the format component; extended derivation in Appendix O). Standard debate drives CR↓\, +SIM↓\, (Section 6.1) via two mechanisms. First, revision pressure removes contradictory steps without replacing them with shared reasoning—subtraction without alignment—so CR drops while SIM also drops. Second, agents adopt each other’s answers without adopting the underlying reasoning, sometimes severely: M3 r1 fails consensus on 1414–15%15\% of questions on both datasets. gdp reverses both via two complementary effects. The format effect alone, before any debate (GDPr0_r0 vs. M4r0_r0 on D1: SIM +0.034+0.034, d=+0.62d=+0.62), shows that much of the apparent misalignment originates from underspecified output format rather than genuine reasoning disagreement. The dominant debate interaction effect (d=+1.51d=+1.51 on D1) operates because Stance statements force agents to share reasoning vocabulary and logical structure with their interlocutors, driving SIM up by +0.077+0.077 on D1 and +0.078+0.078 on D2 from r0 to r1. Format is a precondition: without structured claims, Stance has no targets. RQ2: Beyond CARA, what behavioral indicators support the illusion finding? Answer: Standard debate destabilizes consensus ∼38× 38× more than independent voting on D1 and produces 5%5\% “reasoning collapse” zero-step agents on D2; gdp eliminates both. Two independent behavioral indicators support the illusion finding without depending on agreement-set selection. First, the undefined rate (no majority emerges) jumps from 0.4%0.4\% before debate (M3 r0) to 15.2%15.2\% after (M3 r1) on D1 (38×38×) and 5.6%→14.2%5.6\%→14.2\% on D2: debate itself manufactures disagreement; grounding removes it (M3-GDP r1: 0.4%0.4\%). Second, reasoning collapse—agents emitting zero extractable steps after debate, the behavioral signature of sycophantic convergence—affects 5.0%5.0\% of M3 r1 agreement-set memberships on D2 (N=100N=100) vs. 0.3%0.3\% for M3-GDP r1. gdp’s Claim+Ground requirement prevents both failure modes. RQ3: Which failure modes does GDP structurally eliminate? Answer: gdp eliminates FM1 (complementary reasoning) and FM4 (sycophantic convergence) on both datasets, shifting residual misalignment toward the low-severity baseline FM6 (partial overlap, 6262–64%64\% of remaining low-cara-hyb cases). We examine the 50 lowest-cara-hyb cases per system where each still reaches the correct answer, classifying by deterministic rules over CARA sub-metrics; the full classification rules are in Appendix L. Table 2 groups the six modes into Severe (FM1, FM4), Moderate (FM2, FM3, FM5), and Mild (FM6) tiers. Four findings replicate across the two datasets. First, gdp eliminates sycophantic convergence (FM4): 11→011→0 on D1, 15→015→0 on D2; the Claim+Ground requirement structurally prevents empty answer adoption. Second, gdp eliminates complementary reasoning (FM1): 23→023→0 on D1, 18→018→0 on D2; structured grounding forces convergence on shared medical facts. Third, residual misalignment under gdp shifts to FM6 (partial overlap): 3131/5050 on D1, 3232/5050 on D2 (6262–64%64\%)—a low-severity baseline remaining after severe modes are removed. Fourth, FM3 contradictory premises rises under gdp (2→122→12 on D1, 1→151→15 on D2), consistent with the CR analysis in Section 6.1: the structured format surfaces contradictions that vague free text would hide. A qualitative case study (MQA_0447 hypovolemic shock, where agents misstate SVR direction before self-correcting) is in Appendix L. RQ4: Do the findings generalize across datasets and backbones? Answer: Yes—the gdp effect replicates with Tier-A magnitudes across all three conditions (D1 Qwen d=+1.43d=+1.43; D2 Qwen d=+1.62d=+1.62; D2 Llama d=+1.99d=+1.99), with the system rank ordering identical across the two datasets and two independent backbones. The rank ordering M3-GDPr1≻M3-GDPr0≻M4r0≈M3r0≻M3r1M3-GDP_r_1\! \!M3-GDP_r_0\! \!M4_r_0\!≈\!M3_r_0\! \!M3_r_1 is identical across all three dataset–backbone combinations. The illusion magnitude varies across conditions (d=−0.08,−0.30,−1.32d=-0.08,-0.30,-1.32) for two reasons. First, D1’s fixed 4-option format leaves less room for reasoning to diverge. Second, M3 r1 undefined rates differ (15.2%,14.2%,2.0%15.2\%,14.2\%,2.0\%): a lower undefined rate leaves more hard cases in the sample, strengthening the measured illusion. Is the gain just structural? Because cara-hyb scores step-pair similarity, gdp’s structured Claim/Ground/Stance format could in principle inflate scores mechanically—through shared template vocabulary or cleaner step parsing—rather than through better reasoning alignment. First, the format-only effect (d=+0.62d=+0.62) is less than half the full gdp gain; the dominant component (d=+1.51d=+1.51) is the Stance-mediated debate interaction over content, not template. Second, cara-GT (Section 6.4) shows gdp also aligns agents with the expert Scoring-Points reference (d=+0.32d=+0.32), which is not in gdp format. Third, GPT-4o judges M3-GDP r1 reasoning as more aligned than M3 r1 on the same 50 agreement-set pairs (4.864.86 vs. 4.504.50 on a 5-point scale, Cohen’s d=+0.75d=+0.75, 95%95\% bootstrap CI on Δ [+0.08,+0.65][+0.08,+0.65]); since the LLM judge evaluates holistic logical coherence rather than template structure, this is direct evidence against pure structural confound. Fourth, stripping the literal Claim/Ground/Stance tokens from gdp traces and recomputing cara-hyb attenuates paired Cohen’s dzd_z by only 66–10%10\% across all three conditions (Appendix G). 8 Conclusion We identified the consistency illusion in multi-agent medical QA: standard debate reduces detectable contradictions between agents while simultaneously decreasing the semantic similarity of their reasoning—agents appear to agree more but actually reason less consistently. We introduced cara, a family of metrics for cross-agent reasoning alignment, and applied it on MedQA-USMLE and MedThink-Bench: the illusion is negligible on D1 (d=−0.08d=-0.08) but significantly negative on D2 (Qwen d=−0.30d=-0.30; Llama-3 d=−1.32d=-1.32). We proposed the Grounded Debate Protocol (gdp), a prompt-level intervention requiring Claim+Ground+Stance, which delivers Tier-A alignment gains across all conditions (d=+1.43d=+1.43 to +1.99+1.99), eliminates two failure modes (complementary reasoning and sycophantic convergence) on both datasets and adds no LLM calls. Together, our results challenge the assumption that multi-agent consensus indicates reasoning reliability, and show that protocol-level grounding can close this gap, positioning gdp as a diagnostic and auditing intervention. The structural sensitivity we observe—larger illusion where reasoning has more room to diverge—suggests grounding is most valuable on open-ended clinical tasks and proprietary backbones. Limitations While we evaluate cara and gdp across two medical QA benchmarks (D1 MedQA-USMLE N=500N=500 and D2 MedThink-Bench N=499N=499) and two open-weight backbones (Qwen 2.5 72B and Llama 3.3 70B), with Tier-A effect sizes (d=+1.43d=+1.43 to +1.99+1.99) replicated across all conditions, our analysis has several limitations. cara’s agreement-set definition relies on discrete answer matching, which is well-defined for multiple-choice benchmarks but would require adaptation for free-form clinical generation tasks such as differential diagnosis or treatment-plan drafting; extending the metric to such settings is a natural next step. Our cross-backbone validation covers two open-weight 70B-class models at comparable scales; proprietary models (e.g., GPT-4o, Claude, Gemini) would require an identical-instance experimental setup that is outside the scope of this work. cara-hyb is a step-level semantic-alignment measure designed to operate within the agreement set; the complementary GPT-4o oracle (Appendix F) and human-expert validation (Appendix H) anchor its interpretation, with both showing the predicted monotonicity across cara-hyb terciles. Extending the metric to non-agreement-set settings (e.g., disagreement-pair audit) and to domains beyond clinical reasoning is left for future work. cara and gdp target cross-agent reasoning alignment, a property orthogonal to answer accuracy; GDP is accordingly intended as an alignment-auditing and -improving intervention, and we make no claim that it improves accuracy. Across both backbones and datasets, majority-vote accuracy under GDP does not differ significantly from the standard debate baseline (McNemar tests; Table 6)—consistent with prior evidence that multi-agent debate itself does not reliably improve accuracy (Choi et al., 2025). The Llama 3.3 comparison is based on N=100N=100 and is underpowered to detect small accuracy effects; an adequately powered analysis of the alignment–accuracy relationship is left for future work. We study a single, widely-used debate protocol—the symmetric, single-round design of Du et al. (2024), the canonical multi-agent debate baseline. Whether the consistency illusion and gdp’s correction generalize to other topologies (asymmetric roles, additional rounds, larger agent sets) is an important direction for future work. Ethics Statement This study uses publicly available medical QA benchmarks (MedQA-USMLE and MedThink-Bench) that contain no patient-identifiable information. All LLM inference uses open-weight models (Qwen 2.5 72B, Llama 3.3 70B) running on institutional compute infrastructure. cara-llm validation uses the OpenAI API (GPT-4o) for 80 evaluation pairs. No clinical participants are involved. Three medically-trained annotators rated 50 agent-pair outputs from the agreement set for metric validation; this is an annotation task on model outputs and does not constitute human-subjects research. We emphasize that cara measures reasoning alignment, not reasoning correctness. High cara scores indicate that agents reason consistently, not that they reason correctly. cara should be used as a diagnostic tool alongside accuracy metrics, not as a replacement for clinical validation. Multi-agent medical QA systems evaluated in this work are research prototypes and are not intended for clinical deployment. References H. Afolabi, Z. Afolabi, E. Friel, J. Roberts, A. Ji-Xu, L. Chen, E. Ogbomo, E. Imevbore, P. Eneje, W. El Ouahidi, A. Sohal, A. Kennan, S. Srivastava, A. Vairavan, L. Napitu, and K. McClure (2025) Faithful or just plausible? evaluating faithfulness for medical reasoning in closed-source LLMs. In NeurIPS 2025 Workshop on Generative AI for Health, External Links: Link Cited by: §2. J. Becker, L. B. Kaesberg, A. Stephan, J. P. Wahle, T. Ruas, and B. Gipp (2026) Stay focused: problem drift in multi-agent debate. In Findings of the Association for Computational Linguistics: EACL 2026, V. Demberg, K. Inui, and L. Marquez (Eds.), Rabat, Morocco, p. 5068–5102. External Links: Link, Document Cited by: §2, §4.1. H. K. Choi, J. Zhu, and S. Li (2025) Debate or vote: which yields better decisions in multi-agent large language models?. In Advances in Neural Information Processing Systems 38 (NeurIPS), Note: Spotlight External Links: Link Cited by: §1, §2, §5.1, §6.5, Limitations. Y. Du, S. Li, A. Torralba, J. B. Tenenbaum, and I. Mordatch (2024) Improving factuality and reasoning in language models through multiagent debate. In Proceedings of the 41st International Conference on Machine Learning (ICML), External Links: Link Cited by: Table 10, §2, §4.2, §5.1, Limitations. T. Falke, L. F. R. Ribeiro, P. A. Utama, I. Dagan, and I. Gurevych (2019) Ranking generated summaries by correctness: an interesting but challenging application for natural language inference. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, A. Korhonen, D. Traum, and L. Màrquez (Eds.), Florence, Italy, p. 2214–2220. External Links: Link, Document Cited by: §2. O. Golovneva, M. Chen, S. Poff, M. Corredor, L. Zettlemoyer, M. Galley, and A. Celikyilmaz (2023) ROSCOE: a suite of metrics for scoring step-by-step reasoning. In The Eleventh International Conference on Learning Representations (ICLR), External Links: Link Cited by: §2. A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Guzmán, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kambadur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. Çelebi, P. Alrassy, P. Zhang, P. Li, P. Vasic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Ganapathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Polidoro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot, S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speckbacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchandani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Montgomery, E. Presani, E. Hahn, E. Wood, E. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Jagadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma (2024) The Llama 3 herd of models. arXiv preprint arXiv:2407.21783. External Links: Link Cited by: §5.2. D. Jin, E. Pan, N. Oufattole, W. Weng, H. Fang, and P. Szolovits (2021) What disease does this patient have? A large-scale open domain question answering dataset from medical exams. Applied Sciences 11 (14), p. 6421. Note: Originally arXiv:2009.13081 (2020) External Links: Document, Link Cited by: §5.2. Y. Kim, C. Park, H. Jeong, Y. S. Chan, X. Xu, D. McDuff, C. Breazeal, and H. W. Park (2024) MDAgents: an adaptive collaboration of LLMs for medical decision-making. In Advances in Neural Information Processing Systems 37 (NeurIPS), External Links: Link Cited by: §1, §2. W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica (2023) Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles (SOSP), External Links: Document, Link Cited by: Appendix J. P. Laban, H. Hayashi, Y. Zhou, and J. Neville (2026) LLMs get lost in multi-turn conversation. In The Fourteenth International Conference on Learning Representations (ICLR), Note: Oral External Links: Link Cited by: §2. T. Lanham, A. Chen, A. Radhakrishnan, B. Steiner, C. Denison, D. Hernandez, D. Li, E. Durmus, E. Hubinger, J. Kernion, K. Lukić, K. Nguyen, N. Schiefer, T. Rauber, N. Larson, S. McCandlish, C. Olsson, T. Henighan, E. Perez, and S. R. Bowman (2023) Measuring faithfulness in chain-of-thought reasoning. arXiv preprint arXiv:2307.13702. Note: Anthropic External Links: Link Cited by: §2. M. Laurer, W. van Atteveldt, A. Casas, and K. Welbers (2024) Building efficient universal classifiers with natural language inference. arXiv preprint arXiv:2312.17543. External Links: Link Cited by: Appendix E, §5.3. J. Li, Y. Lai, W. Li, J. Ren, M. Zhang, X. Kang, S. Wang, P. Li, Y. Zhang, W. Ma, and Y. Liu (2025) Agent hospital: a simulacrum of hospital with evolvable medical agents. External Links: 2405.02957, Link Cited by: §2. J. Li, Q. Zhang, Y. Yu, Q. Fu, and D. Ye (2024) More agents is all you need. Transactions on Machine Learning Research. External Links: ISSN 2835-8856, Link Cited by: §2. M. Marelli, L. Bentivogli, M. Baroni, R. Bernardi, S. Menini, and R. Zamparelli (2014) SemEval-2014 task 1: evaluation of compositional distributional semantic models on full sentences through semantic relatedness and textual entailment. In Proceedings of the 8th International Workshop on Semantic Evaluation (SemEval 2014), Dublin, Ireland, p. 1–8. External Links: Link, Document Cited by: §3.2. K. Matton, R. Ness, J. Guttag, and E. Kiciman (2025) Walk the talk? measuring the faithfulness of large language model explanations. In The Thirteenth International Conference on Learning Representations (ICLR), External Links: Link Cited by: §2. P. P. Mishra, M. Arvan, and M. Zalake (2025) TeamMedAgents: enhancing medical decision-making of LLMs through structured teamwork. arXiv preprint arXiv:2508.08115. External Links: Link Cited by: §2. OpenAI (2024) GPT-4o system card. arXiv preprint arXiv:2410.21276. External Links: Link Cited by: Appendix F, §5.3. L. Parcalabescu and A. Frank (2024) On measuring faithfulness or self-consistency of natural language explanations. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 6048–6089. External Links: Link, Document Cited by: §2. P. Pitre, N. Ramakrishnan, and X. Wang (2025) CONSENSAGENT: towards efficient and effective consensus in multi-agent LLM interactions through sycophancy mitigation. In Findings of the Association for Computational Linguistics: ACL 2025, W. Che, J. Nabende, E. Shutova, and M. T. Pilehvar (Eds.), Vienna, Austria, p. 22112–22133. External Links: Link, Document Cited by: §2, §4.1. Qwen, A. Y. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu (2025) Qwen2.5 technical report. arXiv preprint arXiv:2412.15115. External Links: Link Cited by: §5.2. S. Schmidgall, R. Ziaei, C. Harris, J. W. Kim, E. Reis, J. Jopling, and M. Moor (2026) AgentClinic: a multimodal agent benchmark to evaluate AI in simulated clinical environments. npj Digital Medicine. External Links: Document Cited by: §2. X. Sun, Q. Hong, M. Zhang, Y. Li, T. Chen, Z. Huang, G. Liang, W. Tang, S. Xu, X. Ni, J. Pang, P. Wan, and E. Long (2026) Model confrontation and collaboration: A debate intelligence framework for enhancing medical reasoning in large language models. Cell Reports Medicine 7 (1). External Links: Document Cited by: §1, §2. X. Tang, A. Zou, Z. Zhang, Y. Zhao, X. Zhang, A. Cohan, and M. Gerstein (2024) MedAgents: large language models as collaborators for zero-shot medical reasoning. In Findings of the Association for Computational Linguistics: ACL 2024, Bangkok, Thailand. External Links: Link Cited by: §1, §2. D. Tran and D. Kiela (2026) Single-agent LLMs outperform multi-agent systems on multi-hop reasoning under equal thinking token budgets. arXiv preprint arXiv:2604.02460. External Links: Link Cited by: §2. Q. Wang, Z. Wang, Y. Su, H. Tong, and Y. Song (2024) Rethinking the bounds of LLM reasoning: are multi-agent discussions the key?. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), L. Ku, A. Martins, and V. Srikumar (Eds.), Bangkok, Thailand, p. 6106–6131. External Links: Link, Document Cited by: §2. B. Yao, C. Shang, W. Du, J. He, R. Lian, Y. Zhang, H. Su, S. Swamy, and Y. Qi (2025) Peacemaker or troublemaker: how sycophancy shapes multi-agent debate. External Links: 2509.23055, Link Cited by: Appendix O, §2, §4.1. D. Zhang, J. Li, Z. Zeng, and F. Wang (2025a) Jasper and stella: distillation of sota embedding models. External Links: 2412.19048, Link Cited by: Appendix E, §5.3. H. Zhang, Z. Cui, J. Chen, X. Wang, Q. Zhang, Z. Wang, D. Wu, and S. Hu (2025b) Stop overvaluing multi-agent debate – we must rethink evaluation and embrace model heterogeneity. External Links: 2502.08788, Link Cited by: §2. S. Zhou, W. Xie, J. Li, Z. Zhan, M. Song, H. Yang, C. Espinoza, L. Welton, X. Mai, Y. Jin, Z. Xu, Y. Chung, Y. Xing, M. Tsai, E. Schaffer, Y. Shi, N. Liu, Z. Liu, and R. Zhang (2025) Automating expert-level medical reasoning evaluation of large language models. npj Digital Medicine. External Links: Document Cited by: §5.2. Y. Zhu, Z. He, H. Hu, X. Zheng, X. Zhang, Z. Wang, J. Gao, L. Ma, and L. Yu (2025) MedAgentBoard: benchmarking multi-agent collaboration with conventional methods for diverse medical tasks. arXiv preprint arXiv:2505.12371. Note: Accepted at NeurIPS 2025, Datasets & Benchmarks Track External Links: Link Cited by: §1, §2. Appendix A Sample-Size Stability A practical concern for any automated metric is whether scores are stable as the sample grows. Table 3 compares earlier pilot cara-hyb estimates (D1 N=200N=200, D2 pooled N=100N=100) against the full sample (D1 N=500N=500, D2 N=499N=499). All ten cell–round–dataset combinations show deltas within ±0.005± 0.005 (maximum: 0.0050.005 for M3 r0 on D2), with five of ten cells unchanged at the third decimal place. This stability has two implications. First, the metric is well-conditioned at the operational sample sizes used in the paper—further scaling would not change the qualitative conclusions. Second, the pilot estimates were already reliable point estimates rather than artifacts of small-sample noise; the decision to scale up at N=200N=200 accurately anticipated the final N=500N=500 effect sizes. D1 D2 System Round Pilot Δ Pilot Δ M4 r0 0.895 +0.000+0.000 0.894 +0.001+0.001 M3 r0 0.896 −0.001-0.001 0.891 +0.005+0.005 M3 r1 0.891 +0.000+0.000 0.878 +0.003+0.003 M3-GDP r0 0.914 −0.002-0.002 0.913 −0.001-0.001 M3-GDP r1 0.951 +0.001+0.001 0.951 +0.001+0.001 Table 3: Pilot → full-sample cara-hyb stability. All deltas are within ±0.005± 0.005 on both datasets, confirming that cara-hyb is not sensitive to sample fluctuations at this scale. Appendix B Survivorship-Bias Sensitivity M3 r1 excludes the 76/500 (15.2%) of D1 and 71/499 (14.2%) of D2 questions where no majority answer emerged after debate. These exclusions are non-random: harder questions are more likely to produce post-debate disagreement, so the surviving sample is biased toward inherently easier cases with higher baseline agreement. The exclusion mechanism differs by system—M3-GDP retains 99.6%99.6\% of questions on both datasets—so direct comparison of M3 r1 against M3-GDP r1 (or against M4 r0, where M4 retains all D1 questions) involves comparing differently-selected populations. We assess sensitivity through imputation: for the undefined questions, we impute plausible cara-hyb and CR values and recompute the cell mean. Three scenarios are reported in Table 4: observed (current measurement, excluding undefined), neutral (impute the M4 baseline value of that metric), and worst-case (impute cara-hyb =0.50=0.50 and CR=1.0=1.0—both far beyond any value observed in any cell on either dataset). Two conclusions are robust across both datasets. First, the cara-hyb finding is robust: under all imputation scenarios—including the worst case, which assigns cara-hyb =0.50=0.50 (lower than any observed sample), M3 r1’s adjusted score remains at or below M4 r0. The consistency-illusion conclusion (M3 debate does not improve, and on D2 significantly degrades, alignment) therefore does not depend on the survivorship-biased sample. Second, the CR finding is fragile in the opposite direction: under worst-case imputation, M3 r1’s adjusted CR exceeds M4 r0’s on both datasets (0.182>0.1080.182>0.108 on D1; 0.188>0.1610.188>0.161 on D2), so M3 r1’s apparently low observed CR is at least partly an artifact of survivorship. Practically, this means CR should be interpreted alongside the undefined rate, not in isolation, and the consistency-illusion claim rests primarily on the joint movement of CR and SIM in the main cara-hyb result (Section 6.1) rather than on CR alone. D1 (N=500N=500) D2 (N=499N=499) Scenario HYB CR HYB CR M4 baseline 0.895 0.108 0.895 0.161 Observed (M3 r1) 0.892 0.035 0.881 0.054 Neutral impute 0.892 0.046 0.883 0.069 Worst-case 0.832 0.182 0.827 0.188 Table 4: Sensitivity of M3 r1 to imputation of undefined questions. Neutral imputes the M4 baseline value; Worst-case imputes HYB=0.50=0.50 and CR=1.0=1.0. The HYB conclusion (M3 r1 ≤ M4) is robust on both datasets; the CR direction reverses under worst-case on both datasets. Appendix C cara-GT Per-Cell Results D2 includes expert-annotated reasoning steps (Scoring Points) of 22–66 steps per question (mean 3.13.1), authored by clinical experts. cara-GT measures alignment between each agent’s reasoning and the expert reference, computed via the same NLI++embedding pipeline used for cara-hyb but with one agent replaced by the expert reference. Where cara-hyb asks whether agents reason consistently with each other, cara-GT asks whether they reason consistently with experts. cara-GT is undefined when an agent produces zero extractable steps; this affects 3535/2,4952,495 pairwise records (1.4%1.4\%), much cleaner than cara-hyb’s 156156/2,4952,495 error rate (which additionally excludes agents who disagree on the final answer). Table 5 reports per-cell cara-GT-HYB scores and expert-coverage rates on D2 (N=499N=499). gdp improves expert alignment; standard debate degrades it. The pairwise comparisons mirror the inter-agent cara-hyb pattern. GDPr1_r1 vs. M3r1_r1: ΔGT-HYB=+0.011 -HYB=+0.011, d=+0.32d=+0.32, 95%CI=[+0.20,+0.45]95\%\,CI=[+0.20,+0.45] (excludes zero). M3r1_r1 vs. M4r0_r0: Δ=−0.007 =-0.007, d=−0.21d=-0.21, CI=[−0.34,−0.08]CI=[-0.34,-0.08] (excludes zero). The dual evidence—GDP improves alignment with respect to experts, and standard debate degrades it—supports the interpretation that the inter-agent illusion has a substantive directional component: under standard debate, agents drift not merely apart from each other but away from the expert reasoning standard. Expert coverage tells a sharper story. Coverage is the fraction of expert scoring points matched by at least one agent step at the contradiction-filtered alignment threshold. Standard debate loses 8.68.6 percentage points of expert coverage relative to independent voting (M4: 88.8%88.8\% → M3 r1: 80.2%80.2\%, CI=[−11.2,−6.1]CI=[-11.2,-6.1]): roughly one in eleven expert steps is dropped during debate. gdp debate recovers 5.45.4p of that loss (M3 r1: 80.2%80.2\% → M3-GDP r1: 85.6%85.6\%, CI=[+2.7,+8.1]CI=[+2.7,+8.1]). Standard debate drops expert scoring points; gdp debate keeps them. This coverage gap is independent of the inter-agent cara-hyb measurement and provides a second, expert-grounded line of evidence for the same conclusion: gdp produces reasoning that is more consistent with both other agents and external standards. System Round GT-HYB Coverage N M3-GDP r1 0.813 85.6% 499 M3 r0 0.810 87.4% 499 M4 r0 0.809 88.8% 499 M3-GDP r0 0.807 88.2% 499 M3 r1 0.802 80.2% 464 Table 5: cara-GT on D2 (N=499N=499): alignment between agent reasoning and expert Scoring Points. Coverage is the fraction of expert scoring points that match at least one agent step. Appendix D Accuracy and McNemar Tests cara measures reasoning alignment, not answer correctness, but we separately report majority-vote accuracy for context. Accuracy is computed by extracting the majority answer per agent at each round, then comparing to the gold answer with continuity-corrected McNemar tests on paired binary outcomes (correct/incorrect per question) and paired percentile bootstrap CIs (B=10,000B=10,000, seed 4242) on the accuracy delta. Table 6 reports majority-vote accuracy for all dataset–backbone combinations. Findings. Standard debate (M3 r1) does not significantly change accuracy on either Qwen dataset (D1: Δ=+0.4 =+0.4p, p=0.86p=0.86; D2: Δ=0.0 =0.0p, p=0.90p=0.90), consistent with the martingale characterization of debate dynamics: gains from sampling and voting are already realized by the independent baseline. D2 baseline accuracy (∼41% 41\%) is much lower than D1 (∼84% 84\%) because D2 questions have a mean of ∼7.5 7.5 answer options (random baseline ∼13% 13\%), making ∼41% 41\% approximately 3×3× random on this benchmark—i.e., the model is performing non-trivially despite the lower headline number. Llama D2 accuracy (N=100N=100, 49.0%49.0\%) and Qwen D2 (N=499N=499, 41.5%41.5\%) are computed on different sample sizes of the same dataset, so we do not draw cross-backbone accuracy conclusions from this comparison. Dataset Backbone M4 M3 r1 GDP r1 ΔGDP _GDP D1 MedQA (N=500N=500) Qwen 83.8% 84.2% 81.4% −2.4-2.4 ppns^ns D2 MedThink (N=499N=499) Qwen 41.5% 41.5% 39.1% −2.4-2.4 ppns^ns D2 MedThink (N=100N=100) Llama 49.0% 47.0% 43.0% −6.0-6.0 ppns^ns Table 6: Majority-vote accuracy. ΔGDP= _GDP= GDP−r1_r1-M4r0_r0. “ns” = not significant at α=0.05α=0.05. On D1, GDP vs. M3: McNemar p=0.061p=0.061; GDP vs. M4: p=0.097p=0.097; M3 vs. M4: p=0.86p=0.86. On D2 (Qwen), all pairwise McNemar tests yield p>0.15p>0.15. On D2 (Llama, N=100N=100), all pairwise tests yield p>0.28p>0.28. Appendix E cara-hyb Implementation NLI component. DeBERTa-v3-large fine-tuned on MNLI, FEVER, ANLI, LingNLI, and WANLI (Laurer et al., 2024) (304M parameters). We threshold the contradiction probability at τ=0.7τ=0.7 for the hybrid filter. Embedding component. Stella-EN-1.5B-v5 (Zhang et al., 2025a) (1.5B parameters), L2-normalized 10241024-dimensional vectors; cosine similarity reduces to dot product. Step decomposition. Deterministic regex pipeline: primary extraction of numbered lists (e.g., 1., 2)) and bullet points (-, *); sentence-level splitting as fallback for unstructured paragraph responses. The final answer line (e.g., ANSWER: X) is stripped before decomposition; resulting steps shorter than 20 characters are filtered. Runtime. On a single NVIDIA RTX 4090 (24 GB), cara processes ∼9,000 9,000 traces per hour. The full D1 pipeline (N=500N=500, 2,5002,500 pairwise records) completes in ∼8 8 minutes; D2 (N=499N=499, 2,4952,495 records) in ∼8 8 minutes. Appendix F cara-llm Oracle Validation Design. To validate cara-hyb against human-like judgment, we conduct a three-dimensional validation using GPT-4o (OpenAI, 2024) as an alignment judge. We evaluate 8080 agent pairs: 5050 agreement-set pairs (both agents selected the majority answer) and 3030 disagreement pairs (agents selected different answers). The inclusion of disagreement pairs is critical: agreement-set pairs produce a ceiling effect in GPT-4o scores (skewed toward 44–55), and disagreement pairs break this ceiling by introducing genuinely misaligned reasoning (scores 22–33), expanding the effective scoring range. For each pair, GPT-4o assigns a 11–55 alignment score at temperature 0.00.0, repeated three times per pair. We report cara-llm as this rating normalized to [0,1][0,1] via (score−1)/4(score-1)/4; the per-system head-to-head below is given on the raw 11–55 scale for interpretability. Three-dimensional validation results. 1. Moderate pair-level correlation (Spearman ρ=0.55ρ=0.55, n=80n=80): cara-hyb and GPT-4o capture overlapping but non-identical dimensions of alignment. The correlation is moderate rather than high because the two metrics operationalize alignment differently: cara-hyb measures step-level semantic overlap via embeddings, while GPT-4o evaluates holistic logical coherence. 2. Strict stratum monotonicity: when pairs are binned by cara-hyb terciles, mean cara-llm scores are 0.670.67 (low), 0.770.77 (mid), 0.820.82 (high). cara-hyb reliably distinguishes alignment levels even where it does not predict individual GPT-4o scores precisely. 3. Perfect inter-run reliability: all 8080 pairs receive identical GPT-4o scores across three independent runs (inter-run SD =0.00=0.00), confirming deterministic evaluation at temperature 0. Per-system head-to-head. Within the 5050 agreement-set pairs, M3-GDP r1 receives a mean GPT-4o score of 4.864.86 (n=28n=28, SD 0.350.35), vs. 4.504.50 for M3 r1 (n=22n=22, SD 0.580.58); Cohen’s d=+0.75d=+0.75, 95%95\% bootstrap CI on Δ [+0.08,+0.65][+0.08,+0.65] excludes zero (B=10,000B=10,000, seed 4242). This direct head-to-head corroborates the inter-agent cara-hyb ranking and complements the tercile monotonicity reported above. Disagreement-pair divergence. On disagreement pairs (agents with different final answers), cara-hyb assigns scores of 0.830.83–0.950.95 because medical reasoning steps have high embedding similarity even when conclusions differ, while GPT-4o assigns 22–33 out of 55 because it detects the logical incoherence of conflicting conclusions. This divergence is by design: cara-hyb measures reasoning-step alignment within the agreement set, so its high scores on disagreement pairs reflect genuine semantic overlap in medical reasoning vocabulary rather than a metric failure. The complementary roles ground the joint use of cara-hyb (scalable, surface alignment) and cara-llm (oracle, logical coherence) we recommend in deployment. Appendix G Label-Stripped cara-hyb: Structural Confound Check Motivation. cara-hyb computes step-pair embedding cosine and NLI contradictions over reasoning steps extracted from agent responses. gdp requires each step to carry literal Claim/Ground/Stance label tokens, so a reasonable concern is that these template tokens themselves inflate the embedding component of cara-hyb, making part of the gdp-vs-M3 gap a metric artifact of template self-similarity rather than substantive reasoning convergence. The format/debate decomposition (Section 7) and the cara-GT result (Section 6.4) already weigh against this concern; this appendix reports the direct check. Procedure. For each gdp trace (D1 Qwen, D2 Qwen, D2 Llama; both r0r_0 and r1r_1) we remove the literal label tokens CLAIM:, GROUND:, and STANCE: together with their Markdown bold variants (e.g., **CLAIM**), preserving (i) the numbered-step markers (1., 2., …) that M3 traces also use, (i) the Agree/Disagree/Extend content tokens that constitute the agent’s actual stance decision, and (i) all free-text claim, ground, and stance content. We then re-run the unchanged cara-hyb pipeline (DeBERTa-v3 NLI + Stella embeddings, τ=0.7τ=0.7, seed 4242) on the stripped traces; M3 and M4 cells are unchanged. Paired Cohen’s dzd_z is computed at the question level for gdp r1 vs. M3 r1, using the same script for both the original and the stripped condition. Results. Table 7 reports the paired comparison across all three dataset-backbone conditions. Removing the literal label tokens attenuates Cohen’s dzd_z by only 66–10%10\%. All three 95%95\% bootstrap CI lower bounds remain above +0.76+0.76, and mean cara-hyb itself drops by only ∼0.005 0.005 across configurations. Approximately 90%90\% of the alignment effect survives label removal, indicating that the literal template tokens are not the principal source of gdp’s gain on cara-hyb. Config Original dzd_z Stripped dzd_z Stripped 95% CI D1 MedQA (Qwen, n=422n=422) +0.97+0.97 +0.87+0.87 [+0.76,+0.99][+0.76,+0.99] D2 MedThink (Qwen, n=427n=427) +1.16+1.16 +1.09+1.09 [+0.98,+1.21][+0.98,+1.21] D2 MedThink (Llama, n=98n=98) +1.39+1.39 +1.29+1.29 [+0.96,+1.74][+0.96,+1.74] Table 7: Label-stripped cara-hyb confound check. Paired Cohen’s dzd_z between gdp r1 and M3 r1 on the agreement set, before and after removing literal Claim/Ground/Stance tokens from gdp traces. The effect attenuates by 66–10%10\% but remains Tier B–A in all three conditions, with all CI lower bounds above +0.76+0.76. Interpretation and limitations. The label-stripping experiment isolates one specific structural-confound mechanism—that the literal template tokens themselves drive embedding similarity. The data do not support this mechanism: ∼90% 90\% of the effect persists after token removal. Robustness to the NLI model. cara-hyb is also robust to the choice of contradiction detector: substituting a cross-encoder NLI model (cross-encoder/nli-deberta-v3-large) for the default DeBERTa-MNLI on the D2 N=100N=100 set leaves the illusion direction unchanged (M3 r1 << M4 r0) and Cohen’s d nearly identical (−0.36-0.36 vs. −0.38-0.38). Appendix H Human-Expert Validation Setup. Three independent annotators with medical training each rated 5050 agent pairs drawn from the agreement set on the MedThink-Bench (D2) benchmark. The 50 pairs are stratified by cara-hyb tercile (low/mid/high, ∼17 17 pairs each). Each pair was rated on a 5-point alignment scale (5=5=near-identical reasoning, 4=4=substantially aligned, 3=3=partially aligned, 2=2=weakly aligned, 1=1=misaligned reasoning with the same final answer) and a 3-point confidence scale, following the annotation guidelines released with the code repository. Annotators saw only the question and the two agents’ reasoning steps; system identity (M3 vs. M3-GDP) and cara-hyb scores were withheld to prevent bias. Inter-annotator agreement. Pairwise Spearman correlation between annotators is consistently high (Table 8); 100%100\% of pair-level disagreements are bounded to ≤1≤ 1 scale point, confirming task reliability. Annotator pair Spearman ρ Exact agreement Within-1 A1 vs. A2 0.934 76% 100% A1 vs. A3 0.987 98% 100% A2 vs. A3 0.927 74% 100% Table 8: Inter-annotator agreement on the 50-pair human annotation. Correlation with CARA metrics. Aggregate human ratings (mean of three annotators) correlate positively with cara-hyb (Spearman ρ=0.263ρ=0.263), cara-sim (ρ=0.312ρ=0.312), and weakly with the contradiction rate (ρ=0.137ρ=0.137). Each individual annotator shows a similar correlation with cara-hyb (ρ=0.21ρ=0.21–0.260.26), confirming the signal is not driven by any single annotator. Tercile monotonicity. Aggregate ratings increase monotonically across cara-hyb strata: 4.004.00 (low, n=17n=17), 4.654.65 (mid, n=17n=17), 4.714.71 (high, n=16n=16). The same monotonic pattern is observed in the cara-llm GPT-4o oracle terciles (Figure 4), supporting a consistent ordering of cara-hyb-graded alignment across two independent external judges. The low-to-mid gap is larger than the mid-to-high gap, reflecting a real shift in alignment quality at the lower end combined with a ceiling effect at the upper end (more than half of high-stratum pairs receive the maximum rating of 55). Figure 4: Tercile monotonicity across two external validators on the same 50-pair sample stratified by cara-hyb. Both the GPT-4o oracle (cara-llm; Appendix F) and the mean of three medically-trained annotators show monotonic increase from low to mid to high cara-hyb strata. Range restriction. All 50 pairs have cara-hyb∈[0.76,0.98] cara-hyb∈[0.76,0.98] because the metric is conditioned on the agreement set (Section 3.1); the “low” stratum mean is 0.8640.864, far above the corpus-level low end. This narrow range bounds the achievable Spearman correlation against any external judge, paralleling the bounded correlation observed in the cara-llm GPT-4o oracle validation when restricted to the same agreement set. Annotation instructions. Annotators received the following written guidelines before the task (verbatim text is released with our anonymized code repository): Task overview. You will see pairs of AI-generated medical reasoning traces. Both agents chose the same final answer; judge how aligned their reasoning is—not whether the answer is correct. Rating 1 — Reasoning Alignment (1–5): • 5 (Near-identical): same claims, same facts, same logical sequence; minor wording differences only. • 4 (Substantially aligned): same core reasoning and key facts; one may add detail or reorganize. • 3 (Partially aligned): share key steps but diverge on intermediate justifications. • 2 (Weakly aligned): same answer through largely different reasoning; few overlapping steps. • 1 (Misaligned): same answer through fundamentally different or contradictory reasoning. Rating 2 — Confidence (1–3): 33 confident, 22 somewhat confident, 11 low confidence (reasoning too vague to judge). What NOT to judge. Do not evaluate answer correctness, writing quality, response format (structured vs. free text), or step count. Some responses use a Claim/Ground/Stance format; others use free text—judge reasoning content regardless of format. System identity and automated cara scores are withheld to prevent bias. Annotator background and procedure. The three annotators are medically-trained professionals aged 3030–3535 (two female resident physicians and one male postdoctoral fellow), recruited through direct personal contact by the authors and participating as volunteers without compensation. All three annotators rated all 5050 pairs, enabling the pairwise Spearman inter-annotator agreement reported above. Pairs were presented in randomized order, and system identity (M3 vs. M3-GDP), agent round, and cara-hyb/cara-sim scores were withheld throughout. The task does not constitute human-subjects research—annotators rated machine-generated outputs only—and required no institutional review. The verbatim annotation guidelines, blank annotation sheet, and three completed per-annotator scored sheets are released with our anonymized code repository. Appendix I cara Variants and Aggregation Formulas Variant table. The four cara alignment variants used in the paper are summarized below. Variant Range Definition cara-nli −1,0,+1\-1,0,+1\ NLI label: +1+1 entailment, 0 neutral, −1-1 contradiction. cara-sim [−1,+1][-1,+1] Cosine similarity: cos(emb(rik),emb(rjl)) (emb(r_i^k),emb(r_j^l)). cara-hyb [−1,+1][-1,+1] Hybrid: if P(contradiction)>τP(contradiction)>τ then −1-1, else cosine similarity. cara-llm [0,1][0,1] GPT-4o 11–55 alignment rating, normalized to [0,1][0,1] via (score−1)/4(score-1)/4. Table 9: cara alignment variants. cara-hyb is the default metric used throughout; cara-llm serves as a validation oracle. Rescaling. For variants whose range includes negative values (cara-nli, cara-hyb), we define a unit-interval rescaling: CARA[0,1]=CARA+12.CARA_[0,1]= CARA+12. (7) Contradiction Rate. The Contradiction Rate diagnostic is defined in Eq. (6) of Section 3.3. The denominator ∑i∈(q)(|(q)|−1)Ki _i (q)(|S(q)|-1)K_i equals ∑i<j(Ki+Kj) _i<j(K_i+K_j), the total number of best-match operations across all directed agent pairs in the agreement set. CR captures the rate at which best-aligned step pairs across agreeing agents are nonetheless contradictory: a direct indicator of the consistency illusion. Appendix J System Configurations Configuration overview. We evaluate three system configurations under identical compute budgets. Each system uses three agents and aggregates their final answers by majority vote. “Round 0” refers to the independent step in which each agent answers without seeing the others’ responses; “Round 1” refers to the post-debate step in which each agent revises its answer after seeing all other agents’ Round 0 outputs. M4 has only Round 0 (the control condition with no debate); M3 and M3-GDP have both rounds. Majority voting is applied to Round 1 answers for M3 and M3-GDP, and to Round 0 answers for M4. ID System Agents Rounds GDP M3 Du et al. (2024) Debate 3 1 no M4 Independent Vote 3 0 no M3-GDP Debate ++ gdp 3 1 yes Table 10: System configurations. All systems use majority voting for answer aggregation. “Rounds” counts debate rounds; Round 0 (independent) is always present. M4 serves as a no-communication control. Backbone serving environment. Primary trials use Qwen 2.5 72B Instruct AWQ (4-bit quantized) served via vLLM (Kwon et al., 2023) on 2×2×NVIDIA A40 GPUs (48 GB each). The cross-backbone replication uses Llama 3.3 70B Instruct AWQ on a separate pair of A40 GPUs with the same vLLM serving configuration. All experiments use the same decoding hyperparameters: temperature 0.70.7, top-p 0.90.9, and a maximum of 2,0482,048 output tokens per agent response. Standard M3 and M4 prompts. M3 and M4 share a 2-message Round 0 prompt. The system message is verbatim: You are a medical specialist. Answer medical questions with step-by-step reasoning. The user message contains the question text with inline answer options, followed by the suffix: Think step by step and provide your answer in the format: ANSWER: X M3 Round 1 extends this to a 4-message asymmetric prompt: the original system and user messages, the agent’s own Round 0 response as the assistant turn, and a final user turn presenting the other agents’ Round 0 responses with the instruction “Using these as additional information, provide your answer in the format: ANSWER: X.” Each agent sees only the other two agents’ responses in the debate user turn (never its own), preventing self-conditioning artifacts. M3-GDP prompts. M3-GDP preserves the 2-message (Round 0) and 4-message (Round 1) prompt structure of M3 but replaces the system message with format requirements that enforce structured reasoning. The Round 0 system message requires Claim+Ground for each reasoning step, where Ground must name a specific medical fact, mechanism, or guideline supporting the claim. The Round 1 system message adds Stance (Agree/Disagree/Extend) with respect to each other agent’s claims, with the rule that any Disagree must be accompanied by a counter-Ground. Round 1 also includes an explicit anti-sycophancy clause: Do NOT change your answer simply because other agents disagree. Only change if you identify a specific factual error in your own reasoning or receive a more compelling medical ground from another agent. The full GDP Round 0 and Round 1 system messages, along with the debate user template, are released with the code repository. Appendix K GDP Worked Example We compare a standard debate (M3) trace and a gdp (M3-GDP) trace on the same clinical question. Both systems converge on the same correct answer; the contrast is in how the agents arrive there. Clinical question. A patient presents with fasting plasma glucose of 140140 mg/dL on two consecutive measurements. What is the most appropriate first-line management? A. Insulin B. Metformin C. Lifestyle modification only D. Sulfonylurea Gold answer: B. M3 (Standard Debate) — Round 1 Agent A The elevated fasting glucose is consistent with a diagnosis of diabetes. Considering the clinical picture, metformin would be appropriate. I agree with Agent B’s reasoning. Answer: B. Agent B The patient’s clinical picture supports metformin therapy. Given the lab values and the analysis Agent A provided, I support this approach. Answer: B. Voted solution: B ✓ | Both agents converge with no medical citations; agreement is vague. M3-GDP — Round 1 Agent A CLAIM: The fasting glucose of 140140 mg/dL meets ADA diagnostic criteria for T2DM. GROUND: Fasting plasma glucose ≥126≥126 mg/dL on two occasions confirms diabetes (ADA Standards of Care). STANCE: Disagree with Agent B (insulin)—insulin is not first-line for uncomplicated T2DM without DKA. Answer: B. Agent B CLAIM: Metformin is the appropriate first-line agent for newly diagnosed T2DM. GROUND: ADA Standards of Care recommend metformin first-line for adults with T2DM without contraindications, given its efficacy in reducing hepatic glucose production and favorable safety profile. STANCE: Agree with Agent A’s diagnostic claim; reaffirm metformin over insulin given absence of DKA or severe hyperglycemia. Answer: B. Voted solution: B ✓ | Both agents cite ADA criteria and take explicit stances on each other’s claims; agreement is grounded. The contrast is structural: M3 agents agree on the answer without committing to specific medical evidence or explicitly engaging with each other’s claims, whereas M3-GDP agents reach the same answer through named criteria (ADA guidelines) and explicit Stance declarations. This pattern recurs at the dataset scale, producing the cara-hyb and cara-GT gaps reported in Section 6.1 and Appendix C. Appendix L Failure Mode Taxonomy and Case Study Failure-mode codebook. The 50 lowest-cara-hyb correct-answer cases per system are classified using deterministic rules over CARA sub-metrics (CR>0.3>0.3, step-count ratio>2>2, zero-step agents, SIM<0.80<0.80). Code Failure mode FM1 Complementary Reasoning: different but non-contradictory reasoning paths. Low SIM, low CR. FM2 Granularity Mismatch: step-count ratio >2>2. Depth difference, not substantive disagreement. FM3 Contradictory Premises: agents cite contradictory facts but converge on the correct answer. FM4 Sycophantic Convergence: agent adopts the majority answer without substantive reasoning (zero steps). FM5 Terminology Divergence: different medical terms for the same concept. FM6 Partial Overlap: residual baseline misalignment without extreme signals. Table 11: Six failure modes identified by classifying the 50 lowest-cara-hyb correct-answer cases per system. Case study: MQA_0447 (hypovolemic shock). We illustrate FM3 (contradictory premises) with a D1 case in which the correct answer is reached through partially incorrect reasoning. The question describes a 5252-year-old man with variceal bleeding and hypovolemic shock. Two of three gdp agents in the agreement set initially state that systemic vascular resistance (SVR) decreases in hypovolemic shock—a clinically incorrect claim, since SVR increases as a compensatory response to maintain perfusion. Both agents then self-correct within the same response and select the correct treatment (volume resuscitation followed by endoscopic intervention). This case exemplifies the consistency illusion at the individual-fact level: the answer is correct, but the underlying reasoning contains explicit factual contradictions that the metric correctly surfaces. gdp’s structured format makes these contradictions detectable; under standard debate, the same misstatement would more likely be smoothed away in vague free text. Appendix M Undefined Rate The undefined rate—questions where no majority answer emerges—is an independent behavioral indicator that does not depend on agreement-set selection. Table 12 reports per-system rates on both datasets. Standard debate destabilizes consensus. On D1, M4 (independent voting) achieves 100%100\% consensus because agents sample independently from the same distribution: the probability of three-way disagreement among i.i.d. samples is low. Standard debate (M3 r1) drives the undefined rate to 15.2%15.2\% on D1—a 38×38× increase over the pre-debate round (M3 r0, 0.4%0.4\%). On D2, where the harder question distribution gives even M4 a 5.8%5.8\% undefined rate, M3 r1 still pushes this up to 14.2%14.2\% (2.5×2.5× relative). gdp debate maintains consensus at near-M4 levels (0.4%0.4\% on both datasets) while substantially improving reasoning alignment. Reasoning collapse. A related phenomenon is reasoning collapse: agents that produce zero extractable reasoning steps after debate (a behavioral signature of sycophantic convergence). On D2 N=100, M3 r1 yields zero-step agents in 5.0%5.0\% of agreement-set memberships; M3-GDP r1 in only 0.3%0.3\%. gdp’s Claim+Ground requirement prevents this collapse because an agent cannot adopt an answer without producing a named ground supporting it. System / Round D1 (N=500N=500) D2 (N=499N=499) M3 r1 76 (15.2%) 71 (14.2%) M3-GDP r0 3 (0.6%) 26 (5.2%) M3 r0 2 (0.4%) 28 (5.6%) M3-GDP r1 2 (0.4%) 2 (0.4%) M4 r0 0 (0.0%) 29 (5.8%) Table 12: Undefined rate per system–round–dataset. M3 r1 destabilizes consensus on ∼14 14–15%15\% of questions on both datasets; M3-GDP r1 retains essentially all of them. Appendix N Consistency-Illusion Visualization Figure 5 visualizes the (CR, SIM) trajectory of each system on D2 (N=499N=499, Qwen 2.5 72B), making the dual movement of standard debate (CR↓\, + SIM↓\, ) and the repair pattern under gdp (SIM↑\, ) directly observable. The two arrows correspond to the r0→\,→\,r1 transitions under M3 (red) and M3-GDP (green); M4 is shown as a single point for reference. The figure complements the conceptual schematic of Figure 1 with the actual measured coordinates that anchor every claim made in Section 6.1. Figure 5: Empirical (CR, SIM) trajectory on D2 (N=499N=499, Qwen 2.5 72B). M3 (red arrow) moves r0→\,→\,r1 both CR↓\, and SIM↓\, (consistency illusion). gdp (green arrow) moves r0→\,→\,r1 SIM↑\, while keeping CR informative. Appendix O Mechanism Discussion The condensed mechanism account in RQ1 (Section 7) merges two effects under one heading. Here we elaborate each with the original four-component analysis. Contradiction smoothing without reasoning convergence (M3). During debate, agents see each other’s responses and revise their own. When agents disagree on factual claims, revision pressure causes them to soften or remove the conflicting statements, which reduces CR. However, this removal is subtraction without alignment: agents delete contradictory steps without replacing them with shared reasoning anchors that other agents have also committed to. The result is fewer contradictions but also less overlapping reasoning content, so SIM decreases. This pattern is stronger on D2 because the more open answer space (3–10 options, mean ∼7.5 7.5) gives agents more room to diverge in reasoning while still landing on the same final answer. Sycophantic convergence (M3). Agents may adopt each other’s answers without adopting the underlying reasoning—a form of answer-level sycophancy (Yao et al., 2025). This produces answer convergence (the system reports consensus) while reasoning chains remain divergent or become vaguer. The high M3 r1 undefined rate—15.2%15.2\% on D1 and 14.2%14.2\% on D2—shows that for a substantial minority of questions, agents’ positions are sufficiently incoherent that no majority emerges even after revision. The zero-step phenomenon (Appendix M) is the limiting case: an agent produces an answer with no extractable reasoning, indicating that revision has degenerated into pure imitation. Format effect (gdp, pre-debate). The Claim+Ground structure forces agents to decompose reasoning into named, verifiable steps. This produces more comparable reasoning units across agents even before any debate interaction: on D1, GDPr0_r0 vs. M4r0_r0 alone yields a format-effect d=+0.62d=+0.62 with SIM increase of +0.034+0.034. This isolates a structural cause for misalignment in standard debate: a substantial portion of apparent misalignment originates from underspecified output format rather than from genuine reasoning disagreement. Debate interaction effect (gdp, r0→ 1). The Stance mechanism in Round 1 requires each agent to explicitly address every claim made by the other agents. When an agent must state “I Agree with Agent B’s claim that renal function is preserved, because…” or “I Disagree because…,” the resulting response necessarily shares reasoning vocabulary and logical structure with the other agents. This drives SIM upward (+0.077+0.077 on D1, +0.078+0.078 on D2 from r0 to r1 under gdp) at a magnitude 2.4×2.4× larger than the format effect alone (d=+1.51d=+1.51 vs. d=+0.62d=+0.62), indicating that Stance-mediated engagement is the primary mechanism. The format effect is a necessary precondition, however: without structured claims to reference, the Stance mechanism has no well-defined targets. Appendix P Cross-Dataset and Cross-Backbone Discussion Comparison Llama D2 Qwen D2 Dir. (N=100N=100) (N=499N=499) GDPr1_r1 vs. M3r1_r1 +1.99+1.99 +1.62+1.62 ✓ M3r1_r1 vs. M4r0_r0 (illusion) −1.32-1.32 −0.30-0.30 ✓ M3 debate Δ HYB (r0→ 1) −0.038-0.038 −0.014-0.014 ✓ GDP debate Δ HYB (r0→ 1) +0.014+0.014 +0.040+0.040 ✓ Format effect Δ HYB +0.006+0.006 +0.017+0.017 ✓ Table 13: Cross-backbone validation on D2 (referenced from Section 6.3). First two rows give Cohen’s d; remaining rows give Δ cara-hyb. All five comparisons are direction-consistent (“Dir.” ✓). gdp effect generalization. The gdp effect replicates across all tested conditions with Tier A effect sizes: D1 Qwen (N=500N=500): d=+1.43d=+1.43, CI=[+1.29,+1.60]CI=[+1.29,+1.60]; D2 Qwen (N=499N=499): d=+1.62d=+1.62, CI=[+1.48,+1.78]CI=[+1.48,+1.78]; D2 Llama (N=100N=100): d=+1.99d=+1.99, CI=[+1.65,+2.44]CI=[+1.65,+2.44]. The rank ordering of systems is identical across all three conditions: M3-GDPr1≻M3-GDPr0≻M4r0≈M3r0≻M3r1M3-GDP_r_1\! \!M3-GDP_r_0\! \!M4_r_0\!≈\!M3_r_0\! \!M3_r_1 This consistency across two structurally different datasets (fixed 4-option MCQ vs. variable 3–10-option clinical reasoning) and two independent backbone models supports the generalizability of both the consistency-illusion finding and the gdp repair effect. Consistency-illusion magnitude variation. The consistency illusion magnitude varies more across conditions: d=−0.08d=-0.08 (D1, Qwen), d=−0.30d=-0.30 (D2, Qwen), d=−1.32d=-1.32 (D2, Llama). Two factors plausibly drive this variation. First, D1’s fixed 4-option format constrains the space in which reasoning can diverge while answers remain matched, weakening the measurable illusion compared to D2’s mean ∼7.5 7.5-option open answer space. Second, M3 r1 undefined rates differ markedly across conditions (15.2%,14.2%,2.0%15.2\%,14.2\%,2.0\%), and a lower undefined rate preserves more of the worst-alignment cases in the measured sample, strengthening the apparent illusion magnitude. The third dataset (Llama D2) has the lowest undefined rate and the strongest illusion—an interpretive asymmetry we surface honestly via the sensitivity analysis in Appendix B rather than treating the Llama magnitude as a backbone-effect comparison. Implications for deployment. Where the underlying multi-agent question allows substantial reasoning-space variation (open-ended diagnosis, differential reasoning, treatment selection with multiple acceptable plans), the consistency illusion is likely to be larger than on constrained MCQ benchmarks, and the gdp benefit correspondingly more valuable. The D1 vs. D2 magnitude difference is therefore not a defect of the metric or the intervention but a direct reflection of how much room the task gives reasoning to diverge.