Paper deep dive
Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation
Amruta Parulekar, Jinu Lee, Dilek Hakkani-Tür, Hari Sundaram
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 8/1/2026, 2:12:47 AM
Summary
The paper introduces 'Reasoning Consensus,' a framework that ensembles the reasoning structures of multiple Large Language Models (LLMs) by extracting and merging Directed Acyclic Graphs (DAGs) from their chain-of-thought traces. By weighting reasoning steps based on cross-trace attestation (consensus), the method produces an inspectable 'consensus reasoning' graph that outperforms majority-vote baselines and self-consistency on six benchmarks, including statutory interpretation, science, and logic tasks.
Entities (9)
Relation Signals (7)
Reasoning Consensus → uses → Directed Acyclic Graph
confidence 95% · We propose a framework that ensembles the reasoning structure... by weighted merging of Directed Acyclic Graphs (DAGs) extracted from reasoning chains.
Reasoning Consensus → evaluatedon → MuSR-MM
confidence 90% · with a maximum accuracy gain of 3.1% on MuSR-MM
Reasoning Consensus → evaluatedon → SARA
confidence 90% · Across six benchmarks spanning statutory interpretation (SARA)
Reasoning Consensus → evaluatedon → GPQA-Diamond
confidence 90% · graduate-level science (GPQA Diamond)
Reasoning Consensus → evaluatedon → FOLIO
confidence 90% · first-order logical entailment (FOLIO)
Reasoning Consensus → extractsfrom → Chain-of-Thought
confidence 90% · DAGs extracted from reasoning chains... chain-of-thought
Reasoning Consensus → outperforms → Self-Consistency
confidence 90% · On a single model, the framework matches or exceeds self-consistency at the same trace budget
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, users cannot tell which steps are well-supported, which alternatives were seriously considered, or how the final conclusion compares to those the model discarded. We propose a framework that ensembles the reasoning structure, not just the answers, of multiple LLMs by weighted merging of Directed Acyclic Graphs (DAGs) extracted from reasoning chains. We weight each step by how many traces independently attest to it, to return "Consensus Reasoning". Across six benchmarks spanning statutory interpretation, graduate-level science, narrative multi-hop reasoning, and first-order logic, our ensemble outperforms a matched-budget majority-vote baseline, with a maximum accuracy gain of 3.1% on MuSR-MM (narrative multi-hop reasoning). On a single model, the framework matches or exceeds self-consistency at the same trace budget while additionally exposing an inspectable consensus reasoning graph. Ensemble weights correlate with LLM-judge rankings of reasoning quality at Spearman $\rho = 0.30$-$0.51$, and consensus subgraphs are preferred over alternatives leading to the majority-vote answer in 54.4-65.4% of head-to-head comparisons across five of six datasets. We observe that our framework can also be used to analyze diverse reasoning perspectives for a problem.
Tags
Links
- Source: https://arxiv.org/abs/2607.27783v1
- Canonical: https://arxiv.org/abs/2607.27783v1
Trouble viewing inline? Open PDF directly →
Full Text
71,382 characters extracted from source content.
Expand or collapse full text
Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation Amruta Parulekar, Jinu Lee, Dilek Hakkani-Tür, Hari Sundaram University of Illinois Urbana-Champaign amp20, jinulee2, dilek, hs1@illinois.edu Abstract Large Language Models (LLMs) explore problems through chain-of-thought, but this exploration is buried in unstructured prose. On high-stakes tasks, users cannot tell which steps are well-supported, which alternatives were seriously considered, or how the final conclusion compares to those the model discarded. We propose a framework that ensembles the reasoning structure, not just the answers, of multiple LLMs by weighted merging of Directed Acyclic Graphs (DAGs) extracted from reasoning chains. We weight each step by how many traces independently attest to it, to return “Consensus Reasoning”. Across six benchmarks spanning statutory interpretation, graduate-level science, narrative multi-hop reasoning, and first-order logic, our ensemble outperforms a matched-budget majority-vote baseline, with a maximum accuracy gain of 3.1% on MuSR-M (narrative multi-hop reasoning). On a single model, the framework matches or exceeds self-consistency at the same trace budget while additionally exposing an inspectable consensus reasoning graph. Ensemble weights correlate with LLM-judge rankings of reasoning quality at Spearman ρ=0.30ρ=0.30–0.510.51, and consensus subgraphs are preferred over alternatives leading to the majority-vote answer in 54.4–65.4% of head-to-head comparisons across five of six datasets. We observe that our framework can also be used to analyze diverse reasoning perspectives for a problem. 111Our code is available at this link Reasoning Consensus: Structural Ensembling of LLM Reasoning via Weighted DAG Aggregation Amruta Parulekar, Jinu Lee, Dilek Hakkani-Tür, Hari Sundaram University of Illinois Urbana-Champaign amp20, jinulee2, dilek, hs1@illinois.edu 1 Introduction We propose a framework that ensembles the reasoning structure of multiple LLMs to surface diverse, interpretable justifications for complex reasoning tasks. Complex domain-specific reasoning admits multiple legitimate lines of reasoning, and a thorough analysis must consider each Yang et al. (2026). These get lost in the unstructured chain-of-thought (Wei et al., 2023) that LLMs return. Users see a final answer and a free-form rationale, but cannot tell which steps and conclusions are supported or rejected, or which arguments rely on a single uncertain inference. Smaller models compound this by committing to a single solution path, missing alternative justifications that might be better Yang et al. (2025b). The result is brittle reasoning Zeng et al. (2026), that is especially costly in low-resource, high-stakes settings Dahl et al. (2024) where the model might provide the right answer but a wrong supporting reason, misleading users. Existing strategies for improving Chain-of-Thought (CoT) either constrain reasoning diversity or discard the structure that makes reasoning auditable. Self-consistency (Wang et al., 2023) performs majority voting of only the final answer, and does not consider reasoning trace structure. Tree/Graph of Thoughts (Yao et al., 2023; Besta et al., 2024) show more reasoning structure, but remain bounded by the single-model ceiling. Multi-model methods broaden this at the cost of structural auditability: some extend exploration across models but still commit to a single chain (Xu et al., 2025); others aggregate at the level of unstructured prose, making it hard to trace which model contributed what (Wang et al., 2024; Chen et al., 2024); and search-based ensembles require trained process reward models and ultimately return a single best chain (Park et al., 2024). Thus, the structure of disagreement and alternative support is lost. Our framework generalises Self-consistency Wang et al. (2023) from answers to reasoning structure, and from a single model to many. We sample CoT traces from multiple models and ensemble their underlying reasoning graphs. We extract a Directed Acyclic Graph (DAG) from each trace and merge these DAGs while preserving bundles of jointly-supporting premises, so that each merged node carries an “attestation count”. Steps that appear across many traces accumulate weight and steps attested by only one trace are down-weighted automatically, suppressing reasoning step errors. We return the most-supported, inspectable “consensus” subgraphs and conclusions with quantification of support received by each, and other possible answers as well as the amount of support for each. We evaluate our framework on six reasoning benchmarks spanning statutory interpretation (SARA), graduate-level science (GPQA Diamond), narrative multi-hop reasoning (MuSR-M, MuSR-OP, MuSR-TA), and first-order logical entailment (FOLIO). Our accuracy-weighted 4-model ensemble outperforms a matched-budget 20-trace majority-vote baseline on every dataset, with a maximum gain of 3.1% on MuSR-M. On a single model, we match or exceed self-consistency while additionally exposing an inspectable justification graph. Beyond accuracy, consensus subgraphs contain better explanations than single model traces, being preferred over alternatives leading to the majority-vote answer in 54.4–65.4% of head-to-head LLM-judge comparisons across five of six datasets. Our subgraph ranking correlates with the judge at Spearman ρ=0.30ρ=0.30–0.510.51. Qualitatively, we observe that subgraphs can contain differing problem-solving approaches, confirming that reasoning perspective diversity can be preserved through structural aggregation. Our contributions: (1) Structural ensembling of heterogeneous LLM reasoning. We introduce a framework to aggregate reasoning at the DAG level rather than as unstructured prose. This allows users in high-stakes domains to inspect and override individual steps rather than treating the LLM as a black box. (2) Attestation-weighted confidence without auxiliary training. We derive step-level confidence purely from cross-trace agreement across different models. This gives a training-free alternative to Process Reward Models, scaling to low-resource domains where step-level labels are unavailable. (3) Accuracy gains with diverse, inspectable alternatives. Our framework outperforms majority-vote baselines on all six benchmarks and matches or exceeds self-consistency on single-model trace pools, while simultaneously surfacing most-supported subgraphs and preserving competing conclusions that other traces explored. By retaining and not collapsing the disagreement structure, our framework naturally extends to open-ended settings like argument generation and creative writing, where the goal is not a single correct answer but a curated set of distinct, well-supported perspectives. 2 Related Work Recent “reasoning-heavy” models like o1 (OpenAI et al., 2024) and DeepSeek-R1 (Guo et al., 2025) have popularised long-form Chain-of-Thought (CoT) traces (Wei et al., 2023). However, linear CoT collapses logical dependencies more naturally modeled as graphs. To enrich the topology of a single model’s reasoning, Tree-of-Thoughts (Yao et al., 2023) branched the chain to enable backtracking, and Graph-of-Thoughts (Besta et al., 2024) generalised this to thought graphs with aggregation and feedback. While these approaches show more reasoning structure, their diversity remains bounded by what one isolated model can produce. Subsequent methods expanded exploration at the expense of structural auditability. Trace-level aggregation, such as Self-Consistency (Wang et al., 2023), samples CoTs from one model and votes on the final answer. Collaborative Beam Search (Xu et al., 2025) extends this by selecting the best continuation at each step via multi-model consensus. At the response level, Mixture-of-Agents (Wang et al., 2024) and ReConcile (Chen et al., 2024) generate a single response from the critiques of different agents. Finally, structured search methods like LE-MCTS (Park et al., 2024) and MoSA (Yang et al., 2025b) use Monte Carlo tree search over multi-LLM steps, while DER (Hu et al., 2025) trains a router to select the best expert output. These approaches generate diversity, but it ultimately collapses at aggregation time. Through voting, prose synthesis, and reward-guided search, the systems force convergence into a single final chain or answer and alternative paths are discarded. A concurrent line of work avoids this compression by preserving disagreement structure rather than discarding it. ARGORA (Jin et al., 2026) organizes multi-expert discussions into argumentation graphs to identify causal dependencies between competing claims, but relies on constrained interactive debate and causal interventions to evaluate argument fragility. Our framework instead extracts and ensembles DAGs from independent, unconstrained CoT traces across models, preserving the full weighted graph of attested reasoning steps via statistical consensus rather than counterfactual intervention. We never force convergence: when models disagree, both conclusions are retained with their supporting subgraphs. Consensus emerges naturally from independent reasoners. The resulting structure remains auditable at node level. 3 Methodology Our pipeline (Figure 1) operates in five stages—trace extraction (Section 3.1), DAG extraction (Section 3.2), node merging (Section 3.3), ensemble weighting (Section 3.4), and finding consensus (Section 3.5)—to convert multiple free-form reasoning chains into a weighted, auditable graph and obtain the consensus reasoning. A query is given to multiple LLMs, which sample numerous reasoning traces; a Directed Acyclic Graph (DAG) is extracted from each trace, the DAGs are ensembled into a weighted graph and the conclusion with the highest support and its most-supported reasoning trace are extracted. All prompts are in Appendix A. 3.1 Reasoning Trace Extraction We sample multiple stochastic reasoning traces per query from each of several small, open-source LLMs. Each model is constrained to a fixed schema requiring a step-by-step reasoning trace and a final verdict tag. For each generation, we separate the free-form reasoning from the verdict, retaining both for downstream DAG extraction and ensembling. 3.2 DAG Extraction Figure 1: Detailed pipeline. Step 1: Multiple LLMs sample reasoning traces for a query. Step 2: Each trace is decomposed into a DAG with nodes labeled Planning (P), Fact (F), Reasoning (Re), and Conclusion (C). Typed node labels let independently generated DAGs be aligned and merged step-by-step rather than compared as opaque prose. Step 3: Nodes are merged based on semantic similarity. Our hybrid ensembling strategy filters candidate node pairs by type (shown by color) and embedding similarity (threshold τ′τ ), then verifies surviving pairs with an LLM judge. Verified pairs merge into a single node representing the same reasoning step attested by multiple traces. The hybrid avoids both the false positives of similarity-only merging and the O(n2)O(n^2) cost of LLM-only merging. Step 4: The ensemble nodes are weighted based on cross-trace attestation, under simple weighting (each trace casts a unit vote, so weights depend on raw attestation counts) or accuracy-weighted (each vote scaled by its source model’s held-out accuracy). Red edges depict defeaters, gray supporters. Edge numbers are attestation counts. Accuracy weighting (whole-number for clarity, fractional in practice) redistributes mass toward steps attested by stronger models. Step 5: The highest supported answer and its highest supported justification is extracted. The output preserves alternative reasoning paths and step-level support that a single LLM’s chain-of-thought lacks. Each reasoning trace is converted into a fine-grained causal DAG using a small LLM as a structured extractor. To keep the task tractable, the LLM emits nodes together with pre-grouped bundles—each bundle is a list of node ids that must hold together to support a target—so the logical structure is encoded by the premise grouping in bundles. Each node carries an id, a node type, a short textual description, and bundle lists: a support list, whose bundles each independently justify the node, and a defeats list, whose bundles each independently suppress it. Premises within a bundle are conjunctive. Bundles within a list are disjunctive. Logical Reconstruction. An AND/OR DAG G=(V,)G=(V,S) is reconstructed in a single deterministic pass: every support bundle becomes a positive support set, every defeater bundle a negated one, and each node’s gate is assigned by cardinality of its positive sets—Atomic if none, And if one, Or if multiple. Formally, (v)=(S1,δ1),(S2,δ2),…S(v)=\(S_1, _1),(S_2, _2),…\, where each pair (Sk,δk)(S_k, _k) consists of a predecessor set Sk⊆VS_k V jointly supporting v and a polarity flag δk∈0,1 _k∈\0,1\ indicating support (δk=0 _k=0) or defeasibility (δk=1 _k=1). We verify that all node types are permitted, bundle members reference valid nodes, and the graph is acyclic. Traces failing any check are discarded so malformed outputs do not pollute aggregation. DAG Structure. Inspired by Lee et al. (2025), we use four labels: Planning (wherein the LLM decides what to do next), Fact (factual statements from the input context), Reasoning (intermediate inference steps linking facts to an outcome), and Conclusion (the final determination). This structure emphasizes procedural reasoning and is well suited to traces with meta-reasoning about problem-solving. A final Answer node is added after the Conclusion, exactly matching the predicted label. These DAGs provide the inputs to node merging. 3.3 Node Merging Our node-merging strategy filters same node-type candidate pairs by embedding similarity before LLM verification, balancing compute against merge precision. Neither pure approach suffices. An LLM judge, given each candidate pair’s text, labels, and local context yields accurate equivalence decisions but scales as O(n2)O(n^2) in the node count. Embedding similarity, computed as cosine distance between dense vectors of node descriptions, sim(u,v)=u⋅v‖u‖‖v‖sim(u,v)= e_u·e_v\|e_u\|\|e_v\| (1) is cheaper but produces false positives, often merging surface-similar yet substantively distinct nodes. Our hybrid method. We combine the two: a permissive embedding threshold τ′=0.55τ =0.55 prunes the search space, and surviving candidates are verified by the LLM judge. To preserve logical operators, we merge nodes while keeping bundles intact. Bundle-Preserving Merge Operator. Once a clustering π:V→π:V over all source-trace nodes is fixed, we construct the merged ensemble DAG G∗=(V∗,∗)G^*=(V^*,S^*) by mapping each bundle through π and aggregating identical bundles: ∗(c)=⋃v∈π−1(c)⋃(S,δ)∈(v)(π(S),δ)S^*(c)= _v∈π^-1(c) _(S,δ) (v) \ (π(S),δ ) \ (2) where π(S)=π(u):u∈Sπ(S)=\π(u):u∈ S\. Two bundles match across traces iff they reduce to the same cluster-id set under π. Identical bundles are combined, recording how many traces attested the same justification, while distinct bundles become separate support sets, recovering the OR-of-ANDs structure naturally. Gate types are then re-derived, and cycles introduced by clustering are resolved by removing the weakest support set along the cycle. Next, the resulting ensemble DAG is assigned node weights, based on cross-trace consensus. 3.4 Weighted Ensemble of DAGs We weight each merged node to amplify reasoning trajectories attested across multiple traces and suppress hallucinated ones, treating each trace as casting a signed vote on every merged node. Let Supp(v),Def(v)⊆MSupp(v),Def(v) M denote the source traces that contributed supporter and defeater bundles to merged node v, where M is the full set of traces, and let wtw_t be the per-trace weight. We define the positive and negative attestation masses α+(v)=∑t∈Supp(v)wt∑t∈Mwt,α−(v)=∑t∈Def(v)wt∑t∈Mwt.α^+(v)\;=\; _t (v)w_t _t∈ Mw_t, α^-(v)\;=\; _t (v)w_t _t∈ Mw_t. (3) so that α+(v),α−(v)∈[0,1]α^+(v),α^-(v)∈[0,1]. Traces that do not mention v contribute to neither mass but inflate the denominator, thus reducing confidence. Node scores W(v)∈[0,1]W(v)∈[0,1] are computed topologically (predecessors before successors) and combine the predecessor support with the node’s own support. The conjunctive strength of any bundle S is the weakest-link score over its members, ϕ(S)=minu∈SW(u),φ(S)\;=\; _u∈ SW(u), (4) so a conjunction is as strong as its weakest member. The positive and negative aggregate signals are then assembled by gate-aware combination over the disjunctive supporting and defeating bundles: Φ+(v)=1g(v)=Atomicϕ(S)g(v)=And,+(v)=SmaxS∈+(v)ϕ(S)g(v)=Or ^+(v)= cases1&g(v)= Atomic\\ φ(S)&g(v)= And,\,S^+(v)=\S\\\ _S ^+(v)φ(S)&g(v)= Or cases (5) Φ−(v)=0−(v)=∅ϕ(S)−(v)=SmaxS∈−(v)ϕ(S)|−(v)|>1 ^-(v)= cases0&S^-(v)= \\ φ(S)&S^-(v)=\S\\\ _S ^-(v)φ(S)&|S^-(v)|>1 cases (6) The final node score combines these with the attestation masses: W(v)=max(0,α+(v)⋅Φ+(v)−α−(v)⋅Φ−(v)).W(v)\;=\; \! (0,\;\;α^+(v)· ^+(v)\;-\;α^-(v)· ^-(v) ). (7) Simple Weighting. The baseline sets wt=1w_t=1 for all traces, treating consensus as a raw vote count. Accuracy-Weighting. When per-model accuracies are available from a held-out test set, we replace uniform votes with model-quality priors: wt=Waccm(t)w_t=W_acc^m(t), where ∑mWaccm=1 _mW_acc^m=1. Stronger models carry larger votes in both α+(v)α^+(v) and α−(v)α^-(v). With per-node weights that quantify the support to each node computed, the ensemble is parsed to select the highest-weighted conclusion and unfold its highest-weighted supporting reasoning, as described next. 3.5 Finding the Consensus The consensus answer is the candidate conclusion whose answer node carries the highest weight: c⋆=argmaxcW(c).c = _cW(c). (8) Ties are broken by the number of attesting traces. Consensus Subgraph. Given c⋆c , we recursively unfold a proof subgraph (c⋆)P(c ) rooted at c⋆c by picking at each node the support set maximizing ϕ(S)φ(S): • g(v)=Atomicg(v)= Atomic: include v and terminate. • g(v)=Andg(v)= And with unique support set S: include all members and expand each. • If g(v)=Org(v)= Or with support sets S1,S2,…\S_1,S_2,…\: select S⋆=argmaxSϕ(S)S = _Sφ(S), include all members, and recursively expand each. The resulting (c⋆)P(c ) is the strongest supported complete justification tree for c⋆c , with every conjunct required to support every intermediate step preserved. We report it as the consensus reasoning. To surface alternative reasoning chains, we enumerate the top-k subgraphs for each candidate conclusion by varying the support set at each Or node, scoring each subgraph by its total node weight: Ψ()=∑v∈W(v). (P)= _v W(v). (9) Alongside the consensus, we return for each conclusion its top-k justifications and its net support (total node weight across the union of its proof subgraphs) exposing alternative reasoning chains and relative standing of competing conclusions. 3.6 Summary Together, the five stages turn unstructured reasoning from heterogeneous models into a single auditable artifact. Bundle-preserving merging keeps each trace’s logical structure intact, attestation-weighted scoring suppresses hallucinated steps without discarding them, and thus we expose the consensus conclusion, consensus reasoning and the alternatives that remain in contention. The next section evaluates whether these design choices yield measurable accuracy, explainability and diversity gains over majority-vote baselines. 4 Experimental Setup Datasets. We evaluate on six benchmarks spanning four regimes: statutory interpretation, graduate-level science, narrative multi-hop reasoning, and first-order logic. For each, we pick 50 examples for a held-out split (used to estimate per-model accuracy priors) and an evaluation split (222, 148, 200, 206, 200, 153 samples), with a fixed random seed (42) for reproducibility. The entire pipeline runs on a single NVIDIA H200 GPU. (1) SARA Holzenberger et al. (2020) is a 272 question statutory reasoning benchmark over US federal tax law. It pairs a statute excerpt, a case scenario, and a hypothesis judged Entailed or Contradicted. (2) GPQA Diamond Rein et al. (2023) is a 198 question graduate-level multiple-choice benchmark in biology, chemistry, and physics, testing generalization to high-difficulty scientific reasoning. (3–5) MuSR Sprague et al. (2024) is a multi-step soft reasoning benchmark over synthetic narratives, with three domains we evaluate separately: Murder Mysteries (MuSR-M-250), inferring suspect, motive, and opportunity from narrative cues; Object Placements (MuSR-OP-256), tracking belief states across a sequence of actions; Team Allocation (MuSR-TA-250), assigning agents to tasks under capability, preference constraints. (6) FOLIO Han et al. (2024) is a 203 sample first-order logic entailment benchmark with natural-language premises, conclusions and three-way labels with an explicit “insufficient evidence” option. Category Weighting / Baseline SARA GPQA-D MuSR-M MuSR-OP MuSR-TA FOLIO (a) Heterogeneous ensemble (4 models, 5 chains each = 20 chains total) Ensemble (Ours) Simple 0.6225 ± 0.0142 0.2770 ± 0.0461 0.6040 ± 0.0096 0.4835 ± 0.0199 0.2860 ± 0.0195 0.6118 ± 0.0260 Accuracy-based 0.6387 ± 0.0080 0.2784 ± 0.0495 0.5960 ± 0.0074 0.4922 ± 0.0174 0.2890 ± 0.0167 0.6405 ± 0.0257 Baseline 4-model majority vote (20 chains) 0.6279 ± 0.0158 0.2662 ± 0.0421 0.5650 ± 0.0071 0.4874 ± 0.0184 0.2830 ± 0.0057 0.6405 ± 0.0122 Single-model CoT Gemma 3 12B 0.6322 ± 0.0162 0.2478 ± 0.0292 0.5500 ± 0.0068 0.4650 ± 0.0094 0.4150 ± 0.0118 0.6688 ± 0.0197 Llama 3.2 3B 0.4987 ± 0.0323 0.2462 ± 0.0408 0.5166 ± 0.0050 0.4371 ± 0.0130 0.3827 ± 0.0089 0.4532 ± 0.0130 Mistral 7B v0.3 0.5613 ± 0.0126 0.2576 ± 0.0352 0.6034 ± 0.0088 0.4328 ± 0.0038 0.2446 ± 0.0056 0.4797 ± 0.0122 Qwen 2.5 7B 0.7510 ± 0.0207 0.3087 ± 0.0462 0.5854 ± 0.0114 0.5565 ± 0.0115 0.3334 ± 0.0103 0.6518 ± 0.0119 (b) Homogeneous ensemble: Llama 3.2 3B (weak model, 20 chains) Ensemble (Ours) Llama 3.2 3B 0.5838 ± 0.0133 0.2406 ± 0.0304 0.5250 ± 0.0000 0.4524 ± 0.0121 0.4510 ± 0.0279 0.5386 ± 0.0229 Baselines Self-consistency (20 chains) 0.5541 ± 0.0139 0.2324 ± 0.0237 0.5250 ± 0.0000 0.4553 ± 0.0159 0.4420 ± 0.0346 0.5464 ± 0.0136 Single-chain CoT 0.4987 ± 0.0323 0.2462 ± 0.0408 0.5166 ± 0.0050 0.4371 ± 0.0130 0.3827 ± 0.0089 0.4532 ± 0.0130 (c) Homogeneous ensemble: Qwen 2.5 32B (strong model, 20 chains) Ensemble (Ours) Qwen 2.5 32B 0.7730 ± 0.0075 0.3587 ± 0.0244 0.6550 ± 0.0137 0.5281 ± 0.0100 0.5450 ± 0.0061 0.8065 ± 0.0199 Baselines Self-consistency (20 chains) 0.7730 ± 0.0082 0.3540 ± 0.0183 0.6610 ± 0.0096 0.5252 ± 0.0121 0.5460 ± 0.0055 0.8078 ± 0.0164 Single-chain CoT 0.7578 ± 0.0016 0.3555 ± 0.0028 0.6539 ± 0.0048 0.5210 ± 0.0025 0.5306 ± 0.0023 0.7675 ± 0.0069 Table 1: Accuracy of our framework against matched-budget baselines across six benchmarks. (a) Heterogeneous ensemble over four models, comparing our Simple (wt=1w_t=1) and Accuracy-based (wtw_t = source model’s held-out accuracy) weighting schemes to a 4-model majority-vote baseline and the four single-model CoT accuracies. (b, c) Homogeneous ensemble applied to a single weak model (Llama 3.2 3B) and a single strong model (Qwen 2.5 32B), each compared to self-consistency and single-chain CoT. The consensus answer is the answer node that has the highest aggregated weight. Values are mean ± SD over 5 seeds. Bold / underline = best / second-best per column in each block. Takeaway: on mixed models, our accuracy-weighted framework exceeds the majority-vote baseline on every dataset; on a single model, our framework exceeds single-chain CoT and matches or exceeds self-consistency. Models. We sample CoT from four small instruction-tuned LLMs spanning distinct families and pretraining mixtures within a fixed parameter budget (≤ 12B): Qwen 2.5 7B Instruct Qwen et al. (2025), Gemma 3 12B Instruct Team et al. (2025), Llama 3.2 3B Instruct Grattafiori et al. (2024), and Mistral 7B Instruct v0.3 Jiang et al. (2023), each at temperature 0.6. For DAG extraction and LLM-based node merging we use Qwen 2.5 32B Instruct Qwen et al. (2025) as a stronger structured extractor at temperature 0.3. Node similarity uses all-mpnet-base-v2 Reimers and Gurevych (2019), and the LLM judge in analysis is Qwen 3 32B Yang et al. (2025a). All models are open-weight, for reproducibility without proprietary access. Baselines. We compare against majority vote over the same four-model pool at a fixed 20-trace budget, isolating the contribution of structural aggregation: any difference reflects how traces are combined, not the underlying reasoners or sampling budget. Evaluation. We evaluate along three axes. (1) Accuracy: The final answer parsed from the ensemble graph, compared against a majority-vote baseline for both many-model and single-model ensembles. (2) Explainability: The framework returns the most-supported subgraphs, the support level for each conclusion, and the alternative conclusions reached. We additionally quantify if ensemble weights track reasoning quality using Qwen 3 32B as an LLM judge. (3) Diversity: We qualitatively check if alternative subgraphs present different solution perspectives. To motivate the use of multiple model families for obtaining diverse reasoning chains, we conduct an experiment to see if multi-model pools give more semantically diverse correct-answer reasoning trajectories than single-model pools (Appendix C). 5 Experiments and Results This section evaluates our framework. Sections 5.1 and 5.2 report accuracy for multi- and single-model ensembles against majority-vote baselines; Section 5.3 quantifies if ensemble weights track reasoning quality via LLM-judge rankings; Section 5.4 illustrates graphs; Section 5.5 has human evaluation. 5.1 Heterogeneous Ensemble Accuracies Table 1(a) reports our two weighting schemes against a matched-budget 4-model majority-vote baseline (20 chains) and single-model CoT. Different models lead on different datasets.. Strongest-to-weakest gaps within a dataset reach 25.2%, confirming no single model is uniformly best and motivating heterogeneous ensembling. Accuracy-weighted ensembling beats majority voting on every dataset. With the largest gain of 3.1% on MuSR-M (59.6% vs. 56.5%). Simple weighting also beats majority vote on three datasets, indicating gains stem from structural aggregation rather than the accuracy prior alone, though the prior delivers the additional boost for consistent wins. The strongest single model often wins outright. As expected for ensembling, averaging in weaker reasoners introduces noise structural aggregation may not filter. The ensemble’s value here is inspectability that the single best model can not provide. 5.2 Homogeneous Ensemble Accuracies To isolate structural aggregation from cross-model diversity, we apply our framework to single-model traces using a weak reasoner and a stronger reasoner, 20 chains each. Table 1(b,c) compares against self-consistency and single-chain CoT. Weak model: structural aggregation recovers gains. On Llama 3.2 3B, our method exceeds single-chain CoT (max +8.5%+8.5\% on SARA and FOLIO), and matches / exceeds self-consistency (max +3.0%+3.0\% on SARA).222Llama 3.2 3B exhibits near-deterministic majority voting on MuSR-M, making accuracy invariant across five seeds. When individual traces are unreliable, structural aggregation extracts a more trustworthy verdict from cross-trace consensus and exposes the supporting graph that voting discards. Weighting SARA GPQA-D MuSR-M MuSR-OP MuSR-TA FOLIO (a) Spearman ρ between judge ranking and ensemble-weight ranking (Ψ()=∑vW(v) (P)= _vW(v)) Simple 0.490 ± 0.030 0.353 ± 0.034 0.507 ± 0.029 0.451 ± 0.009 0.392 ± 0.033 0.333 ± 0.021 Accuracy 0.503 ± 0.023 0.383 ± 0.034 0.513 ± 0.024 0.449 ± 0.031 0.449 ± 0.020 0.298 ± 0.024 (b) Win rate of consensus subgraph vs. length-matched random alternative (%) Simple 60.0 ± 2.4 54.6 ± 3.9 60.3 ± 4.2 64.3 ± 3.2 60.3 ± 4.9 46.4 ± 2.6 Accuracy 61.6 ± 4.0 54.4 ± 3.8 59.9 ± 3.6 65.4 ± 3.6 61.2 ± 3.8 52.6 ± 4.3 Table 2: Two complementary measures of consensus-subgraph quality, judged by Qwen3-32B. (a) Per-question Spearman correlation between judge-ranked and weight-ranked sampled subgraphs (n=5n=5 per question, 5 shuffled orderings averaged). (b) Per-question win rate of the consensus subgraph against a length-matched random alternative trace leading to the majority-vote answer (ties excluded, ordering swapped). Values are mean ± SD across 5 seeds. Bold marks the higher value per column in each block. Takeaway: ensemble weights correlate positively with judge-rated reasoning quality, and the consensus subgraph is preferred over a random alternative. FOLIO is the only exception under simple weighting, and accuracy weighting recovers preference to 52.6%. Strong model: ensemble matches self-consistency and exceeds single-chain CoT. On Qwen 2.5 32B, our framework is within ±0.6%± 0.6\% of self-consistency on all six benchmarks and matches or exceeds single-chain CoT on every dataset, (max +3.9%+3.9\% on FOLIO). A strong reasoner converges to the right answer in most chains, so any aggregation recovers it. The framework’s contribution at this scale is inspectability, that self-consistency cannot provide. Benefits hold without cross-model diversity. Our framework recovers self-consistency’s accuracy while additionally producing a structured, weighted, auditable graph—demonstrating that its value is not contingent on heterogeneous pools. 5.3 Evaluation of Consensus Graph Quality Ensemble weights correlate positively with LLM-judge rankings of reasoning quality across all datasets and weighting schemes, indicating that the weight signal reflects genuine reasoning quality. Expt A: Subgraph ranking. For each question, we sampled 5 subgraphs from the ensemble, linearized each into a step-by-step chain, emitting nodes in topological order, and asked Qwen3-32B to rank them by reasoning quality (5 shuffled orderings per question, averaged to cancel presentation-order bias). Table 2(a) reports per-question Spearman correlation between judge and weight rankings. Correlation is moderate: weights and LLM-judge measure overlapping but not identical notions of justification quality, and the judge is imperfect. Expt B: Pairwise win rate. We sampled a subgraph leading to the majority-vote answer, length-matched it to the consensus subgraph to remove length bias, linearized both, and asked Qwen3-32B which chain provided better reasoning (each pair evaluated twice with swapped A/B order). Table 2(b) reports the per-question win rate of the consensus subgraph. The consensus trace is preferred over 50% of the time, showing that the weighting procedure captures a meaningful signal for reasoning quality beyond mere plausible justification. 5.4 Qualitative Analysis To support qualitative explainability evaluation, our framework returns three components per query: the top-k most-supported reasoning subgraphs, the support level for each conclusion, and the alternative conclusions reached by the models. Examples are provided in Appendix B, where each shows a consensus graph alongside an alternative subgraph (sampled from the same ensemble DAG) and a random single-trace CoT graph that reached the correct answer. Across these examples, the consensus graph consistently provides the most detailed and well-supported reasoning, the alternative graph offers a different justification for the same conclusion, and the single CoT is at times less thorough. 5.5 Human Evaluation We ran a small human evaluation (Appendix D) to complement our LLM-as-a-Judge results. Ten annotators were each shown 15 correct-answer graph triplets (consensus subgraph, alternative subgraph, single-CoT graph) randomly sampled from MuSR, and asked to pick the graph that best justified the answer. The consensus, single-CoT, alternative, and none options were chosen 41.3%, 33.3%, 20.7%, and 4.7% of the time, respectively, supporting our LLM-judge preferring the consensus reasoning over alternatives. The non-trivial alternative-graph share (20.7%) also suggests that surfacing alternative reasoning chains is useful in its own right. 6 Discussion Our results highlight three properties of structural ensembling. On accuracy, edges attested by only one trace contribute little to the consensus score, so a single model’s spurious thought cannot dominate the aggregate. This is what helps the framework beat majority voting on every dataset and recover gains over self-consistency on the single-weak-model. On interpretability, the output is a weighted graph rather than a free-form rationale, so users can view the consensus reasoning, different possible conclusions, and different justifications from diverse perspectives that support any conclusion—a property no answer-level baseline provides. On diversity, preserving competing conclusions makes query produce multiple supported reasoning paths, and heterogeneous ensembles widen this due to distinct pretraining distributions. This can also be observed qualitatively. These properties suggest a natural extension. In open-ended settings such as argument generation and creative writing, the goal is not to converge on a single answer but to surface a diverse set of well-supported alternatives. Our framework’s existing outputs—competing conclusions with their attested support graphs—map directly onto this use case: rather than report the highest-weighted answer alone, the system can present the top-k alternative conclusions and let the user choose among reasoning paths grounded in different models’ perspectives. This shifts the role of consensus from selection to curation, and we view it as the most promising direction for follow-up work. 7 Conclusion We introduced a framework for ensembling the reasoning structure—not just the answers—of multiple LLMs, by extracting per-trace DAGs, merging them with a logic-preserving operator, and weighting each node by cross-trace attestation. Across six benchmarks, our accuracy-weighted ensemble outperforms matched-budget majority voting on every dataset; applied to a single model, it matches self-consistency. Ensemble weights track LLM-judge rankings of reasoning quality, and consensus subgraphs are consistently preferred over alternatives leading to the majority-vote answer. The resulting weighted Reasoning Consensus graph is auditable at the level of individual reasoning steps, exposes alternative conclusions that competing traces reached, alternative justifications for each conclusion, and naturally extends to open-ended settings where curating diverse perspectives matters. 8 Limitations Our pipeline relies on an extractor model faithfully converting free-form chains of thought into typed, bundle-structured DAGs. Extraction errors would propagate through aggregation. To mitigate this, we use Qwen 2.5 32B, a substantially stronger model than the trace generators, at a low decoding temperature (0.3) for deterministic structured output, and we discard any trace whose extraction fails validation (acyclicity, valid node types, valid bundle references). Cross-trace attestation provides a second filter: a hallucinated step appearing in a single trace receives low weight, so even when individual extractions are imperfect, the aggregated graph remains dominated by steps that multiple independent traces converged on. We use Qwen 3 32B as a judge, and LLM judges exhibit known biases around position, length, and verbosity. We control for the most common ones explicitly: every pairwise comparison is evaluated twice with swapped A/B order, every ranking experiment is repeated five times per question with shuffled presentations and averaged, and alternative subgraphs are length-matched to the consensus subgraph to remove length confounds. To further verify the judge, we ran a human evaluation on randomly sampled graph triplets on the MuSR dataset. Each query requires sampling 20 chains of thought, extracting a DAG per chain, computing pairwise embeddings, and running the LLM judge for node merging — more compute than single-chain CoT or self-consistency. We deliberately target the high-stakes setting (legal, scientific) where the cost of an unsupported answer outweighs the compute cost of our framework, and where users cannot rely on closed APIs. To keep the framework optimized, we restrict all trace-generator models to ≤12≤ 12B parameters, use a permissive embedding threshold (τ′=0.55τ =0.55) to prune most node pairs before the LLM judge sees them, and run the entire pipeline on a single NVIDIA H200 GPU. 9 AI Involvement Disclosure AI assistance was limited to language polishing for grammar and readability. The conception, methodology, experiments, analyses, and interpretations were conducted fully by the authors. References M. Besta, N. Blach, A. Kubicek, R. Gerstenberger, M. Podstawski, L. Gianinazzi, J. Gajda, T. Lehmann, H. Niewiadomski, P. Nyczyk, and T. Hoefler (2024) Graph of thoughts: solving elaborate problems with large language models. Proceedings of the AAAI Conference on Artificial Intelligence 38 (16), p. 17682–17690. External Links: ISSN 2159-5399, Link, Document Cited by: §1, §2. J. C. Chen, S. Saha, and M. Bansal (2024) ReConcile: round-table conference improves reasoning via consensus among diverse llms. External Links: 2309.13007, Link Cited by: §1, §2. M. Dahl, V. Magesh, M. Suzgun, and D. E. Ho (2024) Large legal fictions: profiling legal hallucinations in large language models. Journal of Legal Analysis 16 (1), p. 64–93. External Links: Document, ISSN 1946-5319, Link Cited by: §1. A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Guzmán, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kambadur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. Çelebi, P. Alrassy, P. Zhang, P. Li, P. Vasic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Ganapathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Polidoro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot, S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speckbacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchandani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Montgomery, E. Presani, E. Hahn, E. Wood, E. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Jagadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma (2024) The llama 3 herd of models. External Links: 2407.21783, Link Cited by: §4. D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Xu, H. Ding, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Chen, J. Yuan, J. Tu, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. You, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Zhou, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang (2025) DeepSeek-r1 incentivizes reasoning in llms through reinforcement learning. Nature 645 (8081), p. 633–638. External Links: ISSN 1476-4687, Link, Document Cited by: §2. S. Han, H. Schoelkopf, Y. Zhao, Z. Qi, M. Riddell, W. Zhou, J. Coady, D. Peng, Y. Qiao, L. Benson, L. Sun, A. Wardle-Solano, H. Szabo, E. Zubova, M. Burtell, J. Fan, Y. Liu, B. Wong, M. Sailor, A. Ni, L. Nan, J. Kasai, T. Yu, R. Zhang, A. R. Fabbri, W. Kryscinski, S. Yavuz, Y. Liu, X. V. Lin, S. Joty, Y. Zhou, C. Xiong, R. Ying, A. Cohan, and D. Radev (2024) FOLIO: natural language reasoning with first-order logic. External Links: 2209.00840, Link Cited by: §4. N. Holzenberger, A. Blair-Stanek, and B. V. Durme (2020) A dataset for statutory reasoning in tax law entailment and question answering. External Links: 2005.05257, Link Cited by: §4. J. Hu, Y. Wang, S. Zhang, K. Zhou, G. Chen, Y. Hu, B. Xiao, and M. Tan (2025) Efficient dynamic ensembling for multiple llm experts. External Links: 2412.07448, Link Cited by: §2. A. Q. Jiang, A. Sablayrolles, A. Mensch, C. Bamford, D. S. Chaplot, D. de las Casas, F. Bressand, G. Lengyel, G. Lample, L. Saulnier, L. R. Lavaud, M. Lachaux, P. Stock, T. L. Scao, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed (2023) Mistral 7b. External Links: 2310.06825, Link Cited by: §4. Y. Jin, H. Kim, K. Kim, C. Lee, and S. Shin (2026) ARGORA: orchestrated argumentation for causally grounded llm reasoning and decision making. External Links: 2601.21533, Link Cited by: §2. J. Lee, S. Mukherjee, D. Hakkani-Tur, and J. Hockenmaier (2025) ReasoningFlow: semantic structure of complex reasoning traces. External Links: 2506.02532, Link Cited by: §3.2. OpenAI, :, A. Jaech, A. Kalai, A. Lerer, A. Richardson, A. El-Kishky, A. Low, A. Helyar, A. Madry, A. Beutel, A. Carney, A. Iftimie, A. Karpenko, A. T. Passos, A. Neitz, A. Prokofiev, A. Wei, A. Tam, A. Bennett, A. Kumar, A. Saraiva, A. Vallone, A. Duberstein, A. Kondrich, A. Mishchenko, A. Applebaum, A. Jiang, A. Nair, B. Zoph, B. Ghorbani, B. Rossen, B. Sokolowsky, B. Barak, B. McGrew, B. Minaiev, B. Hao, B. Baker, B. Houghton, B. McKinzie, B. Eastman, C. Lugaresi, C. Bassin, C. Hudson, C. M. Li, C. de Bourcy, C. Voss, C. Shen, C. Zhang, C. Koch, C. Orsinger, C. Hesse, C. Fischer, C. Chan, D. Roberts, D. Kappler, D. Levy, D. Selsam, D. Dohan, D. Farhi, D. Mely, D. Robinson, D. Tsipras, D. Li, D. Oprica, E. Freeman, E. Zhang, E. Wong, E. Proehl, E. Cheung, E. Mitchell, E. Wallace, E. Ritter, E. Mays, F. Wang, F. P. Such, F. Raso, F. Leoni, F. Tsimpourlas, F. Song, F. von Lohmann, F. Sulit, G. Salmon, G. Parascandolo, G. Chabot, G. Zhao, G. Brockman, G. Leclerc, H. Salman, H. Bao, H. Sheng, H. Andrin, H. Bagherinezhad, H. Ren, H. Lightman, H. W. Chung, I. Kivlichan, I. O’Connell, I. Osband, I. C. Gilaberte, I. Akkaya, I. Kostrikov, I. Sutskever, I. Kofman, J. Pachocki, J. Lennon, J. Wei, J. Harb, J. Twore, J. Feng, J. Yu, J. Weng, J. Tang, J. Yu, J. Q. Candela, J. Palermo, J. Parish, J. Heidecke, J. Hallman, J. Rizzo, J. Gordon, J. Uesato, J. Ward, J. Huizinga, J. Wang, K. Chen, K. Xiao, K. Singhal, K. Nguyen, K. Cobbe, K. Shi, K. Wood, K. Rimbach, K. Gu-Lemberg, K. Liu, K. Lu, K. Stone, K. Yu, L. Ahmad, L. Yang, L. Liu, L. Maksin, L. Ho, L. Fedus, L. Weng, L. Li, L. McCallum, L. Held, L. Kuhn, L. Kondraciuk, L. Kaiser, L. Metz, M. Boyd, M. Trebacz, M. Joglekar, M. Chen, M. Tintor, M. Meyer, M. Jones, M. Kaufer, M. Schwarzer, M. Shah, M. Yatbaz, M. Y. Guan, M. Xu, M. Yan, M. Glaese, M. Chen, M. Lampe, M. Malek, M. Wang, M. Fradin, M. McClay, M. Pavlov, M. Wang, M. Wang, M. Murati, M. Bavarian, M. Rohaninejad, N. McAleese, N. Chowdhury, N. Chowdhury, N. Ryder, N. Tezak, N. Brown, O. Nachum, O. Boiko, O. Murk, O. Watkins, P. Chao, P. Ashbourne, P. Izmailov, P. Zhokhov, R. Dias, R. Arora, R. Lin, R. G. Lopes, R. Gaon, R. Miyara, R. Leike, R. Hwang, R. Garg, R. Brown, R. James, R. Shu, R. Cheu, R. Greene, S. Jain, S. Altman, S. Toizer, S. Toyer, S. Miserendino, S. Agarwal, S. Hernandez, S. Baker, S. McKinney, S. Yan, S. Zhao, S. Hu, S. Santurkar, S. R. Chaudhuri, S. Zhang, S. Fu, S. Papay, S. Lin, S. Balaji, S. Sanjeev, S. Sidor, T. Broda, A. Clark, T. Wang, T. Gordon, T. Sanders, T. Patwardhan, T. Sottiaux, T. Degry, T. Dimson, T. Zheng, T. Garipov, T. Stasi, T. Bansal, T. Creech, T. Peterson, T. Eloundou, V. Qi, V. Kosaraju, V. Monaco, V. Pong, V. Fomenko, W. Zheng, W. Zhou, W. McCabe, W. Zaremba, Y. Dubois, Y. Lu, Y. Chen, Y. Cha, Y. Bai, Y. He, Y. Zhang, Y. Wang, Z. Shao, and Z. Li (2024) OpenAI o1 system card. External Links: 2412.16720, Link Cited by: §2. S. Park, X. Liu, Y. Gong, and E. Choi (2024) Ensembling large language models with process reward-guided tree search for better complex reasoning. External Links: 2412.15797, Link Cited by: §1, §2. Qwen, :, A. Yang, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Li, D. Liu, F. Huang, H. Wei, H. Lin, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Lin, K. Dang, K. Lu, K. Bao, K. Yang, L. Yu, M. Li, M. Xue, P. Zhang, Q. Zhu, R. Men, R. Lin, T. Li, T. Tang, T. Xia, X. Ren, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Wan, Y. Liu, Z. Cui, Z. Zhang, and Z. Qiu (2025) Qwen2.5 technical report. External Links: 2412.15115, Link Cited by: §4. N. Reimers and I. Gurevych (2019) Sentence-bert: sentence embeddings using siamese bert-networks. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing, External Links: Link Cited by: Appendix C, §4. D. Rein, B. L. Hou, A. C. Stickland, J. Petty, R. Y. Pang, J. Dirani, J. Michael, and S. R. Bowman (2023) GPQA: a graduate-level google-proof q&a benchmark. External Links: 2311.12022, Link Cited by: §4. Z. Sprague, X. Ye, K. Bostrom, S. Chaudhuri, and G. Durrett (2024) MuSR: testing the limits of chain-of-thought with multistep soft reasoning. External Links: 2310.16049, Link Cited by: §4. G. Team, A. Kamath, J. Ferret, S. Pathak, N. Vieillard, R. Merhej, S. Perrin, T. Matejovicova, A. Ramé, M. Rivière, L. Rouillard, T. Mesnard, G. Cideron, J. Grill, S. Ramos, E. Yvinec, M. Casbon, E. Pot, I. Penchev, G. Liu, F. Visin, K. Kenealy, L. Beyer, X. Zhai, A. Tsitsulin, R. Busa-Fekete, A. Feng, N. Sachdeva, B. Coleman, Y. Gao, B. Mustafa, I. Barr, E. Parisotto, D. Tian, M. Eyal, C. Cherry, J. Peter, D. Sinopalnikov, S. Bhupatiraju, R. Agarwal, M. Kazemi, D. Malkin, R. Kumar, D. Vilar, I. Brusilovsky, J. Luo, A. Steiner, A. Friesen, A. Sharma, A. Sharma, A. M. Gilady, A. Goedeckemeyer, A. Saade, A. Feng, A. Kolesnikov, A. Bendebury, A. Abdagic, A. Vadi, A. György, A. S. Pinto, A. Das, A. Bapna, A. Miech, A. Yang, A. Paterson, A. Shenoy, A. Chakrabarti, B. Piot, B. Wu, B. Shahriari, B. Petrini, C. Chen, C. L. Lan, C. A. Choquette-Choo, C. Carey, C. Brick, D. Deutsch, D. Eisenbud, D. Cattle, D. Cheng, D. Paparas, D. S. Sreepathihalli, D. Reid, D. Tran, D. Zelle, E. Noland, E. Huizenga, E. Kharitonov, F. Liu, G. Amirkhanyan, G. Cameron, H. Hashemi, H. Klimczak-Plucińska, H. Singh, H. Mehta, H. T. Lehri, H. Hazimeh, I. Ballantyne, I. Szpektor, I. Nardini, J. Pouget-Abadie, J. Chan, J. Stanton, J. Wieting, J. Lai, J. Orbay, J. Fernandez, J. Newlan, J. Ji, J. Singh, K. Black, K. Yu, K. Hui, K. Vodrahalli, K. Greff, L. Qiu, M. Valentine, M. Coelho, M. Ritter, M. Hoffman, M. Watson, M. Chaturvedi, M. Moynihan, M. Ma, N. Babar, N. Noy, N. Byrd, N. Roy, N. Momchev, N. Chauhan, N. Sachdeva, O. Bunyan, P. Botarda, P. Caron, P. K. Rubenstein, P. Culliton, P. Schmid, P. G. Sessa, P. Xu, P. Stanczyk, P. Tafti, R. Shivanna, R. Wu, R. Pan, R. Rokni, R. Willoughby, R. Vallu, R. Mullins, S. Jerome, S. Smoot, S. Girgin, S. Iqbal, S. Reddy, S. Sheth, S. Põder, S. Bhatnagar, S. R. Panyam, S. Eiger, S. Zhang, T. Liu, T. Yacovone, T. Liechty, U. Kalra, U. Evci, V. Misra, V. Roseberry, V. Feinberg, V. Kolesnikov, W. Han, W. Kwon, X. Chen, Y. Chow, Y. Zhu, Z. Wei, Z. Egyed, V. Cotruta, M. Giang, P. Kirk, A. Rao, K. Black, N. Babar, J. Lo, E. Moreira, L. G. Martins, O. Sanseviero, L. Gonzalez, Z. Gleicher, T. Warkentin, V. Mirrokni, E. Senter, E. Collins, J. Barral, Z. Ghahramani, R. Hadsell, Y. Matias, D. Sculley, S. Petrov, N. Fiedel, N. Shazeer, O. Vinyals, J. Dean, D. Hassabis, K. Kavukcuoglu, C. Farabet, E. Buchatskaya, J. Alayrac, R. Anil, Dmitry, Lepikhin, S. Borgeaud, O. Bachem, A. Joulin, A. Andreev, C. Hardin, R. Dadashi, and L. Hussenot (2025) Gemma 3 technical report. External Links: 2503.19786, Link Cited by: §4. J. Wang, J. Wang, B. Athiwaratkun, C. Zhang, and J. Zou (2024) Mixture-of-agents enhances large language model capabilities. External Links: 2406.04692, Link Cited by: §1, §2. X. Wang, J. Wei, D. Schuurmans, Q. Le, E. Chi, S. Narang, A. Chowdhery, and D. Zhou (2023) Self-consistency improves chain of thought reasoning in language models. External Links: 2203.11171, Link Cited by: §1, §1, §2. J. Wei, X. Wang, D. Schuurmans, M. Bosma, B. Ichter, F. Xia, E. Chi, Q. Le, and D. Zhou (2023) Chain-of-thought prompting elicits reasoning in large language models. External Links: 2201.11903, Link Cited by: §1, §2. Y. Xu, S. Ren, and J. Zhang (2025) Collaborative beam search: enhancing LLM reasoning via collective consensus. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, C. Christodoulopoulos, T. Chakraborty, C. Rose, and V. Peng (Eds.), Suzhou, China, p. 11398–11410. External Links: Document, ISBN 979-8-89176-332-6, Link Cited by: §1, §2. A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, and K. et al. (2025a) Qwen3 technical report. External Links: 2505.09388, Link Cited by: §4. S. Yang, Y. Li, W. Lam, and Y. Cheng (2025b) Multi-llm collaborative search for complex problem solving. External Links: 2502.18873, Link Cited by: §1, §2. X. Yang, X. Zhang, H. Hu, and F. Ji (2026) Beyond accuracy: evaluating strategy diversity in llm mathematical reasoning. External Links: 2605.09292, Link Cited by: §1. S. Yao, D. Yu, J. Zhao, I. Shafran, T. L. Griffiths, Y. Cao, and K. Narasimhan (2023) Tree of thoughts: deliberate problem solving with large language models. External Links: 2305.10601, Link Cited by: §1, §2. X. Zeng, J. Lin, Y. Yan, F. Guo, L. Shi, J. Wu, and D. Zhou (2026) HalluGuard: demystifying data-driven and reasoning-driven hallucinations in llms. External Links: 2601.18753, Link Cited by: §1. Appendix A Prompts The prompts used in our pipeline are: A.1 Reasoning Trace Extraction A.1.1 Shared System Prompt You are a careful reasoner. For every question, you must: 1. Think step by step inside a single <reasoning>...</reasoning> block. 2. Put ONLY the final answer (no explanation) inside a <answer>...</answer> tag. Do not write anything outside these two tags. Do not abbreviate your reasoning. A.1.2 SARA (Statutory Reasoning) You are reasoning about US federal tax law. STATUTE: statute CASE SCENARIO: case HYPOTHESIS: question TASK: Decide whether the statute, applied to the case scenario, ENTAILS or CONTRADICTS the hypothesis. Walk through: (a) which statutory provisions apply to the case, (b) what facts in the scenario trigger or fail to trigger those provisions, (c) what the provisions imply about the hypothesis. Output exactly: <reasoning>your step-by-step legal analysis</reasoning> <answer>ENTAILED</answer> or <answer>CONTRADICTED</answer> A.1.3 GPQA Diamond (Graduate-Level Science) You are answering a graduate-level science question. Exactly one option is correct. QUESTION: question OPTIONS: A. option_a B. option_b C. option_c D. option_d TASK: Reason carefully through the underlying science. Walk through: (a) what principles or equations the question is testing, (b) how to apply them to the specifics of this problem, (c) why one option follows and the others do not. Output exactly: <reasoning>your step-by-step scientific reasoning</reasoning> <answer>A</answer> (or B, C, or D -- a single capital letter only) A.1.4 MuSR (Narrative Multi-Hop Reasoning) internallinenumbers* You are reasoning about a narrative scenario. Read the narrative carefully, then choose the single correct answer. NARRATIVE: narrative QUESTION: question OPTIONS: A. choice_a B. choice_b ... (up to G) TASK: Reason carefully through the evidence in the narrative. Walk through: (a) what facts the narrative establishes, (b) how those facts bear on each candidate option, (c) why one option is best supported and the others are not. Output exactly: <reasoning>your step-by-step reasoning over the narrative</reasoning> <answer>A</answer> (or B, C, ... up to last_letter -- a single capital letter only) A.1.5 FOLIO (First-Order Logic Entailment) You are reasoning about first-order logical entailment in natural language. PREMISES: premises CONCLUSION: conclusion internallinenumbers* TASK: Decide whether the conclusion is logically entailed by the premises. Use these labels: TRUE -- the conclusion follows necessarily from the premises FALSE -- the negation of the conclusion follows from the premises UNCERTAIN -- neither the conclusion nor its negation follows Walk through: (a) what each premise asserts, (b) what chain of inferences, if any, the conclusion would require, (c) whether the premises are sufficient to establish or refute it. Output exactly: <reasoning>your step-by-step logical analysis</reasoning> <answer>TRUE</answer> or <answer>FALSE</answer> or <answer>UNCERTAIN</answer> A.2 DAG Extraction A.2.1 Shared System Prompt You are a structured-reasoning extractor. Given a reasoning trace and its final answer, you produce a Directed Acyclic Graph (DAG) in JSON that captures the trace’s logical structure. You output ONLY a single JSON object inside <dag>...</dag> tags. No prose. No markdown fences. A.2.2 DAG Extraction User Prompt TASK: Extract a Type-B reasoning DAG from this trace. NODE TYPES (use exactly these strings): - "Planning": meta-reasoning, problem decomposition, deciding what to check - "Fact": factual statements drawn from the question/statute/scenario - "Reasoning": intermediate inferential steps linking facts to outcome - "Conclusion": the final determination REQUIRED: the DAG must contain exactly ONE node of type "Conclusion", whose `text` expresses the final answer below (paraphrase it however the trace did, but it must be the same verdict). All other nodes must transitively support this Conclusion. BUNDLE STRUCTURE: each node has - "id": unique short string id (e.g. "p1", "f1", "r1", "c1") - "type": one of the four node types above - "text": short natural-language description (one sentence) - "support": list of bundles. Each bundle is a list of node ids that JOINTLY (conjunctive AND) justify this node. Different bundles in the list are alternative (disjunctive OR) independent justifications. - "defeats": list of bundles whose joint truth would SUPPRESS this node. Use [] if the trace mentions no counter-considerations. Do NOT invent defeaters. RULES: - Facts may have empty support (they are atomic premises from the question). - Planning/Reasoning/Conclusion nodes MUST have >=1 support bundle. - Every id referenced in a bundle must exist as a node in the graph. - The graph must be acyclic. - Do NOT add information the trace does not contain. ---------------------------------------- QUESTION CONTEXT: question_blob FINAL ANSWER (must be the Conclusion node’s verdict): final_answer REASONING TRACE: reasoning ---------------------------------------- Output ONLY this: <dag>"nodes": [ "id": "...", "type": "Planning|Fact|Reasoning|Conclusion", "text": "...", "support": [["id1","id2"], ["id3"]], "defeats": [] ]</dag> A.3 Node Merging A.3.1 Merge Judge System Prompt You are a strict semantic equivalence judge for reasoning-graph nodes. Two nodes should merge ONLY IF they express the SAME underlying claim, fact, or inference, regardless of wording. Lexical overlap is NOT enough. If they cite different statutes, parties, events, rules, or draw different inferences, they are DIFFERENT. Answer with a single token: YES or NO. A.3.2 Merge Judge User Prompt Decide whether these two reasoning-graph nodes should be merged into one. Both are of type "ntype" (already verified to match). QUESTION CONTEXT (helps interpret references like "the hypothesis"): question_blob NODE A: text_a NODE B: text_b Should A and B be merged because they assert the SAME underlying claim / fact / inference? Reply with exactly one token: YES or NO. Figure 2: Three reasoning graphs reaching the same conclusion (Answer A) for a MuSR Murder Mysteries question. Consensus graph (left): the highest-Ψ proof subgraph of the consensus answer, aggregating attested reasoning across the heterogeneous ensemble. Alternative graph (middle): a different proof subgraph for the same conclusion, sampled from the same ensemble DAG by varying the support set at Or nodes. Single CoT graph (right): a random per-trace DAG from one model whose chain reached the same answer. The consensus graph integrates evidence from multiple traces; the alternative offers a substantively different justification (eliminating Jose’s guilt vs. establishing Bella’s guilt); and the single CoT depends on fewer supporting facts. A.4 LLM-as-a-Judge A.4.1 Ranking Judge System Prompt You are a strict evaluator of reasoning quality. You are given several candidate reasoning chains that all argue toward the SAME final answer. Rank them by the QUALITY of their reasoning and justification: are the steps well-grounded in facts, are the inferences logically sound, are there gaps, leaps, or irrelevant steps. Do NOT reward a chain for reaching the answer -- they all reach the same answer. A chain can reach a correct answer through weak or wrong reasoning; rank that chain low. The number of steps does NOT matter; do not prefer a chain merely because it is longer. A.4.2 Ranking Judge User Prompt You are given n candidate reasoning chains. They ALL conclude the same final answer. Rank them from BEST to WORST reasoning quality. QUESTION CONTEXT: question_blob FINAL ANSWER (every chain reaches this -- it is NOT what you are judging): final_answer chains_block Rank all n chains from best reasoning to worst. Output ONLY the ranking as chain numbers separated by ’>’, best first. For example: 2 > 1 > 3 Do not output anything else. A.4.3 Pairwise Judge System Prompt You are a strict evaluator of reasoning quality. You are given two candidate reasoning chains, each arguing for a final answer to the same question. Decide which chain is the better reasoning and justification: better-grounded steps, sounder inferences, fewer gaps or irrelevant leaps. Chain LENGTH IS IRRELEVANT: a longer chain with more steps is NOT automatically better, and a shorter chain is NOT automatically worse. Do not reward verbosity. Judge only the soundness and grounding of the reasoning, not how many steps it has. Answer with a single token: A or B. A.4.4 Pairwise Judge User Prompt Two candidate reasoning chains each argue for an answer to the same question. Decide which chain is the better reasoning and justification. QUESTION CONTEXT: question_blob --- CHAIN A (concludes: answer_a) --- chain_a --- CHAIN B (concludes: answer_b) --- chain_b Which chain reasons better -- better-grounded steps, sounder inferences, fewer gaps or irrelevant leaps? The number of steps does NOT matter; do not prefer a chain merely because it is longer. Reply with exactly one token: A or B. Appendix B Visualization Warning: This appendix contains examples from MuSR-Murder Mystery subset describing hypothetical murder-related scenarios. Figure 2 depicts an example of consensus, alternative and CoT graphs from MuSR-M that reach the same correct answer. We can observe that the alternative graph solved the murder mystery by eliminating Jose guilty while the consensus graph solves it by asserting Bella’s guilt. This shows us the diverse perspectives taken to solve the same problem by two different graphs extracted from our ensemble DAG. The CoT graph is less detailed and does not take into account various facts that the consensus graph does. Figure 3 also depicts an example of consensus, alternative and CoT graphs from MuSR-M that reach the same correct answer. We can observe that the consensus graph reveals much more intricate reasoning as compared to the alternative or CoT graphs, which is beneficial for users audits. Figure 4 compares a consensus graph with an alternative graph from the MUSR-OP dataset that reach different answers. The alternative graph argues that since tools are usually in a tool shed, they should be searched for in the shed first. The consensus graph argues that since Emma had moved the tools to the backyard, she should look there first. Surfacing multiple reasoning graphs and conclusions allows the users to audit them and decide which answer and justification they agree with. Figure 3: Three reasoning graphs reaching the same conclusion (Answer A) for a MuSR Murder Mysteries question. Consensus graph (left): the highest-Ψ proof subgraph of the consensus answer, aggregating attested reasoning across the heterogeneous ensemble. Alternative graph (middle): a different proof subgraph for the same conclusion, sampled from the same ensemble DAG by varying the support set at Or nodes. Single CoT graph (right): a random per-trace DAG from one model whose chain reached the same answer. The consensus graph integrates evidence from multiple traces and we can observe here that both the alternative graph and CoT graph show less intricate reasoning when compared to the highest weighted consensus graph. Figure 4: Three reasoning graphs reaching the same conclusion (Answer A) for a MuSR Object Placement question. Consensus graph (left): the highest-Ψ proof subgraph of the consensus answer, aggregating attested reasoning across the heterogeneous ensemble. Alternative conclusion graph (middle): a proof subgraph for a different conclusion, sampled from the same ensemble DAG. The consensus graph integrates evidence from multiple traces; the alternative conclusion graph offers a substantively different reasoning and conclusion (tools are usually in a tool shed so search there first vs. they were moved to the backyard so search there first). Appendix C Single vs. multi-model diversity Condition Diversity Correct Traces Qwen-only (8) 0.111±0.0560.111± 0.056 4.893±3.2524.893± 3.252 Gemma-only (8) 0.081±0.0500.081± 0.050 4.317±3.5754.317± 3.575 Llama-only (8) 0.149±0.0700.149± 0.070 3.654±2.6753.654± 2.675 Mistral-only (8) 0.146±0.0710.146± 0.071 4.102±3.4724.102± 3.472 Mixed (2 each) 0.164±0.0640.164± 0.064 4.300±2.9984.300± 2.998 Table 3: Semantic diversity of correct reasoning traces generated by individual model families and a heterogeneous ensemble. Diversity is measured as mean pairwise 1−cosine similarity1-cosine similarity between trace embeddings. Mixed-model setting achieves highest diversity, indicating a broader range of successful reasoning strategies. We use multiple model families in our ensemble to obtain diverse reasoning perspectives that reach the correct answer. To motivate this and verify that heterogeneous ensembles produce more diverse successful trajectories, we sampled eight reasoning traces (for SARA and GPQA-D) either entirely from one model or in a mixed setting with two traces per model (Qwen 2.5 7B, Gemma 3 12B, Llama 3.2 3B, Mistral 7B v0.3). We retained only traces reaching the correct final answer, embedded them with all-MiniLM-L6-v2 Reimers and Gurevych (2019), and computed pairwise diversity as 1−cosine similarity1-cosine similarity. Table 3 reports the mean and standard deviation of diversity, along with the average number of correct traces per question. The mixed-model setting achieves the highest diversity, supporting the claim that combining heterogeneous models yields a broader variety of successful reasoning paths than any single model. Appendix D Human annotation details Annotators were students recruited on a volunteer basis. Instruction: You are provided three reasoning graphs that lead to an answer, select the one that you think justifies the answer in the most logical, correct, coherent way. If none justify the answer well, you can respond with "none". This will be used for research. Warning: contains examples describing hypothetical murder-related scenarios.