Paper deep dive
When Do Institutions Beat Intelligence?
Zhengye Han
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/13/2026, 3:37:15 AM
Summary
The paper investigates when institutional structures outperform raw intelligence in multi-agent systems. It defines four loci of collective failure: access/routing, admission/dependence, state maintenance/incentives, and representation/action. Through controlled artificial ecologies, the authors demonstrate that institutions improve performance by repairing failures in constructing usable public state (e.g., ensuring evidence coverage, validating independent support, maintaining current state). However, institutions lose their advantage when signals are uninformative, when stronger intelligence can perform the transformation directly, or when the resulting state cannot support reliable action.
Entities (10)
Relation Signals (7)
Intelligence Intervention → comparesto → Institutional Intervention
confidence 95% · Our results recast the choice between intelligence and institutions as a diagnosis of where collective reasoning fails.
Institutional Intervention → addresses → Admission and Dependence
confidence 90% · Admission and dependence asks which reports become public evidence... institutions help when they change what is allowed to count as a public reason
Institutional Intervention → addresses → State Maintenance and Incentives
confidence 90% · Maintenance and incentives asks whether public state remains current... institutions help when they repair failures in how a collective constructs usable public state
Institutional Intervention → addresses → Access and Routing
confidence 90% · Institutions help when they repair failures in how a collective constructs usable public state... access and routing
Institutional Intervention → addresses → Representation and Action
confidence 90% · Representation and action asks whether a constructed state remains useful... institutions help when they repair failures in how a collective constructs usable public state
Grounded Validation → improves → Public State Quality
confidence 85% · Grounded validation exceeds syntax-only admission... the gain is rejection of unsupported evidence
LLaMA-3.1-8B → usedin → Access and Routing
confidence 85% · In distributed evidence, the valid five-call 8B institution exceeds the five-call raw 70B system
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:More capable agents do not necessarily form a more capable collective. A multi-agent system may jointly possess sufficient information yet fail because evidence is poorly routed, unreliable reports enter public belief, correlated claims masquerade as independent support, shared state becomes stale or strategically distorted, or useful evidence is exposed through an ineffective action interface. We ask when additional resources should improve the reasoner and when they should instead change the institutional structure through which the collective forms and acts on public information. Drawing on functional distinctions from research on group decision making and distributed cognition, we construct controlled artificial ecologies around four loci of collective failure: access and routing, admission and dependence, state maintenance and incentives, and representation and action. Across these ecologies, we separately vary model capability and institutional structure, pairing positive interventions with matched reasoning baselines and mechanism-breaking controls. The experiments reveal a consistent boundary: institutions help when they repair failures in how a collective constructs usable public state, but lose their advantage when their signals are uninformative or uncheckable, when stronger intelligence can perform the same transformation directly, or when the resulting state cannot support reliable action. Our results recast the choice between intelligence and institutions as a diagnosis of where collective reasoning fails.
Tags
Links
- Source: https://arxiv.org/abs/2608.11357v1
- Canonical: https://arxiv.org/abs/2608.11357v1
Trouble viewing inline? Open PDF directly →
Full Text
49,246 characters extracted from source content.
Expand or collapse full text
When Do Institutions Beat Intelligence? Zhengye Han New York University Department of Electrical and Computer Engineering Brooklyn, New York, United States zh3286@nyu.edu Abstract More capable agents do not necessarily form a more capable collec- tive. A multi-agent system may jointly possess sufficient informa- tion yet fail because evidence is poorly routed, unreliable reports enter public belief, correlated claims masquerade as independent support, shared state becomes stale or strategically distorted, or useful evidence is exposed through an ineffective action interface. We ask when additional resources should improve the reasoner and when they should instead change the institutional structure through which the collective forms and acts on public information. Drawing on functional distinctions from research on group decision making and distributed cognition, we construct controlled artificial ecologies around four loci of collective failure: access and routing, admission and dependence, state maintenance and incentives, and representation and action. Across these ecologies, we separately vary model capability and institutional structure, pairing positive interventions with matched reasoning baselines and mechanism- breaking controls. The experiments reveal a consistent boundary: institutions help when they repair failures in how a collective con- structs usable public state, but lose their advantage when their signals are uninformative or uncheckable, when stronger intelli- gence can perform the same transformation directly, or when the resulting state cannot support reliable action. Our results recast the choice between intelligence and institutions as a diagnosis of where collective reasoning fails. CCS Concepts • Computing methodologies→ Multi-agent systems. Keywords multi-agent systems, institutions, collective intelligence, group deci- sion making, large language model agents, empirical methodology 1 Introduction More capable agents do not necessarily form a more capable collec- tive. A system may jointly contain enough information to solve a task and still fail because that information never becomes jointly accessible, because public belief admits unsupported or dependent evidence, because shared state is not maintained, or because the resulting representation cannot support reliable action. Such fail- ures are not reducible to weakness inside any one agent. They arise along the path by which private observations become public reasons for collective action. This distinction creates a concrete design choice. An intelligence intervention strengthens the reasoner, for example by changing model family, scale, or inference budget. An institutional interven- tion changes the structure through which agents observe, report, verify, update, and act on information. Intelligence expands what a reasoner can compute from the state it receives. Institutions change how the collective constructs that state. Our question is therefore not whether institutions are generally better than intelligence, but which source of collective failure each intervention can actually repair. The ingredients of this distinction are well established in multi- agent systems: protocols, organizations, norms, monitoring, and shared state regulate interaction among autonomous agents [30], while institutional analysis emphasizes how rules-in-use shape col- lective outcomes [19,20]. What is harder to identify empirically is which intervention repaired a failing artificial collective. Contem- porary agent workflows often vary roles, debate, voting, retrieval, memory, model capability, and call count together [4,8,32]. An improvement in end-task performance then shows that the pipeline changed, but does not reveal whether the binding constraint was reasoning capacity, evidence access, evidential validity, enforceabil- ity, or the action interface. This ambiguity matters in both direc- tions: institutional machinery introduces additional computation and failure modes, while additional reasoning is wasted when every reasoner continues to receive the same unusable state. We study this identification problem through controlled artificial ecologies. Rather than treating environments as interchangeable benchmarks, we design them around four locations at which col- lective information processing can fail (Fig. 1). Access and routing asks whether distributed knowledge becomes jointly available. Ad- mission and dependence asks which reports become public evidence and whether repeated reports provide independent support. Main- tenance and incentives asks whether public state remains current and whether deviations can be observed and made consequential. Representation and action asks whether a constructed state remains useful relative to stronger capability and whether the final agent can reliably act on it. These distinctions are motivated by recurrent problems in re- search on human groups, without assuming psychological equiv- alence between human and artificial agents. Hidden-profile and transactive-memory research shows that a group may collectively possess decisive information while failing to retrieve and coordinate it effectively [17,24,25]. Research on social influence shows why repeated agreement need not constitute independent evidence [1,2, 12]. Work on accountability emphasizes that consequences depend on what can be observed and attributed [23,27,28]. Distributed- cognition research, in turn, treats external representations as part of the cognitive system rather than as neutral containers of already- completed reasoning [15,22,34]. We use these literatures to moti- vate distinct functional failure structures, not to claim that artificial agents reproduce human psychological mechanisms. Each ecology then isolates one such failure and asks whether changing capability or changing institutional structure repairs it. arXiv:2608.11357v1 [cs.MA] 11 Aug 2026 Intelligence or Institutions ? Agents with private observations Private observation Access & routing Admission & validation Public-state maintenance Action interface Collective decision Decision Intelligence intervention Institutional intervention Strongermodel More inference Better computation Observe Route VerifyUpdate Improves reasoning over the received state Changes how the collective constructs the state Do we need a stronger reasoner, or a more usable collective decision state? Figure 1: Intelligence and institutions intervene at different levels of collective computation. Intelligence interventions strengthen the reasoning applied to a received state. Institutional interventions change how private observations become decision-relevant public state through access, routing, admission, verification, updating, and the action interface. We ask whether a failing collective requires a stronger reasoner or a more usable decision state. Institutional interventions are paired with stronger-model or call- matched raw baselines, together with mechanism-breaking con- trols under which the proposed institutional explanation should lose value. Call-matched sampling, voting, and debate test whether an apparent gain is merely additional inference. Below-threshold routing tests whether coverage rather than workflow complexity matters. Syntax-only validation separates format from evidential validity. Zero-checkability conditions separate sanctions from en- forceability. Frontier models and learned retrieval test whether capability can substitute for an explicit institution. Fixed-evidence interfaces separate better evidence construction from better exe- cution. The objective is therefore not to rank workflows, but to identify where collective computation fails and why a particular intervention succeeds or becomes redundant. We make three contributions. First, we introduce a function- ally defined taxonomy of artificial collective failures, motivated by group decision research, that separates failures of access, public belief, state maintenance, and action. Second, we develop matched, mechanism-breaking comparisons that distinguish institutional re- pair from additional reasoning and specify conditions under which an institutional explanation should fail. Third, across controlled ecologies, strategic interaction, multiple model families, and natu- ral evidence tasks, we identify both institutional advantage and its boundaries: institutions help when they repair the construction of usable public state, but lose value when their signals are invalid or uncheckable, when stronger capability performs the same transfor- mation directly, or when the resulting state cannot support reliable action. 2 Institutions as Rules for Collective State Construction We separate the task ecology from the rules through which a col- lective turns private observations into a decision-relevant public state. An ecology is E= ⟨ 푁,푆,푂 푖 푖∈푁 ,퐴 푖 푖∈푁 ,푇,푅 푖 푖∈푁 ,퐻 ⟩ , with agents푁, latent task state푆, private observation channels 푂 푖 , reporting or action spaces퐴 푖 , transition dynamics푇, possibly heterogeneous rewards 푅 푖 , and horizon 퐻 . Within an ecology, we represent an institution as 퐼=⟨휌,푉,푈,Φ⟩. The component휌specifies observation and reporting rights: which agents may access, produce, or route which information. Given the resulting reports푟 푡 ,푉specifies which reports are entitled to enter public belief. The admitted evidence ̃ 푟 푡 = 푉(푟 푡 )updates a shared public state푝 푡 = 푈(푝 푡−1 , ̃ 푟 푡 ),andΦdetermines how that state is represented and exposed to the final decision process. A reasoner with capability푚 then acts onΦ(푝 푡 ). 2 This definition makes the institutional boundary explicit. We call휌,푉,푈,Φinstitutional because they are externally specified rules governing collective information rights, admissibility, state maintenance, and action affordances, rather than parameters of an individual reasoner. They may be role- or history-dependent and may condition future interaction on observable behavior. The defi- nition therefore includes routing, verification, public memory, audit, and institutionally constrained action interfaces, while excluding additional sampling or cosmetic reformatting that leaves the infor- mation available to the collective and its action rules unchanged. This operationalization follows the institutional perspective that rules and coordination structures regulate interaction among agents [19, 20, 30]. Institutional and capability interventions. Let푌(푒,퐼,푚)denote system performance on paired instance푒under institution퐼and reasoner capability푚. A capability intervention changes푚while holding the collective information structure fixed; an institutional intervention changes how that information is routed, admitted, updated, or exposed while holding the reasoner fixed. For regime 푟 ∈ +,−, define 휏 푟 퐼 (푚)=E 푒∼E 푌(푒,퐼 푟 ,푚)−푌(푒,퐼 푟 0 ,푚) , where퐼 + contains the proposed informative institutional signal and 퐼 − preserves the workflow while breaking that signal. We define the mechanism-validity interaction Γ 퐼 (푚)=휏 + 퐼 (푚)−휏 − 퐼 (푚). An institutional gain receives mechanistic credit only when perfor- mance tracks the signal the institution claims to exploit: structure alone is insufficient if the gain persists when that signal is broken. To compare structure with additional capability, we use the crossover 퐶 퐼 (푚 푠 ,푚 ℓ )=E 푒∼E 푌(푒,퐼 + ,푚 푠 )−푌(푒,퐼 0 ,푚 ℓ ) , where푚 푠 is a weaker institutionally supported reasoner and푚 ℓ a stronger raw reasoner evaluated on the same paired instances and, where applicable, under matched call budgets. A positive crossover is therefore evidence for institutional advantage only when the corresponding mechanism-validity test is also positive. Diagnostic failure sources. We do not assume that institutional effects admit an additive structural decomposition. Instead, the experiments distinguish five diagnostic sources of success or fail- ure: correction of an active bottleneck (퐵), information lost during state construction (퐿), mismatch between the constructed state and downstream execution (푋), institutional overhead (푂), and substi- tution when a sufficiently capable learned component performs the same transformation directly (푆). Broken or uncheckable signals test퐵; construction diagnostics expose퐿; fixed-evidence interfaces isolate푋; explicit resource accounting measures푂; and stronger raw models or learned constructors probe푆. These are diagnostic labels tied to matched controls, not separately identified causal parameters. 3 Experimental Design Design logic. Every main comparison is paired by seed or question. Controlled structural experiments and semantic validation form the causal core; cross-family and zero-checkability tests provide boundary replication; natural tasks probe transfer, substitution, and execution rather than an aggregate benchmark score. Construction conditions. The locked construction core includes raw inference, call-matched sampling, majority vote, a free-form debate proxy, unvalidated pooling, below-threshold institutions, valid institutions, and diagnostic public-state execution. Validity and admission designs. Static validity studies use a 2×2 design: the same finalizer sees raw or institutionally constructed state under a valid or deliberately broken regime. Controlled se- mantic validation fixes the transcript and changes only whether one unsupported row is admitted. Repeated and natural-task designs. The repeated ecology crosses private-signal capability, audit probability, checkability, and sanc- tion strength under common random numbers. Natural tasks keep selected passage count fixed; learned retrieval tests construction substitution, while ordered and pointer interfaces hold evidence fixed and change only the action representation. Model and scale coverage. The locked hosted core evaluates Llama 3.1 8B and Llama 3.3 70B on 50 held-out seeds per construc- tion ecology. Separate panels test Qwen 235B/Mistral 24B family transfer, validity interactions for Qwen 9B, Qwen 235B, and Mistral 24B, and frontier substitution for Claude Sonnet 4.6, Gemini 2.5 Pro, and DeepSeek V3.2. Sample sizes. The repeated-audit grid uses 100 matched seeds per cell, with 50 for the zero-checkability boundary. HotpotQA uses 20 paired questions per model. MuSiQue retrieval uses 300 questions balanced across 2–4 hops; hosted execution uses 21 or 30 questions crossed with three finalizers. The frozen-history comparison uses 20 simulation seeds and three post-burn-in questions per seed. Statistics. We report paired mean differences and nonparametric 95% bootstrap intervals. Resampling follows the assignment unit: seed for controlled ecologies, question for ordinary natural pan- els, and connected source-component cluster when dependence remains. The structural panels retain arm means because absolute failure is informative; other main panels emphasize paired effects. 4 Distributed Knowledge: Access and Routing Group-psychology motivation. Hidden-profile research separates information that a group possesses from information that enters collective deliberation: when critical facts are split across mem- bers so that no individual’s private view favors the correct option but the group’s pooled view does, discussion still gravitates to- ward information everyone already shares rather than the decisive unique pieces. Across 65 studies, groups discussed substantially more shared than unique information and were far less likely to solve hidden profiles than groups given complete information [17]. Transactive-memory work adds a second distinction: coordination improves when members know who has access to which expertise [25,31]. The functional lesson is not that artificial agents share a human bias. It is that distributed possession is not yet collective availability. Artificial ecology. Figure 2 turns this distinction into a controlled artificial ecology. The task has 12 candidate answers (A–L) and a deck of 16 clue cards, each encoding one logical constraint on the correct answer – e.g., a card of typeeliminated_setwith options 3 Distributed Evidence Bottleneck Worker B Worker A Worker C Worker D Raw reasoning (no coordination)Institutional assembly (coverage guarantee) Stronger reasoner Stronger reasoner Stronger reasoner Stronger reasoner Raw Aggregation (no coverage guarantee) Worker B Worker A Worker C Worker D Insufficient joint coverage Coverage routing ↓ Validation ↓ Shared public state Correct decision Collective knowledge depends on what the decision process can jointly assemble. Capability improves inference; institutions can change access. Figure 2: Distributed evidence as an access bottleneck. Raw reasoning can spend more capability on partial views with- out guaranteeing joint coverage. Here, coverage means that the public decision state contains all complementary evi- dence components required for the task, rather than merely more reports or more computation. Institutional routing and validation instead assemble a jointly sufficient public state by making those complementary components collectively available to the finalizer. K, E states that the answer is neither K nor E. No single card, and no individual worker’s local view, is sufficient to determine the answer. We distinguish three objects: a clue card is the private input a worker is shown; a constraint is the logical restriction it expresses; and a record is the worker’s structured transcription of that card (id, type, options). The institution validates each record against the card the worker actually saw, then aggregates validated records into one public constraint table exposed to the finalizer. Raw 8B and 70B systems receive the same underlying cards with matched calls. An assignment푚×푘gives푚workers푘cards each. 3×2 and 3×4 use option-symbol-balanced routing that allows cards to repeat across workers and leaves coverage incomplete; only the 4×4 assignment uses a disjoint partition that covers all 16 cards exactly once and supplies a jointly sufficient state. A related expertise-routing ecol- ogy uses public credentials to direct domain-specific evidence and breaks the credential signal in the control condition. Result and interpretation. In distributed evidence, the valid five- call 8B institution exceeds the five-call raw 70B system by+0.58 [0.44, 0.72] (Fig. 4). As routing becomes more complete, task accu- racy rises from 0.26 under 3×2, to 0.96 under 3×4, and to 1.00 under the disjoint 4×4 assignment; correspondingly, mean coverage of decisive clues rises from 0.357 to 0.760 to 1.000, while the validated public table uniquely identifies the answer in 0%, 94%, and 100% of trials, respectively – the chain is routing completeness→decisive- evidence coverage→unique public decision state→final accuracy. Separate Qwen/Mistral probes reproduce the qualitative threshold pattern. Expertise-routing validity is tested with the cross-model interactions in Fig. 6a. The conclusion is narrower than “small mod- els beat large models.” A stronger reasoner can make better use of evidence it receives, but cannot infer constraints that never become jointly available. The institution helps by changing access to the decision state. If a capable raw reasoner can directly reconstruct the same transformation, that conclusion should reverse; Section 7 tests exactly that boundary. Public Belief Failure Modes A B Evidence admission Correlated wrong consensus Syntax checks (structure, schema) Syntax-only validation Wrong decision Reasoner Institutional validation (cross-check with records) Admitted to public state (contaminated) · Reasoner Wrong decision Admitted to public state (clean) · · Rejected (unsupported) Grounded validation Format valid ≠ Evidence valid versus Shared misleading cue (common source) Agent A Agent B Agent C Raw pooling (debate) Public belief False consensus Repeated agreement ≠ Independent evidence Shared misleading cue (common source) Agent A Agent B Agent C Public belief Grounded belief Verify against records Figure 3: Two public-belief failure modes. Schema-valid re- ports can remain evidentially unsupported, and repeated reports can share one underlying error source. Grounded verification changes which reports may update public belief; it does not imply that agreement or voting is generally unre- liable. Icons denote representative report sets. 5 Public Belief: Admission and Dependence Group-psychology motivation. Communication can fail after evi- dence has been shared. Human-group studies show that decision makers may evaluate the same information differently depending on prior preferences [13], while work on the wisdom of crowds shows that agreement is informative only under assumptions about the independence and distribution of judgments [1,2,12]. A re- peated claim can therefore increase consensus without adding in- dependent evidence. The corresponding institutional question is epistemic: which reports are entitled to update public belief ? Artificial ecologies. Figure 3 separates admission from evidential dependence. In the faulty-evidence ecology, malformed or logically invalid records enter a common pool. The institution validates re- ports against visible record rules before updating public state; raw and call-matched baselines repeatedly reason over the contami- nated pool. The wrong-consensus ecology creates several reports that repeat one salient but incorrect cue. Voting, repeated sampling, and free-form debate preserve correlation, whereas exact-record verification tests whether claims have independent support. A con- trolled semantic-validation experiment isolates admission from generation: one schema-valid report flips a visible constraint type, and the finalizer receives either syntax-only admission or grounded admission based on the exact visible record. Result and interpretation. Validated construction reaches 1.00 in faulty evidence while call-matched raw reasoning remains at 0.00. Under wrong consensus, the valid 8B institution exceeds the four-call raw 70B baseline by+0.68 [0.55, 0.81] (Fig. 4). In the con- trolled transcript, grounded validation exceeds syntax-only admis- sion by+0.50 [0.30, 0.70], while clean minus grounded is only+0.067 [−0.100, 0.233] (Fig. 6b). Because clean and grounded-filtered fi- nalizer prompts are identical, the gain is rejection of unsupported 4 Raw 8B Stronger raw Broken / incomplete Valid institution Valid institution - stronger raw 0.00.51.0 Δ +0.58 (a) Distributed evidence 0.00.51.0 Accuracy / paired accuracy difference Δ +1.00 (b) Faulty evidence 0.00.51.0 Δ +0.68 (c) Wrong consensus Figure 4: Structural information bottlenecks survive stronger reasoning over the same unusable state. The three ecologies share four directly labeled method rows: raw 8B, a call-matched stronger raw reasoner, the ecology-specific broken or incomplete rule, and the valid 8B institution. Points are arm accuracies with seed-bootstrap 95% intervals (푛=50); the separated bottom row reports the paired valid-institution minus stronger-raw effect. These are call-matched frozen ecologies, not a universal 8B-over-70B or dollar-cost claim. evidence rather than cleaner prose, extra sampling, or an answer- bearing hint. These experiments distinguish social agreement from evidential warrant. Institutions help when they change what is allowed to count as a public reason, not when they merely cause the same evidence to be reconsidered more often. The broken-rule and syntax- only controls are therefore central: structure without a valid admis- sion signal receives no causal credit. Checkable Enforcement No checkability Checkable audit Early round Later roundPublic state Early round Later round Public state Agent Private signal Report distorted Sanction Agent Private signal Report distorted Sanction Inaccurate public state •Violation not observable •Sanction cannot target behavior Agent Private signal Report distorted Audit (Checkable Signal) Reject/ penalty Agent Private signal Report accurate Audit (Checkable Signal) Reject/ penalty Improving public state •Checkable audit •Observable consequence A sanction without observability is an unenforceable declaration. Figure 5: Checkability makes enforcement behaviorally active. The main panel depicts the simulated repeated- reporting mechanism: sanctions cannot condition on un- observable violations, whereas checkable audit can connect deviations to future consequences. 6 Accountability Through Time: Maintenance and Incentives Group-psychology motivation. Collective knowledge must re- main current and strategically credible. Dynamic situation man- agement requires teams to update shared representations as the task changes [14]. Accountability research also warns that merely demanding justification does not uniformly improve judgment: effects depend on what is observable, who evaluates, and how consequences are attached [23,27,28]. A sanction that cannot be connected to an observable violation is a declared norm, not an enforceable institution. Artificial ecologies. Figure 5 separates observability from conse- quence. The dynamic-state environment gives agents chronological records in which a later observation supersedes an earlier one. The institution maintains a current public state from a visible update relation; the broken regime degrades the visibility and reliability of that relation. The repeated-reporting ecology introduces misaligned agents that may report a private capture option rather than their signal. Audit probability determines whether reports are inspected, checkability determines whether inspection can identify a fabrica- tion, and sanctions alter future incentives. Private-signal capability supplies the intelligence comparison. A separate hosted-policy pilot asks whether language-model reporters exhibit the same adaptive behavior. Result and interpretation. Across Qwen 9B, Mistral 24B, and Qwen 235B, dynamic-state valid-minus-broken interactions are weak at 0.15, 0.15, and 0.30; only the Qwen 235B interval excludes zero (Fig. 6a). This heterogeneous boundary shows that public mem- ory does not receive a generic benefit from added structure. In the adaptive simulator, a 10% audit at capability 0.55 raises accuracy from 0.265 to 0.535, reduces private capture from 0.676 to 0.398, and beats unaudited capability 0.65 by+0.135 [0.096, 0.175]. At zero checkability, every sanction has exactly 0.00 accuracy gain under common random numbers (Fig. 6c). The mechanism is not punishment by itself, but observable evidence that connects behav- ior to future consequence. 5 0.00.51.0 Validity interaction: valid - broken Public state (weak) Expertise routing Strategic reporting (a) Valid rules help across model families Qwen 9BMistral 24BQwen 235B 0.00.40.8 Paired accuracy difference Grounded - syntax Clean - grounded +0.50 +0.07 (b) Evidence must be admissible 0.000.150.30 Paired accuracy difference Institutional gain Institution - stronger No checkability +0.271 +0.135 0.000 (c) Checkability is necessary Figure 6: Institutional effects track valid and behaviorally active signals. (a) Valid-rule gain minus broken-rule gain across three model families; the public-state row is the deliberately weak boundary. (b) Holding the transcript and finalizer fixed, grounded rejection exceeds syntax-only admission and remains close to the clean control (arm accuracies 0.23, 0.73, and 0.80). (c) In the adaptive simulator, checkable audit improves accuracy and can exceed the next unaudited capability, whereas four sanction levels have exact-zero gain when violations are unobservable. Whiskers are paired 95% bootstrap intervals under each experiment’s registered seed set; simulator adaptation is not a hosted-agent claim. The hosted pilot marks the behavioral limit: across 2,400 Llama and Mistral decisions, both families reported honestly in every con- dition, so the predeclared behavioral gate stopped scaling (Table 1). The selector’s 0.022 regret on 16 disjoint simulator settings is a prospective within-ecology design result, not a universal policy. Evidence, Substitution, and Execution 1. Evidence construction 2.Substitution 3. Execution interface Question Candidate records · A. Sparse retrieval B. Graph-based public-state construction C. Learned reranker · 1 2 3 · Flat list, no explicit relations Structured, linked public state Ranked, cohesive set Same underlying task Baseline construction Institutional construction Capability substitution Sparse state → Model Graph state →same model Raw state → strong model Same finalizer Same finalizer Sparse-evidence baseline Structured-state advantage Structured-state advantage Less reliable action Morereliable action · Graph interface Ordered interface Hold evidence fixed (identical accepted records) 1 · 2 3 Figure 7: Construction, substitution, and execution are dis- tinct. Learned components or stronger models may substitute for explicit institutional construction, while fixed-evidence comparisons isolate whether the resulting state is executable. 7 From Public Evidence to Public Action Distributed-cognition motivation. Collective reasoning does not end once relevant evidence has been collected. Accounts of dis- tributed cognition treat agents and external representations as parts of one functional system [15,18], and work on representation shows that two states can contain the same information while imposing very different operations on the decision maker [22,34]. For an artificial collective, this creates three distinct questions. Can an institution construct a better public evidence state? Can a stronger model or learned component perform the same transformation without the institution? And, even when the evidence is better, can the final agent reliably act on the way that evidence is represented? Can the institution construct a better public state? We first test whether explicit structure improves which evidence reaches the fi- nalizer. HotpotQA [33] compares an observable title-link graph with equal-size lexical selection at similar input-token counts. MuSiQue [29] provides a harder construction test: on 300 balanced 2–4 hop examples, we compare an entity graph with BM25, dense retrieval, hybrid retrieval, and an all-passage learned reranker. Gold support labels are used only after selection for retrieval diagnostics. Graph routing transfers across three HotpotQA finalizers and improves MuSiQue support recall over equal-passage BM25 by+0.135 [0.098, 0.172]. Thus a hand-specified relational structure can recover evi- dence that sparse lexical retrieval misses. When does capability substitute for explicit structure? A useful construction rule need not remain useful as capability improves. On MuSiQue, the learned reranker exceeds the hand-specified graph by +0.070 [0.036, 0.104], showing that learned evidence selection can substitute for explicit graph construction. We test the same prin- ciple at the model level with a separate frontier panel. There, the 6 0.00.51.0 Valid - broken DeepSeek Gemini Claude (a) Public state 0.00.51.0 Valid - broken Expertise routing 0.00.51.0 Valid - broken Strategic reporting −0.20.00.2 Paired difference Graph - BM25 (support) Reranker - graph (support) Graph - BM25 (answer) Public state - ordered +0.135 +0.070 +0.079 -0.067 (b) From construction to action No auditAudit 0.25 0.4 0.5 0.6 0.7 0.8 Accuracy 0.47 0.52 Raw records 0.80 Routed public state 0.75 (c) Supply x interface Figure 8: Capability, learned construction, and the action interface bound institutional advantage. (a) Frontier-model valid- minus-broken interactions test whether capable raw reasoners substitute for transparent institutional transformations. (b) Construction and execution are separated directly: retrieval comparisons measure support recovery and answer transfer, while public-state versus ordered passages holds evidence fixed and changes only the interface. (c) Frozen histories cross upstream audit with the downstream public-state interface. Whiskers denote paired 95% bootstrap intervals. raw model must reconstruct transparent dynamic-state, expertise- routing, and reporting transformations without receiving the cor- responding institutional state. DeepSeek retains large valid-minus- broken interactions, Gemini partially substitutes, and Claude pro- duces zero incremental interaction in all three tested ecologies (Fig. 8a). These are not failed replications. They identify a boundary on institutional value: when a capable reasoner can reliably perform the same transparent transformation from raw context, keeping that transformation as permanent institutional machinery becomes unnecessary. Can the finalizer act on the constructed state? Better evidence construction is still not sufficient for better decisions. Although the graph improves MuSiQue support recall over BM25, its hosted answer advantage over BM25 is imprecise. We therefore hold the selected evidence fixed and change only how it is exposed to the finalizer. The same passages are presented either through the graph- style public state or through an ordered or passage-pointer inter- face. Under this fixed-evidence comparison, the ordered interface reverses the graph comparison (Fig. 8b). The evidence is present in both conditions; what changes is the operation required of the final- izer. This identifies the action interface as a possible downstream bottleneck: a public state can contain better evidence without mak- ing the remaining decision easier to execute. Where does the end-to-end gain come from? Frozen natural histo- ries finally separate upstream evidence supply from downstream representation. The same frozen visible-record pool is evaluated under a 2×2 design that crosses upstream audit with either a raw- record or public-state finalizer interface. At the raw interface, a 0.25 checkable audit raises accuracy from 0.467 to 0.800, a paired supply effect of+0.333 [0.217, 0.450]. Holding the accepted records fixed, however, the downstream public-state transformation has no detectable simple effect either without audit (+0.050 [−0.067, 0.150]) or under audit (−0.050 [−0.150, 0.050]). The full audit-plus- public-state pipeline gains+0.283 [0.150, 0.417] over no-audit raw execution (Fig. 8c). The detectable improvement is therefore asso- ciated with changing which evidence reaches the finalizer, rather than with reformatting the same accepted evidence alone. Only the finalizer is hosted in this experiment; the adaptive reporters are simulated. 8 When Do Institutions Beat Intelligence? The ecologies represent different failures, but their positive, null, and reversal results support three recurring empirical conditions. The failure is social-structural rather than merely computational. Institutional value persists when stronger or call-matched reason- ing receives the same missing, contaminated, correlated, stale, or strategically distorted state. It disappears when a capable raw model or learned constructor performs the same transformation. Institu- tional advantage is therefore relative to a capability frontier, not a permanent property of a workflow. The institutional signal is valid and behaviorally active. A rule earns causal credit when performance follows the signal it claims to use and the gain disappears when the signal is broken or made uncheckable. Grounded rather than syntactic validation, valid rather than broken credentials, and nonzero rather than zero checkability satisfy this requirement. Labels, sanctions, and structure have no independent force. 7 The public state is nonredundant and executable. An institution can improve an intermediate representation without improving the final decision. Fixed-evidence comparisons and formatting nulls show that construction and use must be evaluated separately. The relevant object is the entire path from private observation to action, not the apparent quality of a public artifact in isolation. These conditions are cross-ecology regularities, not a sequen- tial certification algorithm or a universal theorem. They imply a three-way design choice. Buy intelligence when capability directly performs the required transformation. Build an institution when the active structural failure survives scaling, the rule observes it reli- ably, and the resulting state supports action under explicit resource accounting. Redesign the interface, or choose neither intervention, when the rule is uninformative, redundant, or produces a state the finalizer cannot use. Table 1: Boundary, null, and substitution evidence. Each probe limits where institutional advantage should be ex- pected. ProbeEvidence and implication Dynamic-state validityInteractions are 0.15, 0.15, and 0.30; only Qwen 235B excludes zero. Public-state structure is not generically beneficial. Zero checkabilityAll four sanction levels yield exactly 0.00 gain. Sanctions require observable violations. Hosted strategic pilot0 misreports in 2,400 decisions. The interface was exer- cised, but hosted strategic adaptation was not observed. Frontier substitutionClaude shows 0.00 incremental interaction in all three tested ecologies. Transparent institutional rules can be- come capability-substitutable. Learned substitutionReranker minus graph support recall is+0.070 [0.036, 0.104]. Hand-specified construction is not permanently privileged. Graph answer transferGraph minus BM25 is+0.079 [−0.143, 0.302]. Better sup- port recovery need not yield a detectable answer gain. Fixed-evidence interfaceGraph minus ordered accuracy is−0.067 [−0.128, −0.010]. Representation can reverse performance with evidence held fixed. Frozen-pool interfacePublic-state effects are+0.050 [−0.067, 0.150] without audit and−0.050 [−0.150, 0.050] with audit. Reformat- ting the same accepted evidence has no detectable benefit. 9 Related Work Multi-agent institutions. Electronic institutions and normative multi-agent systems treat roles, protocols, admissible interactions, monitoring, and enforcement as system-level rules for coordinat- ing autonomous agents [3,9,10]. Institutional analysis similarly studies how rules-in-use structure collective outcomes [19,20]. Re- cent work extends this perspective to LLM collectives, showing that governance topology, market rules, and executable governance mechanisms can substantially change collective behavior [6,11,26]. These studies establish that institutional design can matter for artifi- cial collectives. Our focus is the comparative identification problem: when does a structural intervention repair a binding collective failure relative to additional capability, and when can stronger or learned reasoning substitute for the same transformation? We therefore pair institutional gains with matched capability baselines, mechanism-breaking controls, and explicit substitution tests. Collective information processing. Hidden profiles, transactive memory, social influence, accountability, and external cognition explain how collective performance can diverge from individual competence [2,17,25,28,34]. Rather than cite these traditions as analogy alone, we use each to motivate a separable experimen- tal variable: coverage, evidential dependence, checkability, state maintenance, or action affordance. LLM multi-agent systems. Debate, role specialization, voting, re- trieval, memory, and tool-mediated coordination are widely used in LLM-agent systems [4,8,32]. Recent work also begins to sepa- rate organization, coordination, and collaboration protocol as inde- pendently configurable design dimensions [5]. Agent benchmarks assess broad competence or interactive behavior [16,21], while cooperative-AI work emphasizes coordination and common ground [7]. Our contribution is not another fixed workflow, organizational architecture, or aggregate benchmark. Instead, every positive struc- tural intervention is paired with a stronger or call-matched reasoner and a condition under which its proposed mechanism should fail or become redundant. 10 Limitations and Scope Several causal ecologies are deliberately synthetic, and transcript- level corruption provides causal precision rather than open-domain fact checking. The psychological literature supplies construct prove- nance, not evidence that language-model agents share human cog- nitive mechanisms. Adaptive strategic behavior remains simulation evidence: the hosted audit pilot elicited no strategic variation, and the frozen-history study hosts the finalizer rather than the reporters. Natural execution panels are smaller than the offline retrieval audit. Results depend on provider implementations and model versions, and the selector is prospective only within one strategic ecology. Call matching does not equal matching tokens, latency, monetary cost, or engineering burden; we use explicit resource accounting without claiming timeless provider-price rankings. Finally, the di- agnostic decomposition organizes evidence but does not constitute an identified structural theory. 11 Conclusion A collective can contain enough intelligence and enough informa- tion yet fail to transform private evidence into public action. Across ecologies of access, admission, maintenance, incentives, representa- tion, and execution, institutions outperform additional intelligence when they repair a structural bottleneck that capability alone does not remove. The nulls and reversals define the same phenomenon: unobservable violations make sanctions inert, capable models and learned retrieval absorb transparent institutional functions, and unusable interfaces erase gains from better evidence. The practical question is therefore not whether institutions or intelligence are generally superior. It is where collective computation breaks, and which intervention changes that location most directly. References [1]Abdullah Almaatouq, M. Amin Rahimian, Jason W. Burton, and Abdullah J. Alhajri. 2022. The Distribution of Initial Estimates Moderates the Effect of Social Influence on the Wisdom of the Crowd. Scientific Reports 12 (2022). doi:10.1038/ s41598-022-20551-7 8 [2]Joshua Becker, Devon Brackbill, and Damon Centola. 2017. Network Dynamics of Social Influence in the Wisdom of Crowds. Proceedings of the National Academy of Sciences 114, 26 (2017), E5070–E5076. doi:10.1073/pnas.1615978114 [3] Guido Boella, Leendert Van Der Torre, and Harko Verhagen. 2006. Introduction to normative multiagent systems. Computational & Mathematical Organization Theory 12, 2 (2006), 71–79. [4]Chi-Min Chan, Weize Chen, Yusheng Su, Jianxuan Yu, Wei Xue, Shanghang Zhang, and Jie Fu. 2024. ChatEval: Towards Better LLM-Based Evaluators through Multi-Agent Debate. In International Conference on Learning Representations. [5]Huan Chen, Xiang Song, Jian Jin, Pan Ren, and Liang-Jie Zhang. 2026. Toward an Organizational Science of Multi-Agent LLM Systems: Decoupling Who, How, and Which Algorithm. arXiv preprint arXiv:2607.25446 (2026). [6] Maxim Chupilkin. 2026. Artificial Institutions: How Institutional Design Shapes LLM Simulations. arXiv preprint arXiv:2608.04020 (2026). [7]Allan Dafoe et al.2021. Cooperative AI: Machines Must Learn to Find Common Ground. Nature 593 (2021), 33–36. doi:10.1038/d41586-021-01170-0 [8] Yilun Du, Shuang Li, Antonio Torralba, Joshua B. Tenenbaum, and Igor Mor- datch. 2024. Improving Factuality and Reasoning in Language Models through Multiagent Debate. In Proceedings of the 41st International Conference on Machine Learning. [9]Marc Esteva, David De La Cruz, and Carles Sierra. 2002. ISLANDER: an electronic institutions editor. In Proceedings of the first international joint conference on Autonomous agents and multiagent systems: part 3. 1045–1052. [10]Marc Esteva, Juan-Antonio Rodriguez-Aguilar, Carles Sierra, Pere Garcia, and Josep L Arcos. 2001. On the formal specification of electronic institutions. In Agent Mediated Electronic Commerce: The European AgentLink Perspective. Springer, 126– 147. [11]Chao Fei, Hongcheng Guo, and Yanghua Xiao. 2026. When Agents Evolve, Institutions Follow. arXiv preprint arXiv:2604.27691 (2026). [12]Vincenz Frey and Arnout van de Rijt. 2021. Social Influence Undermines the Wisdom of the Crowd in Sequential Decision Making. Management Science 67, 7 (2021), 4273–4286. doi:10.1287/mnsc.2020.3713 [13] Tobias Greitemeyer and Stefan Schulz-Hardt. 2003. Preference-Consistent Evalu- ation of Information in the Hidden Profile Paradigm. Journal of Personality and Social Psychology 84, 2 (2003), 322–339. doi:10.1037/0022-3514.84.2.322 [14]Jean-Michel Hoc. 2000. Cognitive Aspects of Dynamic Situation Management. In Proceedings of the Human Factors and Ergonomics Society Annual Meeting, Vol. 44. 156–156. doi:10.1177/154193120004400141 [15] Edwin Hutchins. 1995. Cognition in the Wild. MIT Press. [16] Xiao Liu et al.2024. AgentBench: Evaluating LLMs as Agents. In International Conference on Learning Representations. [17] Li Lu, Y. Connie Yuan, and Poppy Lauretta McLeod. 2012. Twenty-Five Years of Hidden Profiles in Group Decision Making: A Meta-Analysis. Personality and Social Psychology Review 16, 1 (2012), 54–75. doi:10.1177/1088868311417243 [18]Kourken Michaelian and John Sutton. 2013. Distributed Cognition and Memory Research: History and Current Directions. Review of Philosophy and Psychology 4 (2013), 1–24. doi:10.1007/s13164-013-0131-x [19]Elinor Ostrom. 1990. Governing the Commons: The Evolution of Institutions for Collective Action. Cambridge University Press. doi:10.1017/CBO9780511807763 [20]Elinor Ostrom. 2009. A General Framework for Analyzing Sustainability of Social- Ecological Systems. Science 325, 5939 (2009), 419–422. doi:10.1126/science.1172133 [21]Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative Agents: Interactive Simulacra of Human Behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology. doi:10.1145/3586183.3606763 [22] Mike Scaife and Yvonne Rogers. 1996. External Cognition: How Do Graphical Representations Work? International Journal of Human-Computer Studies 45, 2 (1996), 185–213. doi:10.1006/ijhc.1996.0048 [23]Thomas Schillemans. 2022. Accountability and the Quality of Regulatory Judg- ment Processes. Public Performance & Management Review 45 (2022), 473–498. doi:10.1080/15309576.2022.2040034 [24]Garold Stasser and William Titus. 1985. Pooling of Unshared Information in Group Decision Making: Biased Information Sampling During Discussion. Journal of Personality and Social Psychology 48, 6 (1985), 1467–1478. doi:10.1037/0022- 3514.48.6.1467 [25]Garold Stasser, Sandra I. Vaughan, and Dennis D. Stewart. 2000. Pooling Unshared Information: The Benefits of Knowing How Access to Information Is Distributed among Group Members. Organizational Behavior and Human Decision Processes 82, 1 (2000), 102–116. doi:10.1006/obhd.2000.2890 [26]Marcantonio Bracale Syrnikov, Federico Pierucci, Marcello Galisai, Matteo Prandi, Piercosma Bisconti, Francesco Giarrusso, Olga Sorokoletova, Vincenzo Suriani, and Daniele Nardi. 2026. Institutional AI: Governing LLM collusion in multi-agent cournot markets via public governance graphs. arXiv preprint arXiv:2601.11369 (2026). [27]Philip E. Tetlock and Richard Boettger. 1989. Accountability: A Social Magnifier of the Dilution Effect. Journal of Personality and Social Psychology 57, 3 (1989), 388–398. doi:10.1037/0022-3514.57.3.388 [28] Philip E. Tetlock, Linda Skitka, and Richard Boettger. 1989. Social and Cogni- tive Strategies for Coping with Accountability: Conformity, Complexity, and Bolstering. Journal of Personality and Social Psychology 57, 4 (1989), 632–640. doi:10.1037/0022-3514.57.4.632 [29]Harsh Trivedi, Niranjan Balasubramanian, Tushar Khot, and Ashish Sabharwal. 2022. MuSiQue: Multihop Questions via Single-Hop Question Composition. Transactions of the Association for Computational Linguistics 10 (2022), 539–554. doi:10.1162/tacl_a_00475 [30]Paul Valckenaers, John Sauter, Carles Sierra, and Juan A. Rodríguez-Aguilar. 2007. Applications and Environments for Multi-Agent Systems. Autonomous Agents and Multi-Agent Systems 14 (2007), 61–85. doi:10.1007/s10458-006-9002-5 [31]Wendy P. van Ginkel and Daan van Knippenberg. 2009. Knowledge about the Distribution of Information and Group Decision Making: When and Why Does It Work? Organizational Behavior and Human Decision Processes 108, 2 (2009), 218–229. doi:10.1016/j.obhdp.2008.10.003 [32]Qingyun Wu, Gagan Bansal, Jieyu Zhang, Yiran Wu, Beibin Li, Erkang Zhu, Li Jiang, Xiaoyun Zhang, Shaokun Zhang, Jiale Liu, Ahmed Hassan Awadallah, Ryen W. White, Doug Burger, and Chi Wang. 2024. AutoGen: Enabling Next-Gen LLM Applications via Multi-Agent Conversation. In Conference on Language Modeling. [33]Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William W. Cohen, Ruslan Salakhutdinov, and Christopher D. Manning. 2018. HotpotQA: A Dataset for Diverse, Explainable Multi-Hop Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing. 2369–2380. doi:10.18653/v1/D18-1259 [34] Jiajie Zhang and Donald A. Norman. 1994. Representations in Distributed Cogni- tive Tasks. Cognitive Science 18, 1 (1994), 87–122. doi:10.1207/s15516709cog1801_3 9