Paper deep dive
Artificial collectives of specialists and generalists excel at different tasks
John Meluso, Laurent Hébert-Dufresne, Christoph Riedl, H. Oliver Gao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/9/2026, 6:23:54 AM
Summary
This study investigates how agent interpretive abilities and bounded rationality influence collective performance in multi-agent systems. It finds that specialists (sparse, centralized networks) and generalists (dense, decentralized networks) excel at different task qualities. Generalists outperform on generate, choose, and coordinate tasks, while specialists excel at negotiate tasks. Bounded rationality moderates these dynamics: loose computational bounds favor specialists, tight bounds favor generalists, and moderate bounds reveal a performance-convergence speed trade-off. The findings advocate for matching network topology to task demands and computational limits.
Entities (10)
Relation Signals (10)
Specialists â correspondsto â Sparse centralized networks
confidence 95% · Collectives of specialists correspond to sparse, centralized networks
Generalists â correspondsto â Dense decentralized networks
confidence 95% · collectives of generalists correspond to dense, decentralized ones
Bounded rationality â moderates â Network topology effects
confidence 94% · Rationality bounds then moderate these relationships.
Loose rationality bounds â favors â Specialists
confidence 93% · At loose bounds, specialists outperform generalists
Tight rationality bounds â favors â Generalists
confidence 93% · At tight bounds, generalists outperform specialists
Specialists â excelsat â Negotiate tasks
confidence 92% · collectives of specialists with a few generalist mediators perform better on tasks that involve negotiating
Generalists â excelsat â Generate tasks
confidence 92% · collectives of generalists to perform better on tasks that involve generating
Generalists â excelsat â Choose tasks
confidence 92% · perform better on tasks that involve generating, choosing, and coordinating
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Collective artificial intelligence, where multiple agents work on shared tasks, holds potential to solve expansive problems in fields from medicine to collective governance. But while prescriptive engineering solutions abound, we lack descriptive scientific understanding of artificial collectives, and therefore principles for how to design resource efficient multi-agent systems. Through systematic experiments with optimizing agents, we characterize how agent interpretive abilities, rationality bounds, and task qualities interact to shape collective performance. Agents range from specialists, with narrow interpretive abilities, to generalists, with broad ones. Collectives of specialists correspond to sparse, centralized networks, while collectives of generalists correspond to dense, decentralized ones. We show that interpretive network properties have small performance effects on average (0.07 standard deviations of performance). However, for specific task qualities, these effects are 4.5 times larger (0.33 sd) and can reach much higher for certain task qualities (1.84 sd). This leads collectives of generalists to perform better on tasks that involve generating, choosing, and coordinating, while collectives of specialists with a few generalist mediators perform better on tasks that involve negotiating. Rationality bounds then moderate these relationships. At loose bounds, specialists outperform generalists through more effective sampling of high-dimensional decision spaces. At tight bounds, generalists outperform specialists through better gradient estimation. A fundamental trade-off between performance and convergence speed emerges at moderate bounds. These findings suggest that multi-agent design could benefit from matching interpretive networks to both task demands and agents' computational limits, with implications for the efficiency and energy costs of multi-agent systems.
Tags
Links
- Source: https://arxiv.org/abs/2606.20877v1
- Canonical: https://arxiv.org/abs/2606.20877v1
Trouble viewing inline? Open PDF directly â
Full Text
54,387 characters extracted from source content.
Expand or collapse full text
Artificial collectives of specialists and generalists excel at different tasks John Meluso 1 , Laurent HĂ©bert-Dufresne 2,3 , Christoph Riedl 4 and H. Oliver Gao 1 1 Cornell University 2 University of Vermont 3 Santa Fe Institute 4 Northeastern University This manuscript was compiled on June 23, 2026 Abstract Collective artificial intelligenceâwhere multiple agents work on shared tasksâholds potential to solve expansive problems in fields from medicine to collective governance. But while prescriptive engineering solutions abound, we lack descriptive scientific understanding of artificial collectives, and therefore principles for how to design resource efficient multi-agent systems. Through systematic experiments with optimizing agents, we characterize how agent interpretive abilities, rationality bounds, and task qualities interact to shape collective performance. Agents range from specialists, with narrow interpretive abilities, to generalists, with broad ones. Collectives of specialists correspond to sparse, centralized networks, while collectives of generalists correspond to dense, decentralized ones. We show that interpretive network properties have small performance effects on average (0.07 standard deviations of performance). However, for specific task qualities, these effects are 4.5 times larger (0.33 sd) and can reach much higher for certain task qualities (1.84 sd). This leads collectives of generalists to perform better on tasks that involve generating, choosing, and coordinating, while collectives of specialists with a few generalist mediators perform better on tasks that involve negotiating. Rationality bounds then moderate these relationships. At loose bounds, specialists outperform generalists through more effective sampling of high-dimensional decision spaces. At tight bounds, generalists outperform specialists through better gradient estimation. A fundamental trade-off between performance and convergence speed emerges at moderate bounds. These findings suggest that effective multi-agent design could benefit from matching interpretive networks to both task demands and agentsâ computational limits, with likely implications for the efficiency and energy costs of multi-agent systems. Keywords: collective intelligence, multi-agent systems, bounded rationality, network science, distributed problem-solving, artificial general intelligence Corresponding author: John Meluso E-mail address: jam627@cornell.edu This document is licensed under Creative Commons C BY 4.0. Introduction Collective intelligence in systems of artificial agents is a grow- ing scientific frontier with significant implications for problem- solving across domains [1,2]. These multi-agent systems demon- strate remarkable capabilities in strategy games [3], scientific discovery [4], and spatial navigation [5]. Their capabilities hold potential to assist humans with challenging medical decisions [6], engineering optimization [7,8], collective governance [9,10], and many other complex tasks. At the same time, multi-agent systems are substantial contributors to growing datacenter re- source consumption and carbon emissions [11â15]. Together, these pros and cons make it essential to answer: How should we design multi-agent systems to take advantage of their benefits while mitigating their drawbacks? Answering such questions requires not just prescriptive engineering of solutions but a descriptive science of artificial systems [5,16]. This paper systematically in- vestigates two understudied design qualities of artificial systems: agent interpretive abilities [17,18] and bounded rationality [19]. The first underexplored quality of these systems is the spec- trum of agent interpretive abilities (Fig. 1) [17,18]. Interpretive abilities enable agents to encode and decode messages [20]. In turn, this allows agents to form mental models of others because Interpretive Abilities Narrow Broad (Generalists) (Specialists) Can Interpret Most Actions Can Interpret Few Actions Figure 1. Interpretive abilities correspond to network ties. Agents can have many different interpretive abilities. Varying how many in- terpretive abilities an agent has implies a conceptual spectrum, from a narrow set of abilities (can interpret few actions, like specialists) to a broad set (can interpret varied actions, like generalists). With respect to a group, having greater interpretive abilities corresponds to more network ties while having fewer corresponds to fewer network ties. Creative Commons C BY 4.0June 23, 2026Artificial collectives of specialists and generalists excel at different tasks1â10 arXiv:2606.20877v1 [cs.MA] 18 Jun 2026 Artificial collectives of specialists and generalists excel at different tasksMeluso et al. Multi-Agent SystemsSimulations Analyses 30 Collective Tasks that map decisions to performance 4 Task Qualities Searchable decision variable fraction per model time step: Unbounded (100%) Loose (10%) Moderate (1%) Tight (0.1%) 18 Interpretive Networks CompleteRandomPreferential Attachment Small World e.g. 4 Group Sizes 4 Rationality Bounds 24168 250 Trials per Task Q1. How do interpretive abilities affect performance? Q2. How does bounded rationality moderate interpretive network effects? Generate Choose Coordinate Negotiate Difficulty Figure 2. Methodological overview. We systematically varied multi-agent system design properties including group size (2, 4, 8, 16 agents), interpretive networks (18 topologies including complete, random, small-world, and preferential attachment graphs), and agent rationality bounds (searchable fraction of decision variable domain of±0.1%,1%,10%,100%per time step). Each multi-agent system sought to maximize performance on one of 30 collective tasks (mathematical objective functions mapping states to performance). Task qualities varied significantly as measured by 4 difficulty measures (generating useful new solutions, choosing the best option, coordinating decisions, and negotiating with competing objectives). We measured convergence performance across 250 trials per task. Analysis addressed two research questions: (Q1) What interpretive network properties affect performance across task types? We used linear regressions to quantify network property and task quality effects. (Q2) How does bounded rationality moderate interpretive network effects? We compared performance differences across network densities for systems with varying rationality bounds. Equal computational resources (32 decision variables) were maintained across all group sizes. agents can productively understand othersâ actions and share information in interpretable ways [21,22]. However, agent inter- pretive abilities vary substantially. Agents with narrow interpre- tive abilities can interpret few othersâ actions and are themselves interpretable by few. In contrast, agents with broad interpretive abilities can interpret many othersâ actions and can share infor- mation to make themselves interpretable by many. This range of interpretive abilities corresponds to networks among agents in which ties describe who can effectively interpret whom [23], just as social perceptiveness shapes ties in human groups [24,25]. Such abilities also map onto the specialist-generalist spectrum of skills studied in human collective intelligence research [26,27]. Specialists develop expertise in a narrow set of skills at the cost of interpretive breadth, while generalists develop a broad set of interpretive abilities but less depth in any domain. Google DeepMindâs AlphaStar exemplifies this interpretive spectrum [28]. Using about 900 co-trained agents, the system achieved grandmaster-level performance in StarCraft I, a chal- lenging strategy game. Training patterns shaped each agentâs in- terpretive abilities: specialized âexploiterâ agents trained against limited strategies, learning to interpret only those approaches and thus developing narrow interpretive abilities. Meanwhile, generalist âmainâ agents trained against diverse opponents, de- veloping broad interpretive abilities that achieved grandmaster- level performance. AlphaStarâs success depended on collabora- tion among agents with different interpretive abilitiesânarrow exploiters and broad main agents. Such differences in interpre- tive abilities correspond to networks among agents where ties represent who can effectively interpret whom. Similar inter- pretive heterogeneity appears throughout multi-agent systems research [4, 29â31]. Network topologies influence group performance in human collective intelligence research, with a groupâs balance of in- terpretive abilities shaping effectiveness [25,32â38]. Network topologies also influence performance on tasks with different qualities such as generate, choose, negotiate, and coordinate tasks [39â42]. Whether these dynamics of human groups trans- fer to artificial systems remains a largely open question in col- lective intelligence research, though. Despite growing deployment of multi-agent systems, we lack scientific understanding of how interpretive networks between artificial agents influence collective performance across different tasks. Current approaches to multi-agent systems use various broadcasting and targeted interaction topologies [17,23] but seldom study the appropriateness of interpretive network topolo- gies for specific tasks. Additionally, current approaches rarely ex- 2â10 Meluso et al.Artificial collectives of specialists and generalists excel at different tasks amine how network design should be informed by agentsâ inher- ent computational limitsâwhat Herbert Simon called bounded rationality [19]. This leaves two critical knowledge gaps. First, while the effects of network topology on group performance are well-established in human research, we lack descriptive evidence of how inter- pretive network topologies affect collective performance among artificial agents in controlled settings. Second, and more fun- damentally, we have limited understanding of how interaction between network topologies and agentsâ computational limits affect collective performance. These gaps pose fundamental scientific challenges because, just as insufficient knowledge of physics would limit effective bridge design, lacking a science of artificial collective intelligence limits our ability to design effec- tive multi-agent systems. Our research addresses two exploratory questions: 1. What interpretive network properties most improve multi- agent group performance, both overall and for specific task qualities? 2. How does bounded rationality moderate the relationship between interpretive network density and collective perfor- mance? To answer these questions, we designed an experiment that systematically varies key properties of multi-agent systems, mea- suring their performance across a battery of 30 varied mathemat- ical representations of tasks (see Fig. 2). Instead of using specific AI implementations (like particular large language models or other machine learning implementations), we model agents ab- stractly as optimizers that iteratively search constrained state spaces. This abstraction captures how researchers and practi- tioners predominantly design agentic systems to maximize per- formance by searching state spaces within computational limits [43â45] (see Methods). Tasks are represented by objective func- tions that map agentsâ collective states to performance across high-dimensional landscapes. The landscapes vary along four non-mutually exclusive qualities derived from organizational psychology [38,39,46]: generate, choose, coordinate, and nego- tiate, each capturing a distinct source of collective difficulty (see Methods). Our results show that collective performance depends on ap- propriately matching interpretive network topology with both task qualities and rationality bounds. Task qualities produce effects 4.5 times larger than network properties alone. Inter- pretive network density (the proportion of agent pairs that can effectively understand one another), decentralization, and path length can have effects 20 times larger (nearly 2 standard devia- tions) but vary significantly by task. We also find that bounded rationality has mixed effects, favoring groups of specialists in some cases (loosely bounded, +6.7%) and generalists in others (tightly bounded, +7.3%). Between the extremes, moderately bounded searches pose a trade-off between speed and quality with conditions that will require practitioner testing and greater scientific investigation. Together, these findings suggest that efficient multi-agent system designs must be specific to task qualities and agent computational limits, rather than universal. Results Different tasks benefit from different interpretive network properties To study our first research question, we examined how network properties influence group performance both overall and for specific task qualities (Fig. 3). Like others [38,41], we used mul- tivariate ordinary least squares regressions with robust standard errors to isolate the independent performance effects of each network property while controlling for task qualities and agent rationality bounds (see Methods). Fig. 3A shows these effects as a heatmap: the first column shows each network propertyâs average performance (the main effects) across all task types. Subsequent columns show the combined performance of each network property and task quality (task-conditional effects). For brevity, we refer to âgenerate tasksâ to mean tasks with high gen- erate difficulty; similarly for choose, coordinate, and negotiate tasks. Fig. 3B illustrates how specific network topologies per- form on coordinate and negotiate tasks as functions of network density, decentralization, and path length. All values represent statistically significant effects (í < 0.05) in standard deviations above or below the dataset mean. The main performance effects of network properties are uni- formly small with an average magnitude of 0.07 standard devia- tions. The main effect of network density is only +0.07 standard deviations, while decentralization (mean eigenvector central- ity) shows -0.13 standard deviations, and average shortest path length shows -0.18 standard deviations. These modest main ef- fects suggest that network properties do not uniformly improve or impair artificial collective performance across all task qual- ities, similar to findings on human groups [25]. Rather, their influence depends on the type of task being solved. This task-dependence becomes evident when examining con- ditional effects. The average magnitude of the task conditional effects is 0.33 standard deviations, 4.5 times larger than for the main effects. For generate tasksâwhich require finding novel states across complex landscapesânetwork density provides a substantial benefit of +0.87 standard deviations, while decentral- ization contributes +0.77 standard deviations. Shorter average path lengths also help (effect of -0.26 on path length, meaning shorter paths improve performance). 3â10 Artificial collectives of specialists and generalists excel at different tasksMeluso et al. -2.0-1.00.01.02.0 Effect Size (St. Devs. of Performance) Main EffectGenerateChooseCoordinateNegotiate Path Length (avg. path len.) Decentralization (eig. cent. avg.) Max. Distance (diameter) Neighbor Connect. (nnd avg.) Intermediaries (bet. cent. avg.) Varied Intermediaries (bet. cent. std.) Neighbor Connect. Var. (nnd std.) Centralization (eig. cent. std.) Connect. Variation (deg. cent. std.) Triangle Density (clust. coeff.) Connection Homophily (deg. assort.) Network Density Network Property (with metric) -0.18-0.26-0.44-0.940.24 -0.130.771.101.84-0.68 -0.11-0.14-0.28-0.140.05 -0.07-0.010.070.12-0.18 -0.05-0.080.020.15-0.10 -0.05-0.28-0.25-0.270.05 -0.05-0.12-0.09-0.01-0.07 -0.05-0.32-0.35-0.520.17 -0.020.040.050.14-0.10 0.040.190.180.33-0.13 0.070.200.180.150.03 0.070.870.971.60-0.55 Task Conditional Effects 00.250.50.751 0.30 0.31 0.32 0.33 0.34 0.35 0.36 Decentralization (eig. cent. avg.) Coordinate Random (high probab.) Hypercube 00.250.50.751 Negotiate Random (low probab.) Complete 00.250.50.751 Network Density 1.0 1.5 2.0 2.5 3.0 Path Length (avg. path len.) Star 00.250.50.751 Network Density Tree Random (high probab.) -1.00-0.75-0.50-0.250.000.250.500.751.00 Performs Worse Performs Better (task conditional effects vs. complete graph reference) Effect Size (St. Devs. of Performance) AB Figure 3. Performance effects of interpretive network properties. (A) A heatmap showing the performance of different network properties in terms of standard deviations above or below the mean of all simulation runs. Values show network property effects, both across tasks qualities (the main effects) and limited to specific task qualities (the task conditional effects). (B) The performance of specific network topologies corroborate the large effects seen for decentralized and dense networks with short paths between nodes (task conditional effects for teams ofí = 8). All values shown are statistically significant multivariate regression results (í < 0.05). Higher graph density corresponds to greater average agent interpretive abilities. Choose and coordinate tasks exhibit similar network property benefits but with even larger effect magnitudes. For choose tasks, which involve selecting among multiple alternatives, density provides +0.97 standard deviations benefit and decentralization adds +1.10 standard deviations. Coordinate tasks, which re- quire synchronizing interdependent actions, show the strongest effects: density contributes +1.60 standard deviations and de- centralization contributes +1.84 standard deviationsâ23 times larger and 14 times larger than their respective main effects. Shorter path lengths benefit both task types as well, with ef- fects of -0.44 and -0.94 standard deviations respectively. These patterns indicate that generate, choose, and coordinate task per- formances improve when interpretive abilities are abundant and widely distributed, as they are for decentralized interpretive networks with efficient communication paths. Negotiate tasks reveal a contrasting pattern. In these tasks, agents must balance conflicting assessments of performance. Unlike other task qualities, negotiate tasks show negative effects for density (-0.55 standard deviations) and decentralization (- 0.68 standard deviations), while benefiting from longer average paths (+0.24 standard deviations). This suggests that negotiate tasks benefit from sparser, more centralized topologies in which a few generalists mediate between specialists. Fig. 3B corroborates these patterns through specific network topologies. For coordinate tasks (left column), dense, decentral- ized graphs with short path lengths fall further to the right (like the complete, hypercube, and high-probability random graphs). Most of these graphs show strong positive performance effects (large green circles). In contrast, negotiate tasks (right column) show a different pattern: sparser networks and those with longer pathsâincluding tree structures and low-probability random graphsâappear in regions corresponding to better performance. These patterns also replicate across other optimizers with vary- ing effect magnitudes (SI Appendix S1.1). A noteworthy exception to these trends is the complete graph which, despite its high density, decentralization, and short paths, 4â10 Meluso et al.Artificial collectives of specialists and generalists excel at different tasks underperforms on coordinate tasks and overperforms on nego- tiate tasks. This aligns with convergent findings in collective human intelligence on the inferior performance on efficient net- works [25,40,47,48]. However, several mechanisms related to information diffusion speed [40] and search interference [38] might cause this phenomenon. We revisit these mechanisms in the following section when examining how rationality bounds moderate network effects. Task-specific patterns likely emerge because collective tasks are network games [38,49] that groups âplayâ by collectively searching a state space. A groupâs task performance depends on the interdependent actions of individuals which correspond to network topologies. Tasks aided by state diversity benefit from sparser connectivity that preserves independent search, while tasks emphasizing information aggregation benefit from denser connectivity. This matches evidence in human groups that net- work topology has opposite effects on information diffusion and exploration [40,50]. Coordinate tasks illustrate the aggregation case: dense interpretive ties help agents interpret neighborsâ ac- tions and make productive adjustments. Negotiate tasks then illustrate the diversity case: dense interpretive ties may create premature convergence toward suboptimal compromise posi- tions before agents can adequately explore mutually-beneficial alternatives. Loose bounds favor specialists, tight bounds favor general- ists To answer our second research question, we examined how net- work density affects multi-agent performance across different bounds of rationality (Fig. 4). These rationality bounds con- strain agentsâ abilities to search the state space by limiting the domain of each decision variable per time step, thereby modeling Simonâs concept of bounded rationality [19]. We refer to these three regimes as loosely bounded (±10%), moderately bounded (±1%), and tightly bounded (±0.1%) search conditions. We cal- culated mean convergence performance and convergence time for each network topology and rationality bound. Examining convergence outcomes across rationality bounds shows that the relationship between network topology, convergence timing, and performance quality differs markedly across rationality bounds with three distinct regimes. First, when rationality bounds are loose (±10%), specialist- focused topologies substantially outperform generalist topolo- gies on both performance and speed (Fig. 4A). The empty graphâ representing pure specialists with no interpretive connectionsâ achieves 6.7% higher performance than the complete graph (0.956 vs. 0.896) while converging 23% faster (11 vs. 14 steps). Between these density extrema, topology averages cluster near the performance of the complete graph, with lower densities slightly outperforming higher densities. This relationship reverses when rationality is tightly bounded (±0.1%; Fig. 4C). Here, generalist-focused topologies excel: the complete graph outperforms the empty graph by 7.3% (0.674 11121314 Convergence Time (steps) 0.89 0.90 0.91 0.92 0.93 0.94 0.95 0.96 Convergence Performance empty complete star random (low probab.) A Performs Better Loosely Bounded (±10.0%) 60657075 Convergence Time (steps) 0.88 0.90 0.92 0.94 empty complete star random (low probab.) B Moderately Bounded (±1.0%) 92949698 Convergence Time (steps) 0.63 0.64 0.65 0.66 0.67 0.68 empty complete star wheel C Tightly Bounded (±0.1%) 0.0 0.2 0.4 0.6 0.8 1.0 Network Density Figure 4. Network property effects on convergence timing and performance vary across rationality bounds. Each panel shows convergence performance versus convergence time for teams of 8 agents, with points colored by network density (gray = specialists, green = generalists). (A) At loose rationality bounds (±10%), specialists achieve both faster convergence and superior performance. (B) At moderate bounds (±1%), a Pareto frontier emerges: specialists converge slower but achieve better performance. (C) At tightly bounded rationality (±0.1%), this reverses: generalists achieve both fastest convergence and best performance. Groups of 2, 4, and 16 agents show similar results. Hence, optimal network topologies may depend on the relative capacity of agents to search a state space. 5â10 Artificial collectives of specialists and generalists excel at different tasksMeluso et al. vs. 0.628) while converging 6% faster (91 vs. 97 steps). At this bound, many topologies approach or reach the 100-step time limit before fully converging, meaning these performance values represent the best states achieved within available computational steps rather than true convergence values. Still, topologies with density as low as 0.5 (such as wheel graphs) achieve 99.7% of the complete graphâs performance with half the interpretive ties, though with more time required. The moderately bounded case (±1%) exhibits a qualitatively distinct pattern (Fig. 4B): a speed-quality trade-off. Specialist networks converge 27% slower than generalist networks (76 vs. 59 steps) but achieve 7% better performance (0.937 vs. 0.873). The empty graph stands as an outlier, achieving the highest per- formance but at substantially greater time cost than all other configurations, which cluster more closely to the complete graph in both timing and performance. This creates a discontinuous Pareto frontier where design decisions involve a choice between topologies of expensive-but-excellent specialists and faster-but- adequate generalists, rather than a smooth continuum of inter- mediate options with increasing density. What drives these outcomes mechanistically? Rationality bounds determine the local landscape each agent optimizes over during each turn. At tight bounds (±0.1%), agents effectively optimize over smooth, convex regions. Each additional inter- pretable neighbor improves an agentâs gradient estimation by incorporating more dimensions of the groupâs objective func- tion. This provides more accurate directional information for hill-climbing, so generalist topologies (with dense interpretive ties) outperform specialist topologies (with sparse interpretive ties) because agents optimize faster. With loose bounds (±10%), agents often optimize over rugged, non-convex regions contain- ing many local optima. Each additional interpretable neigh- bor adds decision variablesâand thus dimensions [51]âto an agentâs optimization task, creating exponentially larger deci- sion spaces. Given the fixed computational budgets imposed by bounded rationality, these high-dimensional spaces become increasingly difficult to sample. Specialist topologies therefore outperform generalist topologies because agents sample more effectively in lower-dimensional spaces. These patterns replicate with two other optimizers (L-BFGS-B & Dual Annealing, SI Appendix S1.2). However, Random Walk agents show weaker effects at tight bounds and reversed patterns at loose bounds (also SI Appendix S1.2). Unlike the other opti- mizers, which search strategically, Random Walk agents evaluate single random samples. So, groups of Random Walk agents with more interpretive ties appear to improve their evaluation accu- racy without creating the high-dimensional sampling challenges that strategic optimizers face. This suggests that our mechanistic effects may require goal-directed optimization under bounded rationality instead of random, undirected optimization. This progression extends both Simonâs bounded rationality concept [19] and Marchâs exploration-exploitation framework [52] into multi-agent systems, suggesting that appropriate net- work topologies can partially compensate for limited computa- tional capacity in complex decision spaces. Discussion An interpretive network describes which individuals in a group can understand each otherâs actions. In these networks, individu- als with few interpretive abilities correspond to specialists while those with many function as generalists. Our results suggest that topological properties of these networks shape the collec- tive performance of multi-agent systems depending on both task qualities and agentsâ rationality bounds. Across task qualities, interpretive network properties have small performance effects (average 0.07 standard deviations). But with respect to specific task qualities, network effects are 4.5 times larger (average 0.33 standard deviations). Some tasks fa- vor decentralized topologies (generate, choose, coordinate tasks) while others favor centralized ones (negotiate tasks). Bounded rationality mediates these relationships, though. When agents have loose rationality bounds relative to a decision space, spe- cialist topologies perform best because they sample state spaces more effectively; with tight bounds, generalists outperform be- cause they optimize faster. A speed-quality trade-off lies between these extremes. These findings support a nuanced theory of collective artifi- cial intelligence. The benefits of human coordination [53,54] rely upon how well social perceptiveness topologically balances information efficiency and information diversity [25,40,42,55]. Similarly, our findings show that for most collaborative tasks (generate, choose, and coordinate), agentic systems benefit from interpretive abilities that produce dense, decentralized networks which also balance efficiency and diversity [38], though negoti- ate tasks benefit from centralized topologies with network inter- mediaries instead. Our results also yield practical recommendations for crafting efficient multi-agent systems across application domains. Cur- rent approaches to multi-agent system design often select agent abilities and network topologies without systematically consid- ering how they interact with each other, with task qualities, or with agent rationality bounds [17,23]. Our findings suggest this approach risks substantial performance losses. Several design principles based on interpretive networks may overcome these challenges. Like human multidisciplinary teams [38], dense topologies where most agents have broad interpretive abilities provide robust performance, particularly across generate, choose, and coordinate tasks. Centralized configurations with a few 6â10 Meluso et al.Artificial collectives of specialists and generalists excel at different tasks coordinating agents become preferable when tasks require rec- onciling competing objectives (negotiate tasks). AlphaStarâs centralized league topology is consistent with these recommen- dations [28]. Generalist main agents trained widely against specialized exploiters, and the system achieved grandmaster- level performance. However, StarCraft itself involves generate, choose, and negotiate qualities, underscoring the difficulty of designing topologies for single task qualities. Empirical studies of LLM multi-agent systems also reach compatible conclusions, finding that coordination benefits depend on task structure and degrade when individual agents are already capable [56]. Regarding agent rationality bounds: when agents can thor- oughly search decision spaces (loose rationality bounds), sparse interpretive topologies (such as star graphs) may help agents sample state spaces more thoroughly while preventing agents from interfering with each otherâs optimization processes. Con- versely, when individual compute is severely constrained relative to the task (tight rationality bounds), fully connected topologies may be unnecessary. Moderately connected topologies (density â„ 0.5, like a wheel graph) achieve near-optimal performance with substantially fewer interpretive ties (Fig. 4C). At moderate rationality bounds, though, a fundamental trade-off emerges between performance and convergence time. System designers must weigh these competing objectives based on their specific requirements and constraints. In terms of computational efficiency, our results suggest sig- nificant opportunities to optimize resource use and performance through appropriate multi-agent topologies. Pairing topologies with task qualities could yield large gains. Decentralized sys- tems show nearly 2 standard deviation improvements on co- ordinate tasks, while specialist topologies achieve 6.7% better performance in 23% less time than generalist topologies at loose rationality bounds. These improvements hold potential to re- duce societal needs for compute time, for more powerful agents, and for greater numbers of agents. Given the substantial energy and computational costs of multi-agent systems [11â14], under- standing which interpretive network properties provide benefits and costs should be an engineering priority. Furthermore, these findings hold implications for human-AI collaboration, a major focus of current research [1]. To date, the many complications of human-computer interaction mean mixed human-AI groups often underperform the better of the two working alone [57]. The interpretive network topologies that benefit artificial collectives may inform optimal configurations for hybrid teams as well, potentially bridging the performance characteristics of human and artificial collective intelligence. Interpretive networks could provide a common language for designing mixed topologies of humans and artificial agents, posi- tioning each according to their respective interpretive strengths. Finally, our methodology enables future research on coordi- nation processes [54], dynamic network adaptation [58], and heterogeneous group composition [59] in artificial collectives. By bridging human and artificial collective intelligence research, this work advances our understanding of how interpretive net- work properties shape emergent intelligence across organiza- tional scales and contexts. Several limitations warrant future investigation. Our model employed static interpretive network topologies, while adaptive networks that evolve based on performance or task qualities often yield additional benefits [58,60]. We also explored only homogeneous agent optimization abilities, whereas heteroge- neous groups with differentiated specializations and rationality bounds are likely to provide further insights and better perfor- mance [26,59]. While our findings point to interpretive network properties as key determinants of performance, factors like ini- tial condition sensitivity and system size likely contribute to these patterns, meriting further investigation. Although par- tially validated by large language model findings [56], validation with other types of agents and in specific domains like scien- tific discovery, systems engineering, and autonomous systems represents important next steps toward practical application. Methods We designed a systematic experiment with multi-agent systems to inves- tigate how agent interpretive abilities, rationality bounds, and task quali- ties influence collective performance. To this end, we carefully designed properties of our multi-agent systems, simulation experiments, and sta- tistical analyses (Fig. 2). Code accompanying this paper can be found at https://github.com/meluso/mas-interpretive-networks-code/. Data can be found on Zenodo at https://doi.org/10.5281/zenodo.19682737. Multi-Agent Systems For our multi-agent systems, we made design decisions at the group level and the agent level. Groups At the group level, we represented collective configurations of inter- pretive abilities through network topologies and objective functions. Network topologies ranged from completely disconnected agents with no interpretive abilities (empty graphs) to fully connected agents with the ability to interpret all other actions (complete graphs), plus 16 in- termediate topologies including small world, random, and preferential attachment networks (SI Appendix S2.1). Including many topologies en- abled us to significantly vary values of network properties with known effects elsewhere and therefore to measure their relative performance effects here (see SI Appendix S2.2 for network properties and corre- sponding metrics). We also varied group sizes (í = 2, 4, 8, 16 agents) to control for effects of group size on network properties and individ- ual degrees of freedom while holding system computational resources constant. This was done by assigning each agent a set of variables to 7â10 Artificial collectives of specialists and generalists excel at different tasksMeluso et al. optimize that is inverse in size to the group size (íŁ = 16, 8, 4, or 2 variables, respectively), totaling 32 variables per system in all cases. Multi-agent systems were assigned task objective functionsâthat define collective performance mappings. Each agentíoperates with a neighborhood objective functioní í that matchesââs mathematical form but incorporates only information the agent can interpret through its network connections: its own assigned variables plus those of its connected neighbors. For example, if the group taskâaverages all agentsâ variables, then each agent evaluatesí í as the average of only its own variables and its neighborsâ variables. This design ensures that interpretive network topology directly determines which information agents use to evaluate states. Agents At the agent level, we abstractly modeled artificial agents as optimizers that iteratively search constrained state spaces. We use classical opti- mization algorithms as scientific representations of artificial agentsâ specifically the Nelder-Mead simplex algorithm (shown in the main Results) plus the L-BFGS-B algorithm, a simulated annealing algorithm, and a random walk algorithm for comparison (SI Appendix S1) [61]. We chose this approach because it captures a fundamental prop- erty shared across many multi-agent systems while providing scientific tractability. Artificial agents generally must search for high-quality states within computational and design constraintsâwhether through gradient descent, beam search, policy networks, or tree traversal, they are bounded by finite resources and processing capacities [19,62,63]. Classical optimizers model this functional constraint directly: agents evaluate candidate states and select among them within defined limits. This abstraction does not require modeling the specific implementa- tion details of particular AI architectures, enabling us to isolate the effects of interpretive network properties on collective performance in a scientifically rigorous way. In our multi-agent algorithm, agents iteratively optimize their as- signed variables at each time step, incorporating information from network neighbors before producing new states (see Algorithm). This design enables us to study how network properties mediate collective search under bounded rationality while maintaining relevance to con- temporary and future agentic systems that operate under computational bounds. We operationalized bounded rationality through four rationality bounds (í = ±0.1%,±1%,±10%,±100%) representing the fraction of each decision variableâs domain that an agent can search per time step. These bounds span a range of realistic constraint levels: Bounds of ±0.1%model tightly constrained agents that search within a small frac- tion of the decision space. Bounds of±1%and±10%model moderately and loosely constrained agents, respectively. Bounds of±100%model unbounded agents, unconstrained in their ability to search the full de- cision space. This parameterization facilitates systematic investigation of how collective performance depends on both network topology and the degree of individual rationality bounds. The mechanisms we identifyâhow interpretive abilities shape in- formation access and influences collective search behaviorâapply broadly to systems where agents evaluate state quality within con- straints, though specific quantitative effects may vary with implemen- tation details. Algorithm During each simulation trial, agents iteratively optimize their assigned variables within their rationality bounds. The procedure follows this sequence: 1. Each agentí â íŒobserves the current state (variable valuesí„ í ) of connected neighborsí â íœ í as determined by the interpretive network topology. 2. Agentíuses its optimizer to generate a candidate stateí„ âČ í within its rationality bounds, based on its objective function í í (í„ í , í„ í íâíœ í ). 3.After all agents generate candidate states simultaneously, agents calculate a new objective evaluationí âČ í using the updated states í„ âČ í íâíœ í of all connected neighbors. 4. The group objectiveâis computed using all agentsâ new states (í„ âČ 1 , ... , í„ âČ í ), so â(í„ âČ í íâíŒ ). 5.The process repeats until convergence (change inâ < 10 â4 for 3 consecutive steps) or maximum iterations (100 steps). This synchronous update procedure ensures that all agents optimize simultaneously based on the information they can interpret as defined. The formal algorithmic procedure is detailed in SI Appendix S4. Simulation Experiment To systematically examine the effects of different properties, we ran simulation experiments with multi-agent systems, varying the group size, network type, and rationality bounds as described above. Multi- agent systems were assigned 30 mathematically defined optimization tasks that vary in difficulty along four task quality dimensions from collective intelligence research [38, 39, 54, 55]: âą Generate: Finding novel high-quality states âą Choose: Selecting among alternatives âą Coordinate: Synchronizing interdependent actions âą Negotiate: Balancing competing objectives Tasks range from simple functions like averaging to complex opti- mization landscapes like the Ackley function (shown in Fig. 2). All objective functions are normalized to[0, 1]to enable comparison across tasks, with 1 representing optimal performance. Detailed task defini- tions and difficulty measurement methods are provided in SI Appendix S3. The experimental design crossed 4 group sizesĂ18 network topolo- giesĂ30 tasksĂ4 rationality bounds, with 250 independent trials per parameter combination, yielding about 1.7 million simulation trials. Statistical Analysis We estimated the effects of network properties and topologies on group performance through ordinary least squares regressions with robust standard errors (type HC2). To isolate the independent contributions of network properties from the confounding effects of multicollinearity (see SI Appendix S5 for correlation analysis), we conducted separate regression analyses for individual network properties (each with a cor- responding network metric) and for specific network topologies. We used data that includes all 13 network metrics, which required excluding disconnected topologies where path-based metrics (shortest 8â10 Meluso et al.Artificial collectives of specialists and generalists excel at different tasks path length, diameter) are undefined. This filtering retained topolo- gies with valid values for all network properties, providing the most comprehensive characterization of network property effects. Prior to regression, all continuous variables were normalized using MinMax scaling to the range [0,1]. Each regression included controls for agent rationality bounds (log 10 (agent_steplim)) and the four task qualities (generate, choose, negotiate, coordinate). We conducted two variants of each regression: one with main effects only, and one that included interaction terms between the primary predictor and each task quality (task-conditional effects). This approach allowed us to estimate both the overall effect of each network property or network topology and how that effect varies across different task types. For network property analyses, we ran separate regressions with each of 13 network metrics as the primary predictor: degree centrality (mean and standard deviation), betweenness centrality (mean and stan- dard deviation), eigenvector centrality (mean and standard deviation), nearest neighbor degree (mean and standard deviation), clustering co- efficient, density, degree assortativity, shortest path length (mean), and graph diameter. These analyses used data from all group sizes (2, 4, 8, and 16 agents) with normalized group size included as a control variable. Model fitting used scikit-learnâs LinearRegression [64], with robust standard errors calculated using the HC2 estimator [65]. For topology-specific analyses, we ran separate regressions for each of the 18 network topologies, using dummy variables to indicate topology with the complete graph as the reference category. The normalized variables enable interpretation of regression coef- ficients as standardized effects: a coefficient of 1.0 indicates that a one standard deviation increase in the predictor corresponds to a one standard deviation change in convergence performance. The intercept represents expected performance when all predictors are at their mean values. For visualization purposes, we filtered results of Fig. 3A to include only statistically significant effects (í < 0.05). Task-conditional ef- fects shown in the figure represent the cumulative effect of a network property on a specific task quality, calculated as the sum of the net- work propertyâs main effect, the task qualityâs main effect, and their interaction effect, all from the model including interactions. This cumu- lative effect captures the total predicted impact of the network property when that specific task quality is at its maximum value while other task qualities are at their mean values. Analyses for Fig. 3B used only groups of size 8. We did this to help visualize the relationships between performance and three network properties (density, decentralization, and average shortest path length) while holding these network metrics constant for each topology. References [1] Riedl C, De Cremer D (2025) AI for collective intelligence. Collec- tive Intelligence 4(2):26339137251328909. [2]Sehwag UM, McAvoy A, Plotkin JB (2025) Collective artificial in- telligence and evolutionary dynamics. Proceedings of the National Academy of Sciences 122(25):e2505860122. [3]Silver D, et al. (2018) A general reinforcement learning algorithm that masters chess, shogi, and Go through self-play. Science 362(6419):1140â1144. [4]Strieth-Kalthoff F, et al. (2024) Delocalized, asynchronous, closed-loop discovery of organic laser emitters. Science 384(6697):eadk9227. [5]Rahwan I, et al. (2019) Machine behaviour. Nature 568(7753):477â 486. [6] Kim Y, et al. (2024) A Demonstration of Adaptive Collaboration of Large Language Models for Medical Decision-Making. [7]Chen Q, Heydari B (2024) The SoS conductor: Orchestrating re- sources with iterative agent-based reinforcement learning. Systems Engineering p. sys.21747. [8]Nitti A, de Tullio MD, Federico I, Carbone G (2025) A collective intelligence model for swarm robotics applications. Nature Com- munications 16(1):6572. [9]Barfuss W, et al. (2025) Collective cooperative intelligence. Pro- ceedings of the National Academy of Sciences 122(25):e2319948121. [10]Tacchetti A, et al. (2025) Deep mechanism design: Learning so- cial and economic policies for human benefit. Proceedings of the National Academy of Sciences 122(25):e2319949121. [11] (2024) Big techâs great AI power grab. The Economist. [12] IEA (2024) Electricity 2024, (IEA, Paris), Technical report. [13] Leppert R (2025) What we know about energy use at U.S. data centers amid the AI boom. [14]IEA (2025) Energy and AI, (IEA, Paris), World Energy Outlook Special Report. [15]Aczel M, et al. (2026) The Environmental Cost of AIâs Energy Use: Carbon, Water and Land Footprints, (United Nations University Institute for Water, Environment and Health (UNU INWEH)), Technical report. [16] Simon HA (1996) The Sciences of the Artificial. (MIT Press, Cam- bridge, MA). [17]Gronauer S, Diepold K (2022) Multi-agent deep reinforcement learning: A survey. Artif Intell Rev 55(2):895â943. [18]Oroojlooy A, Hajinezhad D (2023) A review of cooperative multi- agent deep reinforcement learning. Appl Intell 53(11):13677â 13722. [19] Simon HA (1957) Models of Man; Social and Rational, Models of Man; Social and Rational. (Wiley, Oxford, England), p. xiv, 287. [20]Shannon CE (1948) The Mathematical Theory of Communication. (University of Illinois Press) Vol. 27, p. 656. [21]Westby S, Riedl C (2023) Collective intelligence in human-AI teams: A bayesian theory of mind approach. Proceedings of the AAAI Conference on Artificial Intelligence 37(5):6119â6127. [22]Kelley S, Cremer D, Riedl C (2025) Personalized AI Scaffolds Synergistic Multi-Turn Collaboration in Creative Work. [23]Horling B, Lesser V (2004) A survey of multi-agent organizational paradigms. The Knowledge Engineering Review 19(4):281â316. [24]Amelkin V, Askarisichani O, Kim YJ, Malone TW, Singh AK (2018) Dynamics of collective performance in collaboration networks. PLOS ONE 13(10):e0204547. [25] Centola D (2022) The network science of collective intelligence. Trends in Cognitive Sciences 26(11):923â941. 9â10 Artificial collectives of specialists and generalists excel at different tasksMeluso et al. [26] Hong L, Page SE (2004) Groups of diverse problem solvers can outperform groups of high-ability problem solvers. Proceedings of the National Academy of Sciences of the United States of America 101(46):16385 LPâ16389. [27]Goldstone RL, Andrade-Lotero EJ, Hawkins RD, Roberts ME (2024) The Emergence of Specialized Roles Within Groups. Topics in Cognitive Science 16(2):257â281. [28]Vinyals O, et al. (2019) Grandmaster level in StarCraft I using multi-agent reinforcement learning. Nature 575(7782):350â354. [29]Tan M (1993) Multi-agent reinforcement learning: Independent vs. cooperative agents in Proceedings of the Tenth International Conference on Machine Learning. p. 330â337. [30] Stone P, Veloso M (2000) Multiagent Systems: A Survey from a Machine Learning Perspective. Autonomous Robots 8(3):345â383. [31]Shoham Y, Leyton-Brown K (2008) Multiagent Systems: Algorith- mic, Game-Theoretic, and Logical Foundations. (Cambridge Uni- versity Press). [32]Bavelas A (1950) Communication patterns in task-oriented groups. The journal of the acoustical society of America 22(6):725â730. [33]Bavelas A, Hastorf AH, Gross AE, Kite WR (1965) Experiments on the alteration of group structure. Journal of Experimental Social Psychology 1(1):55â70. [34]Grant RM (1996) Toward a knowledge-based theory of the firm. Strategic Management Journal 17(S2):109â122. [35]Bunderson JS, Sutcliffe KM (2002) Comparing Alternative Con- ceptualizations of Functional Diversity in Management Teams: Process and Performance Effects. Academy of Management Jour- nal 45(5):875â893. [36]Rulke DL, Galaskiewicz J (2000) Distribution of knowledge, group network structure, and group performance. Management Science 46(5):612â625. [37]Postrel S (2002) Islands of shared knowledge: Specialization and mutual understanding in problem-solving teams. Organization Science 13(3):303â320. [38]Meluso J, HĂ©bert-Dufresne L (2023) Multidisciplinary learning through collective performance favors decentralization. Proceed- ings of the National Academy of Sciences 120(34):e2303568120. [39]McGrath JE (1984) Groups: Interaction and Performance. (Prentice-Hall Englewood Cliffs, NJ) Vol. 14. [40]Lazer D, Friedman A (2007) The network structure of exploration and exploitation. Administrative Science Quarterly 52(4):667â694. [41]Mason W, Watts DJ (2012) Collaborative learning in networks. PNAS 109(3):764â769. [42]Barkoczi D, Galesic M (2016) Social learning strategies modify the effect of network structure on group performance. Nature Communications 7(1):13109â13109. [43]Russell S, Norvig P (2020) Artificial Intelligence: A Modern Ap- proach. (Pearson, London), 4. ed edition. [44]Silver D, Singh S, Precup D, Sutton RS (2021) Reward is enough. Artificial Intelligence 299:103535. [45]Gershman SJ, Horvitz EJ, Tenenbaum JB (2015) Computational ra- tionality: A converging paradigm for intelligence in brains, minds, and machines. Science 349(6245):273â278. [46] Hu XE, Whiting ME, Gandhi L, Watts DJ, Almaatouq A (2025) The Task Space: An Integrative Framework for Team Research. [47]Derex M, Boyd R (2016) Partial connectivity increases cultural accumulation within groups. Proceedings of the National Academy of Sciences 113(11):2982â2987. [48]Brackbill D, Centola D (2020) Impact of network structure on collective learning: An experimental study in a data science com- petition. PLOS ONE 15(9):e0237978. [49]Galeotti A, Goyal S, Jackson MO, Vega-Redondo F, Yariv L (2010) Network Games. The Review of Economic Studies 77(1):218â244. [50]Shore J, Bernstein E, Lazer D (2015) Facts and figuring: An exper- imental investigation of network structure and performance in information and solution spaces. Organization Science 26(5):1432â 1446. [51]Bellman R (1961) Adaptive Control Processes: A Guided Tour. (Princeton University Press). [52]March JG (1991) Exploration and exploitation in organizational learning. Organization Science 2(1):71â87. [53]Riedl C, Woolley AW (2017) Teams vs. Crowds: A field test of the relative contribution of incentives, member ability, and emergent collaboration to crowd-based problem solving performance. AMD 3(4):382â403. [54]Riedl C, Kim YJ, Gupta P, Malone TW, Woolley AW (2021) Quantifying collective intelligence in human groups. PNAS 118(21):e2005737118. [55]Woolley AW, Chabris CF, Pentland A, Hashmi N, Malone TW (2010) Evidence for a collective intelligence factor in the perfor- mance of human groups. Science 330(6004):686â688. [56] Kim Y, et al. (2025) Towards a Science of Scaling Agent Systems. [57]Vaccaro M, Almaatouq A, Malone T (2024) When combinations of humans and AI are useful: A systematic review and meta-analysis. Nature Human Behaviour p. 1â11. [58] Galesic M, et al. (2023) Beyond collective intelligence: Col- lective adaptation. Journal of The Royal Society Interface 20(200):20220736. [59]Page SE (2019) The Diversity Bonus: How Great Teams Pay off in the Knowledge Economy. (Princeton University Press). [60] Rubenstein M, Cornejo A, Nagpal R (2014) Programmable self- assembly in a thousand-robot swarm. Science 345(6198):795â799. [61]Virtanen P, et al. (2020) SciPy 1.0: Fundamental algorithms for scientific computing in python. Nature Methods 17(3):261â272. [62]Russell SJ, Subramanian D (1994) Provably Bounded-Optimal Agents. Journal of Artificial Intelligence Research 2:575â609. [63]Gottwald S, Braun DA (2019) Systems of Bounded Rational Agents with Information-Theoretic Constraints. Neural Computation 31(2):440â476. [64] Pedregosa F, et al. (2011) Scikit-learn: Machine learning in python. Journal of Machine Learning Research 12:2825â2830. [65] Seabold S, Perktold J (2010) Statsmodels: Econometric and sta- tistical modeling with python in Proceedings of the 9th Python in Science Conference. (Austin, TX), Vol. 57, p. 10â25080. 10â10