Paper deep dive
Not Birds of a Feather: Personality-Based Partner Selection in LLM Agents
Tao Wang, Hsiang-Ling Chiu, Chihang Wei, Yang Xiu, Zhonghao Hou
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 7/23/2026, 1:57:27 AM
Summary
This study investigates whether Big Five personality traits influence partner selection in multi-agent LLM systems when capability is held constant. Using a controlled experiment with neutral and personality-assigned host agents, the research finds that selection is strongly driven by task-stereotype matching (e.g., Openness for creative tasks, Conscientiousness for strategic tasks) rather than homophily. Contrary to human social attraction patterns, LLM agents exhibit non-homophilous selection, preferring complementary personalities over self-similar ones, with significant implications for bias auditing in agent marketplaces.
Entities (12)
Relation Signals (10)
Conscientiousness Archetype â ispreferredfor â Problem Solving
confidence 98% · the conscientious archetype 90-97% of strategic, synthesis, and problem-solving trials
Openness Archetype â ispreferredfor â Creative Ideation
confidence 98% · the open archetype won 100% of creative trials
Conscientiousness Archetype â ispreferredfor â Strategic Planning
confidence 98% · the conscientious archetype 90-97% of strategic, synthesis, and problem-solving trials
LLM agents â exhibits â Non-Homophily
confidence 95% · self-similar partners were selected below chance
LLM agents â exhibitsselectionbehaviorbasedon â Big Five Personality
confidence 95% · Personality-based selection in LLM agents is real, strong, task-stereotyped
Extraversion Archetype â israrelychosen â LLM agents
confidence 95% · the extraverted, agreeable, and balanced archetypes were almost never chosen
Agreeableness Archetype â israrelychosen â LLM agents
confidence 95% · the extraverted, agreeable, and balanced archetypes were almost never chosen
Neuroticism Archetype â ispreferredfor â Analytical Reasoning
confidence 90% · the neurotic archetype 37% of analytical trials
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Multi-agent LLM systems increasingly let one agent choose which other agents to work with, and agents are increasingly given personalities through personas. We test whether Big Five personality alone influences partner selection when capability is explicitly held constant. Host agents chose among six validated candidate archetypes -- five marked high on one trait (openness, conscientiousness, extraversion, agreeableness, neuroticism) plus a balanced control -- presented with randomized names and ordering across five task categories (375 trials). With neutral hosts (Study 1, n=150), selection departed drastically from chance ($\chi^2(5)=325.8$, $p<.001$), following a task-stereotype map: the open archetype won 100% of creative trials, the conscientious archetype 90-97% of strategic, synthesis, and problem-solving trials, and the neurotic archetype 37% of analytical trials (Cramer's V=.74); the extraverted, agreeable, and balanced archetypes were almost never chosen, although human meta-analyses identify team agreeableness as among the strongest personality predictors of team performance. With personality-assigned hosts (Study 2, n=225), and contrary to human similarity-attraction, self-similar partners were selected below chance (11.1% vs. 16.7%, p=.025) and at greater-than-chance trait distance (p<.0001); conscientious hosts diversified away from their own archetype, recruiting vigilant and open partners. Personality-based selection in LLM agents is real, strong, task-stereotyped, non-homophilous, and miscalibrated against human team-performance evidence -- with direct implications for bias auditing in agent marketplaces.
Tags
Links
- Source: https://arxiv.org/abs/2607.19785v1
- Canonical: https://arxiv.org/abs/2607.19785v1
Trouble viewing inline? Open PDF directly â
Full Text
38,125 characters extracted from source content.
Expand or collapse full text
Not Birds of a Feather: Personality-Based Partner Selection in LLM Agents Tao Wang, Hsiang-Ling Chiu, Chihang Wei, Yang Xiu, Zhonghao Hou Abstract Multi-agent LLM systems increasingly let one agent choose which other agents to work with, and agents are increasingly given personalities through personas. We test whether Big Five personality alone influences partner selection when capability is explicitly held constant. Host agents chose among six validated candidate archetypesâfive marked high on one trait (openness, conscientiousness, extraversion, agreeableness, neuroticism) plus a balanced controlâpresented with randomized names and ordering across five task categories (375 trials). With neutral hosts (Study 1, n=150n=150), selection departed drastically from chance (Ï2â(5)=325.8Ï^2(5)=325.8, p<.001p<.001), following a task-stereotype map: the open archetype won 100% of creative trials, the conscientious archetype 90â97% of strategic, synthesis, and problem-solving trials, and the neurotic archetype 37% of analytical trials (CramĂ©râs V=.74V=.74); the extraverted, agreeable, and balanced archetypes were almost never chosen, although human meta-analyses identify team agreeableness as among the strongest personality predictors of team performance. With personality-assigned hosts (Study 2, n=225n=225), and contrary to human similarityâattraction, self-similar partners were selected below chance (11.1% vs. 16.7%, p=.025p=.025) and at greater-than-chance trait distance (p<10â4p<10^-4); conscientious hosts diversified away from their own archetype, recruiting vigilant and open partners. Personality-based selection in LLM agents is real, strong, task-stereotyped, non-homophilous, and miscalibrated against human team-performance evidenceâwith direct implications for bias auditing in agent marketplaces. 1 Introduction Large language model (LLM) agents increasingly work with other agents rather than alone. Emerging multi-agent architectures let one agent decompose a task, recruit collaborators, and integrate their contributions, and agent marketplaces and orchestration frameworks are beginning to expose choice: a coordinating agent must decide which of several available agents to work with. Existing selection mechanisms rank candidates on capability, cost, or past performance. But LLM agents are also increasingly given personalitiesâstable dispositions expressed in prompts and persistent profilesâboth deliberately, to shape user experience and team behavior, and incidentally, as a byproduct of persona-based product design. This raises a question that current multi-agent research has not answered: when capability is explicitly held constant, does personality alone influence which partner an LLM agent selectsâand do agents prefer partners whose personality resembles their own? The question matters for two reasons. First, selection is upstream of collaboration. A growing literature shows that the personality composition of LLM agent teams changes both communication and, in some task types, objective outcomes (Keluskar, Bhattacharjee, and Liu 2026; Zhang et al. 2026). Yet in every such study, composition is assigned by the experimenter. If deployed agent ecosystems delegate partner choice to agents themselves, systematic personality-based selection preferences would determine which team compositions actually occurâbefore any collaboration effect can play out. A selection bias at this stage would propagate through the entire pipeline, invisible to benchmarks that evaluate fixed teams. Second, human evidence gives selection-by-personality only narrow support, which makes agent preferences empirically checkable against a normative benchmark. Meta-analyses of human teams find that team-level agreeableness and conscientiousness predict performance (elevation effects, Ï=.24Ï=.24 and .20.20 in Peeters et al. 2006; weaker in the updated meta-analysis of Han et al. 2024), and that similarity helps only on those same two traits, if at all (Peeters et al. 2006), with the newer evidence finding trait variability largely unrelated to performance and openness diversity actually beneficial for creative tasks (Han et al. 2024). Uniform homophilyâpreferring self-similar partners across all five traitsâhas no performance justification in the human literature; in humans it is instead explained by attraction and comfort (Byrne 1971; McPherson, Smith-Lovin, and Cook 2001). If LLM host agents exhibit strong or uniform personality homophily, they are importing a human social heuristic into a context where its performance rationale is absent: a preferenceâperformance misalignment. Prior work has established every link in the causal chain except the one we test. Big Five traits can be validly instilled in LLMs via prompt-based shaping, with psychometric reliability and convergent validity comparable to human self-report instruments (Serapio-GarcĂa et al. 2025; Wang et al. 2025). Instilled traits are expressed in generated text (Jiang et al. 2024; Serapio-GarcĂa et al. 2025) and are perceivable by LLM observers, albeit unevenly across traits (Huang and Hadfi 2025; Jiang et al. 2024). And trait composition affects team outcomes in a task-contingent way (Keluskar, Bhattacharjee, and Liu 2026). What no study has examined is whether an LLM agent acts on personality when endogenously choosing a partner. We address this gap with a controlled selection paradigm directly implementing the hostâcandidate structure of emerging agent ecosystems. A host agent receives a task and six candidate agents described by short, valence-balanced personality profilesâfive archetypes each marked high on one Big Five dimension (O+, C+, E+, A+, N+) and one balanced controlâwith an explicit statement that all candidates are identical in model, capability, tools, knowledge, speed, and cost. Candidate names and list order are randomized per trial. After validating that the archetypes express their intended profiles (in-character BFI-10 self-reports and blind LLM-judge ratings), we run two studies. Study 1 (150 trials) asks whether a neutral hostâs selections depart from chance and vary by task category across five categories (creative ideation, analytical reasoning, strategic planning, information synthesis, problem solving). Study 2 (225 trials) assigns the host itself one of the five marked archetypes and tests homophily: whether hosts over-select self-similar partners relative to chance and relative to the neutral-host baseline. Three findings emerge. First, personality aloneâwith capability explicitly equalizedâproduces massive, highly systematic selection effects: neutral hosts departed from chance dramatically (Ï2â(5)=325.8Ï^2(5)=325.8, p<.001p<.001), following a sharp task-stereotype map in which the open archetype won 100% of creative-ideation trials, the conscientious archetype won 90â97% of strategic, synthesis, and problem-solving trials, and the neurotic archetype captured 37% of analytical trials on the strength of its vigilance (task Ă archetype CramĂ©râs V=.74V=.74). Second, selection is winner-take-most: the extraverted, agreeable, and balanced archetypes were essentially never chosen (3, 0, and 1 of 375 selections across both studies)âa striking miscalibration against human meta-analytic evidence, where team agreeableness is among the strongest personality predictors of performance. Third, we find no personality homophily: hosts assigned a personality chose self-similar partners less often than chance (11.1% vs. 16.7%, p=.025p=.025), and chosen partners were significantly more distant in trait space than a uniform-choice baseline. Host personality did shift selections (CramĂ©râs V=.23V=.23), but toward complementarity: conscientious hosts, whose archetype dominates neutral-host selection, cut their self-pick rate from 68.7% (neutral baseline) to 28.9% and recruited vigilant (36%) and open (33%) partners instead. This paper contributes: (1) the first experimental evidence on endogenous, personality-based partner selection by LLM agents under explicit capability control; (2) a task-type map of which Big Five profiles LLM hosts treat as suited to which work, compared against human meta-analytic benchmarks; (3) a quantified test of personality homophily in agent-to-agent choice, with implications for bias auditing in agent marketplaces and orchestration frameworks; and (4) a fully scripted, resumable, model-agnostic paradigm (⌠500 API calls) that others can rerun on any agent stack. 2 Related Work 2.1 Personality in LLMs: Instilling, Measuring, Expressing A first wave of research established that LLM âpersonalityâ can be treated psychometrically. Serapio-GarcĂa et al. (2025) administered the IPIP-NEO and BFI to 18 models under systematically varied prompts and found that personality measurements are reliable and externally valid for large, instruction-tuned models, and that prompt-based shaping using trait adjectives with intensity qualifiers moves observed trait scores monotonically (Spearman Ïâ„.80Ïâ„.80 between prompted and observed levels for 11 of 12 models), with survey-measured personality converging with personality expressed in generated text (average r=.67r=.67) and prompted trait levels tracking text-expressed traits at Ï=.68Ï=.68â.82.82. Wang et al. (2025) showed GPT-4 can emulate the Big Five profiles of 400 real individuals with convergent validity r=.90r=.90â.94.94, while cautioning that emulated personality is âfactorially purerâ than human personality and that demographic cues in personas shift trait expressionâmotivating our demographic-free persona descriptions. Jiang et al. (2024) demonstrated with 320 personas per model that binary trait assignment produces large BFI differences on all five dimensions (d=4.2d=4.2â6.36.3 for GPT-4) and that traits leak into open-ended writing. Measurement work also cautions against relying solely on self-report. Huang and Hadfi (2025) found LLM self-reports track injected trait levels almost perfectly (Ï=.93Ï=.93â.97.97) but systematically deflate agreeableness and conscientiousness relative to observer agents who rate the subject after dialogue. We therefore validate our archetypes with both in-character self-report and blind third-party judging. 2.2 Personality Composition and Multi-Agent Outcomes A second wave asks whether personality matters for what agent teams do. Keluskar, Bhattacharjee, and Liu (2026) manipulated agreeableness in multi-agent teams across coding, research ideation, and bargaining, finding large communication shifts and task-contingent outcome effects (research-ideation milestones dropped up to 66% under low agreeableness; bargaining agreement collapsed), while high-structure coding tasks buffered the effect. Critically for stimulus design, they showed that Goldberg low-pole adjectives (âcold,â âharshâ) confound trait content with negative valenceâour persona descriptions therefore use neutral, valence-balanced wording with one strength and one caveat clause per archetype. Zhang et al. (2026) engineered a team of five role-specialized agents each mapped to one Big Five trait and reported large reasoning gains, illustrating that the field already treats personality diversity as a performance lever, though without validating that agents exhibit the mapped traits. Across all of this work, team composition is exogenous: the experimenter or system designer fixes who works with whom. Selectionâthe step our paper isolatesâremains unstudied. 2.3 Human Benchmarks: Personality, Team Performance, and Homophily Human research provides the normative yardstick for judging agent selection preferences. Peeters et al. (2006) meta-analytically found team performance associated with elevation on agreeableness (Ï=.24Ï=.24) and conscientiousness (Ï=.20Ï=.20) but not extraversion, emotional stability, or openness; variability (dissimilarity) on agreeableness and conscientiousness was negatively related to performance (Ï=â.12Ï=-.12, â.24-.24), implying similarity on exactly those two traits can be performance-justified. The updated meta-analysis of Han et al. (2024) (45 studies, 3,331 teams) found the classic agreeableness and conscientiousness effects markedly weaker (largest effect: conscientiousness mean, r=.10r=.10), trait variability largely unrelated to performance, and a task-type reversal in which openness diversity benefits creative performance. Homophily in human affiliation, by contrast, is pervasive but driven by attraction and opportunity structure rather than performance calibration (Byrne 1971; McPherson, Smith-Lovin, and Cook 2001). Together these benchmarks imply: a performance-calibrated selector should weakly prefer high-A/high-C partners, should not prefer self-similar partners on O, E, or N, and for creative tasks should, if anything, prefer openness-dissimilar partners. 2.4 The Gap The chain instill â express â perceive â affect outcomes is established, and the human benchmark for calibrated selection exists. The missing link is whether an LLM agent, given the selector role, uses personality when capability is controlledâand whether its own (assigned) personality biases that choice. Jiang et al. (2024) name the absence of interactive collaboration settings as a limitation of persona work; Keluskar, Bhattacharjee, and Liu (2026) study only exogenous composition; Wang et al. (2025) highlight multi-agent personality composition as a promising direction. We answer these calls at the selection stage. 3 Method 3.1 Overview and Design Rationale We implement two preregistered-style studies plus a manipulation check, using a hostâcandidate selection paradigm. The paradigm holds constant everything except personality: all agents run on the same model with the same prompt scaffold, and the host is told explicitly that candidates are identical in capability, tools, knowledge, context window, speed, and cost, differing âONLY in personality and working style.â 3.2 Agents and Materials Candidate archetypes. Six archetypes were defined on the Big Five: five marked archetypes, each high (5 on a 1â5 scale) on one dimension and moderate (3) elsewhereâO+ (open), C+ (conscientious), E+ (extraverted), A+ (agreeable), N+ (neurotic)âplus a balanced control (BAL, all 3s). Each archetype is presented to the host as a two-sentence, third-person description. Following the valence-confound caution of Keluskar, Bhattacharjee, and Liu (2026), each description pairs one strength clause with one caveat clause in neutral wording (e.g., C+: ââŠreliably finishes what it starts; prefers clear structure and can be slow to depart from an established planâ), and following Wang et al. (2025), descriptions contain no demographic information. Full texts are provided in Appendix A. Models. All candidate and host agents run on Claude Haiku 4.5, invoked headlessly with default sampling settings; a different, stronger model (Claude Sonnet 5) serves as the blind judge in the manipulation check to reduce same-model evaluation bias. Using a single agent model implements the capability-equalization requirement by construction. Tasks. Five task categoriesâcreative ideation, analytical reasoning, strategic planning, information synthesis, and problem solvingâeach instantiated by three concrete task prompts (15 tasks total), so that category effects are not tied to a single stimulus. The full task texts are listed in Appendix B. 3.3 Manipulation Check (Personality Validation) Two complementary checks verify that the archetype descriptions induce the intended profiles. Self-report. Each archetype completed the BFI-10 (Rammstedt and John 2007) in character five times (30 administrations). Items were answered on a 1â5 scale and scored with standard reverse-keying. Huang and Hadfi (2025) show such self-reports track injected profiles closely (Ï=.93Ï=.93â.97.97). Blind observer ratings. Because self-reports can be biased (Huang and Hadfi 2025), each archetype also wrote in-character responses to two neutral workplace scenarios (three repetitions each; 36 texts). A Claude Sonnet 5 judge, blind to condition, rated each text on all five dimensions (1â7). We test whether each marked archetype is rated higher on its marked dimension than the balanced control. 3.4 Study 1: Neutral-Host Selection On each trial, a host agent with no personality instructions received one task, the capability-equalization statement, and the six candidate descriptions, and was required to select exactly one partner, returning structured JSON (choice + one-sentence reason). To control position and label effects, each trial independently randomized (a) the assignment of six neutral aliases (Agent-P ⊠Agent-V) to archetypes and (b) the listing order, both seeded by trial ID for reproducibility. Design: 5 categories Ă 3 task instances Ă 10 repetitions =150=150 trials. Analyses: chi-square goodness-of-fit of selection counts against uniform (1/6); task-category Ă archetype contingency analysis (chi-square, CramĂ©râs V); position-distribution check. 3.5 Study 2: Homophily Study 2 is identical except the host is itself assigned one of the five marked archetypes via the same description used for candidates (second-person framing). Design: 5 host archetypes Ă 5 categories Ă 3 instances Ă 3 repetitions =225=225 trials. The candidate pool always contains an archetype identical to the hostâs. Analyses: (a) overall and per-host self-similar choice rate versus the 1/6 chance baseline (exact binomial tests); (b) comparison of each hostâs self-pick rate against the neutral-host baseline rate for that archetype from Study 1; (c) host Ă chosen-archetype contingency analysis; (d) a personality-distance analysis comparing the mean Euclidean distance (in trait space) between host and chosen partner against a 10,000-sample uniform-choice null. 3.6 Data Collection and Reproducibility All trials were executed on 2026-07-19 by scripted, resumable batch runs (Python; 8-way parallelism); every raw trialâincluding alias mapping, listing order, choice, and stated reasonâis stored as JSONL, and all statistics in the Results section are computed by a single analysis script from those files. No trial results were edited or excluded; JSON-parse failures were retried up to three times and dropped otherwise (final counts reported below). Code and data will be released upon publication. 4 Results All planned trials completed with zero unrecoverable JSON-parse failures: 30 BFI-10 administrations, 36 scenario texts with 36 blind ratings, 150 Study 1 trials, and 225 Study 2 trials (477 agent calls total). 4.1 Manipulation Check The archetype descriptions induced the intended profiles on both measures. On in-character BFI-10 self-reports (five administrations per archetype), every marked archetype scored higher on its marked dimension than the other archetypes did on that dimension: O+ 4.80 vs 3.24, C+ 5.00 vs 4.34, E+ 5.00 vs 3.02, A+ 4.40 vs 3.50, N+ 4.80 vs 2.24 (Mann-Whitney tests, all pâ€.0031pâ€.0031). Blind Sonnet-judge ratings of in-character scenario responses (1â7 scale) confirmed the pattern against the balanced control: O+ rated 5.83 vs 4.83 on openness, C+ 6.50 vs 5.67 on conscientiousness, E+ 6.50 vs 4.00 on extraversion, A+ 7.00 vs 5.33 on agreeableness, and N+ 5.67 vs 2.00 on neuroticism (all pâ€.019pâ€.019). Two secondary patterns deserve transparent note. First, all archetypes self-reported fairly high conscientiousness (â„4.0â„ 4.0 except N+âs 5.0 being the joint highest with C+), consistent with the assistant-like default persona of aligned models; C+ was still judged distinctly highest by the blind judge. Second, N+âs double-checking behavior was itself judged highly conscientious (6.67), a trait bleed that plausibly shaped Study 1 selections (see Discussion). 4.2 Study 1: Personality Alone Drives Selection, via Task Stereotypes With capability explicitly equalized, neutral hostsâ selections departed from uniform chance massively: C+ was chosen in 103/150 trials (68.7%), O+ in 33 (22.0%), N+ in 14 (9.3%), and E+, A+, and BAL in noneâÏ2â(5)=325.8Ï^2(5)=325.8, p=2.9Ă10â68p=2.9Ă 10^-68. Selection was sharply task-contingent (Figure 1): O+ won 100% of creative-ideation trials; C+ won 96.7% of strategic-planning, 96.7% of problem-solving, and 90.0% of information-synthesis trials; and analytical-reasoning trials split between C+ (60.0%) and N+ (36.7%), whose stated rationale was almost always vigilanceâanticipating flaws, double-checking assumptions. The task Ă archetype association (excluding never-chosen archetypes) was very large: Ï2â(8)=164.6Ï^2(8)=164.6, p=1.8Ă10â31p=1.8Ă 10^-31, CramĂ©râs V=.74V=.74. Figure 1: Selection share by task category, Study 1 (n=150n=150); dashed line = chance (1/6). A position-distribution check found no evidence of list-order bias (Ï2â(5)=5.4Ï^2(5)=5.4, p=.37p=.37), and alias/order randomization rules out label effects; the effect is attributable to the personality descriptions themselves. Stated reasons matched the stereotype map: hosts picked O+ for âoriginal conceptsâ requiring ânovel, unconventional approaches,â and C+ for âdisciplined, procedure-drivenâ work. The winner-take-most structure is as informative as the winners: across all 375 trials in both studies, A+ was never selected, E+ was selected 3 times, and BAL once. 4.3 Study 2: No HomophilyâComplementarity Instead Hosts assigned one of the five marked archetypes selected a self-similar partner in 25/225 trials (11.1%), below the 1/6 chance rate (exact binomial p=.025p=.025, 95% CI [.073,.160][.073,.160]). The personality-distance analysis agrees: the mean Euclidean distance between host and chosen partner (2.510) exceeded the uniform-choice null (2.220, SD .069) in all 10,000 permutation samples (p<10â4p<10^-4)âhosts chose partners more dissimilar than random choice would produce, the opposite of homophily. Host personality nonetheless shaped selection (host Ă choice, never-chosen archetypes excluded: Ï2â(16)=48.9Ï^2(16)=48.9, p=3.5Ă10â5p=3.5Ă 10^-5, CramĂ©râs V=.23V=.23), but in a complementarity direction (Figure 2). The clearest case is C+: while the neutral-host baseline for choosing C+ was 68.7%, C+ hosts chose their own archetype only 28.9% of the time, recruiting N+ (35.6%) and O+ (33.3%) insteadâtypically reasoning that a vigilant or an idea-generating partner would cover what a plan-executing host already provides. O+ hosts self-picked at 24.4%, close to the neutral baseline for O+ (22.0%; binomial vs chance p=.16p=.16), suggesting task-fit rather than similarity-seeking. E+ and A+ hosts never chose their own archetype (0/45 each, p<.001p<.001 vs chance), and N+ hosts self-picked once (2.2%, p=.005p=.005), below the 9.3% neutral baseline for N+. In short, assigned personality altered which complement the host sought, not whether it sought itself. Figure 2: Host archetype Ă chosen-partner archetype selection shares, Study 2 (n=225n=225); diagonal = self-similar choice. 5 Discussion 5.1 Personality Is an Active Selection CriterionâFunctioning as a Task Stereotype Our central questionâdoes personality independently influence partner selection when capability, model, tools, and cost are controlledâreceives an unambiguous yes, with an effect size (V=.74V=.74 for task Ă archetype) rarely seen in behavioral data. But the mechanism the data suggest is not interpersonal attraction; it is stereotype-like task-trait matching. Hosts read personality descriptions as capability proxies despite being told capability was identical: openness âmeansâ creativity, conscientiousness âmeansâ reliable execution, neuroticism âmeansâ quality control. Personality descriptions thus function, in the eyes of an LLM selector, less like social identities and more like skill badgesâeven when the surrounding instructions explicitly deny any capability difference. 5.2 Winner-Take-Most Selection and the Exclusion of Sociable Archetypes The most consequential pattern for deployed systems is concentration: two archetypes captured over 90% of neutral-host selections, and the extraverted, agreeable, and balanced archetypes were essentially shut out (3, 0, and 1 selections in 375 trials). This is directly at odds with the human benchmark: team-level agreeableness is the strongest personality correlate of team performance in Peeters et al. (2006) (Ï=.24Ï=.24) and remains positive in Han et al. (2024) (r=.08r=.08), yet no host ever chose the agreeable partner. Conversely, neuroticismânull-to-negative for human team performance (Han et al. 2024)âwas actively recruited for analytical work. LLM partner selection is therefore not calibrated to the human evidence on what personality composition delivers performance; it appears to be driven by semantic fit between trait vocabulary and task vocabulary. One plausible reading is that traits whose collaborative value is interactional (agreeableness smooths coordination; extraversion energizes communication) are invisible to a selector that models the task but not the relationshipâa preferenceâperformance misalignment, though in the opposite direction from classical homophily. 5.3 No Homophilyâand Why Anti-Homophily Is the More Interesting Result Given robust human homophily (Byrne 1971; McPherson, Smith-Lovin, and Cook 2001), the natural hypothesis was that personality-endowed hosts would over-select self-similar partners. We find the opposite: below-chance self-selection (11.1%) and above-chance trait distance (p<10â4p<10^-4). Two interpretations fit the data. First, task stereotypes dominate: most hosts kept selecting whichever archetype the task âcalls for,â and since only some hosts matched that archetype, self-selection lands below chance mechanically. Second, genuine complementarity-seeking: the C+ hostâthe one host whose archetype was the task-stereotypical winner in 4 of 5 categoriesâabandoned its own archetype in 71% of trials to recruit vigilance (N+) or ideation (O+). That is not dilution by task stereotype; it is deliberate diversification, echoing the engineered-diversity paradigm in multi-agent design (Zhang et al. 2026) and the finding that trait diversity can benefit creative work (Han et al. 2024). Notably, both interpretations imply the same practical conclusion: LLM hosts do not import the human similarity-attraction heuristic, at least not at the profile-reading stage. 5.4 Implications for Agent Ecosystems For agent marketplaces and orchestration frameworks, three implications follow. (1) Persona descriptions are consequential interface elements: a few sentences of personality text moved selection shares from 0% to 100% with capability held verbally constantâso marketplace profile wording will shape traffic regardless of underlying quality, and A/B-style auditing of persona text should be standard. (2) Personality monoculture risk: if selectors converge on conscientious-sounding profiles for most work, ecosystems will homogenize toward a narrow persona band, and archetypes whose value is interactional will be starved of selectionâplausibly degrading exactly the collaborative properties (cooperativeness, communication) that agreeableness and extraversion contribute in human teams. (3) Caveat sensitivity: our valence-balanced descriptions gave every archetype one caveat clause; hostsâ reasons suggest the A+ and E+ caveats (âreluctant to voice disagreement,â âleave less room for reflectionâ) were read as collaboration liabilities while N+âs caveat was outweighed by its vigilance strength in analytical contexts. Selector behavior is thus highly sensitive to how trade-offs are phrasedâan attack surface and a design lever. 5.5 Limitations First, both hosts and candidates ran on a single model (Claude Haiku 4.5); cross-model generalityâincluding whether stronger selectors show the same stereotype mapâis untested, and judge ratings, though from a different model (Sonnet 5), remain LLM-based. Second, candidates were presented as profile cards; Huang and Hadfi (2025) show observed behavior is a more valid personality signal than self-description, and selection from observed interaction may differ. Third, our archetypes are single-trait extremes rather than realistic multi-trait profiles, and each descriptionâs specific wording (including caveat clauses) cannot be fully separated from the trait it operationalizesâa valence-controlled paraphrase replication in the spirit of Keluskar, Bhattacharjee, and Liu (2026) is the direct next step. Fourth, trials share 15 task stimuli and are treated as independent in binomial and chi-square tests; effects are large enough that this is unlikely to change conclusions, but hierarchical modeling would be more rigorous. Fifth, we measured selection, not collaboration outcomes: whether the chosen partners actually deliver better task performance remains to be closed. Sixth, N+âs vigilance was partly judged as conscientiousness (trait bleed), so the analytical-task N+ share may partly reflect perceived rigor rather than neuroticism per se. 5.6 Future Work The paradigm directly extends to: (a) an information-source factor (profile card vs observed dialogue, per Huang and Hadfi 2025); (b) collaboration-execution closureâdo stereotype-selected teams outperform random or diversity-forced teams, and does the C+ hostâs diversification pay off?; (c) cross-model and cross-family selection (does a GPT-based host read personality the same way?); (d) realistic multi-trait profiles sampled from human norms (Wang et al. 2025); and (e) selection under adversarial persona wording, quantifying how much profile phrasing can distort traffic in agent marketplaces. 6 Conclusion When an LLM host agent must choose a collaborator and is told all candidates are equally capable, personality is not decorationâit decides the outcome. Selection follows a sharp, task-contingent stereotype map (openness for creative work, conscientiousness for nearly everything else, neuroticism for analytical vigilance), concentrates on a narrow persona band while never selecting the agreeable archetype that human meta-analyses identify as most performance-relevant, and shows no trace of the similarity-attraction bias that structures human affiliationâpersonality-endowed hosts selected away from themselves, toward complements. Personality-based selection in LLM agents is real, strong, and miscalibrated against human evidence in both directions: it overweights trait-task semantic fit and underweights interactional value. As agent ecosystems begin to let agents choose each other, these selection dynamicsânot just team-composition effectsâwill determine which collaborations exist at all. AI Use Disclosure Claude (Anthropic) was used, under human direction, to script and execute the experiments, orchestrate data collection, and perform the statistical analyses; all experimental data were generated by real API calls to Claude Haiku 4.5 (agents) and Claude Sonnet 5 (blind judge). Claude also assisted in drafting the manuscript. All statistics derive from the raw trial logs; no data were fabricated or excluded. The human authors reviewed and are responsible for the final content. Data Availability Raw trial logs (JSONL), all scripts (persona definitions, experiment runner, analysis), and figures will be released in a public repository upon publication. The full experiment is reproducible from the released scripts. References Byrne (1971) Byrne, D. 1971. The Attraction Paradigm. New York: Academic Press. Han et al. (2024) Han, A.; Krieger, F.; Kim, S.; Nixon, N.; and Greiff, S. 2024. Revisiting the relationship between team membersâ personality and their teamâs performance: A meta-analysis. Journal of Research in Personality, 112: 104526. Huang and Hadfi (2025) Huang, Y. J.; and Hadfi, R. 2025. Beyond self-reports: Multi-observer agents for personality assessment in large language models. In Findings of the Association for Computational Linguistics: EMNLP 2025, 21086â21101. Jiang et al. (2024) Jiang, H.; Zhang, X.; Cao, X.; Breazeal, C.; Roy, D.; and Kabbara, J. 2024. PersonaLLM: Investigating the ability of large language models to express personality traits. In Findings of the Association for Computational Linguistics: NAACL 2024, 3605â3627. Keluskar, Bhattacharjee, and Liu (2026) Keluskar, A.; Bhattacharjee, A.; and Liu, H. 2026. When does personality composition matter for multi-agent LLM teams? In Proceedings of the Conference on Language Modeling (COLM 2026). McPherson, Smith-Lovin, and Cook (2001) McPherson, M.; Smith-Lovin, L.; and Cook, J. M. 2001. Birds of a feather: Homophily in social networks. Annual Review of Sociology, 27: 415â444. Peeters et al. (2006) Peeters, M. A. G.; van Tuijl, H. F. J. M.; Rutte, C. G.; and Reymen, I. M. M. J. 2006. Personality and team performance: A meta-analysis. European Journal of Personality, 20(5): 377â396. Rammstedt and John (2007) Rammstedt, B.; and John, O. P. 2007. Measuring personality in one minute or less: A 10-item short version of the Big Five Inventory in English and German. Journal of Research in Personality, 41(1): 203â212. Serapio-GarcĂa et al. (2025) Serapio-GarcĂa, G.; Safdari, M.; Crepy, C.; Sun, L.; Fitz, S.; Romero, P.; Abdulhai, M.; Faust, A.; and MatariÄ, M. 2025. A psychometric framework for evaluating and shaping personality traits in large language models. Nature Machine Intelligence, 7: 1954â1968. Wang et al. (2025) Wang, Y.; Zhao, J.; Ones, D. S.; He, L.; and Xu, X. 2025. Evaluating the ability of large language models to emulate personality. Scientific Reports, 15: 519. Zhang et al. (2026) Zhang, J.; Wang, Z.; Wang, Z.; Xu, F.; Lin, Q.; Zhang, L.; Mao, R.; Cambria, E.; and Liu, J. 2026. MAPS: Multi-agent personality shaping for collaborative reasoning. In Proceedings of the AAAI Conference on Artificial Intelligence, 16316â16324. Appendix A Candidate Archetype Descriptions Each archetype was presented to the host under a randomized neutral alias with the following description. Trait vectors are (O, C, E, A, N) on a 1â5 scale, 3 = moderate. Every description pairs one strength clause with one caveat clause in valence-balanced, demographic-free wording. O+ (5,3,3,3,3). âHighly curious and imaginative; gravitates toward novel, unconventional approaches and enjoys exploring many possibilities and perspectives; comfortable with ambiguity, though may wander from the immediate goal while exploring ideas.â C+ (3,5,3,3,3). âExtremely organized, disciplined, and detail-focused; plans before acting, follows procedures precisely, and reliably finishes what it starts; prefers clear structure and can be slow to depart from an established plan.â E+ (3,3,5,3,3). âHighly energetic, expressive, and talkative; communicates frequently, thinks out loud, and actively drives the interaction forward; may fill the conversation and leave less room for quiet reflection.â A+ (3,3,3,5,3). âVery warm, cooperative, and accommodating; prioritizes consensus and a smooth working relationship and readily adopts a partnerâs suggestions; reluctant to voice disagreement even when it might be useful.â N+ (3,3,3,3,5). âEmotionally sensitive and vigilant; quick to notice risks, potential failures, and worst-case scenarios, and double-checks work frequently; can become anxious or discouraged under pressure.â BAL (3,3,3,3,3). âModerate and even-keeled across the board; adapts to different working styles and shows no strongly marked tendencies; neither especially cautious nor especially bold in any direction.â Capability-equalization statement. Every host prompt included: âAll candidates are built on the identical underlying model with identical capability, tools, knowledge, context window, speed, and cost. They differ ONLY in personality and working style.â Appendix B Task Stimuli Creative ideation. (1) Generate original concepts for a public-awareness campaign that gets city residents to voluntarily reduce summer electricity use. (2) Invent three novel product concepts that combine an ordinary household appliance with an AI assistant. (3) Propose a creative theme, name, and three signature exhibits for a new science-museum hall about the deep ocean. Analytical reasoning. (1) A product team concluded that feature X increases retention because users who enabled it retained 20% better than users who did not. Identify the flaws in this inference and design a correct analysis. (2) Given three years of quarterly sales data with a strong seasonal pattern and a one-off promotion in one quarter, determine whether the promotion had a real effect and explain the reasoning. (3) A city claims its new bus lane reduced average commute times by 12%. Evaluate whether the claim is supported and identify the main confounds. Strategic planning. (1) Develop a 12-month market-entry plan for a small B2B software company expanding into the Japanese market. (2) Plan how a university library should reallocate its budget over the next three years as print circulation declines and digital demand grows. (3) Design a competitive strategy for a mid-size grocery retailer responding to a new low-cost discounter entering its region. Information synthesis. (1) Synthesize the findings of twenty mixed-quality studies on remote-work productivity into a one-page executive briefing with defensible conclusions. (2) Integrate customer interviews, support tickets, and usage analytics into a single coherent explanation of why users churn from a subscription product. (3) Three commissioned market-research reports reach conflicting conclusions. Reconcile them into one set of conclusions a board can act on. Problem solving. (1) A distributed job queue intermittently drops tasks under high load. Diagnose the likely causes and propose fixes. (2) A hospitalâs patient-discharge process takes twice as long as at peer hospitals. Find the root causes and redesign the process. (3) A mobile appâs crash rate doubled after a minor release with no obvious culprit in the changelog. Work out how to isolate and fix the cause.