Paper deep dive
Toward Human-Centered Multi-Agent Systems: Integrating Cognition, Culture, Values, and Cooperation in AI Agents
Safia Baloch, Rahemeen Khan
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 88%
Last extracted: 7/9/2026, 3:29:29 AM
Summary
This survey argues that contemporary AI and multi-agent systems prioritize task optimization over human-centered capabilities. It proposes a unified framework integrating cognition, culture, values, and cooperation to enable socially adaptive, culturally aware, and value-aligned autonomous agents.
Entities (8)
Relation Signals (9)
Large Language Models (LLMs) â enable â Multi-Agent Systems
confidence 95% ¡ The emergence of large language model (LLM)-based agents and multi-agent systems has enabled a shift from narrow task automation to more autonomous decision-making.
Human-Centered Agents â require â Cognition
confidence 90% ¡ We define human-centered multi-agent systems as systems in which agents are designed not only to solve tasks, but also to reason and interact in ways that account for human cognition...
Human-Centered Agents â require â Culture
confidence 90% ¡ ...account for human cognition, culture, values, and cooperative dynamics.
Human-Centered Agents â require â Values
confidence 90% ¡ ...account for human cognition, culture, values, and cooperative dynamics.
Human-Centered Agents â require â Cooperation
confidence 90% ¡ ...account for human cognition, culture, values, and cooperative dynamics.
Bounded Rationality â contrastswith â Classical Decision Theory
confidence 85% ¡ Classical decision theory often assumes rational actors who optimize expected utility... Simon (1957) introduced bounded rationality to describe how humans make satisfactory rather than optimal decisions...
Preference Learning â supports â Value Alignment
confidence 85% ¡ Preference learning has become one of the most influential approaches to alignment in foundation models.
Theory of Mind â underpins â Human-Agent Collaboration
confidence 85% ¡ ToM underlies persuasion, explanation, coordination, and conflict resolution. In collaborative settings, humans continuously model what others know...
Shared Mental Models â facilitate â Multi-Agent Coordination
confidence 80% ¡ Humans also collaborate through trust, role awareness, negotiation, and shared mental models.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The emergence of large language model (LLM)-based agents and multi-agent systems has enabled a shift from narrow task automation to more autonomous decision-making. Despite progress in language generation, planning, tool use, and coordination, most agents still treat intelligence as prediction, optimization, and task completion. Human environments are social and normative, where people reason under bounded rationality, communicate in culturally situated language, and make decisions guided by values, beliefs, trust, and social norms. This survey argues that future AI agents, especially those acting on behalf of humans, must move beyond task competence toward human-centered capabilities. We review research across six areas: (1) evolution of intelligent agents, (2) human cognition and decision-making, (3) language, culture, and social context, (4) human values and belief systems, (5) human-agent collaboration, and (6) multi-agent coordination and modeling of human characteristics. We synthesize work from cognitive science, sociolinguistics, computational social science, and AI alignment, along with recent advances in LLM agents, cultural alignment benchmarks, preference learning, explainability, and agent societies. We identify a key gap: existing systems do not provide a unified framework integrating cognition, culture, values, and social behavior into autonomous agents. We conclude with directions for building culturally aware, value-aligned, cognitively grounded, and cooperative multi-agent systems.
Tags
Links
- Source: https://arxiv.org/abs/2606.08274v1
- Canonical: https://arxiv.org/abs/2606.08274v1
Trouble viewing inline? Open PDF directly â
Full Text
55,521 characters extracted from source content.
Expand or collapse full text
Toward Human-Centered Multi-Agent Systems: Integrating Cognition, Culture, Values, and Cooperation in AI Agents Safia Baloch safia.baloch@giki.edu.pk Rahemeen Khan rahemeen_ahmed@yahoo.com Abstract The emergence of large language model (LLM)-based agents and multi-agent systems has accelerated a transition in artificial intelli- gence from narrow task automation to increas- ingly autonomous decision-making. Yet, de- spite remarkable progress in language genera- tion, planning, tool use, and coordination, most contemporary agents still treat intelligence pri- marily as prediction, optimization, and task completion. Human environments, however, are fundamentally social and normative: peo- ple reason under bounded rationality, communi- cate through culturally situated language, make decisions through values and beliefs, and col- laborate through trust, shared mental models, and social norms. This survey argues that the next generation of AI agents especially those acting on behalf of humans must move beyond task competence toward human-centered capa- bilities. We synthesize research across six intercon- nected areas: (1) the evolution of intelligent agents, (2) human cognition and decision- making, (3) language, culture, and social con- text, (4) human values and belief systems, (5) human-agent collaboration, and (6) multi- agent coordination and computational model- ing of human characteristics. We review foun- dational theories from cognitive science, soci- olinguistics, computational social science, and AI alignment, alongside recent work on LLM agents, cultural alignment benchmarks, prefer- ence learning, explainability, and agent soci- eties. Based on this synthesis, we identify a central research gap: while substantial progress has been made in each dimension individually, there remains no widely adopted framework that jointly integrates cognition, culture, values, and social behavior into autonomous agents capable of authentic human-centered collabora- tion. We conclude by outlining future research directions for building culturally aware, value- aligned, cognitively grounded, and cooperative multi-agent systems. 1 Introduction Artificial intelligence has entered an era in which agents are no longer limited to narrow task execu- tion. The rise of foundation models and LLMs has enabled systems that can perceive instructions, rea- son over long contexts, call tools, simulate social behavior, and coordinate with other agents (Brown et al., 2020; Wang et al., 2025; Guo et al., 2024; Chen et al., 2025c). This shift has fueled an emerg- ing paradigm sometimes described as agentic AI: systems capable not merely of responding, but of acting. At the same time, a conceptual limitation has become increasingly visible. Much of current AI still models intelligence as a combination of predic- tion, optimization, and procedural task completion (Russell, 2019; Ji et al., 2024). In contrast, human intelligence is inseparable from social context. Hu- man reasoning is often bounded rather than optimal (Simon, 1957); judgments are shaped by heuristics and biases (Tversky and Kahneman, 1974; Kahne- man, 2011); communication depends on pragmat- ics, social identity, and culture (Hall, 1976; Hofst- ede, 2001); and decisions are influenced not only by goals but by values, norms, and beliefs (Rus- sell, 2019; Ji et al., 2024). Humans also collabo- rate through trust, role awareness, negotiation, and shared mental models (Salas et al., 2005; Doshi- Velez and Kim, 2017). These observations matter because AI agents are increasingly deployed in roles traditionally occu- pied by humans: assistants, advisors, mediators, teammates, and autonomous participants in orga- nizational workflows. In such settings, success cannot be measured only by accuracy or efficiency. Agents acting on behalf of humans must be able to preserve and reflect human-centered properties: culturally appropriate language, context-sensitive reasoning, value-aligned decision making, and so- cially adaptive collaboration. 1 arXiv:2606.08274v1 [cs.MA] 6 Jun 2026 This survey develops the argument that future multi-agent systems should not merely be intelli- gent; they should be human-centered. We define human-centered multi-agent systems as systems in which agents are designed not only to solve tasks, but also to reason and interact in ways that account for human cognition, culture, values, and coopera- tive dynamics. The survey is organized around the following progression: â˘We first review the evolution of intelligent agents, from rule-based systems to LLM-based autonomous and multi-agent architectures. â˘We then examine human cognition and decision-making, focusing on bounded rational- ity, dual-process reasoning, common sense, and theory of mind. ⢠Next, we discuss language, culture, and social context, including cross-cultural psychology, so- ciolinguistics, and culturally adaptive language technologies. â˘We then consider human values and belief sys- tems, drawing on AI alignment, preference learn- ing, and machine ethics. â˘We connect these attributes to human-agent collaboration, emphasizing trust, explainability, and shared mental models. â˘Finally, we analyze multi-agent systems and computational modeling of human character- istics, highlighting the gap between coordination- centric architectures and richer human-centered representations. The central claim of this paper is straightforward: existing AI and multi-agent systems primarily focus on intelligence as prediction, optimization, and task completion, whereas agents acting on behalf of humans in cooperative environments require deeper computational representations of human cognition, culture, values, and social be- havior. While significant progress has been made in each of these topics independently, the literature still lacks unified frameworks that integrate them into robust autonomous agents. 2 A Unified View of Human-Centered Agents Before reviewing the literature in detail, it is use- ful to articulate the conceptual space this survey addresses. Figure 1 presents a unified perspective on human-centered agents. The framework places a multi-agent infrastructure at the base, but argues that effective participation in socially embedded environments requires four additional layers: cog- nition, culture, values, and cooperation. The figure is intentionally layered rather than modular. It emphasizes that human-centered be- havior is not a single capability added to a task- oriented agent; it emerges from the interaction of multiple dimensions. Cognition without culture yields abstract reasoning detached from context. Culture without values risks superficial adaptation. Values without cooperation produce rigid agents that may be aligned in isolation but ineffective in teams. The broader research challenge, therefore, is integration. 3 Evolution of Intelligent Agents 3.1 From Rule-Based Automation to Learning Systems The history of intelligent agents can be read as a se- quence of increasing representational and decision- making flexibility. Early AI systems were primarily rule-based. Classical expert systems and symbolic agents relied on hand-crafted knowledge bases and inference rules to solve well-defined problems (Nilsson, 1980; Russell and Norvig, 2021). These systems excelled in constrained domains but suf- fered from brittleness, poor scalability, and limited adaptability. Machine learning shifted the field away from manual rule-writing toward data-driven generaliza- tion (Bishop, 2006). Statistical NLP, reinforcement learning, and deep learning substantially improved performance across perception, classification, lan- guage understanding, and control. Yet most such systems still optimized narrowly defined objectives, often without explicit representations of social con- text or human preferences. 3.2Foundation Models and the Re-emergence of Agents The introduction of large-scale pretrained language models marked a major reconfiguration of the agent discussion. Models such as GPT-3 demonstrated remarkable in-context learning and language gen- eration abilities (Brown et al., 2020). Subsequent work revealed that LLMs can support planning, reflection, tool use, role specialization, and long- horizon workflows, leading to a rapid expansion 2 Multi-Agent System Infrastructure communication protocols, memory, planning, coordination, tool use, distributed decision-making Cognition ⢠bounded rationality ⢠dual-process reasoning ⢠common-sense reasoning ⢠mental models ⢠theory of mind Culture & Social Context ⢠sociolinguistics ⢠cultural norms ⢠contextual communication ⢠regional variation ⢠social identity Values & Beliefs ⢠value alignment ⢠preference learning ⢠pluralistic values ⢠moral reasoning ⢠belief representation Cooperation ⢠trust ⢠explainability ⢠shared mental models ⢠teamwork ⢠negotiation Human-Centered Agent Behavior context-aware, culturally adaptive, value-aligned, and socially cooperative decision-making Figure 1: A unified conceptual framework for human-centered multi-agent systems. Contemporary agents typically rely on a multi-agent infrastructure for planning, communication, and task execution. This survey argues that agents acting on behalf of humans must additionally model cognition, culture, values, and cooperation in an integrated manner. of research on autonomous and multi-agent LLM systems (Wang et al., 2025; Xi et al., 2023; Guo et al., 2024; Chen et al., 2025c). Recent surveys have converged on the idea that LLM-based agents can be decomposed into compo- nents such as profiling, memory, planning, action, and feedback (Wang et al., 2025; Xi et al., 2023). In multi-agent settings, research has explored role assignment, collaboration structures, communica- tion protocols, and collective problem solving (Guo et al., 2024; Tran et al., 2025). These architectures expand the operational scope of AI from single-turn response generation to persistent action in complex environments. 3.3 The Remaining Limitation Despite these advances, the dominant design logic of many current agents remains task-centric. Agents may communicate fluently, invoke tools effectively, and coordinate procedurally, yet still fail to capture key dimensions of human behavior. They often exhibit weak grounding in social norms, inconsistent persona maintenance, limited cultural adaptation, and narrow notions of alignment (Cao et al., 2023; AlKhamissi et al., 2024; Wang et al., 2024; Baltaji et al., 2024). In other words, modern agents are increasingly autonomous, but not yet deeply human-centered. 3.4 Research Gap The evolution of intelligent agents reveals a widen- ing gap between capability and human representa- tion. Contemporary agents have gained powerful mechanisms for task decomposition, communica- tion, and action, but they still lack unified represen- tations of how humans reason, what humans value, and how humans behave across cultural and social settings. This gap motivates the rest of the survey. 4 Human Cognition and Decision-Making If agents are expected to act on behalf of humans, they must do more than optimize utility functions or rank likely next actions. They must approximate how humans reason under uncertainty, constraints, and social context. 4.1 Rationality and Its Limits Classical decision theory often assumes rational actors who optimize expected utility. While such models remain foundational, decades of cognitive research show that human reasoning deviates sys- tematically from strict rationality. Simon (1957) 3 introduced bounded rationality to describe how humans make satisfactory rather than optimal deci- sions under limited time, information, and cogni- tive resources. This insight is highly relevant for agent design: human-centered agents should not be modeled as perfect optimizers detached from realistic decision conditions. Tversky and Kahneman (1974) and Kahneman (2011) further demonstrated that human judgment is shaped by heuristics and biases. People rely on mental shortcuts such as availability and represen- tativeness, which can be efficient but also error- prone. The implication is not that AI should re- produce human biases uncritically, but that human decision-making is not reducible to clean optimiza- tion. Agents intended to support or imitate human reasoning must account for this hybrid structure of efficiency, boundedness, and context dependence. 4.2 Dual-Process Accounts of Reasoning Dual-process theories distinguish between fast, intuitive reasoning (often called System 1) and slow, deliberative reasoning (System 2) (Kahne- man, 2011; Evans, 2008). This distinction has become increasingly relevant in AI. Many LLM systems appear strong at rapid pattern completion and associative inference, yet struggle with deliber- ate, grounded, or multi-step reasoning unless scaf- folded by prompting, decomposition, or external tools (Wei et al., 2022; Kojima et al., 2022; Huang and Chang, 2023). This parallel has motivated interest in cognitive hybrids that combine neural language models with more structured reasoning or cognitive architec- tures. For example, recent work explores integrat- ing ACT-R-inspired mechanisms into LLM sys- tems to ground perception, memory, and deliberate decision-making (Wu et al., 2024). Such work re- flects a broader recognition that human-like reason- ing may require architectures capable of shifting between intuitive and reflective modes. 4.3 Common-Sense Reasoning Human reasoning is deeply supported by com- mon sense: tacit world knowledge about causality, temporality, physical interaction, social routines, and practical inference. Common-sense reason- ing has long been recognized as a core AI chal- lenge (McCarthy, 1959). While LLMs often dis- play impressive common-sense behavior in lan- guage tasks, their competence remains uneven and benchmark-dependent (Huang and Chang, 2023). In human-centered agent settings, common sense is not merely a benchmark ability; it is a requirement for believable, safe, and contextually appropriate action. 4.4 Theory of Mind and Mental Models A particularly important aspect of social cognition is Theory of Mind (ToM): the ability to reason about the beliefs, desires, intentions, and knowledge states of others (Premack and Woodruff, 1978). ToM underlies persuasion, explanation, coordi- nation, and conflict resolution. In collaborative settings, humans continuously model what others know, what they expect, and how they are likely to react. Recent work has begun evaluating ToM-like ca- pabilities in LLMs and surveying related bench- marks (Nguyen, 2025; Chen et al., 2025b). Find- ings suggest that LLMs can perform well on some stylized tasks, yet remain non-robust and sensitive to prompt design, dataset artifacts, and superficial cues. This distinction is crucial: high benchmark performance does not necessarily imply stable, de- ployable social reasoning. For multi-agent and human-facing systems, ToM must be reliable, con- textual, and interaction-aware. Relatedly, the literature on mental models in teamwork emphasizes that effective collaboration relies on partially shared internal representations of goals, roles, constraints, and plans (Salas et al., 2005). Human-centered agents should not only in- fer othersâ mental states, but also maintain and up- date interoperable representations of what is jointly understood. 4.5 Implications for Agent Design Taken together, the literature suggests that repre- senting human reasoning in AI requires at least five ingredients: 1.bounded rather than perfectly rational decision processes, 2.the ability to combine fast heuristics with slow deliberation, 3. common-sense world knowledge, 4.ToM-like reasoning about othersâ beliefs and intentions, and 5.dynamic mental models of shared tasks and so- cial interactions. 4 Current agents often implement fragments of this picture, but rarely the whole. The result is a recurring mismatch: agents can complete tasks but still behave in ways that feel inhuman, socially clumsy, or insensitive to context. 5 Language, Culture, and Social Context One of the most distinctive and underdeveloped aspects of human-centered agents is cultural and sociolinguistic adaptation. Human communication is never purely semantic. Meaning is mediated by culture, audience, norms, and interactional expec- tations. 5.1 Culture as a Determinant of Behavior Cross-cultural psychology provides foundational evidence that values, communication styles, and behavioral expectations vary systematically across societies. Hofstede (2001) characterized differ- ences along dimensions such as individualism vs. collectivism, power distance, uncertainty avoid- ance, masculinity/femininity, long-term orientation, and indulgence/restraint. Hall (1976) further dis- tinguished between high-context and low-context communication, highlighting how different cul- tures encode meaning explicitly or implicitly. These frameworks are, of course, abstractions and have been critiqued when used simplistically. Nevertheless, they remain influential reference points for understanding cultural variability. For AI, the key lesson is that the âsameâ instruction, explanation, or negotiation strategy may be inter- preted differently across cultural settings. 5.2 Sociolinguistics and Pragmatics Sociolinguistics shows that language varies as a function of region, class, gender, institutional con- text, and community norms. Meaning arises not only from literal content but from style, stance, pre- supposition, and politeness. As a result, agents that are fluent but socially generic may still be culturally inappropriate or pragmatically ineffective. This issue becomes especially important in col- laborative environments where language performs social work: signaling respect, maintaining face, negotiating disagreement, or expressing uncer- tainty. Human-centered agents must therefore go beyond grammatical fluency toward pragmatic and sociocultural competence. 5.3 Cultural Alignment in LLMs Recent NLP research has begun to investigate whether LLMs reflect cultural diversity or dispro- portionately encode dominant linguistic and ideo- logical patterns. Cao et al. (2023) show that Chat- GPT aligns more strongly with American cultural patterns and adapts less effectively to other con- texts, while English prompting can flatten cross- cultural differences. AlKhamissi et al. (2024) find that cultural alignment improves when models are prompted in the dominant language of a culture and when pretraining better reflects relevant lan- guage mixtures. Wang et al. (2024) introduce CDE- val, a benchmark measuring cultural dimensions in LLM behavior, and show that models vary sub- stantially across domains and dimensions. Masoud et al. (2025) similarly evaluate cultural alignment through Hofstede-inspired analysis and report un- even adaptation across regions and languages. Together, these studies establish an important point: current models can appear globally deploy- able while remaining culturally asymmetric. They often reflect a narrow center of gravity rather than genuinely pluralistic social understanding. 5.4 Context-Aware Dialogue and Social Intelligence Another line of work examines context-aware and socially aware dialogue systems. While recent models are far more fluent than earlier dialogue systems, contextual adaptation often remains shal- low: models can maintain local coherence but may miss deeper social cues, interactional histories, or culturally embedded expectations. This gap is am- plified in high-stakes settings such as healthcare, education, administration, and conflict mediation, where how something is said may be as important as what is said. 5.5 Implications for Human-Centered Agents Culturally and socially aware agents require more than multilingual capacity. They need: ⢠representations of social and cultural context, â˘adaptive communication strategies sensitive to audience norms, â˘mechanisms to avoid collapsing diverse value systems into a single default style, and â˘evaluation frameworks that measure not only cor- rectness, but cultural appropriateness and social fit. 5 This is one of the clearest gaps in current agent research: language generation has advanced rapidly, but genuine cultural understanding remains limited. 6 Human Values and Belief Systems An agent acting on behalf of humans must not only understand what people say; it must also reason about what they care about. This brings the survey into the domain of alignment, ethics, preference learning, and computational representations of be- lief. 6.1 From Objective Optimization to Alignment As AI systems become more capable, alignment has emerged as a central research concern. Broadly speaking, alignment asks whether system behavior accords with human intentions and values (Russell, 2019; Ji et al., 2024). This issue becomes particu- larly important when agents make autonomous de- cisions under uncertainty or operate over extended horizons. A key challenge is that many AI systems opti- mize measurable proxies rather than human goals themselves. Reward misspecification, distribution shift, and objective gaming can all produce behav- ior that is technically competent but normatively undesirable (Amodei et al., 2016; Ji et al., 2024). In multi-agent contexts, this challenge becomes dy- namic: interaction among agents may amplify or distort misalignment (Carichon et al., 2025). 6.2 Preference Learning Preference learning has become one of the most influential approaches to alignment in foundation models. Early work on inverse reinforcement learn- ing sought to infer reward functions from observed behavior (Ng and Russell, 2000). More recent ap- proaches include reinforcement learning from hu- man feedback (RLHF) (Ouyang et al., 2022), di- rect preference optimization (DPO), constitutional methods, and broader forms of human preference modeling. Recent surveys provide a detailed taxonomy of how preferences are sourced, represented, and opti- mized in LLM alignment (Jiang et al., 2024; Gao et al., 2024). These works make clear that pref- erence alignment is not a single technique but an ecosystem of choices involving feedback collec- tion, reward modeling, policy optimization, and evaluation. 6.3 Pluralistic and Personalized Alignment A particularly important development for your topic is the move from generic âhuman preferenceâ to pluralistic and personalized alignment. Human values are not homogeneous. They vary across individuals, communities, and cultures. Recent surveys on pluralistic and personalized alignment argue that one-size-fits-all alignment is insufficient for real-world deployment (Xie et al., 2025; Guan et al., 2025). Benchmarks such as PERSONA em- phasize the challenge of representing diverse and potentially conflicting user value profiles (Castri- cato et al., 2025). This literature is directly relevant to your re- search framework. If agents are to act on behalf of humans in cooperative environments, they must preserve not only generalized norms but also user- specific and culturally situated valuesâwhile re- maining within ethical boundaries. 6.4 Belief Representation and Moral Reasoning Modeling values also requires modeling beliefs. Human behavior depends on what people perceive to be true, what they regard as important, and which norms they think apply in a given context. Some recent work proposes valueâbeliefânorm reasoning as a means of improving opinion alignment and persona-sensitive prediction (Do et al., 2025). This is promising because it moves beyond static per- sona labels toward explanatory structures linking values, beliefs, and decisions. Computational ethics and machine ethics further ask how moral concerns can be represented algo- rithmically (Wallach and Allen, 2009). Although the field remains philosophically and technically unresolved, its relevance for agent systems is un- deniable. Agents deployed in socially embedded roles must make choices that affect autonomy, fair- ness, confidentiality, harm, respect, and legitimacy. 6.5 Implications for Human-Centered Agents The literature suggests that value-aware agents need: 1. methods for learning and updating human pref- erences, 2.support for pluralism rather than a single univer- salized user model, 3.explicit mechanisms for linking beliefs, values, and decisions, 6 4.safeguards against misalignment in interactive and multi-agent settings, and 5.evaluation standards that capture normative ade- quacy rather than only task success. This remains one of the largest unresolved ar- eas in human-centered agent design. Current sys- tems often optimize a measurable objective, but real-world human representation demands richer normative modeling. 7 Human-Agent Collaboration If agents are to become teammates rather than tools, they must collaborate in ways that humans perceive as intelligible, trustworthy, and adaptive. Human- agent collaboration is therefore not a peripheral application area; it is a test of whether human- centered design has succeeded. 7.1 From Automation to Teaming Traditional AI systems often operated as automa- tion tools: they executed subtasks delegated by hu- mans. Human-agent collaboration moves beyond this paradigm toward joint activity, where humans and agents contribute complementary strengths. Recent overviews from NLP and HCI emphasize that the key research question is no longer whether AI can perform tasks alone, but how AI and hu- mans can work together most effectively (Wu et al., 2025; Yang et al., 2024). This reframing matters because collaboration re- quires more than capability. A strong autonomous solver may still be a poor teammate if it fails to communicate uncertainty, infer preferences, or co- ordinate interactively. 7.2 Trust and Explainability Trust is central to collaboration. People are more likely to rely on systems that are transparent, pre- dictable, and responsive. Explainable AI has long been motivated by the need to render system behav- ior understandable (Doshi-Velez and Kim, 2017). In the LLM era, this concern has expanded into a rapidly growing research area on explainable and transparent language models (Cambria et al., 2024; Palikhe et al., 2025). However, explainability should not be reduced to post-hoc verbalization. In collaborative set- tings, useful explanations must be relevant to the userâs goals, calibrated to expertise, and actionable. This is especially important in high-stakes domains where over-trust and under-trust can both be harm- ful. 7.3 Shared Mental Models and Mixed-Initiative Interaction The teamwork literature stresses the importance of shared mental models: collaborators need an overlapping understanding of task structure, roles, and expectations (Salas et al., 2005). In human- agent interaction, this implies that agents should be capable of communicating plans, updating goals, and negotiating task boundaries. Recent work on decision-oriented dialogue for- malizes scenarios in which AI assistants collabo- rate with humans to reach complex decisions, show- ing that current language models still underperform human assistants despite fluent dialogue (Lin et al., 2024). Similarly, work on reinforcement learning- based human-agent collaboration for complex tasks shows that strategic human intervention can sub- stantially improve outcomes when agents are not fully reliable (Feng et al., 2024). These studies collectively suggest that collaboration should be treated as an optimization target in its own right, not merely as a side-effect of agent competence. 7.4 What Makes an Agent a Teammate? An effective teammate typically exhibits: â˘awareness of the humanâs goals, limitations, and preferences, ⢠timely and context-appropriate communication, ⢠transparency about confidence and rationale, ⢠adaptive turn-taking and initiative management, ⢠willingness to defer, ask, or clarify when neces- sary. Most current agents only partially satisfy these criteria. They are often impressive assistants but inconsistent teammates. 8 Multi-Agent Systems and Social Coordination Human-centeredness becomes even more challeng- ing when moving from individual agents to soci- eties of agents. Multi-agent systems (MAS) have a long history in AI, but LLM-based multi-agent sys- tems have revitalized the field by enabling natural- language communication, role specialization, and emergent collaboration (Wooldridge, 2009; Guo et al., 2024; Chen et al., 2025c; Tran et al., 2025). 7 8.1 Classical Foundations Classical MAS research studied distributed prob- lem solving, coordination, communication pro- tocols, negotiation, and collective behavior (Wooldridge, 2009). Many applications involved autonomous entities with local information that needed to coordinate under partial observability or distributed objectives. These foundations remain highly relevant. Mod- ern LLM-based MAS often reinvent classical MAS patterns centralized orchestration, peer-to-peer communication, role differentiation, and iterative deliberation under a new language-based interface. 8.2 LLM-Based Multi-Agent Systems Recent surveys show that LLM-based MAS are being applied to software development, planning, simulation, social science experiments, information seeking, and collaborative reasoning (Guo et al., 2024; Chen et al., 2025c; Tran et al., 2025). Their appeal lies in combining the flexibility of language- mediated interaction with the modularity of dis- tributed agents. Yet these systems reveal several problems related to your research agenda: â˘persona drift and instability in multi-agent dis- cussions, ⢠conformity and peer-pressure-like effects, ⢠weak preservation of cultural roles or identities, â˘insufficient representation of value conflicts, and â˘limited models of trust and social accountability. For instance, recent work on persona incon- stancy in multi-agent collaboration finds that even when agents are instructed to maintain cultural posi- tions, they may conform or shift inconsistently dur- ing discussion (Baltaji et al., 2024). This suggests that current multi-agent coordination mechanisms are socially thin: they coordinate discourse, but not necessarily identity, value, or norm-consistent behavior. 8.3 The Problem of Multi-Agent Misalignment Another emerging theme is that alignment in MAS cannot be treated as a static property of isolated agents. Carichon et al. (2025) argue that alignment in multi-agent systems is dynamic and interaction- dependent: even individually aligned agents may become misaligned through social organization, competition, or emergent incentives. This obser- vation is particularly important for your proposal because it links alignment directly to social struc- ture. 8.4 Toward Socially Grounded Coordination Culturally and cognitively aware agents require coordination mechanisms that go beyond message passing. They need ways to represent: 1. socially meaningful roles, 2. culturally appropriate communication norms, 3. trust relationships, 4. collective and individual value constraints, and 5. dynamic negotiation over conflicting prefer- ences. This is where many current systems remain un- derdeveloped. They coordinate effectively at the task level, but not yet at the human-social level. 9 Computational Modeling of Human Characteristics The themes reviewed so far are conceptually rich, but they must eventually be implemented. This raises a core engineering question: how can human cognition, social behavior, and value systems be represented computationally? 9.1 Cognitive Architectures Cognitive architectures such as ACT-R and Soar were developed to model human reasoning, mem- ory, and problem solving at a computational level (Anderson et al., 1997; Newell, 1990; Lebiere and Anderson, 2014). Their significance for present- day agent research lies in their explicitness. Unlike purely statistical models, cognitive architectures provide interpretable mechanisms for memory re- trieval, production rules, goal management, and procedural control. Although they predate modern LLMs, these architectures are increasingly relevant in neuro- symbolic and hybrid agent design. Integrating them with LLMs offers one path toward modeling slower, more structured cognition on top of fluent language generation (Wu et al., 2024). 8 9.2 Agent-Based Modeling and Social Simulation Agent-based modeling (ABM) provides another im- portant implementation paradigm. Instead of mod- eling cognition in isolation, ABM simulates popu- lations of interacting agents whose local rules gen- erate emergent macro-level behavior (Bonabeau, 2002). This tradition has long been used in eco- nomics, epidemiology, political science, and com- putational social science. The recent rise of LLM-driven generative agents has renewed interest in the use of language models for simulating believable human-like social behav- ior (Park et al., 2023; Wang et al., 2023). Such systems show how memory, reflection, planning, and interaction can produce emergent social phe- nomena. However, they also reveal a key limi- tation: believable behavior does not necessarily imply faithful modeling of culture, values, or psy- chological diversity. 9.3 Persona and User Modeling User and persona modeling aim to represent sta- ble or evolving characteristics of individuals, such as preferences, demographics, habits, beliefs, or communication styles. Recent work has explored trainable role-playing agents (Shao et al., 2023), persona prompting and its limitations (Hu and Col- lier, 2024; Giorgi et al., 2024; Lutz et al., 2025), dynamic persona updating (Chen et al., 2025a), and recommendation-oriented persona modeling (Shi et al., 2025). This line of research is highly relevant to your motivation because it addresses the technical chal- lenge of representing individual human differences. However, existing approaches often remain shal- low or fragile. Persona labels can produce stereo- typed outputs, and prompt-based personas may be inconsistent across tasks or interaction histories. Richer user modeling must therefore move toward interaction-grounded, belief- and value-aware rep- resentations rather than static descriptors alone. 9.4 Digital Humans and Human-Like Agents Several recent systems explicitly aim to build human-like generative agents or digital characters (Park et al., 2023; Wang et al., 2023; Shao et al., 2023). These efforts are important because they make the problem concrete: what kinds of mem- ory, reflection, emotion, need states, and persona structures make an artificial agent behave more believably like a person? Yet believability and human-centered represen- tativeness are not identical. A convincing conver- sational persona may still fail to preserve cultural nuance, user values, or normative intent. Thus, the future challenge is not only to simulate human-like behavior, but to simulate human-centered behavior grounded in real social and ethical constraints. 10 Comparison of Traditional and Human-Centered Agents Table 1 summarizes the central contrast developed in this survey. 11 Illustrative Application Scenario To make the discussion concrete, consider a multi- cultural healthcare support environment in which multiple AI agents assist clinicians, patients, and administrators. A task-centric system might perform well on information retrieval, appointment scheduling, or treatment recommendation. Yet such a system could still fail in important ways: it may commu- nicate too directly in one cultural setting and too vaguely in another; it may ignore patient beliefs about treatment, family involvement, or authority; it may explain decisions in a technically correct but socially ineffective manner; or it may coordinate with other agents in a way that satisfies a global optimization objective while violating local value expectations. A human-centered system would need to: â˘represent patient preferences and culturally shaped expectations, ⢠adapt language and explanation style to audience and context, â˘communicate uncertainty to clinicians transpar- ently, ⢠preserve confidentiality and ethical constraints, â˘coordinate among agents while maintaining shared mental models and trust. Thisexampleillustrateswhyhuman- centeredness is not an optional layer of per- sonalization. In many real domains, it is integral to competence. 9 DimensionTraditional / Task-Centric AgentsHuman-Centered Agents Primary ObjectiveTask completion, optimization, predictionContext-sensitive representation and action on behalf of humans Decision-MakingUtility-driven or heuristic task planningBounded rationality, reflective reasoning, context- aware judgment Reasoning ModelLogical or statistical inference Hybrid cognition: common sense, dual-process elements, mental models Language UseFluent,oftengeneric,instruction- following Pragmatic, audience-aware, culturally adaptive communication Cultural AwarenessMinimal or implicit Explicit modeling of norms, regional context, and sociolinguistic variation Values and BeliefsOften absent or collapsed into generic safety alignment Preference-aware, pluralistic, belief-sensitive, norm-constrained decision-making Human InteractionUser as requester or supervisorHuman as teammate, stakeholder, or represented principal ExplainabilityOptional add-onCore requirement for trust, coordination, and ac- countability Coordination in MASTask-level communication and synchro- nization Socially grounded coordination including roles, trust, negotiation, and norm awareness User ModelingStatic profiles or noneDynamic representations of preferences, beliefs, history, and social context EvaluationAccuracy, efficiency, benchmark perfor- mance Human-centered success: appropriateness, trust- worthiness, cultural fit, collaboration quality Table 1: Conceptual comparison between traditional task-centric agents and the human-centered agent perspective advocated in this survey. 12Open Challenges and Future Research Directions The literature points toward several urgent research directions. 12.1 Unified Cognitive-Social Agent Architectures Current systems typically specialize in one dimen- sion: cognitive modeling, cultural adaptation, align- ment, or collaboration. Future work should aim at unified architectures in which memory, planning, social context, value constraints, and communica- tion policies are modeled jointly rather than bolted together. 12.2 Representing Cultural Context Without Stereotyping Cultural adaptation is necessary, but simplistic cul- tural labeling can lead to overgeneralization or stereotyping. Research is needed on probabilis- tic, interaction-sensitive, and self-updating models of culture that treat users as individuals embedded in social contexts rather than as fixed demographic templates. 12.3 Pluralistic and Dynamic Value Alignment Alignment research increasingly recognizes value pluralism, but scalable implementations remain dif- ficult. Future agents must balance universal safety requirements with individual and community- specific preference adaptation, including mecha- nisms for revising value models over time. 12.4 Robust Social Reasoning Theory of Mind-like capabilities, negotiation strategies, and shared mental model tracking are promising, but current evaluations remain mostly benchmark-centric. Future work needs interactive, ecologically valid tests of social reasoning in long- horizon human-agent and agent-agent collabora- tion. 12.5 Trust-Aware Multi-Agent Coordination Trust is not just a human-agent issue. In multi- agent societies, trust calibration, reputation, and ac- countability may affect whether collaborative struc- tures remain aligned and stable. This suggests the need for MAS frameworks that explicitly model trust, role legitimacy, and social responsibility. 12.6 Evaluation Beyond Task Accuracy Human-centered agents require new evaluation cri- teria. Important metrics likely include: ⢠cultural appropriateness, ⢠value conformity under pluralistic settings, ⢠explanation usefulness, ⢠collaboration quality, 10 ⢠persona consistency, ⢠robustness of social reasoning, ⢠long-term trustworthiness. Benchmarks like CDEval are a useful start (Wang et al., 2024), but the field still lacks compre- hensive evaluation frameworks that cover all major human-centered dimensions. 12.7 Ethics, Governance, and Societal Oversight Finally, human-centered agent research cannot re- main purely technical. Agents acting in socially embedded roles raise questions of accountability, consent, representation, fairness, governance, and institutional legitimacy. Building such systems re- sponsibly requires interdisciplinary collaboration across AI, HCI, psychology, sociology, linguistics, and ethics. 13 Discussion A coherent pattern emerges across the literature sur- veyed here. Advances in LLMs and autonomous agents have made it possible to build systems that appear increasingly intelligent in traditional com- putational terms. They can reason, converse, plan, retrieve information, invoke tools, and collaborate in distributed workflows. However, this progress has exposed a deeper limitation: intelligence alone does not yield effective participation in human en- vironments. Human environments are not merely informa- tional they are social, cultural, and normative. Peo- ple do not decide only by calculating utility; they reason under constraints, emotion, uncertainty, and social expectation. They do not communicate only through literal semantics; they speak through con- text, audience, and culture. They do not collaborate simply by dividing labor; they cooperate through trust, mutual modeling, explanation, and negotiated norms. And when they act on behalf of others, they must preserve values, beliefs, and responsibilities. The literature has begun to address each of these dimensions: ⢠cognitive science provides theories of bounded rationality, dual-process reasoning, common sense, and theory of mind; â˘cultural and sociolinguistic research clarifies how communication and behavior vary across con- texts; â˘alignment and preference learning provide tools for modeling values and human intentions; â˘HCI and collaborative AI emphasize trust, ex- plainability, and teaming; â˘MAS research provides infrastructures for dis- tributed interaction; ⢠cognitive architectures, ABM, persona model- ing, and generative agents offer implementation pathways. Yet the key problem remains integration. Exist- ing systems typically combine strong task compe- tence with shallow representations of the human dimensions in which they operate. As a result, they can be competent but socially brittle, fluent but cul- turally narrow, aligned in the aggregate but not in the particular, collaborative in protocol but not in understanding. This survey therefore positions the next frontier of AI agents not simply as greater autonomy, but as human-centered autonomy: autonomy constrained and informed by cognition, culture, values, and cooperation. 14 Conclusion This survey has argued that the evolution of AI agents from rule-based systems to LLM-based au- tonomous and multi-agent architecturesâhas cre- ated an opportunity and a necessity. The oppor- tunity lies in building systems that can partici- pate meaningfully in complex human environments. The necessity lies in confronting the fact that such participation requires more than language genera- tion and task optimization. We reviewed six major literatures: the evolu- tion of intelligent agents, human cognition and decision-making, language and culture, human val- ues and belief systems, human-agent collaboration, and multi-agent coordination with computational models of human behavior. Across these areas, the same conclusion recurs: the most important proper- ties of human-centered interaction are still weakly represented in current agents. The central research gap can therefore be stated as follows: Existing AI and multi-agent systems primarily model intelligence as predic- tion, optimization, and task comple- tion. However, agents intended to act 11 on behalf of humans in cooperative en- vironments require deeper computa- tional representations of human cog- nition, cultural context, value systems, and social behavior. Despite substan- tial advances in individual areas such as cognitive modeling, cultural intel- ligence, and AI alignment, there re- mains no widely adopted framework that integrates these dimensions into autonomous agents capable of authen- tic human-centered collaboration. Addressing this gap is likely to define a major part of the next stage of agent research. The future of multi-agent systems is not simply more agents or more autonomy. It is agents that can reason with, communicate like, and collaborate on behalf of humans in ways that are contextually appropri- ate, culturally aware, value-aligned, and socially trustworthy. References Badr AlKhamissi, Muhammad ElNokrashy, Mai Alkhamissi, and Mona Diab. 2024. Investigating cultural alignment of large language models. In Pro- ceedings of ACL. Dario Amodei, Chris Olah, Jacob Steinhardt, Paul Christiano, John Schulman, and Dan ManĂŠ. 2016. Concrete problems in ai safety.arXiv preprint arXiv:1606.06565. John R. Anderson, Michael Matessa, and Christian Lebiere. 1997. Act-r: A theory of higher level cog- nition and its relation to visual attention. Human- Computer Interaction, 12(4):439â462. Razan Baltaji, Babak Hemmatian, and Lav Varshney. 2024. Conformity, confabulation, and impersonation: Persona inconstancy in multi-agent llm collaboration. Proceedings of the 2nd Workshop on Cross-Cultural Considerations in NLP. Christopher M. Bishop. 2006. Pattern Recognition and Machine Learning. Springer. Eric Bonabeau. 2002. Agent-based modeling: Meth- ods and techniques for simulating human systems. Proceedings of the National Academy of Sciences, 99:7280â7287. Tom B. Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared Kaplan, Prafulla Dhariwal, and 1 oth- ers. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems. Erik Cambria, Lorenzo Malandri, Fabio Mercorio, Navid Nobani, and Andrea Seveso. 2024. Xai meets llms: A survey of the relation between explain- able ai and large language models. arXiv preprint arXiv:2407.15248. Yong Cao, Li Zhou, Seolhwa Lee, Laura Cabello, Min Chen, and Daniel Hershcovich. 2023. Assessing cross-cultural alignment between chatgpt and human societies: An empirical study. In Proceedings of the First Workshop on Cross-Cultural Considerations in NLP. Florian Carichon, Aditi Khandelwal, Marylou Fauchard, and Golnoosh Farnadi. 2025. The coming crisis of multi-agent misalignment: Ai alignment must be a dynamic and social process. arXiv preprint arXiv:2506.01080. Louis Castricato, Nathan Lile, Rafael Rafailov, Jan- Philipp Fränken, and Chelsea Finn. 2025. Persona: A reproducible testbed for pluralistic alignment. In Proceedings of COLING. Aili Chen, Chengyu Du, Jiangjie Chen, Jinghan Xu, Yikai Zhang, Siyu Yuan, Zulong Chen, Liangyue Li, and Yanghua Xiao. 2025a. Deeper insight into your user: Directed persona refinement for dynamic persona modeling. In Proceedings of ACL. Ruirui Chen, Weifeng Jiang, Chengwei Qin, and Che- ston Tan. 2025b.Theory of mind in large lan- guage models: Assessment and enhancement. arXiv preprint arXiv:2505.00026. Shuaihang Chen, Yuanxing Liu, Wei Han, Weinan Zhang, and Ting Liu. 2025c.A survey on llm-based multi-agent system: Recent advances and new frontiers in application. arXiv preprint arXiv:2412.17481. Xuan Long Do, Kenji Kawaguchi, Min-Yen Kan, and Nancy F. Chen. 2025. Aligning large language mod- els with human opinions through persona selection and valueâbeliefânorm reasoning. In Proceedings of COLING. Finale Doshi-Velez and Been Kim. 2017. Towards a rigorous science of interpretable machine learning. arXiv preprint arXiv:1702.08608. Jonathan St. B. T. Evans. 2008. Dual-processing ac- counts of reasoning, judgment, and social cognition. Annual Review of Psychology, 59:255â278. Xueyang Feng, Zhi-Yuan Chen, Yujia Qin, Yankai Lin, Xu Chen, Zhiyuan Liu, and Ji-Rong Wen. 2024. Large language model-based human-agent collab- oration for complex task solving. In Findings of EMNLP. Bofei Gao, Feifan Song, Yibo Miao, Zefan Cai, Zhe Yang, Liang Chen, Helan Hu, Runxin Xu, and 1 oth- ers. 2024. Towards a unified view of preference learning for large language models: A survey. arXiv preprint arXiv:2409.02795. 12 Salvatore Giorgi, Tingting Liu, Ankit Aich, Kelsey Jane Isman, Garrick Sherman, Zachary Fried, Joao Sedoc, Lyle Ungar, and Brenda Curtis. 2024. Modeling human subjectivity in llms using explicit and implicit human factors in personas. In Findings of EMNLP. Jian Guan, Junfei Wu, Jia-Nan Li, Chuanqi Cheng, and Wei Wu. 2025.A survey on personalized alignment â the missing piece for large language models in real-world applications. arXiv preprint arXiv:2503.17003. Taicheng Guo, Xiuying Chen, Yaqi Wang, Ruidi Chang, Shichao Pei, Nitesh V. Chawla, Olaf Wiest, and Xi- angliang Zhang. 2024. Large language model based multi-agents: A survey of progress and challenges. arXiv preprint arXiv:2402.01680. Edward T. Hall. 1976. Beyond Culture. Anchor Books. Geert Hofstede. 2001. Cultureâs Consequences: Com- paring Values, Behaviors, Institutions and Organiza- tions Across Nations. Sage. Tiancheng Hu and Nigel Collier. 2024. Quantifying the persona effect in llm simulations. In Proceedings of ACL. Jie Huang and Kevin Chen-Chuan Chang. 2023. To- wards reasoning in large language models: A survey. Findings of ACL. Jiaming Ji, Tianyi Qiu, Boyuan Chen, Borong Zhang, Hantao Lou, Kaile Wang, Yawen Duan, and 1 others. 2024. Ai alignment: A comprehensive survey. arXiv preprint arXiv:2310.19852. Ruili Jiang, Kehai Chen, Xuefeng Bai, Zhixuan He, Juntao Li, Muyun Yang, Tiejun Zhao, Liqiang Nie, and Min Zhang. 2024. A survey on human preference learning for large language models. arXiv preprint arXiv:2406.11191. Daniel Kahneman. 2011. Thinking, Fast and Slow. Far- rar, Straus and Giroux. Takeshi Kojima, Shixiang Gu, Machel Reid, Yutaka Matsuo, and Yusuke Iwasawa. 2022. Large lan- guage models are zero-shot reasoners. arXiv preprint arXiv:2205.11916. Christian Lebiere and John R. Anderson. 2014. The act- r cognitive architecture: Principles and applications. In The Oxford Handbook of Cognitive Engineering. Oxford University Press. Jessy Lin, Nicholas Tomlin, Jacob Andreas, and Jason Eisner. 2024. Decision-oriented dialogue for human- ai collaboration. Transactions of the Association for Computational Linguistics, 12:892â911. Marlene Lutz, Indira Sen, Georg Ahnert, Elisa Rogers, and Markus Strohmaier. 2025. The prompt makes the person(a): A systematic evaluation of sociodemo- graphic persona prompting for large language models. In Findings of EMNLP. Reem I. Masoud, Ziquan Liu, Martin Ferianc, Philip Treleaven, and Miguel Rodrigues. 2025. Cultural alignment in large language models: An explanatory analysis based on hofstedeâs cultural dimensions. In Proceedings of COLING. John McCarthy. 1959. Programs with common sense. Proceedings of the Teddington Conference on the Mechanization of Thought Processes. Allen Newell. 1990. Unified Theories of Cognition. Harvard University Press. Andrew Y. Ng and Stuart Russell. 2000. Algorithms for inverse reinforcement learning. In Proceedings of ICML. Hieu Minh Nguyen. 2025. A survey of theory of mind in large language models: Evaluations, representations, and safety risks. arXiv preprint arXiv:2502.06470. Nils J. Nilsson. 1980. Principles of Artificial Intelli- gence. Springer. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, and 1 others. 2022. Training language models to follow instructions with human feedback. Advances in Neural Information Processing Systems. Avash Palikhe, Zhenyu Yu, Zichong Wang, and Wenbin Zhang. 2025. Towards transparent ai: A survey on explainable large language models. arXiv preprint arXiv:2506.21812. Joon Sung Park, Joseph C. OâBrien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. 2023. Generative agents: Interactive sim- ulacra of human behavior. In Proceedings of UIST. David Premack and Guy Woodruff. 1978. Does the chimpanzee have a theory of mind? Behavioral and Brain Sciences, 1(4):515â526. Stuart Russell. 2019. Human Compatible: Artificial Intelligence and the Problem of Control. Viking. Stuart Russell and Peter Norvig. 2021. Artificial Intelli- gence: A Modern Approach, 4 edition. Pearson. Eduardo Salas, Dana E. Sims, and C. Shawn Burke. 2005. Is there a âbig fiveâ in teamwork?Small Group Research, 36(5):555â599. Yunfan Shao, Linyang Li, Junqi Dai, and Xipeng Qiu. 2023. Character-llm: A trainable agent for role- playing. In Proceedings of EMNLP. Yunxiao Shi, Wujiang Xu, Zeqi Zhang, Xing Zi, Qiang Wu, and Min Xu. 2025. Personax: A recommen- dation agent-oriented user modeling framework for long behavior sequence. In Findings of ACL. Herbert A. Simon. 1957. Models of Man: Social and Rational. Wiley. 13 Khanh-Tung Tran, Dung Dao, Minh-Duong Nguyen, Quoc-Viet Pham, Barry OâSullivan, and Hoang D. Nguyen. 2025.Multi-agent collaboration mech- anisms:A survey of llms.arXiv preprint arXiv:2501.06322. Amos Tversky and Daniel Kahneman. 1974. Judgment under uncertainty: Heuristics and biases. Science, 185(4157):1124â1131. Wendell Wallach and Colin Allen. 2009. Moral Ma- chines: Teaching Robots Right from Wrong. Oxford University Press. Lei Wang, Chen Ma, Xueyang Feng, Zeyu Zhang, Hao Yang, Jingsen Zhang, Zhiyuan Chen, Jiakai Tang, Xu Chen, Yankai Lin, Wayne Xin Zhao, Zhewei Wei, and Ji-Rong Wen. 2025. A survey on large language model based autonomous agents. Frontiers of Com- puter Science. Yuhang Wang, Yanxu Zhu, Chao Kong, Shuyu Wei, Xi- aoyuan Yi, Xing Xie, and Jitao Sang. 2024. Cdeval: A benchmark for measuring the cultural dimensions of large language models. In Proceedings of the 2nd Workshop on Cross-Cultural Considerations in NLP. Zhilin Wang, Yu Ying Chiu, and Yu Cheung Chiu. 2023. Humanoid agents: Platform for simulating human- like generative agents. In Proceedings of EMNLP System Demonstrations. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, Ed H. Chi, Quoc V. Le, and Denny Zhou. 2022. Chain-of-thought prompt- ing elicits reasoning in large language models. In Advances in Neural Information Processing Systems. Michael Wooldridge. 2009. An Introduction to MultiA- gent Systems. Wiley. Sherry Wu, Diyi Yang, Joseph Chang, Marti A. Hearst, and Kyle Lo. 2025. Human-ai collaboration: How ais augment human teammates. In Proceedings of ACL Tutorial Abstracts. Siyu Wu, Alessandro Oltramari, Jonathan Francis, C. Lee Giles, and Frank E. Ritter. 2024. Cognitive llms: Towards integrating cognitive architectures and large language models for manufacturing decision- making. arXiv preprint arXiv:2408.09176. Zhiheng Xi, Wenxiang Chen, Xin Guo, Wei He, Yiwen Ding, Boyang Hong, Ming Zhang, Junzhe Wang, Senjie Jin, Enyu Zhou, and 1 others. 2023. The rise and potential of large language model based agents: A survey. arXiv preprint arXiv:2309.07864. Zhouhang Xie, Junda Wu, Yiran Shen, Yu Xia, Xintong Li, Aaron Chang, Ryan Rossi, Sachin Kumar, and 1 others. 2025. A survey on personalized and plural- istic preference alignment in large language models. arXiv preprint arXiv:2504.07070. Diyi Yang, Tongshuang Wu, and Marti A. Hearst. 2024. Human-ai interaction in the age of llms. In Proceed- ings of NAACL Tutorial Abstracts. 14