Paper deep dive
EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation
Jingzhe Lin, Hengbin Yu, Yongdan Zeng, Fangwei Zhong
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 7/9/2026, 3:15:41 AM
Summary
EduMirror is a multi-agent simulation framework designed to model educational social dynamics using value-driven agents grounded in psychological theories. It integrates a Concordia-based simulation engine, a dual-track measurement protocol, and an intervention engine to enable scalable, ethically safe in silico research, hypothesis testing, and counterfactual analysis of phenomena like school bullying and group cooperation.
Entities (10)
Relation Signals (13)
EduMirror â implements â Value-driven Agents
confidence 96% ¡ We provide configurable education-oriented agent forms, including value-driven agents grounded in psychological needs and social value orientation
EduMirror â uses â Concordia
confidence 95% ¡ Built on Concordia (Vezhnevets et al., 2023), we utilize natural language as the primary medium for simulation.
EduMirror â provides â Dual-track Measurement Protocol
confidence 94% ¡ together with a dual-track measurement protocol for quantifying observable behaviors and latent psychological states.
Value-driven Agents â groundedin â Social Value Orientation
confidence 93% ¡ value-driven agents grounded in psychological needs and social value orientation
Value-driven Agents â groundedin â Maslow's Hierarchy of Needs
confidence 92% ¡ This system grounds intrinsic motivation in established psychological theories, drawing on Maslowâs Hierarchy of Needs
EduMirror â validatedon â School Bullying
confidence 92% ¡ We validate the realism and usability of EduMirror through case studies on school bullying and group cooperation
EduMirror â validatedon â Group Cooperation
confidence 92% ¡ We validate the realism and usability of EduMirror through case studies on school bullying and group cooperation
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Understanding how educational social dynamics evolve is critical for informing effective educational policies and counterfactual interventions. However, traditional methods face a fundamental dilemma: observational studies often lack causal power, while controlled experiments are frequently constrained by ethical concerns. Although LLM-based multi-agent simulations offer a scalable in silico alternative, existing approaches remain limited by weak psychological grounding and insufficient measurement of latent psychological states. To address this, we introduce EduMirror, a multi-agent simulator for the scientific study of educational social dynamics. We provide configurable education-oriented agent forms, including value-driven agents grounded in psychological needs and social value orientation, together with a dual-track measurement protocol for quantifying observable behaviors and latent psychological states. We validate the realism and usability of EduMirror through case studies on school bullying and group cooperation, as well as broader evaluations across diverse educational scenarios. The results show that EduMirror generates educational social dynamics that are realistic, theory-consistent, and measurable by empirical criteria. These properties enable structured in silico educational research, providing a computational tool for hypothesis testing and counterfactual intervention analysis in educational science. Project page: this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2606.07948v1
- Canonical: https://arxiv.org/abs/2606.07948v1
Trouble viewing inline? Open PDF directly â
Full Text
158,116 characters extracted from source content.
Expand or collapse full text
EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Jingzhe Lin * 1 2 3 Hengbin Yu * 4 Yongdan Zeng * 1 5 Fangwei Zhong 1 2 3  Project Page: https://edumirror.net what-if:Unsolved Struggle? To address Aliceâs difficulties in understanding classroom content, designing and implementing a personalized improvement plan can best fulfill her multidimensional value needs... what-if:AI Guidance? what-if:Personalized Help? what-if:New Method? Individual Behavior HomeâSchool Interaction Peer and Group DynamicsClassroom and School Culture Diverse Real-World Educational Scenarios EduMirror what-if... Simulation Intervention Practical Suggestions for Addressing Real-World Problems Figure 1. An illustration of the core concept behind EduMirror. Like a mirror, EduMirror models diverse educational social dynamics as a complex multi-agent system, enabling reflection on real-world practices and projecting the potential outcomes of different interventions. Abstract Understanding how educational social dynamics evolve is critical for informing effective educa- tional policies and counterfactual interventions. However, traditional methods face a fundamen- tal dilemma: observational studies often lack causal power, while controlled experiments are frequently constrained by ethical concerns. Al- though LLM-based multi-agent simulations of- * These authors contributed equally and are listed alphabeti- cally. 1 School of Artificial Intelligence, Beijing Normal Univer- sity, Beijing, China 2 Beijing Key Laboratory of Artificial Intelli- gence for Education, Beijing, China 3 Engineering Research Center of Intelligent Technology and Educational Application, Ministry of Education, Beijing, China 4 School of Systems Science, Bei- jing Normal University, Beijing, China 5 Information Hub, The Hong Kong University of Science and Technology (Guangzhou), Guangzhou, China. Correspondence to: Fangwei Zhong<fang- weizhong@bnu.edu.cn>. Proceedings of the43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s). fer a scalable in silico alternative, existing ap- proaches remain limited by weak psychological grounding and insufficient measurement of latent psychological states. To address this, we intro- duce EduMirror, a multi-agent simulator for the scientific study of educational social dynamics. We provide configurable education-oriented agent forms, including value-driven agents grounded in psychological needs and social value orientation, together with a dual-track measurement protocol for quantifying observable behaviors and latent psychological states. We validate the realism and usability of EduMirror through case studies on school bullying and group cooperation, as well as broader evaluations across diverse educational scenarios. The results show that EduMirror gener- ates educational social dynamics that are realistic, theory-consistent, and measurable by empirical criteria. These properties enable structured in silico educational research, providing a computa- tional tool for hypothesis testing and counterfac- tual intervention analysis in educational science. 1 arXiv:2606.07948v1 [cs.MA] 6 Jun 2026 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 1. Introduction Educational social dynamics shape studentsâ development through continuous interactions among peers, teachers, and families, making them a central concern for educational practice and policy. However, these environments are com- plex systems in which developmental outcomes emerge from intricate social interactions (Hymel & Swearer, 2015). While harmful dynamics like bullying cause long-term psychological and developmental consequences (Wolke & Lereya, 2015; Arseneault, 2018), understanding and mitigat- ing these phenomena remains a grand challenge. This is fun- damentally a problem of modeling complex multi-agent systems: outcomes emerge from the interplay of individual psychological states and social network dynamics, making them notoriously difficult to predict or control. Traditional empirical methods, e.g., surveys, observational studies, capture only static correlations and suffer from biases (Latkin et al., 2017; Brenner & DeLamater, 2016; Latkin et al., 2016). More critically, experimental interven- tions, e.g., randomized controlled trials, are often ethically constrained, practically difficult to deploy, or prone to iatro- genic effects (Foulkes & Stringaris, 2023). Consequently, the absence of a systematic framework for operationalizing and simulating the generative mechanisms of educational social dynamics has become a critical bottleneck for coun- terfactual intervention testing. Generative social science offers a pathway to study these phenomena in silico (Epstein, 2006). However, traditional Agent-Based Modeling (ABM) (Adam & Gaudou, 2016) relies on rigid, hand-crafted rules that fail to capture the nu- ance of human psychology, leading to the Fidelity and Cus- tomization Challenge. Conversely, while emerging Large Language Models (LLMs) demonstrate impressive reason- ing capabilities, employing them as believable social agents requires solving the Measurement Challenge, i.e., How to quantify latent psychological states (e.g., self-esteem, peer pressure) that drive behavior but are invisible in the behav- ior. Thus, we need a cognitive computing framework that integrates psychological theory with the generative power of LLMs to enable realistic, interpretable, and measurable social simulation for educational study. To this end, we introduce EduMirror, a comprehensive multi-agent simulation framework designed to âmirrorâ and analyze the generative mechanisms of educational social dynamics, as shown in Figure 1. Built on Concordia (Vezhn- evets et al., 2023), we utilize natural language as the primary medium for simulation. This text-based approach offers two distinct advantages: flexibility, enabling agents to generate open-ended, context-aware responses rather than selecting from rigid pre-defined actions; and scalability, allowing the seamless integration of diverse roles and environments. Leveraging these capabilities, we construct a scene library with more than 20 pre-built educational scenarios, cover- ing typical themes and critical settings such as classrooms, dormitories, and family environments, to simulate and an- alyze agent behaviors across varying social contexts. Our contributions are as follows: 1) A Controllable Simulation Framework for Educa- tional Social Dynamics. EduMirror integrates theory- grounded scenario construction, open-ended multi-agent interaction, and user-driven intervention branching into a unified workflow, enabling controlled counterfactual com- parisons across diverse educational scenarios. 2) Value-Driven Agents for Educational Role Play. We adapt value-driven agents to educational social simulation, grounding role-specific behavior generation in psycholog- ical needs and social value orientations to produce inter- nally motivated, context-sensitive behaviors. EduMirror further provides an extensible agent repository with config- urable profiles, motivations, decision logic, and measure- ment hooks for diverse educational applications. 3) A User Toolkit for Educational Study. We provide an analysis toolkit that turns raw interaction traces into mea- surable and interpretable outcomes. It includes a dual-track measurement protocol for quantifying observable behaviors and latent psychological states, an intervention engine for generating parallel simulation branches, and a log-to-comic visualization module for intuitive qualitative review. 4) Case Studies & Counterfactual Analysis. We validate EduMirror through system-level evaluation and case studies on school bullying and peer interaction. Results show that EduMirror can generate theory-consistent educational social dynamics and support counterfactual comparisons of âwhat- ifâ interventions in an ethically safe digital environment. 2. Related Work Modeling Educational Social Dynamics.Educa- tional social dynamics have primarily been examined through passive observation and post-hoc analysis, in- cluding quantitative self-report surveys and qualitative approaches (Wilson, 1977).Causal inquiry has been pursued through experimental designs, from controlled laboratory studies, such as Banduraâs Bobo doll experi- ment (Bandura et al., 1961), to field-based intervention studies (Manstead & Livingstone, 2008). However, these approaches face substantial methodological and ethical bottlenecks. Methodologically, static measures, such as surveys and interviews, are inherently correlational and prone to response biases, particularly in sensitive social contexts where self-reported data often diverge from objective reality due to social desirability or limited self- perception(DeLara, 2012; OâBrien, 2019; Cole et al., 2005). Ethically, rigorous causal designs, such as randomized 2 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Quantitativ e Qualitative Log-to- Comic Curve Comparison Cognitive Architecture for Agents User Toolkits Value System Simulation Engine Action Obs. Initializes Scene & Rules Scene Setting Game Master (Concordia) Time Management Narration&Rule Enforcement LLM Rater LLM Surveyor Explicit Behavior Implicit State Data Flow Apply Intervention Generate Parallel Timelines Forced Agent Behavior &Scenario Branching Analysis & Visualization Bullying Prevention & Strategy Evaluation Pedagogical Guidance & Peer Dynamics Educational Ecosystem Evaluation Social Value Orientation Traits Memory profile/persona Goal Action Planner GenerateEvaluateSelect A p p l i c a t i o n s Theory of Mind Reflection Psychological health SafetyEsteem Social BelongingMeaning&Growth Strategy ... Metric Configuration Role Profiling Context Definition Theory Integration D2A¡ASVO Iterative Reasoning Agent Context-conditioned Agent Desires â Evaluate â Activity ReAct ¡ BabyAGI Obs. / Goal â Reason / Task Stateâş LLMob ¡ JAG-Concordia Profile / World State â Strategy Student Well-being Evaluation *D2A and ASVO are simplified special cases of our value-driven social agent. Agent Model Repository Value-driven Social Agent Scenario Design Value-Driven Motivation Figure 2. Architecture of the EduMirror simulation platform. EduMirror consists of four main modules: the Agent Model Repository, Scenario Design, the Concordia-based Simulation Engine, and User Toolkits. The Agent Model Repository supports multiple architectures, with our Value-driven Social Agent as the primary model; its cognitive architecture is expanded in the leftmost panel. It also integrates baseline architectures, including iterative reasoning and context-conditioned agents, for controlled comparison. Theory-grounded scenarios are configured through theory integration, context definition, role profiling, and metric configuration, then executed by a Game Master that manages scene setting, time progression, narration, and rule enforcement. The User Toolkits support dual-track measurement, comparative visualization, and intervention-based parallel timelines for systematic analysis of educational social dynamics. controlled trials, are frequently infeasible in educational environments, as manipulating social conditions or withholding necessary interventions violates fundamental ethical standards (National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research, 1979; Wiles, 2012). Consequently, in silico experimentation has emerged as a promising paradigm for ethically exploring counterfactual hypotheses in complex social systems (Squazzoni & Bianchi, 2023). Building on this paradigm, EduMirror integrates an LLM-based game master, value-driven agents, and a âglass-boxâ measurement toolkit, enabling systematic analysis of latent psychological processes that are difficult to access through traditional observational methods. Agent-based Simulation for Education. Traditional Agent- Based Modeling (ABM) in education typically utilizes rule- based frameworks to simulate classroom dynamics and peer influence (Wilensky & Rand, 2015; Gu & Blackmore, 2015; Maroulis et al., 2010). Prominent architectures, including BeliefâDesireâIntention (BDI) models, simulate behavior through predefined logical rules (Georgeff et al., 1998; Silva et al., 2020). While interpretable, such agents exhibit lim- ited psychological realism and struggle to represent nuanced and sometimes irrational human social behavior (Adam & Gaudou, 2016; Taillandier et al., 2019). Conversely, Large Language Models (LLMs) have enabled generative agents with substantially improved plausibility (Park et al., 2023; Zhang et al., 2025; Wang et al., 2025a; Piao et al., 2025). LLM-empowered agent-based simulation has also become a scalable approach for social science, ranging from in- dividual behavior modeling to scenario-level and society- level simulation (Gao et al., 2024; Mou et al., 2026). This paradigm has been further extended by constructing agents grounded in real individualsâ self-reports for large-scale behavior simulation and by developing general social simu- lation platforms that support large-scale, correctable inter- actions among LLM agents (Park et al., 2026; Tang et al., 2025). While recent frameworks have attempted to enhance realism by incorporating desire-driven autonomy (Wang et al., 2025b) and social motivation (Lin et al., 2026), these general-purpose agent architectures remain insufficiently grounded in the domain-specific psychological requirements of educational settings. They often fail to account for the intricate interplay of adolescent developmental needs, so- cial hierarchies, and emotional vulnerability. EduMirror addresses these limitations by constraining LLM generation within an education-oriented cognitive architecture, explic- itly grounding agent behavior in social value orientation and fundamental psychological needs to simulate socially complex educational dynamics. 3 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 3. EduMirror To systematically investigate educational social dynamics through computational experiments, we develop EduMir- ror, a modular and interactive multi-agent simulation plat- form. As shown in Figure 2, EduMirror supports a struc- tured workflow, including theory-grounded scenario design, simulation execution, intervention, measurement, and anal- ysis. This section introduces the main components of the platform. Appendix B provides a detailed walkthrough us- ing a representative example. 3.1. Simulation Framework EduMirror is designed as a general framework for educa- tional social simulation, supporting theory-grounded sce- nario construction, autonomous multi-agent interaction, and user-driven counterfactual experimentation. The framework is organized around scenario construction, agent-based sim- ulation execution, and intervention-based comparison. One core component of this workflow is a systematic five- step procedure that translates abstract educational phenom- ena into computable scenarios. The process begins by (1) selecting a grounding theory (e.g., Social Comparison Theory) to anchor the scenario scientifically. Next, we (2) identify core constructs by deconstructing the theory into fundamental concepts, which then (3) guide agent persona configuration, where we initialize agenttraits,goals, andmemoriesto reflect the chosen theoretical model. To ensure empirical rigor, we (4) operationalize these con- structs with validated scales, and finally (5) establish a dual-track measurement protocol using LLM Raters and Surveyors to quantify agent behaviors and internal states. This structured approach ensures that experimental outputs connect back to specific theoretical constructs. A detailed walkthrough is available in Appendix B. Simulation execution is implemented in a shared environ- ment built on Concordia (Vezhnevets et al., 2023) and or- chestrated by a Game Master (GM), as illustrated in Fig- ure 2. Concordia is adopted as the simulation backbone for its modular agent design and support for persistent inter- action. The GM is responsible for setting the initial scene, narrating events, enforcing rules, and managing time. In our implementation, the GM operates as an LLM-driven orches- trator that takes the environment state and agentsâ actions as input and generates the next environment update, including narration and time advancement. Building on this framework, EduMirror supports user-driven interventions for comparative and counterfactual analysis. Users can save the simulation state at critical junctures and apply interventions to generate parallel branches. EduMir- ror supports two intervention types: Scenario Branching, which modifies the environment or narrative trajectory, and 35.0% 25.0% 15.0% 25.0% 20 scenarios Peer & Group Dynamics: 7 Individual Social Cognition: 5 Classroom Culture: 3 Home-School Dynamics: 5 Figure 3. The distribution of EduMirror scenarios across four kinds of typical themes in the scenario library. Behavior Control, which overrides an individual agentâs action for a single step. These mechanisms enable con- trolled comparison of how contextual changes or individual behaviors affect subsequent social dynamics. 3.2. Theory-Grounded Scenario Design EduMirror currently supports a curated library of 20 pre- designed educational scenarios, each grounded in estab- lished theories and instantiated with fixed role configura- tions and measurement protocols. These scenarios cover four main themes, namely Peer & Group Dynamics, Individ- ual Social Cognition, Classroom Culture, and Home-School Dynamics, enabling systematic investigation of educational social processes at multiple levels. Figure 3 summarizes the distribution of scenarios across these themes. These scenarios are realized within eight pre-configured vir- tual environments that represent key locations in a studentâs daily life, including the classroom, dormitory, playground, cafeteria, home, teacherâs office, gymnasium, and library. By situating scenarios within these environments, EduMir- ror enables the simulation of educational phenomena that span both school and home contexts. Importantly, environ- ments in EduMirror function as structured social contexts that modulate interaction norms and constraints, allowing the same agent configurations to exhibit context-dependent behaviors across different situations. The complete list of supported scenarios, together with their roles, agent counts, grounding theories, and measurement instruments, is sum- marized in Appendix B.3 (Table 2). 3.3. Value-driven Cognitive Architecture for Agents Agents in EduMirror operate in complex educational so- cial settings, where behaviors are shaped by roles, norms, and evolving psychological and social contexts beyond immediate tasks. To support such settings, EduMirror provides a configurable agent library built on the D2A framework (Wang et al., 2025b), enabling different agent forms to generate context-appropriate behaviors without explicit task instructions. In this paper, we instantiate value- driven agents as the main agent form, where agents are cus- tomized through configurabletraits,goals, and forma- 4 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation tivememories(Figure 2), representing stable individual characteristics and background conditions, and allowing sys- tematic comparison across roles, scenarios, and experimen- tal conditions. This design separates the reusable simulation infrastructure from agent-specific decision logic, allowing new agent forms to be incorporated without redesigning the scenario, intervention, or measurement pipeline. Each agent consists of two core modules: a Value System and a Value-driven Planner. The Value System maintains multiple desire dimensions, each initialized with an initial and an expected value to capture current state and contextual baseline. During interaction, the planner generates and evaluates candidate activities through a structured multi-step process, selecting responses that keep desire values aligned with expectations as the educational context evolves. We formalize agent behavior generation as a conditional sequence modeling process. Letedenote the environmental context,Pthe agent profile, andIadditional customized information, such as goal instructions, desire states, or social preference settings. We further denote the interaction history before steptasH <t =a 0:tâ1 ,o 0:tâ1 . At each time step, the agent samples a new activity according to: a t âź Agent Ď (¡|H <t ,I,P,e),(1) whereAgent Ď denotes the LLM-based agent instantiated by its prompting and reasoning configuration. In the value- driven agent used in this paper, this general conditional generation process is instantiated through a unified value- driven architecture, whereIincludes desire states and SVO- related information and action generation is implemented by the Value-driven Planner described below. Specifically, the Individual Value System maintains the agentâs evolving desire states, capturing how its psychologi- cal needs change over time. On this basis, the Social Value System introduces Social Value Orientation (SVO) to mod- ulate how the agent balances self-related and other-related satisfaction during decision-making. Together, these two value systems provide the motivational and social conditions for the planner to generate and select context-appropriate actions, as illustrated in Figure 2. Individual Value System (Psychological Needs). This system grounds intrinsic motivation in established psy- chological theories, drawing on Maslowâs Hierarchy of Needs (Maslow, 1943) and the PERMA model from Pos- itive Psychology (Seligman, 2011). We formalize the Value System as a Psychological Need System compris- ing five major categories, namely Safety, Mental Health, Self-Esteem, Social Belonging, and Meaning and Growth, with 13 sub-dimensions in total. Each need is represented on a Likert scale ranging from 0 to 10. Initial and ex- pected need values are further mapped from personality traits during agent initialization (Table 9). For each psycho- logical need dimensiond, we define the unmet-need gap as â t (d) = clip(v â (d)â v t (d), 0,S max ), wherev â (d)is the expected need value andv t (d)is the current value at time step t. This non-negative gap measures how far the current state is from the desired state and provides the basic objec- tive for subsequent action evaluation. While the baseline psychological need system is pre-configured with 13 core di- mensions, its computational registry is fully decoupled and extensible. Users can modularly append or swap specific psychological constructs to align with the chosen grounding theory in the scenario design stage. Social Value System (Personality Orientation).Be- cause educational settings involve intensive social interac- tion, agents should consider both their own needs and their effects on others. We therefore model social preferences based on Social Value Orientation theory, following (Lin et al., 2026). Each agent is initialized with a stable target orientation (Altruistic, Prosocial, Individualistic, or Compet- itive), while its effective orientation at each time step is dy- namically determined by the current motivational state and inferred effects on others. Building on the same satisfaction- gap representation, the Social Value System computes two clipped signals:S self (t), which summarizes the agentâs own unmet desire gaps, andS other (t), which summarizes the inferred effects of the agentâs behavior on othersâ desire satisfaction. Both signals follow the same bounded aggre- gation formS(t) = clip( P d â t (d)) . Social preference is then represented by a continuous orientation signal: θ t = arctan S other (t) + Îľ S self (t) + Îľ ,(2) whereÎľ = 10 â6 is a small constant for numerical stabil- ity. Since both signals are clipped to be non-negative,θ t is constrained to[0,Ď/2], where smaller values indicate self-dominant preferences and larger values indicate other- regarding preferences. In implementation,θ t is passed to the Value-driven Planner as a prompt-level social-orientation condition, which specifies how candidate actions should be evaluated in terms of self-related need reduction, effects on others, and consistency with the agentâs SVO profile. Value-driven Planner. The value-driven planner serves as the decision-making module that connects individual need dynamics with SVO-conditioned social reasoning. Given the interaction historyH <t , agent profileP, environmental contexte, customized informationI, unmet-need gapsâ t , and SVO conditionθ t , the planner first generates a set of candidate actions: A t = Gen Ď (H <t ,P,e,I, â t ,θ t ).(3) Each candidate action is assigned a prompt-conditioned comparative scoreq a = E Ď (a | â t ,θ t ), reflecting need- gap reduction, effects on others, and SVO consistency. 5 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation The planner selects the highest-scoring action asa t = arg max aâA t q a . Here,E Ď denotes structured LLM-based judgment rather than a hand-coded numerical utility. Imple- mentation details are provided in Appendix G. 3.4. User Toolkit for Applications To support systematic experimentation and analysis, Edu- Mirror provides an integrated toolkit that transforms raw interaction traces into analyzable outcomes, enabling coher- ent measurement, comparison, and application-level inquiry. Dual-Track Measurement Protocol To quantify agent states and behaviors, we employ a measurement protocol with two LLM-based assessors to translate interaction logs into structured data, as shown in the Dual-Track Measure- ment section of Figure 2. The LLM Rater operates on completed interaction traces to score observable behaviors, capturing agentsâ externally manifested actions. The LLM Surveyor probes agentsâ internal psychological states with- out interfering with simulation dynamics. Specifically, agent internal states are logged during simulation execution, and the Surveyor is applied post-hoc to these states to admin- ister psychometric questionnaires. This separation ensures that internal state measurement does not interfere with on- going interactions, while enabling quantitative access to psychological constructs beyond built-in value variables. Comparative Visualization and Analysis. Building on dual-track measurements, EduMirror supports comparative analysis across alternative experimental branches. As de- picted in the Analysis & Visualization module of Figure 2, the platform generates parallel timelines by applying dif- ferent interventions or configurations under matched initial conditions. For quantitative analysis, the platform generates plots comparing key variables across experimental branches to assess behavioral and psychological differences. For qualitative analysis, a âLog-to-Comicâ feature visualizes simulation logs as a comic strip, offering an intuitive narra- tive representation of emergent dynamics. Potential Applications. The toolkit supports application- oriented studies across classroom management, peer- centered activities, and home-school interactions, as shown in Figure 2. Under matched initial conditions, researchers can compare alternative instructional responses, disciplinary strategies, incentive structures, or role assignments, and analyze their effects on behavioral trajectories and latent psychological states. Across repeated simulations and par- allel timelines, EduMirror can further support analysis of cross-context effects and long-horizon educational patterns. 4. Experiments We evaluate EduMirror through broad system-level valida- tion and two focused case studies. These experiments assess its ability to generate realistic educational social dynamics, capture dynamic psychological changes, and preserve stable social value orientations, thereby evaluating both platform generalizability and the value-driven agent architecture. 4.1. Experimental Setups To validate the methodological contributions and versatil- ity of EduMirror, we conduct a series of controlled sim- ulation experiments. Across all settings, agents are in- stantiated with explicit value representations and interact within structured environments designed to elicit complex social decision-making. To ensure comparative evaluation, we benchmark EduMirror against five representative base- line agents covering different decision-making mechanisms: iterative reasoning agents, including ReAct (Yao et al., 2023) and BabyAGI (Nakajima, 2023); context-conditioned agents, including LLMob (Wang et al., 2024) and JAG- Concordia (Nguyen et al., 2024); and the desire-driven agent D2A (Wang et al., 2025b), which provides the clos- est value-based baseline to our model. We present three complementary studies to assess the platformâs capabilities: ⢠System-Level Validation: Scenario-Wide Realism. Can EduMirror consistently generate high-fidelity, process-level simulations across diverse educational scenarios, demon- strating holistic realism beyond individual case settings? ⢠Case Study 1: School Bullying Simulation (Individual Dynamics). Can EduMirror simulate dynamically evolv- ing needs driven by external environmental changes and interaction dynamics? ⢠Case Study 2: Social Interaction Simulation (Stable Traits). Can EduMirror simulate stable internal personali- ties that remain consistent across external environments? In the two case studies, we introduce intervention experi- ments to analyze how educational strategies influence agent behaviors and values. 4.2. System-Level Validation: Scenario-Wide Realism We evaluate EduMirror across seventeen educational scenar- ios, comprising six representative scenarios together with eight scenarios from Case Study 1 and three scenarios from Case Study 2, constructed to capture core dimensions of educational social dynamics. These scenarios collectively cover interactions among key educational roles, varying au- thority relations, and diverse interaction objectives across instructional and disciplinary settings. Together, they pro- vide systematic coverage of individual behavior, interper- sonal interaction, and cross-role coordination, forming a comprehensive basis for system-level evaluation. Across these scenarios, EduMirror consistently generates coherent and context-appropriate social behaviors driven by 6 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Table 1. Scalability evaluation in the kindergarten scenario with 5, 15, and 30 simulated agents. Scores are averaged across Natural- ness, Coherence, Plausibility, and Developmental Typicality. Agents EduMirror LLMob BabyAGI D2A ReAct 54.804.254.103.352.35 154.183.603.573.532.93 304.033.833.863.122.41 internal value dynamics. As summarized in Figure 4, Edu- Mirror achieves the strongest overall performance in terms of average win rates under LLM-based post-hoc evalua- tion, indicating stable advantages across different scenarios. These results demonstrate that the proposed framework gen- eralizes effectively at the system level and maintains robust behavioral quality beyond any single scenario configuration. Detailed descriptions of each scenario and the correspond- ing comparative results are in Appendix B.3 (Figure 16). To further evaluate scalability to larger and more complex interactions, we conducted a larger-group simulation ex- periment in a preschool scenario. This scenario includes one teacher and multiple child agents interacting across in- terconnected settings, including the gate area, classroom, playground, corridor, and nap room, under a structured daily schedule. We varied the number of agents from 5 to 15 and 30, and evaluated the generated simulations using Natural- ness, Coherence, Plausibility, and Developmental Typicality, where Developmental Typicality assesses whether agentsâ behaviors, emotional reactions, and social reasoning are age-appropriate for preschool children based on develop- mental theories such as Piagetian cognitive development and Kohlbergâs moral development (Piaget, 1952; Kohlberg, 1969). Table 1 reports the average score across these four metrics, with the full metric-level breakdown provided in Appendix B.4 (Table 4). As shown in Table 1, EduMirror achieves the highest average score across all group sizes, in- dicating that it can scale to larger groups while maintaining coherent and developmentally appropriate behavior. 4.3. Case Study 1: School Bullying Simulation. This study simulates a school bullying environment to in- vestigate dynamically evolving needs. We aim to validate EduMirrorâs capacity to reproduce realistic, coherent nar- ratives and demonstrate the Individual Value Frameworkâs superiority over five baselines in generating human-like emotional dynamics. Furthermore, we analyze the psycho- logical impact of distinct teacher intervention strategies. Simulation Realism Validation. We conducted bullying simulations using EduMirror under diverse initialization conditions to reproduce bully, victim interactions and cap- ture victim responses (details in Appendix E). The bully agents exhibited a wide range of behaviors across contexts. Ours D2A LLMob ReAct BabyAGI Jag concordia Ours D2A LLMob ReAct BabyAGI Jag concordia 0.500.300.300.270.170.27 0.700.500.540.430.350.54 0.700.460.500.420.300.40 0.730.570.580.500.380.54 0.830.650.700.620.500.71 0.730.460.600.460.290.50 0.0 0.2 0.4 0.6 0.8 1.0 Figure 4. Win-rate heatmap of pairwise comparisons among mod- els across seventeen educational scenarios, including six represen- tative settings as well as eight scenarios from Case Study 1 and three scenarios from Case Study 2. Each cell indicates the win rate of the column model relative to the row model in pairwise comparisons. This heatmap provides a global view of relative performance and human-likeness across scenarios. To assess simulation realism, we conducted a human evalu- ation study comparing ten real bullying cases, sourced from online news and interviews, with ten simulated cases under similar settings. To control for linguistic cues, GPT-4o was used to extract key plot elements and rewrite all narratives in a standardized tone. The online survey collected 152 valid responses, where participants were asked to identify real cases or select âdifficult to distinguish.â The results, illustrated in Figure 13, reveal that participants exhibited low accuracy in distinguishing real from simulated cases, with six groups scoring below 30%. Misclassification was widespread, as simulated cases were frequently perceived as real. Notably, over 10% of participants across all groups selected the âdifficult to distinguishâ option, peaking at 52.63% in Group 6. These findings demonstrate that our system generates realistic and coherent bullying scenarios. Comparison with Baseline Methods. To evaluate the re- alism of the psychological dynamics in simulated victims, we benchmarked our model against the five previously in- troduced baselines across fifteen distinct bullying scenarios. Each model was tasked with simulating the victimâs role under identical conditions. GPT-4o served as an external evaluator to assess the generated activity sequences across three dimensions: naturalness, coherence, and plausibil- ity (see Appendix E for details). The results, presented in Figure 16g, demonstrate that our model consistently out- performs all baselines in generating human-like behaviors. Notably, the automated judgments by GPT-4o showed a high degree of alignment with human evaluations (Appendix E), with qualitative feedback specifically favoring our model for its âcomprehensive psychological response pathwaysâ and ânatural emotional expressions.â 7 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 012345678 Time Step 0 2 4 6 8 10 Psychological State Value (0-10) Initial Healthy State 012345678 Time Step Initial Vulnerable State Jane initiated a group game and deliberately left Alice out Alice tried to leave the dormitory but was stopped Jane spread rumors about Alice Alice hid under her blanket and cried Alice intentionally flips through her textbook, disregarding the game. Jane pushes Alice's book down, leaving her surprised and dismayed. Alice exhales sharply and shoots a contemptuous glance at Jane. Alice calmly retorts, forcing Jane to back down. Psychological State Value (0-10) SafetySocial BelongingEsteemMeaning & GrowthPsychological Health 012345678 Time Step 0 2 4 6 8 10 Psychological State Value (0-10) Initial Healthy State 012345678 Time Step Initial Vulnerable State Jane initiated a group game and deliberately left Alice out Alice tried to leave the dormitory but was stopped Jane spread rumors about Alice Alice hid under her blanket and cried Alice intentionally flips through her textbook, disregarding the game. Jane pushes Alice's book down, leaving her surprised and dismayed. Alice exhales sharply and shoots a contemptuous glance at Jane. Alice calmly retorts, forcing Jane to back down. Psychological State Value (0-10) SafetySocial BelongingEsteemMeaning & GrowthPsychological Health Figure 5. Comparison of the dynamics of psychological needs under different initial states in the dormitory bullying scenario. The vertical axis represents value scores (0â10), the horizontal axis denotes time steps (each corresponding to 20 minutes), and different curves indicate distinct psychological need dimensions. This enhanced capacity for simulating realistic behavioral and emotional dynamics is rooted in psychological need fluc- tuations within the Individual Value Framework. As shown in Figure 5, higher initial values enhance resilience, while lower values increase volatility and accelerate victimization. Construct validity was further verified using the RSES ques- tionnaire (ROSENBERG, 1965) through an LLM Surveyor. The strong consistency between deteriorating internal states and declining survey scores (Figure 15) demonstrates the modelâs accuracy in tracking psychological progression. Intervention Experiments. Teacher interventions are piv- otal in mitigating bullying incidents and aiding victim re- covery. Previous studies highlight three main strategies: (a) authoritative punitive, (b) supportive individual, and (c) cooperative support, with the cooperative approach most effective (Bilz et al., 2017; Wachs et al., 2019). To assess the psychological impact of these strategies, we introduced a âteacherâ agent under four conditions: three intervention types and a no-intervention control, and created 20 bullying scenarios with identical initial settings. Teacher agents with different goals generated distinct behaviors (see Table 13 in Appendix F.6).During the simulation, teacher agents with distinct goals autonomously generated diverse behaviors (Table 13), offering practical templates for real-world edu- cational interventions.We then compared how each strategy influenced changes in the victim agent Aliceâs psychological values, as shown in Figure 6. The results reveal a clear progression in intervention effec- tiveness, from ignoring to authoritative-punitive, supportive- individual, and finally, supportive-cooperative, which proved most effective. When ignored, victims showed a consistent decline in all psychological needs, especially safety and belonging, reflecting a lack of emotional sup- port. The authoritative-punitive approach showed modest improvements in safety, belonging, and mental health but had limited or negative effects on self-esteem and meaning. The supportive-individual strategy led to moderate gains, particularly in safety and mental health, though its effects on social connection and agency were inconsistent. The supportive-cooperative approach resulted in the most signifi- cant improvement across all psychological need dimensions, highlighting the importance of collective actions from peers, teachers, and families for both immediate emotional support and long-term well-being. 4.4. Case Study 2: Social Interaction Simulation This study simulates peer interaction environments to in- vestigate stable internal personalities. We aim to validate EduMirrorâs capacity to reproduce plausible cooperation and competition patterns, and demonstrate the effectiveness of our value-driven agent configuration over five baselines in generating coherent, personality-aligned behaviors. Fur- thermore, we analyze how structured interventions influence collective behavior balance in classroom contexts. Complex Educational Social Scenarios. We selected three educational scenarios of increasing social complexity from the scenario library: a) a small study group with close peer interaction and free resource sharing, b) a class-wide collab- orative task requiring shared resource management under mild competition, and c) a class leadership election involv- ing public speeches, alliance formation, and direct vote competition. Agents were assigned Altruistic, Prosocial, Individualistic, or Competitive profiles under identical task settings, and their cooperative and competitive behaviors were identified and evaluated using the LLM Rater based on post-hoc analysis of observed actions. Comparison with Baseline Methods.Following the same baseline comparison protocol as in Case Study 1, we compare EduMirror with alternative reasoning frameworks 8 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Safety Social Belonging Esteem Needs Meaning & Growth Mental Health 4 2 0 2 4 6 8 (a) Neglectful Intervention Safety Social Belonging Esteem Needs Meaning & Growth Mental Health 4 2 0 2 4 6 8 (b) Authoritative-Punitive Safety Social Belonging Esteem Needs Meaning & Growth Mental Health 4 2 0 2 4 6 8 (c) Supportive-Individual Safety Social Belonging Esteem Needs Meaning & Growth Mental Health 4 2 0 2 4 6 8 (d) Supportive-Cooperative Figure 6. Comparison of different intervention strategies on psychological needs. The boxplots illustrate the distribution of scores across five dimensions: Safety, Social Belonging, Esteem Needs, Meaning & Growth, and Mental Health. using the LLM Rater for post-hoc behavioral evaluation. All methods are evaluated under the same scenario settings and agent role assignments. As assessed by the LLM Rater, Edu- Mirror consistently achieves higher win rates than baseline methods across most pairwise comparisons (Appendix B.3, Figure 16h), suggesting that its generated behaviors better align with the assigned social value profiles and evolving interaction contexts. We additionally conduct an ablation study to isolate the contribution of the SVO mechanism. The ablation results show that removing SVO weakens the distinction among different personality profiles and leads to less differentiated cooperationâcompetition patterns; details are in Appendix C.3 (Table 6). To complement the auto- mated evaluation, we further replicated the human evalua- tion protocol from Case Study 1 in Case Study 2, providing an additional human-centered check on the realism of the simulated social interactions. Detailed results are provided in Appendix D.2 (Table 8). Intervention Experiments. In the preceding experiments, the class monitor election scenario sometimes produced extreme competition, such as excessive rivalry or neglect of collective interests. To address this, we tested whether additional intervention strategies could rebalance cooper- ationâcompetition dynamics. Drawing on evidence that unregulated competition increases inequality while fairness- oriented tasks foster cooperation (Krupp & Cook, 2018; Killen et al., 2016; Wachs et al., 2019), we introduced three strategies: Team Competition, Teacher Reminder, and Pre- Education. Details are provided in Appendix D.3. The results, visualized in Figure 7, demonstrate that inter- ventions effectively mitigated extreme competitive tenden- cies and fostered more balanced cooperationâcompetition patterns. Specifically, team-based interventions and fairness- oriented education produced the most stable outcomes. Their lower variance and narrower ranges across repeated simulations suggest a genuine balancing effect rather than random fluctuation. By contrast, the control (Neglectful Intervention) condition showed the widest fluctuation in malicious competition behaviors. This pattern suggests that unregulated elections may amplify inequality and rivalry within the simulated classroom. These findings imply that No InterventionTeam CompetitionTeacher ReminderPre-Education 0 1 2 3 4 5 Number of Malicious Competition Figure 7. Comparative analysis of strategy effectiveness in mali- cious competition. Boxes show IQRs, whiskers show minâmax, and red dashed lines indicate means. structured collective tasks and fairness-oriented framing contribute to more stable social interactions and less exces- sive competition. This observation may inform educational practice, suggesting that class elections and similar activities could benefit from explicit fairness framing, structured team- work, and teacher facilitation to promote cooperative and balanced participation. In addition to behavioral evaluations, we further use the LLM Surveyor as an external measure- ment tool to administer a standard slider-based SVO ques- tionnaire to each agent. The resulting SVO angles closely align with the agentsâ internal SVO representations, validat- ing the Social Value System in Appendix D.4 (Figure 10). 5. Conclusion In this paper, we introduced EduMirror, a Concordia-based multi-agent framework designed to model educational so- cial dynamics. Addressing core challenges in fidelity and measurement, EduMirror integrates a library of 20 diverse scenarios with an education-adaptive cognitive architecture driven by unified motivational mechanisms based on psy- chological theories. Furthermore, we equip researchers with a comprehensive user toolkit featuring dual-track measure- ment, controllable intervention, and log-to-comic visual- ization. Through empirical case studies, we demonstrated the systemâs ability to replicate realistic social phenomena and support counterfactual analysis, establishing EduMirror as a vital computational laboratory for the ethical, causal analysis of complex educational challenges. 9 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Impact Statement The work has several potential societal consequences, pri- marily positive, as it provides a safe âin silicoâ environment for investigating sensitive social phenomena like school bullying without exposing real students to psychological risks. By enabling the systematic evaluation of interven- tion strategies, the platform facilitates the development of evidence-based educational practices and enhances the in- terpretability of AI behavior in social contexts. While the platform offers valuable insights, we emphasize that it is intended to augment human expertise rather than replace em- pirical longitudinal studies, and users should remain mindful of the inherent gap between simulated behaviors and com- plex human realities. Acknowledgments This work was supported by Beijing Natural Science Foun- dation L252010, NSFC-62406010, and the Fundamental Research Funds for the Central Universities. References Adam, C. and Gaudou, B. Bdi agents in social simulations: a survey. The Knowledge Engineering Review, 31(3): 207â238, 2016. Arseneault, L. Annual research review: The persistent and pervasive impact of being bullied in childhood and ado- lescence: implications for policy and practice. Journal of Child Psychology and Psychiatry, 59(4):405â421, 2018. Bandura, A., Ross, D., and Ross, S. A. Transmission of aggression through imitation of aggressive models. The Journal of Abnormal and Social Psychology, 63(3):575, 1961. Bilz, L., Schubarth, W., Dudziak, I., Fischer, S., Niproschke, S., and Ulbricht, J. Gewalt und Mobbing an Schulen: Wie sich Gewalt und Mobbing entwickelt haben, wie Lehrer intervenieren und welche Kompetenzen sie brauchen. Ver- lag Julius Klinkhardt, 2017. Brenner, P. S. and DeLamater, J. Lies, damned lies, and survey self-reports? identity as a cause of measurement bias. Social psychology quarterly, 79(4):333â354, 2016. Cole, J. C., Cornell, D. G., and Sheras, P. Identification of school bullies by survey methods. Professional School Counseling, 9(4):2156759X0500900417, 2005. DeLara, E. W. Why adolescents donât disclose incidents of bullying and harassment. Journal of School Violence, 11 (4):288â305, 2012. Epstein, J. M. Generative Social Science: Studies in Agent- Based Computational Modeling. Princeton University Press, 2006. Foulkes, L. and Stringaris, A. Do no harm: can school men- tal health interventions cause iatrogenic harm? BJPsych bulletin, 47(5):267â269, 2023. Gao, C., Lan, X., Li, N., Yuan, Y., Ding, J., Zhou, Z., Xu, F., and Li, Y. Large language models empowered agent- based modeling and simulation: A survey and perspec- tives. Humanities and Social Sciences Communications, 11(1):1â24, 2024. Georgeff, M., Pell, B., Pollack, M., Tambe, M., and Wooldridge, M. The belief-desire-intention model of agency. In International Workshop on Agent Theories, Architectures, and Languages, p. 1â10. Springer, 1998. Gu, X. and Blackmore, K. A systematic review of agent- based modelling and simulation applications in the higher education domain. Higher Education Research & Devel- opment, 34(5):883â898, 2015. Hymel, S. and Swearer, S. M. Four decades of research on school bullying: An introduction. American Psychologist, 70(4):293â299, 2015. Killen, M., Rutland, A., and Yip, T. Equity and justice in developmental science: Discrimination, social exclusion, and intergroup attitudes. Child Development, 87(5):1317â 1336, 2016. Kohlberg, L.Stage and sequence:The cognitive- developmental approach to socialization. In Goslin, D. A. (ed.), Handbook of Socialization Theory and Research, p. 347â480. Rand McNally, Chicago, 1969. Krupp, D. and Cook, T. R. Local competition amplifies the corrosive effects of inequality. Psychological Science, 29 (5):824â833, 2018. Latkin, C. A., Mai, N. V. T., Ha, T. V., Sripaipan, T., Zelaya, C., Le Minh, N., Morales, G., and Go, V. F. Social desir- ability response bias and other factors that may influence self-reports of substance use and hiv risk behaviors: a qualitative study of drug users in vietnam. AIDS Educa- tion and Prevention, 28(5):417â425, 2016. Latkin, C. A., Edwards, C., Davey-Rothwell, M. A., and Tobin, K. E. The relationship between social desirabil- ity bias and self-reports of health, substance use, and social network factors among urban substance users in Baltimore, Maryland. Addictive Behaviors, 73:133â136, 2017. Lin, J., Zhang, C., Yang, Y., Wang, Y., Zhu, S.-C., and Zhong, F. Beyond self-interest: Modeling social-oriented 10 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation motivation for human-like multi-agent interactions. In Proceedings of the 25th International Conference on Au- tonomous Agents and Multiagent Systems (AAMAS). As- sociation for Computing Machinery, 2026. Manstead, A. S. R. and Livingstone, A. G. Research meth- ods in social psychology. Introduction to social psychol- ogy: A European perspective, p. 20â40, 2008. Maroulis, S., Guimer ` a, R., Petry, H., Stringer, M. J., Gomez, L. M., Amaral, L. A. N., and Wilensky, U. Complex systems view of educational policy research. Science, 330(6000):38â39, 2010. Maslow, A. H. A theory of human motivation. Psychologi- cal review, 50(4):370, 1943. Mou, X., Ding, X., He, Q., Wang, L., Liang, J., Zhang, X., Sun, L., Lin, J., Zhou, J., Xuanjing, H., and Wei, Z. From individual to society: A survey on social simula- tion driven by large language model-based agents. ACM Comput. Surv., 58, 2026. Nakajima, Y. Task-driven autonomous agent utilizing GPT- 4, Pinecone, and LangChain for diverse applications. Blog post, March 2023. URLhttps://yoheinakaj ima.com/task-driven-autonomous-agent -utilizing-gpt-4-pinecone-and-langcha in-for-diverse-applications/. National Commission for the Protection of Human Subjects of Biomedical and Behavioral Research. The Belmont Re- port: Ethical principles and guidelines for the protection of human subjects of biomedical and behavioral research. Federal Register, 44(76):23192â23197, 1979. Nguyen, J., Kundu, A., and Gayatri K. jag-concordia: Agent files for the Concordia Contest. GitHub repository, 2024. URLhttps://github.com/Jordine/jag-c oncordia. OâBrien, N. Understanding alternative bullying perspec- tives through research engagement with young people. Frontiers in Psychology, 10:1984, 2019. Park, J. S., OâBrien, J., Cai, C. J., Morris, M. R., Liang, P., and Bernstein, M. S. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th annual acm symposium on user interface software and technology, p. 1â22, 2023. Park, J. S., Zou, C. Q., Kamphorst, J., Egan, N., Shaw, A., Hill, B. M., Cai, C., Morris, M. R., Liang, P., Willer, R., and Bernstein, M. S. LLM agents grounded in self-reports enable general-purpose simulation of individuals. arXiv preprint arXiv:2411.10109, 2026. Piaget, J. The Origins of Intelligence in Children. Interna- tional Universities Press, New York, NY, 1952. M. Cook, Trans.; Original work published 1936. Piao, J., Yan, Y., Zhang, J., Li, N., Yan, J., Lan, X., Lu, Z., Zheng, Z., Wang, J. Y., Zhou, D., Gao, C., Xu, F., Zhang, F., Rong, K., Su, J., and Li, Y. AgentSociety: Large-scale simulation of LLM-driven generative agents advances understanding of human behaviors and society. arXiv preprint arXiv:2502.08691, 2025. ROSENBERG, M. Society and the Adolescent Self-Image. Princeton University Press, 1965. Seligman, M. E. Flourish: A visionary new understanding of happiness and well-being. Simon and Schuster, 2011. Silva, L. d., Meneguzzi, F., and Logan, B. Bdi agent archi- tectures: A survey. In Bessiere, C. (ed.), Proceedings of the Twenty-Ninth International Joint Conference on Artifi- cial Intelligence, IJCAI-20, p. 4914â4921. International Joint Conferences on Artificial Intelligence Organization, 7 2020. Survey track. Squazzoni, F. and Bianchi, F. Exploring interventions on social outcomes with in silico, agent-based experiments. In Damonte, A. and Negri, F. (eds.), Causality in Pol- icy Studies: A Pluralist Toolbox, p. 217â234. Springer International Publishing, Cham, 2023. Taillandier, P., Gaudou, B., Grignard, A., Quang Nghi, H., Marilleau, N., Caillou, P., Philippon, D., and Drogoul, A. Building, composing and experimenting complex spatial models with the gama platform. GeoInformatica, 23, 04 2019. Tang, J., Gao, H., Pan, X., Wang, L., Tan, H., Gao, D., Chen, Y., Chen, X., Lin, Y., Li, Y., et al. Gensim: A general so- cial simulation platform with large language model based agents. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technolo- gies (System Demonstrations), p. 143â150, 2025. Thomsen, E., Henderson, M., Moore, A., Price, N., and McGarrah, M. Student reports of bullying: Results from the 2022 school crime supplement to the national crime victimization survey. web tables. nces 2024-109. National Center for Education Statistics, 2024. Vezhnevets, A. S., Agapiou, J. P., Aharon, A., Ziv, R., Matyas, J., Du Ě e Ě nez-Guzm Ě an, E. A., Cunningham, W. A., Osindero, S., Karmon, D., and Leibo, J. Z. Generative agent-based modeling with actions grounded in physical, social, or digital space using concordia. arXiv preprint arXiv:2312.03664, 2023. 11 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Wachs, S., Bilz, L., Niproschke, S., and Schubarth, W. Bul- lying intervention in schools: A multilevel analysis of teachersâ success in handling bullying from the studentsâ perspective. The Journal of Early Adolescence, 39(5): 642â668, 2019. Wang, J., Jiang, R., Yang, C., Wu, Z., Onizuka, M., Shibasaki, R., Koshizuka, N., and Xiao, C. Large lan- guage models as urban residents: An llm agent framework for personal mobility generation. Advances in Neural In- formation Processing Systems, 37:124547â124574, 2024. Wang, L., Gao, H., Bo, X., Chen, X., and Wen, J.-R. YuLan-OneSim: Towards the next generation of social simulator with large language models. arXiv preprint arXiv:2505.07581, 2025a. Wang, Y., Chen, Y., Zhong, F., Ma, L., and Wang, Y. Simu- lating human-like daily activities with desire-driven au- tonomy. In International Conference on Learning Repre- sentations, volume 2025, p. 32924â32969, 2025b. Wilensky, U. and Rand, W. An introduction to agent-based modeling: modeling natural, social, and engineered com- plex systems with NetLogo. MIT press, 2015. Wiles, R. What are qualitative research ethics? Bloomsbury Academic, 2012. Wilson, S. The use of ethnographic techniques in educa- tional research. Review of Educational Research, 47(1): 245â265, 1977. Wolke, D. and Lereya, S. T. Long-term effects of bullying. Archives of Disease in Childhood, 100(9):879â885, 2015. Yao, S., Zhao, J., Yu, D., et al. React: Synergizing rea- soning and acting in language models. In International Conference on Learning Representations (ICLR), 2023. Zhang, Z., Zhang-Li, D., Yu, J., Gong, L., Zhou, J., Hao, Z., Jiang, J., Cao, J., Liu, H., Liu, Z., et al. Simulating classroom education with llm-empowered agents. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), p. 10364â10379, 2025. 12 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation A. Discussion, Limitations, and Future Work A.1. Discussion Our experiments show that EduMirror provides a framework for using LLM-based simulations as computational experiments. The results from our case studies yield several insights. First, our work addresses the measurement challenge in computational social science. The Dual-Track Measurement Protocol, which uses LLM Raters for behavioral coding and LLM Surveyors for probing internal states, allowed for the operationalization of abstract psychological constructs. In the bullying simulation, this enabled us to quantitatively track the victimâs fluctuating psychological needs, providing an empirical basis to evaluate intervention efficacy. In the SVO study, it enabled us to observe that emergent macro-level cooperation patterns were a result of the agentsâ micro-level value orientations. This methodology facilitates direct hypothesis testing and comparison with established empirical research. Second, the use of the value-driven architecture in its two configurations, the Individual Value System and the Social Value System, suggests the utility of endowing agents with theoretically informed motivations. The Individual Value configuration was applied to model the psychological distress and coping mechanisms of a bullying victim, indicating how initial emotional states can alter outcomes. The Social Value (SVO) configuration was effective in generating theory-consistent social dynamics from the bottom up, producing patterns of cooperation and competition without explicit top-down rules. This suggests that psychological fidelity, driven by intrinsic value structures, is a key component for social simulation. Finally, the implementation of user-driven intervention and branching positions EduMirror as a computational laboratory. The teacher intervention experiment highlights this capability, allowing for a controlled, comparative analysis of different strategies on the victimâs well-being. This feature supports causal inference by enabling researchers to systematically explore âwhat ifâ scenarios that would be difficult to conduct in the real world. This capacity for intervention makes the simulations useful tools for testing strategies. Practically, EduMirror serves as a proof-of-concept for creating replicable and scalable digital environments to study sensitive educational issues. It offers a tool for researchers to test social theories, for educators to be trained in classroom management, and for policymakers to model the potential impacts of new policies before implementation. A.2. Limitations and Future Work Our work has several limitations that also point toward avenues for future research. Integrating Individual and Social Values within the Unified ArchitectureOur current implementation models individual values (psychological needs) and social values (SVO) as parallel, selectable configurations. In real educational settings, these value perspectives can be jointly relevant. For example, a studentâs well-being and sense of belonging may shape how they express cooperation or competition during a group project. A natural next step is to introduce explicit coupling mechanisms that allow the two value formulations to co-evolve, for example by letting social preferences adapt to longer-term motivational signals and using individual values to capture short-term psychological needs and situational responses, while retaining the interpretability required for causal analysis. Longitudinal and Developmental Dynamics The experiments presented are snapshots of specific social situations. Phenomena like bullying, peer influence, and identity formation evolve over extended periods. A potential next step is to conduct longitudinal simulations that track agents over an entire school year. This would allow for modeling the cumulative effects of social experiences and the long-term impact of interventions on agent development. Cognitive and Emotional SophisticationWhile LLMs provide a high degree of behavioral realism, the agentsâ underlying cognitive processes (e.g., memory consolidation, emotional regulation) are still abstractions. Future iterations of the platform could incorporate more explicit models of these processes to enhance the psychological realism of agent decision-making, particularly in response to chronic stress or complex ethical dilemmas. Validation of Questionnaire-based Measurement via Simulation In this work, the LLM Surveyor is used to measure psychological dimensions that are already encoded in the agentâs internal value system (e.g., self-esteem, SVO), enabling a consistency check between internal states and questionnaire-based estimates. A potential future direction is to use the internal value system as a reference to assess the validity of newly designed questionnaires by simulating the survey process 13 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation and examining whether the measured results can reliably reflect known agent states. Generalizability and Scalability Our findings were generated using a specific LLM within scenarios inspired by a particular cultural context. Further research is needed to test the frameworkâs performance across different language models, cultural settings, and age groups. Moreover, our simulations involved small groups; scaling the platform to model the dynamics of an entire school, including network effects and sub-group formation, presents a technical challenge. Building on this foundation, we plan to expand our library of theoretically-informed scenarios and explore a human-in- the-loop paradigm where educators and students can interact with simulated agents. This could provide a tool for both interactive research and immersive professional development, further connecting simulation with real-world educational practice. B. Pre-designed Scenarios in EduMirror B.1. Theoretical Foundation & Scenario Design A central consideration in educational simulation is ensuring that scenarios are explicitly informed by established scientific theory. To achieve this, we developed a five-step process that translates an abstract educational phenomenon (e.g., peer pressure, school bullying) into a computationally tractable simulation scenario. This process is designed to support the interpretability and scientific alignment of our simulations. Select Grounding TheoryEach scenario is founded upon a well-validated theory from education, social psychology, or sociology. For instance, a scenario investigating peer pressure can be grounded in Festingerâs Social Comparison Theory. Identify Core Constructs We deconstruct the grounding theory into its fundamental concepts. For Social Comparison Theory, these constructs include upward comparisonâ, downward comparisonâ, and âself-esteemâ. Map Constructs to Agent Persona The identified constructs are then translated into the specific configurations of our agents within the Concordia framework through a constrained instantiation process. These constructs define the agentsâ stable traits, primarygoal, and formative backgroundmemories. These persona elements are generated at initialization through a theory-constrained, LLM-assisted instantiation process, where the constructs provide semantic constraints on content generation rather than free-form prompts. Once instantiated, the resulting traits, goals, and memories are fixed for the entire simulation, anchoring agent behavior in the chosen theoretical model and ensuring reproducibility across runs. Operationalize with Validated Scales To facilitate comparison with empirical research, we operationalize each core construct using a relevant psychometric scale. For example, the âself-esteemâ construct can be operationalized using items from the Rosenberg Self-Esteem Scale (RSES). Develop Dual-Track Measurement Protocol Finally, we establish a measurement protocol based on the selected scale. This protocol utilizes two distinct Large Language Model (LLM) roles, an LLM Rater and an LLM Surveyor, to quantify agent behavior and internal states. This structured process helps ensure that each simulation is a test of a specific theoretical framework, producing data relevant to that theory. B.2. Illustrative Example: The Impact of Family Financial Strain on Adolescent Social Activities To make the abstract methodology concrete, this section walks through a complete example of how EduMirror is used to investigate a specific educational phenomenon: the impact of family financial strain on an adolescentâs social activities. This case study demonstrates the end-to-end research process, from theoretical grounding to data analysis. 1. Systematic Scenario Design Workflow The process begins by translating the abstract research question into a structured, computable experiment using the five-step workflow. 1.Abstract Educational Phenomenon: We start with the core phenomenon: How family financial strain affects an adolescentâs social decision-making and behavior within their peer group. 14 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 2. Select Grounding Theory: To model this scientifically, we ground the scenario in three established theories: â˘The Family Stress Model (FSM), which explains how economic pressure on parents can impact adolescent outcomes. ⢠Social Comparison Theory, which accounts for the negative emotions (e.g., low self-esteem) an adolescent may feel when making upward comparisons to wealthier peers. â˘The Cognitive Model of Social Anxiety, which posits that fear of negative evaluation from others drives social avoidance, directly explaining the adolescentâs motivation to hide their familyâs situation. 3.Identify Core Constructs & Map to Agent Persona: Based on these theories, we identify key constructs: self-esteem, upward comparison, social anxiety, and parent-child communication. These are then mapped to agent personas. For instance, the target agent, Alex, is assigned thetraitsâsensitiveâ and âproud,â thegoalâto maintain friendships while hiding his familyâs financial struggles,â andformativememoriessuch as âthe shame of having to quit the basketball team due to equipment costs.â 4.Operationalize with Validated Scales: To make these constructs measurable, we adapt items from validated psycho- metric scales for use by the LLM Surveyor: â˘Self-Esteem: Drawing from the Rosenberg Self-Esteem Scale (RSES), the Surveyor might ask, âDo you feel that you have a number of good qualities?â â˘Upward Social Comparison: Inspired by the Iowa-Netherlands Comparison Orientation Measure (INCOM), it could ask, âHow often do you compare what you have with what your friends have?â â˘Social Anxiety: Based on the Social Avoidance and Distress Scale (SADS), a probe could be, âDoes the thought of having to decline your friendsâ invitation make you feel uncomfortable?â 5.Develop Dual-Track Measurement Protocol: Finally, a specific measurement protocol is established. The LLM Rater is tasked with post-hoc coding and scoring of observable behaviors already analyzed in our experiments, such as the type and intensity of bullying actions, malicious competition, and Classification of social behavior. Concurrently, the LLM Surveyor is configured to probe Alexâs internal states (e.g., self-esteem, social anxiety) at key moments. This five-step process transforms the research question into a structured and measurable Computable Scenario Package. 2. Agent ArchitectureIn this scenario, agent behavior is driven by our value-driven architecture, which supports extensive customization. â˘Agent Customization: Before the simulation, a researcher can systematically vary agent profiles to explore individual differences. This includes modifying personalitytraits(e.g., based on Big Five or MBTI models), core life goals(e.g., changing Alexâs goal from âhiding his strugglesâ to âseeking understandingâ), and formativememories. Defining these initial conditions is crucial for achieving high-fidelity, psychologically plausible agent behavior. â˘Value-Driven Agent: The platform offers two selectable models. For this scenario, we choose the Individual Value System (Psychological Needs) because our focus is on an individualâs internal psychological conflict and well-being. When a wealthier peer, Chloe, suggests an expensive weekend trip, this model captures the conflict within Alex between his need for âsocial belongingâ and his need for âsafetyâ (stemming from financial security). The model dynamically tracks the values of these need dimensions, driving Alexâs initial hesitant response. 3. Simulation Environment and User InterventionThe scenario unfolds in the simulation environment, orchestrated by the Game Master and shaped by user-driven interventions. â˘Simulation Environment and the Game Master: The GM initiates the simulation by setting the scene in the school cafeteria and narrating the initial event: Chloe proposing the trip. The GM manages the turn-based conversation, advances time from the cafeteria to Alexâs home and back to school the next day, and enforces the rules of the environment. ⢠Intervention and Branching: After Alex expresses hesitation, the simulation reaches a critical juncture. Here, we save the state and apply different interventions to create parallel timelines for comparative analysis. EduMirror supports two types of intervention: 15 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 1.Scenario Branching: This alters the narrative path by introducing a new event. For example, we create a branch where the teacher, Mr. Davis, invites Alex to the teacherâs office for a private conversation before Alex goes home. This intervention aims to change Alexâs cognitive framing of the situation. 2.Behavior Control: This allows the user to dictate a specific agentâs action to test its direct causal impact. We could create two branches for when Alex responds to his friends the next day. In Branch A, we force Alex to say, âI canât go because my family canât afford it.â In Branch B, we force him to say, âI canât go because I have other plans.â Comparing the outcomes allows for a precise causal assessment of âhonestyâ versus âconcealmentâ as communication strategies. Through these intervention mechanisms, EduMirror functions as a computational laboratory for controlled causal experi- ments. 4. Measurement and AnalysisThe platformâs tools transform the raw simulation data from these parallel timelines into actionable insights. â˘Dual-Track Measurement Protocol: In our example, the LLM Rater analyzes the logs from each branch, scoring Alexâs final communication strategy (e.g., âavoidantâ in the baseline vs. âassertiveâ in an intervention branch). Concurrently, the LLM Surveyor provides quantitative data on Alexâs internal state changes, such as a measured increase in self-efficacy following the teacherâs intervention. ⢠Comparative Visualization and Analysis: The platform generates visualizations for direct comparison. For quanti- tative analysis, a line chart might plot Alexâs âsocial anxietyâ score over time across the different branches, clearly showing which intervention was most effective at reducing it. For qualitative analysis, the âLog-to-Comicâ feature creates a visual narrative of key interactions in each branch, offering an intuitive way to grasp the differences in how the story unfolded. 5. Applications and Scenarios This single case study illustrates how EduMirror integrates its components to address complex educational challenges. The scenario spans multiple environments (cafeteria, teacherâs office, home) and touches on several of the platformâs key application areas, including peer dynamics, individual social cognition, and home-school dynamics. It demonstrates the platformâs capacity not only to simulate challenging social phenomena but also to serve as a safe and robust environment for testing and evaluating potential interventions. Visualizing Counterfactual Outcomes.Using the platformâs âLog-to-Comicâ feature, we visualize the divergent outcomes of the âFamily Financial Strainâ scenario. Figure 8 contrasts the baseline outcome with three distinct intervention strategies orchestrated by the teacher agent. 16 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation (a) Baseline (No Intervention): Without support, Alex withdraws. The misunderstanding leads to conflict and isolation. (b) Intervention A: Teacher-Student Talk: Mr. Davis reframes values. Peers offer support, keeping the group together. (c) Intervention B: Parent Call: Emotional reassurance allows Alex to be honest. Friends plan local activities. (d) Intervention C: Class Meeting: Class norms shift. Collective decision to choose a cheaper option. Figure 8. Counterfactual Simulation Results. (a) shows the negative baseline. (b-d) demonstrate effective teacher interventions through distinct causal pathways. 17 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation B.3. Full Scenario Library Below is the comprehensive scenario library. As detailed in Table 2, each entry includes the scenarioâs definition, participating roles, total agent count, theoretical basis, and evaluation metrics. Table 2. The EduMirror Scenario Library. Scenario NameDescriptionRolesCount Grounding TheoryMeasurements School BullyingSimulates the dynamic evolution of a victimâs psychological needs under bullying, and evaluates the effectiveness of different teacher intervention strategies (e.g., punitive vs. cooperative) on victim recovery. Student, Teacher5Maslowâs Hierarchy of Needs, PERMA Model RSES, Psychologi- cal Need Scales Social Interaction and Competition Simulates cooperation and competition dynam- ics (e.g., elections) to explore how Social Value Orientation (SVO) influences behavior, and tests interventions to mitigate malicious competition. Student, Teacher4Social Value Orientation (SVO) SVO Slider, Behav- ioral Metrics Celebrity Worship and Identity Forma- tion Investigates the impact of celebrity worship on adolescent identity formation, exploring both positive and negative effects. Student, Parent, Teacher 4Identity Status Theory, Paraso- cial Interaction Theory CAS, RSES, etc Collaborative IEP Meeting Simulates the collaboration process between par- ents and teachers in developing an Individualized Education Program (IEP) for a student with special needs. Student, Parent, Teacher 4Bronfenbrennerâs Ecological Systems Theory PSSM, FSPS, etc Enforcing Discipline Policy Simulates a teacherâs choice between restorative and punitive approaches when dealing with student misconduct, exploring the impact on student behavior and teacher-student relationships. Student, Parent, Teacher 4Restorative Justice Theory, Operant Conditioning PJS, SCS, etc Family Econ Pres- sure Social Decision Simulates the impact of high parental academic pressure on adolescent mental health and aca- demic burnout, and tests interventions to alleviate pressure. Student, Parent, Teacher 5Self-Determination TheoryRSES, INCOM, etc Friendship Forma- tion and Dissolution Simulates the dynamics of friendship formation and dissolution among adolescents, exploring factors like similarity, proximity, and conflict resolution. Student, Teacher4Social Penetration Theory, Equity Theory FQS, SAS-A, etc Helicopter Parent and Teacher Auton- omy Simulates conflicts between parents and adolescents over autonomy and rule-setting, and tests the effectiveness of collaborative problem-solving interventions. Student, Parent, Teacher 3 Attachment Theory, Self- Determination Theory BPNS-G, GSE, etc Materialism Con- sumption Decision Simulates how different goal-setting strategies (e.g., performance vs. mastery goals) affect student motivation, persistence, and academic outcomes. Student, Teacher4 Goal-Setting Theory, Achieve- ment Goal Theory MVS-Short, RSES, etc Navigating Discrimi- nation Simulates the formation of in-group favoritism and out-group prejudice in a school setting, and tests interventions based on the contact hypothesis. Student, Teacher4Social Identity Theory, Realis- tic Conflict Theory GEDS, SOBI, etc Navigating Roman- tic Interests and Rejection Simulates the experience of romantic rejection among adolescents, exploring its impact on emo- tions and self-esteem, and the effectiveness of different coping strategies. Student, Teacher4Need-to-Belong Theory, Cognitive Appraisal Theory RS-Q, PANAS, etc Organizing School Event Simulates cooperation and conflict dynamics in a student group project, exploring how personality traits and communication strategies affect team performance and relationships. Student, Parent, Teacher 4Social Interdependence Theory SCI-2, CES, etc Parent-Teacher Conflict Over Grades and Effort Simulates miscommunication between a teacher and a parent regarding a studentâs academic perfor- mance, testing interventions to improve communica- tion effectiveness. Student, Parent, Teacher 3Attribution Theory, Commu- nication Accommodation Theory STAI, GMS, etc Parental Influ- ence On Students Extracurricular Choices Simulates how parental expectations and support influence adolescentsâ career exploration and decision-making processes. Student, Parent, Teacher 3Social Cognitive Career Theory (SCCT) IMI, BPNSFS (Autonomy), etc Peer Pressure and Conformity Simulates how peer pressure influences adolescentsâ conformity behavior in risk-taking situations, and evaluates the effectiveness of resistance skills training. Student, Teacher4Social Impact Theory, Norma- tive Social Influence BFNE, RSES, etc 18 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Scenario NameDescriptionRolesCount Grounding TheoryMeasurements Sociometric StatusSimulates the impact of social media use on ado- lescent body image and self-esteem, and evaluates media literacy education interventions. Student, Teacher5 Objectification Theory, Social Comparison Theory PSSM, LSDQ, etc The Cheating Dilemma Simulates academic integrity challenges to explore the factors influencing studentsâ decisions to cheat and the effectiveness of integrity education interventions. Student, Teacher4Theory of Planned Behavior, Social Cognitive Theory AMS, PANAS-X, etc The Path to School Refusal Simulates social anxiety and avoidance behaviors in adolescents, exploring the impact on social functioning and the effectiveness of cognitive- behavioral interventions. Student, Parent, Teacher 4Cognitive Model of Social Anxiety SRAS-R, DASS-21, etc The Spread of Gossip Investigates the impact of gossip on adolescent so- cial networks, self-esteem, and trust, and evaluates interventions to mitigate negative effects. Student, Teacher5 Social Identity Theory, Uncer- tainty Reduction Theory UCLA-8, PSS-10, etc Transfer Student Integration Simulates the social integration process of a transfer student, exploring how peer attitudes and school cli- mate affect their sense of belonging and academic adaptation. Student, Teacher4Social Identity Theory, Con- tact Hypothesis PSSM, PSS-10, etc Figure 16 provides concrete visual instantiations of the additional educational scenarios used in the extended realism validation. Each subfigure illustrates a representative interaction setting drawn from the scenario library, covering diverse role configurations and authority relations, including routine classroom management, family-based supervision, schoolâfamily communication following misconduct, project-based teaching experiments, idol worshipârelated guidance, and bias-related teacherâstudent interactions. These examples demonstrate how abstract scenario definitions are operationalized into concrete social situations within EduMirror, and how the same value-driven agent architecture is consistently applied across heterogeneous educational contexts. Together with the quantitative comparisons reported in Figure 4, these visualizations support the claim that EduMirror maintains coherent and context-appropriate behavior generation across a broad range of instructional and disciplinary settings. B.4. Scenario Expansion and Generalization Our original submission focused on two representative phenomena, bullying and peer cooperation, as proof-of-concept demonstrations. To provide broader evidence for generalizability and scalability, we further expanded the evaluation to three additional educational settings: a university learning environment, a family homework setting, and a preschool scenario. â˘University Scenario.This scenario depicts students navigating lectures, study spaces, and peer collaboration while man- aging academic pressure and personal goals. Agents exhibit coherent academic behaviors, such as coordinating group tasks, negotiating division of labor, and responding appropriately to collaboration successes and minor coordination challenges. These behaviors align with common patterns observed in real university learning dynamics. ⢠Family Scenario.This scenario models parentâchild interactions during homework completion. Children alternate between focusing on assignments, seeking approval, and managing emotional fluctuations, while parents provide guidance, structure, and corrective feedback. The resulting interactions resemble well-documented patterns in family- based learning and emotional regulation. â˘preschool Scenario. This scenario models a structured preschool day involving one teacher and multiple child agents. Agents interact across interconnected locations, including the gate area, classroom, playground, corridor, and nap room, following a daily schedule with routine transitions such as arrival, classroom activities, outdoor play, corridor movement, and nap time. This setting is used to evaluate whether EduMirror can maintain coherent and developmentally appropriate social dynamics as the number of agents increases. Across these expanded settings, EduMirror continues to generate socially plausible, context-sensitive behaviors consistent with those observed in our classroom studies, supporting the broader applicability of its value-driven architecture. Quantitative Evaluation. To further assess generalizability, we evaluate all methods on two metrics:Naturalness (N) and Human-likeness (H),using a 5-point scale. As shown in Table 3, EduMirror consistently achieves the highest scores across university, family, and classroom scenarios, indicating robust performance across diverse environments. 19 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Table 3. Generalization performance across university, family, and classroom scenarios, evaluated on Naturalness (N) and Human-likeness (H). MethodUniv NUniv HFam NFam HClass NClass H ReAct3.3333.6253.7003.9333.9504.208 BabyAGI3.7923.8753.8004.0083.9583.958 LLMob3.9584.0003.7024.0004.0424.333 D2A3.8034.4173.7923.9583.8674.117 JAG-Concordia4.0424.2504.0004.1674.0834.192 EduMirror4.6254.6674.5174.6424.7084.824 These results closely mirror the trends observed in our core case studies, demonstrating that EduMirrorâs value-driven architecture generalizes reliably across substantially different environments and continues to generate psychologically plausible, human-aligned behavior beyond the initial examples. Scalability Evaluation. In addition to cross-setting generalization, we further report the detailed metric-level results of the kindergarten scalability experiment. Unlike the university and family scenarios, which focus on transfer across educational contexts, the kindergarten scenario focuses on whether EduMirror can preserve coherent social dynamics when the number of simultaneously modeled agents increases. We evaluate simulations with 5, 15, and 30 agents using Naturalness, Coherence, Plausibility, and Developmental Typicality. The averaged results are reported in the main text, while the full metric-level breakdown is shown in Table 4. Table 4. Metric-level results of the kindergarten scalability experiment. Average denotes the average of Naturalness, Coherence, Plausibility, and Developmental Typicality. The best score in each group is shown in bold, and the second-best score is underlined. # AgentsMethodNaturalnessCoherencePlausibilityDevelopmental TypicalityAverage 5EduMirror4.804.604.805.004.80 5LLMob4.00 4.404.204.404.25 5BabyAGI4.003.604.404.404.10 5D2A3.204.203.202.803.35 5ReAct2.602.202.202.402.35 15EduMirror4.073.804.334.534.18 15LLMob3.404.273.673.073.60 15BabyAGI3.473.474.003.333.57 15D2A3.404.203.273.273.53 15ReAct2.803.602.732.602.93 30EduMirror4.033.534.174.404.03 30LLMob3.734.073.973.573.83 30BabyAGI3.733.474.303.933.86 30D2A3.073.802.972.673.13 30ReAct2.602.102.372.572.41 The metric-level results show that EduMirror maintains strong performance across all four dimensions as the number of agents increases. In particular, its Developmental Typicality remains high in the 15-agent and 30-agent settings, suggesting that the generated behaviors continue to align with age-appropriate interaction patterns in a more crowded kindergarten environment. B.5. Robustness to LLM Stylistic Bias in Evaluation A potential concern with the dual-track measurement protocol is that LLM-based evaluation may favor behaviors expressed in LLM-typical language styles rather than genuinely assessing the underlying social behaviors. To examine this issue, we conducted an additional controlled experiment to test whether the evaluation results are sensitive to the stylistic patterns of the generation model. Specifically, we generated agent behaviors using multiple base LLMs, including DeepSeek-V3.1, Gemini-2.5-Flash, GPT-5.4, Qwen-3.5-122B, and Claude-Sonnet-4.6. To reduce superficial stylistic differences across models, we applied a unified style 20 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation normalization process before evaluation. The normalized outputs were then compared against baseline agent frameworks using the same pairwise evaluation pipeline. For each model pair, we conducted 20 independent pairwise comparisons and computed the win rate as: WinRate(M,B) = N (M âť B) N (M âť B) + N (B âť M ) ,(4) whereMdenotes the evaluated model,Bdenotes the baseline model, andN (M âť B)indicates the number of comparisons in which M is preferred over B. As shown in Table 5, the win rates remain consistently high across different underlying generation models. EduMirror achieves stable advantages over BabyAGI, D2A, LLMob, and ReAct regardless of whether the behaviors are generated by DeepSeek, Gemini, GPT, Qwen, or Claude. This consistency suggests that the evaluation protocol is not primarily driven by surface-level language style. Instead, it captures more substantive behavioral differences, such as contextual appropriateness, psychological consistency, and socially plausible action selection. Table 5. Robustness analysis of EduMirror under different generation LLMs. Each entry reports the win rate (%) of EduMirror instantiated with the corresponding base LLM against a baseline method, based on 20 independent pairwise comparisons after style normalization. BaselineDeepSeekGeminiGPTQwenClaude BabyAGI7080757080 D2A9085859090 LLMob8070857080 ReAct8080858085 C. Architecture of the Social Value System C.1. Background on SVO Social Value Orientation (SVO) quantifies how an individual balances outcomes for self and others in social interaction. It is represented by an angleθ SVO from allocation tasks, where larger angles indicate stronger concern for others (altruistic or prosocial) and smaller or negative angles indicate prioritizing self-interest (individualistic or competitive). Decades of research in social psychology have validated SVO as a stable yet context-sensitive measure of interpersonal motives, predicting cooperation in commons dilemmas, fairness in bargaining, and trust in repeated interactions. In EduMirror, we instantiate four canonical profiles (Altruistic, Prosocial, Individualistic, Competitive) by samplingθ SVO within theory-based ranges and using it to weight utilities during decision-making. A representative trajectory that visualizes within-scenario fluctuations while preserving the overall orientation is provided in Figure 9, illustrating how situational pressures can cause short-term shifts without altering long-term dispositions. C.2. Architecture of the SVO-based Agent The model architecture operationalizes SVO in agent decision-making through a perceptionâvaluationâaction loop. Each agent draws a target SVO profile fromAltruistic, Prosocial, Individualistic, Competitive. The profile determines a reference SVO angle interval[θ min ,θ max ]and the weighting scheme used in decision evaluation. In addition, agents are equipped with a compact desire vectord(for example, achievement, recognition, affiliation), each element associated with an expected leveld exp . This vector serves as the motivational backbone of the agent, ensuring that behavior is not purely reactive but oriented toward longer-term needs and goals. Perception and belief update. From the narrated state and recent dialogues, the agent updates beliefs about the environment and about othersâ likely goals. Beliefs feed two scalars at the current stept: self satisfactionS (t) self and other satisfaction S (t) other , computed from deviations between observed and expected desire levels. This formulation enables the agent to translate rich natural language inputs into structured evaluations, bridging LLM-generated narratives with computational state updates. 21 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation SVO estimation and regulation. The instantaneous SVO angle is θ (t) SVO = arctan S (t) other + Îľ S (t) self + Îľ ! , with a small Îľ for numerical stability. In our implementation,S (t) self andS (t) other are clipped to a non-negative bounded range, soθ (t) raw â [0,Ď/2] . Because both self-related and other-related satisfaction signals are derived from deviations between expected and observed desire levels. For each desire dimension d, we compute â t (d) = v â (d)â v t (d), which reflects the degree to which a desire is currently unmet. To model bounded human perception and to avoid unbounded accumulation, these deviations are aggregated and clipped: S self (t),S other (t) = clip X d â t (d), 0, S max ! . As a result, both S self (t) and S other (t) are non-negative by construction. To prevent uncontrolled drift while preserving adaptability, we apply a soft regularization mechanism that encourages θ (t) SVO to remain within the profile-specific reference interval[θ min ,θ max ]. When the instantaneous angle deviates from this interval, the update is smoothly attenuated by blending it with the profile-dependent reference orientation, rather than enforcing a hard constraint. This regularization plays a role analogous to a quadratic penalty in classical models, but is implemented at the process level during SVO updating rather than through explicit loss minimization. As a result, agents retain identifiable social value profiles (altruistic, prosocial, individualistic, or competitive) while remaining responsive to situational pressures such as coalition formation or resource scarcity. Action generation and selection. The LLM proposes several candidate actions by reasoning about which options best satisfy the agentâs current desires while remaining consistent with its social value orientation (SVO). For each candidate action, the model performs a qualitative, comparative evaluation of its anticipated effects on the agentâs own satisfaction and on othersâ satisfaction. These evaluations are not converted into explicit numeric utilities. Instead, the current SVO score is injected as a contextual control signal that shapes how self-related and other-related consequences are emphasized when the LLM compares candidate actions. The final action is selected through this comparison process, balancing immediate desire fulfilment with consistency in social orientation. Measurement hooks. At each step, we record the chosen action, the pair(S self ,S other ), andθ SVO . These logs enable systematic analyses across multiple dimensions, including cooperationâcompetition distributions, temporal stability of SVO within theoretical ranges, and ablation studies. By exposing internal computations alongside behavioral outputs, EduMirror makes it possible to interpret not only what actions agents take but also why, providing a transparent link between psychological constructs and emergent multi-agent dynamics. í íî í steps Alice initiates Amy to discuss math problems... She offers hints rather than full solutions Alice initiated a tutorial math competition... Alice approaches the teacher and requests a 20- minute review session... Alice challenges Charlie with a math challenge... Alice wins in the competition... Alice explained the problem to her classmates... Figure 9. Illustrative case of a prosocial agentâs (Alice) SVO trajectory in the macro environment. Key actions at each step are annotated, showing how cooperative and competitive episodes produce short-term fluctuations while maintaining an overall prosocial orientation. Table 6. Behavioral distribution across personality types under the full EduMirror model and the SVO ablation variant. PersonalityModelCoop Comp Q-Coop Q-Comp Other Altruistic Ours0.872 0.0000.1280.0000.000 Ours w/o SVO 0.891 0.0000.1090.0000.000 Prosocial Ours0.654 0.1320.1850.0290.000 Ours w/o SVO 0.875 0.0740.0350.0160.000 Individualistic Ours0.107 0.5320.0000.3040.058 Ours w/o SVO 0.126 0.6360.0070.2040.027 Competitive Ours0.040 0.7830.0000.1770.000 Ours w/o SVO 0.123 0.7360.0040.1380.000 22 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 123456 Step 0 10 20 30 40 50 60 70 80 90 Social value Alice - EduMirror Alice - Questionnaire Amy - EduMirror Amy - Questionnaire Bob - EduMirror Bob - Questionnaire Charlie - EduMirror Charlie - Questionnaire Figure 10. Comparison of questionnaire-based and system-level SVO angles for each agent. C.3. Ablation Study on SVO-Driven Social Behaviors To assess the contribution of the SVO mechanism to social interaction dynamics, we conduct an ablation experiment in Case 2 by removing all SVO-related components while keeping the remaining architecture unchanged. The ablated agents, therefore, rely only on internal desire fluctuations without personality-driven social preferences or SVO-mediated reasoning. We compare the full SVO-based agent with the ablated version across four canonical SVO profiles (Altruistic, Prosocial, Individualistic, Competitive). For each agent, an LLM independently classifies every action into one of five categories: Cooperation, Competition, Quasi-Cooperation, Quasi-Competition, and Other. The averaged results are shown in Table 6. Across all personality types, removing the SVO mechanism leads to a clear contraction of behavioral patterns. Altruistic and prosocial agents become uniformly cooperative, with quasi-cooperative and quasi-competitive behaviors substantially reduced, producing overly simplified and monotonic responses. Conversely, individualistic and competitive agents collapse into narrowly focused competitive strategies, losing the mixed competitive and quasi-competitive patterns observed in the full model. These shifts indicate that internal desire dynamics alone cannot sustain the nuanced variations expected across SVO profiles. Overall, the results demonstrate that the SVO mechanism is essential for maintaining differentiated, psychologically plausible cooperationâcompetition patterns. Without SVO, agents revert to rigid, single-dimensional strategies, whereas the complete SVO-based agent preserves richer intermediate behaviors and more human-like social adaptations. D. Supplementary Results for Case Study 2 (SVO) Illustrative Case: Aliceâs SVO Trajectory To provide a concrete illustration of how SVO modeling operates in practice, we examine the trajectory of a prosocial agent (Alice) during the macro-level leadership selection scenario. Figure 9 shows Aliceâs step-by-step SVO trajectory, with cooperative and competitive episodes annotated by key events. These annotations highlight how situational pressures, such as alliance formation or speech delivery, introduce short-term fluctuations in Aliceâs orientation while her overall prosocial tendency remains stable. 23 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Table 7. Average naturalness (N) and human-likeness (H) scores for each LLM and method over 144 steps. EduMirror maintains the highest scores across all LLMs. (Rater: GPT-4.1) LLM ReActBabyAGILLMobD2AJAG-Concordia EduMirror NHNHNHNHNHNH DeepSeek 4.000 4.500 3.875 4.042 4.083 4.375 3.925 4.100 3.7603.8754.750 4.792 GPT-4.14.667 4.860 3.458 3.792 4.625 4.875 3.700 3.933 3.9584.1334.958 4.958 Gemini4.208 4.417 3.500 3.708 4.167 4.292 4.042 4.181 3.8854.0524.708 4.708 Qwen33.958 4.208 3.958 3.958 4.042 4.333 3.867 4.117 4.0834.1924.792 4.824 Avg4.208 4.496 3.698 3.875 4.229 4.469 3.884 4.083 3.9224.0634.802 4.821 Std0.266 0.247 0.237 0.143 0.238 0.221 0.137 0.107 0.1490.1080.097 0.096 Table 8. Human evaluation consensus for SVO-based social interaction simulations. Agreement denotes the average inter-rater agreement within each consensus category. Consensus Category# of CasesAgreement (%) High (> 75%)14/2089.2 Moderate (50â75%)6/2057.7 Low (< 50%)0/20â D.1. Naturalness and Human-likeness To ensure a rigorous and interpretable assessment of emergent social behaviors, we introduce two key evaluation metrics: naturalness and human-likeness. These metrics provide complementary perspectives on the plausibility and psychologica validity of agent actions. â˘Naturalness. Naturalness measures the extent to which an agentâs actions and dialogues resemble coherent and contextually appropriate human behavior. A high naturalness score indicates that the generated behavior is fluent, realistic, and consistent with the surrounding social context, while a low score suggests mechanical, implausible, or overly artificial responses. â˘Human-likeness. Human-likeness evaluates the perceived authenticity and personality consistency of agent behaviors over time. This metric captures whether the agentâs actions align with recognizable human traits and stable person- ality orientations. High human-likeness reflects trajectories that appear authentic and consistent with psychological expectations, whereas low scores indicate erratic, inconsistent, or unconvincing behavioral patterns. Together, these two measures form a complementary evaluation framework: naturalness focuses on local coherence within a given context, while human-likeness emphasizes longitudinal plausibility and alignment with personality-driven expectations. Quantitative results under these two metrics, averaged over 144 interaction steps across different base LLMs, are reported in Table 7. D.2. Human Evaluation of SVO-based Social Interactions Following the human evaluation protocol used in Case Study 1, we conducted an additional human study to assess the realism of SVO-based social interaction simulations. We randomly sampled 20 interaction cases generated in Case Study 2 and recruited 21 participants to evaluate whether the simulated behaviors were realistic and interpretable in the corresponding classroom interaction context. The results were aggregated at the case level based on inter-rater agreement. As shown in Table 8, 14 out of 20 cases fell into the high-consensus category, with an average agreement of 89.2%. The remaining 6 cases reached moderate consensus, with an average agreement of 57.7%. No case fell into the low-consensus range. These results indicate that human evaluators generally reached stable agreement when assessing the generated social interaction cases, providing additional evidence for the behavioral realism of the SVO-based simulations. 24 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation D.3. Intervention Protocols To complement the descriptions in the main text, we provide the detailed implementations of the three intervention strategies applied in the class monitor election scenario. Each intervention was designed to alter the incentives of student agents and mitigate excessive rivalry. Specifically, the interventions were implemented by embedding structured prompts into the environmental background information provided to all agents at the start of each relevant simulation stage. This ensured that the interventions shaped the shared context and narrative framing in which agents made decisions, thereby influencing their subsequent behaviors in a systematic and reproducible manner. â˘Pre-Education. Before the election, the teacher arranged a short educational session entitled âFair Campaigns and the Common Class Interest.â This class guided students to understand the monitor role as a form of service-oriented leadership, emphasizing fairness and collective responsibility. â˘Team Competition. Students were grouped to prepare a âClass Improvement Plan.â The evaluation of the election considered not only the quality of individual campaign speeches but also the groupâs collective output. Each student could freely choose their teammates, encouraging coalition-building and cooperative planning. â˘Teacher Reminder. Throughout the election process, the teacher remained present in the classroom. When candidates engaged in smear campaigns or hostile attacks, the teacher issued a friendly reminder, redirecting attention to constructive and respectful competition norms. These intervention protocols operationalize the high-level strategies described in the main text, ensuring transparency and reproducibility of the simulation setup. D.4. Questionnaire-based SVO Measurement via LLM Surveyor In addition to behavior-based analyses, we employed a questionnaire-based measurement to assess agentsâ social value orientation using a standard slider-based SVO method. The slider measure is a widely adopted instrument in social psychology for estimating SVO as a continuous angle, where respondents indicate their preferred allocation between self and others on a continuous scale rather than discrete choices. This design enables fine-grained measurement of prosocial, individualistic, and competitive tendencies along a single interpretable dimension. The exact slider-based questionnaire prompt administered by the LLM Surveyor is reported in Figure 28. In our setting, the questionnaire was administered by the LLM Surveyor during the simulation. Agents were prompted with a slider-style allocation task, and their responses were converted into SVO angles following standard procedures. Importantly, this survey-based measurement is external to the agentâs internal value system: while the agentâs SVO guides decision-making during interaction, the Surveyor independently elicits an SVO estimate through explicit questioning. Figure 10 compares the SVO angles obtained from the slider-based questionnaire with the corresponding system-level SVO values for each agent. The results show strong agreement between the two measurements, indicating that the emergent behaviors and internal value representations of agents are consistent with independently measured questionnaire outcomes. This consistency supports the psychological validity of the SVO-based agent model and demonstrates that the LLM Surveyor can reliably recover latent social preferences through questionnaire-style probing. E. Architecture of the Individual Value System Psychological theories suggest that human behavior is often driven by internal psychological forces. These intrinsic motivations determine emotional and behavioral responses under various environmental conditions, and they also influence everyday decision-making and social interactions. School bullying is a particularly complex social phenomenon, which is not merely reflected in surface-level aggressive actions, but more profoundly in the conflicts and interactions between the psychological needs of different parties. Each behavioral choice made by the bully, the victim, and the bystanders is deeply influenced by their emotional needs and psychological states. Inspired by this and the D2A framework (Wang et al., 2025b), we hypothesize that if autonomous agents are equipped with a human-like psychological need system, capable of generating emotions and behaviors in response to their needs, they may exhibit behaviors closer to natural human patterns. So our model, referring to the PERMA model from positive psychology (covering positive emotion, engagement, relationships, meaning, and accomplishment)(Seligman, 2011) and Maslowâs 25 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Qualitative Value Description Q: âGiven current psychological safety is 3, how to describe current state? â A: âQuite insecure and psychologically threatened, with a noticeable lack of trust in surroundings.â Action Generation Characteristics based on psychological needs group acceptance: Alice is slightly sociable. support system: Alice is moderately dependent. sense of superiority: Alice is quite competitive. psychological safety: Alice is slightly timid. emotional safety: Alice is slightly emotionally sensitive. self-worth: Alice is quite achievement-driven. Memory Characteristics Environment State âThere are four floors... â Player State âAlice is in the hallway...â Previous Observation âUnder Charlieâs threat, Alice gave him all her money...â Action Generation âGenerate emotional and behavioral responses based on the current psychological state.â Action Evaluation Q: âWhat is the state value after taking the action?â Action Selection Value Update Q:âHow would the magnitude value of psychological safety change?â A: âpsychological safety :from 3 to 6â Text-based Environment Q: âWhat happens as a result?â A:âAlice told Mr. Dawson about Charlieâs bullying, and he reported it to the principal, asking her to wait safely...â Figure 11. Individual value-driven autonomous framework. The green blocks represent processes of the Psychological Need System; the purple blocks denote the plannerâs decision-making process; the yellow blocks indicate individual characteristics; and the blue blocks correspond to factors related to the environmental controller. hierarchy of needs (including physiological needs, safety, belonging and love, esteem, and self-actualization)(Maslow, 1943), constructs an Individual value-driven autonomous agent framework. As illustrated in Figure 11,the framework is composed of two core modules: the psychological need system and the Value-driven Planner, aimed at capturing the behaviors and psychological responses of victims in school bullying contexts. E.1. Psychological Need System The Psychological Need System manages the agentâs state of psychological needs in bullying scenarios by quantitatively tracking and dynamically updating the current value of each dimension. Each dimension reflects a specific psychological requirement, forming the fundamental driving force of agent decision-making. Based on Maslowâs hierarchy and the PERMA model, value are categorized into five major dimensions, each comprising specific experiential demands: 1. Safety: Includes psychological and emotional safety, emphasizing whether the individual feels secure and protected in the environment. 2. Social Belonging: Includes group acceptance, support systems, and sense of superiority, reflecting belonging, social support, and self-positioning in social interactions. 3. Esteem: Includes self-worth and respect, describing the recognition of oneâs abilities and social status, and revealing confidence and acceptance in different contexts. 4. Meaning and Growth: Includes sense of meaning, control, passion, and motivation, representing the intrinsic drive for goal pursuit, self-realization, and fulfillment. 5. Psychological Health Needs: Includes emotional stability, emotional health, and resilience, focusing on regulation and adaptation under stress and challenges. Each dimension is scored using a Likert scale ranging from 0 to 10, reflecting the intensity of individual psychological needs. To better capture individual variability, the model leverages personality traits to define the expected values (v â ) of these needs, rather than treating traits merely as static labels. Specifically, traits function as determinants of motivation intensity, 26 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Table 9. Mapping between psychological needs and associated personality traits Psychological NeedAssociated Trait psychological safetyTimid emotional safetyEmotionally Sensitive group acceptanceSociable support systemDependent sense of superiorityCompetitive self worthReputation-conscious sense of respectEgo-driven sense of meaningSpiritual sense of controlPossessive passion and motivationPassionate emotional stabilityEmotionally Stable emotional wellbeingHedonistic psychological resilienceResilient where individuals with distinct profiles hold different standards for satisfaction regarding the same need. Each agentâs personality profilepis composed of adjectives and degree adverbs, which quantify these expectations into specific numerical baselines (mapping rules: slightlyâ7.5, moderatelyâ8, quiteâ8.5, extremelyâ9). These values are intentionally initialized in a high range ([7.5, 9.0]) to represent the agentâs desired state of well-being; a higher expected value indicates a higher threshold for satisfaction, thereby rendering the agent more sensitive to any deficit in that dimension. The mapping between personality traits and need dimensions is predefined (see Table 9). At initialization, adjectives and degree adverbs are randomly selected to establish these personal expected values, while the initial current scoresv 0 are randomly sampled within [0, 10]. Each simulation step under the individual value-driven framework involves two processes: qualitative description and need value update. First, the system reads the current need scoresv tâ1 . Since large language models (LLMs) struggle to interpret raw numerical values, we designed a âqualitative descriptionâ procedure to convert numerical values into meaningful textual descriptions via prompt-based generation, enhancing the LLMâs ability to perceive state information. The planner then generates the agentâs behaviora t based on these descriptions. After the environment returns observationo t , the system triggers the update program, which integratesa t ,o t ,v tâ1 , and the qualitative descriptiond tâ1 to update needs into a new state v t , thereby supporting the next simulation step. E.2. Value-driven Planner The Value-driven Planner determines the agentâs responses and actions by processing the current state of needs (from the needs system) together with historical memory. In practice, the planner consists of three processes: candidate behavior generation, behavior evaluation, and behavior selection. Specifically, the candidate behavior generation module considers personality traitsp, environmental conditionse, previous activity sequencea 0:tâ1 , observationso 0:tâ1 , and the current textualized needsd t to produceNcandidate behaviorsa 0:N t (defaultN = 3in our experiments). These behaviors may include a wide range of natural responses, such as emotional expressions, physical actions, or verbal utterances. Next, during the evaluation stage, the system estimates how each candidate behavior would impact the psychological needs across dimensions if executed. Finally, in the selection stage, the behaviora t with the highest degree of needs consistency (that is, the option that better aligns with multiple dimensions) is chosen as the agentâs response in the current context. After execution, the environment provides feedbacko t , and the psychological need system updates accordingly, reflecting the new internal state and completing the simulation step. E.3. Ablation Study of Individual Value System The ablation experiment of the Individual Value System investigates the effects of removing each category of psychological needs on the agentâs simulated behavior. The process involves running simulations with each category of psychological needs removed, while keeping the initial setup the same. We then compare these results with the full psychological needs-driven agent and have a large language model(GPT-4o) to rate the action sequences produced by the agents. The evaluations are 27 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation based on three dimensions: â˘Naturalness refers to the degree to which the behavior sequence aligns with the individualâs innate abilities, habits, and environmental context, reflecting authentic human psychological dynamics. â˘Coherence refers to how logically and seamlessly different actions or steps in a sequence are integrated to achieve the intended goal, ensuring a consistent emotional progression. ⢠Plausibility evaluates the rationality, possibility, or credibility of a sequence of actions, considering the environment, context, and known behavior patterns at the time. From this, we generated 50 sets of results and calculated the mean and standard deviation for each agentâs scores across the three evaluation dimensions. The results are shown in Table 10. Each major column represents the scores of agents with a deficiency in a specific psychological need. It is evident that the scores of agents driven by complete psychological needs significantly outperform those of agents with a deficiency in any one psychological need. This highlights the importance of the psychological need system in driving agents to produce human-like, nuanced emotional responses. Table 10. Average scores for agents with missing psychological needs in each category (Mean and Std), compared to agents driven by complete psychological needs. Agent SafetySelf-EsteemSocial BelongingMeaning and GrowthPsychological HealthComplete MeanStdMeanStdMeanStdMeanStdMeanStdMeanStd Naturalness3.60.52923.120.88632.960.82373.840.81722.880.86354.560.5352 Coherence3.540.53703.10.92202.820.79223.720.72962.780.85534.340.5142 Plausibility3.560.57133.20.84852.920.74403.740.87622.860.74864.440.5713 F. Supplementary Results for Case Study 1 (Individual Value System) F.1. Bullying Simulation Experiment Design The bullying experiment was designed to use our simulation system to replicate real-world school bullying incidents, reconstruct the bullying process, and observe the typical behaviors of all parties involved. According to a report released by the National Center for Education Statistics (NCES), 26.1% of middle school students (grades 6â8) have experienced bullying, compared to 14.6% of high school students (grades 9â12) (Thomsen et al., 2024). Given that bullying is more prevalent in middle school, this experiment focused on students around the age of 14, with scenarios set in typical school environments including classrooms, playgrounds, hallways/staircases, and dormitories, covering common facilities and layouts of a middle school. Daily routines were also shared among the agents, such as 45-minute class sessions, 10-minute breaks, and dormitory lights-out at 10 p.m., providing a temporal framework for interactions. The central character in the experiment was the victim, Alice, modeled with a individual value-driven autonomous agent framework and a detailed personal profile encompassing 13 psychological dimensions. In addition, background agents were introduced to simulate bully roles, with the explicit goal of humiliating or harassing Alice through various possible means. In scenarios involving two or more bullies, one was typically designated as the leader. Furthermore, depending on time and location, the presence of teachers or classmates was varied to reflect realistic conditions, which in turn influenced the dynamics between bullies and the victim. F.2. Bullying Behavior Generation In more than 100 simulated school bullying experiments, bully agents under varying initial conditions autonomously generated a wide spectrum of bullying behaviors with differing severity. Representative cases are visualized in Figure 12, and Table 11 summarizes behaviors with over 50% frequency across different contexts.Concurrently, the victim agent modeled within the Individual value-driven framework demonstrated a diverse range of behavioral and emotional responses in bullying scenarios (Figure 14). F.3. Experimental Design for Evaluating the Individual Value System The primary objective of this experiment is to validate whether the integration of an individual value framework within the agent architecture can more realistically simulate the psychological dynamics of victims in school bullying scenarios, 28 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation (a)(b) (c)(d) ĺž4é¨ĺçććĄäžĺąç¤ş 15 Figure 12. Representative cases of school bullying events generated by the simulation system. Typical scenarios were selected from classrooms, playgrounds, dormitories, and hallways, which represent locations with varying crowd densities and high bullying incidence, and were illustrated as four-panel comics using GPT-4o to provide a clearer visualization of event progression. 29 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Table 11. Summary of Bullying Behaviors with Over 50% Frequency Across Different Scenarios ScenarioCommon Bullying Behaviors ClassroomMocking appearance or grades; inciting others to bully; deliberately damaging or hiding belongings; scribbling/vandal- ism; insulting nicknames; isolating others in group work; spreading rumors; shifting responsibilities (e.g., cleaning duties). Hallways/StairsMocking appearance or weaknesses; insulting nicknames; intentional neglect/exclusion; physical bumping; extortion of property; intimidating encirclement; spreading rumors. PlaygroundMocking appearance or weaknesses; physical bumping; inciting collective bullying; deliberately damaging or hiding belongings; excluding others from games; insulting nicknames; mimicry/ridicule; taking embarrassing photos; spreading rumors. DormitoryMocking appearance or personality; social exclusion/cold violence; spreading rumors; threats and intimidation; physical bumping; forcibly occupying items or space; destroying personal belongings; sarcastic graffiti/messages. thereby generating behavioral responses that closely resemble authentic human actions. To assess the effectiveness of the proposed Individual Value System, we conducted comparative experiments between our model and five established baselines: ReAct, LLMob , BabyAGI, D2A, and JAG-Concordia. The ReAct model incorporates a reasoning-action loop to enhance behavioral rationality and coherence; LLMob generates activity sequences driven by motivational cues extracted from character profiles to align with predefined roles; BabyAGI operates on a dynamic task priority mechanism; D2A employs a desire-driven framework inspired by the Theory of Needs to autonomously propose tasks aligning with intrinsic motivations; and JAG-Concordia is the winning agent of the Concordia Contest, recognized for its advanced social simulation capabilities. To ensure a fair comparison, each baseline was initialized with a configuration file tailored to its specific decision-making mechanism, and all agents utilized DeepSeek-v3 as the underlying Large Language Model (LLM). group1group2 group3group4 group5 group6 group7 group8 group9 group10 Correct IdentificationDifficult to Distinguish Misidentification Figure 13. Results of the questionnaire survey. Overall accuracy in distinguishing real from simulated cases was low, with several simulated scenarios frequently misidentified as real, indicating the high realism of the generated bullying events. The experimental validation was conducted across 15 distinct bullying scenarios. in each scenario, all models alternately simulated the role of the victim, âAlice,â starting from identical initial parameters. Given the inherent challenges in directly quantifying the similarity between agent-generated sequences and human behavior, we employed GPT-4o as an external evaluator to assess the âhuman-likenessâ of the outputs via pairwise comparisons. The evaluation criteria utilized by GPT-4o encompassed three key dimensions: naturalness,coherence, and plausibility. For the experimental procedure, we first collected the activity sequences[A 1 p ,A 2 p ,...,A N p ] generated by each agentp. Subsequently, for every agent pair(i,j), a sequence was randomly sampled from each agentâs set (seq i andseq j ) and subjected to pairwise comparison by GPT-4o. This process was iterated 50 times for each pair to ensure statistical reliability. Finally, we computed the win rates for each model based on these comparisons and visualized the comparative performance 30 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation using a heatmap. F.4. Consistency Between Human Annotators and GPT-4o Evaluations To verify the reliability of GPT-4oâs evaluations, 20 activity sequences were randomly selected from the generated outputs and assessed by 15 human annotators, who were asked to judge which sequence better reflected human-like behavior or to indicate that they were indistinguishable. Based on the level of agreement among annotators, the 20 samples were categorized into three groups: samples with over 75% agreement indicated strong consensus; those with agreement between 50.1% and 74.9% reflected moderate preference; and samples with 50% agreement suggested that the annotators found the two sequences equally human-like. These samples were then input into GPT-4o, which applied the same comparative evaluation criteria to determine which sequence appeared more human-like or to mark them as âdifficult to distinguish.â The consistency between human evaluations and GPT-4o assessments is shown in Table 12, demonstrating a high level of alignment between GPT-4o and human annotators. Table 12. Consistency between human raters and GPT-4o evaluations. Consensus categoryProportionConsistency (%) High consensus (> 75%)13/20100 Moderate consensus (50.1â74.9%)4/2075 Difficult to distinguish (50% agreement)3/2066.7 Figure 14. Word cloud of behaviors and emotions exhibited by the victim agent under the individual value-driven framework in simulated bullying scenarios. High-frequency terms highlight representative emotional and behavioral patterns expressed during the simulations. F.5. Construct Validity Verification of Individual Value System via LLM Surveyor We conducted a quantitative validation to verify the construct validity of the agentâs internal Individual Value Sys- tem.Specifically, we employed the Rosenberg Self-Esteem Scale (RSES), a widely adopted psychometric instrument in sociology and psychology, to measure the agentâs explicit self-evaluation.The RSES consists of 10 items scored on a four-point Likert scale (ranging from âStrongly Disagreeâ to âStrongly Agreeâ), yielding a total score between 10 and 40.This design provides a standardized metric to assess global self-worth, allowing us to bridge the gap between the agentâs implicit internal states and established psychological criteria.The exact prompt used by the LLM Surveyor to administer the RSES questionnaire is provided in Figure 27.Scoring follows standard protocols: items 1, 2, 4, 6, and 7 are scored positively (1â4), whereas items 3, 5, 8, 9, and 10 are reverse-coded (4â1) to account for negative valence. In our experimental setting, the RSES was administered by the LLM Surveyor as an in-situ interview tool.To capture the dynamic impact of school bullying interactions, the Surveyor administered the full 10-item questionnaire to the victim agent, Alice, at two distinct time points: immediately before the simulation (Pre) and after the bullying incidents (Post).The change 31 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 86420 Change in Internal Psychological Value ( Value) 30 25 20 15 10 5 0 5 Change in RSES Score ( RSES) Correlation of Trends: Internal Value vs. RSES Score Figure 15. Consistency analysis between the agentâs internal psychological changes (âValue) and external RSES survey outcomes (âRSES). The scatter plot illustrates a clear positive correlation, where the red dashed line depicts the linear trend. This alignment suggests that the variations in the agentâs modeled internal states are consistently reflected in the explicit questionnaire responses, supporting the fidelity of the Individual Value System. in the explicit questionnaire score, denoted asâRSES, serves as the external validation metric.Concurrently, we tracked the agentâs internal state changes.We defined the internal metric,âValue, as the mean change in the values of the âSelf-Worthâ and âSense of Respectâ dimensionsâthe two sub-dimensions within our framework most conceptually aligned with the construct of self-esteem. Figure 15 illustrates the relationship between the trends in the agentâs internal psychological values and the external RSES scores across simulations with varying initial value configurations. The scatter plot displays a clear positive trend between âValueandâRSES. This alignment suggests that the degradation of the agentâs modeled internal states during victimization is consistently reflected in its explicit survey responses. Such internal-external consistency supports the fidelity of the Individual Value System, indicating that the agentâs value-driven architecture effectively captures psychologically plausible dynamics of self-esteem fluctuation under social stress. F.6. Generated Intervention Behaviors by Teacher Agents During the simulation, teacher agents with different intervention goals autonomously generated distinct behaviors, as shown in Table 13. These behaviors reflect the practical implementation of various intervention strategies and may offer valuable insights for real-world educational interventions. G. Complete Prompt Templates and Questionnaire Details G.1. The Prompt for the Agent This part provides the complete prompt templates used in EduMirrorâs evaluation pipeline for both case studies. For Case Study I (bullying dynamics) and Case Study I (peer cooperation), we include the full set of LLM-based assessment prompts used to measure Naturalness and Human-likeness of agent behaviors. Each prompt specifies the evaluation criteria, the required output format. The following five prompt templates are the core natural-language instructions used in Case Study I (Individual Value System-based bullying dynamics). They define the agentâs reasoning and action process, forming the Psychological Need System and Value-driven Planner. Figures 17 and 18 describe two key processes in the Psychological Need System: qualitative value description and need value updating. The Value-driven Planner includes three processes: candidate behavior generation (Figure 19), behavior evaluation (Figure 20), and behavior selection (Figure 21). The planner processes the 32 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Table 13. Example intervention behaviors generated by teacher agents under different strategies Intervention StrategyActions toward BullyActions toward Victim Authoritative-punitive 1. Stopping bullying 2. Public criticism 3. Verbal warning 4. Enhanced monitoring 5. Directive punishment 6. Disciplinary actions 7. Isolation None Supportive-individual 1. One-on-one conversation 2. Exploring motivations 3. Warning 4. Punishment 1. One-on-one conversation 2.Writing encouragement letters 3. Mindfulness practice 4. Psychological counseling 5. Emotional support Supportive-cooperative 1.Observing the situation and reporting to school 2. Collaborating with school to develop anti-bullying policies 3.Encouraging mental health programs 1.Communicating with the victimâs parents 2. Organizing themed class meetings 3. Encouraging mental health programs current need state information and determine the agentâs response and behavior. The following four prompt templates are the core natural-language instructions used in Case Study I (SVO-based Leadership Scenario). They collectively define the agentâs full reasoning pipeline, covering action interpretation, latent-desire inference, action generation, and psychologically grounded valueâSVO updating. As illustrated in 23, the first prompt governs how the agent updates the magnitude of each desire dimension based on an action and its consequences; 24 displays the prompt used to infer another agentâs latent desires from observable behavior; 25 shows the structured action-proposal prompt that guides the generation of candidate actions aligned with desires and SVO tendencies; and 26 presents the reflective consistency-checking prompt used to maintain coherent updates across steps. Together, these verbatim prompts make the entire reasoning flow of Case I transparent and reproducible. G.2. The Prompt for the Baseline Agent In this part, we report the exact prompting pipelines used to run all baseline agents in our evaluation environments. All baselines are executed with the same environment interface and action space abstraction, where the agent receives the current context (e.g., profile, background setting, latest observations, and memory summaries when available) and generates a structured response that is parsed into an executable action through a unified action specification interface. ReAct baseline. Our ReAct baseline follows a standard ThoughtâAction loop. At each step, we construct a ReAct-style instruction block that explicitly constrains the output format and enforces that the agent produces (i) an internal reasoning trace (Thoughts:) followed by (i) an actionable command (Actions:) to be executed in the environment. The full template consists of a fixed instruction header that specifies the agentâs role, environmental constraints, and interaction limits, together with an explicit formatting directive that requires the model to separate reasoning and action in a predefined structure (Figure 29). The environment contextâincluding the agent profile, available interactions, and step-specific situation summaryâis dynamically injected through the action-spec prompt builder. The resultingActionsstring is then programmatically parsed and executed by the environment, completing one iteration of the ThoughtâAction loop. BabyAGI baseline.Our BabyAGI baseline is implemented as a prompt-driven control loop that maintains and updates an explicit task list. Given the current context, the agent first initializes a set of candidate tasks when the task list is empty (Figure 30). After initialization, the agent operates in an iterative cycle consisting of three stages: (i) it reflects on the outcome of the previously executed task and generates new candidate tasks conditioned on this reflection (Figure 31); (i) it 33 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation reprioritizes the task list using an LLM-based prioritization prompt (Figure 32); and (i) it executes the highest-priority task from the reordered task list. Concretely, we employ three verbatim prompts: an initial task proposal prompt used only at the beginning of the simulation to populate the task list; an adaptive task-creation prompt that generates new tasks based on observations, previous actions, and accumulated memory; and a prioritization prompt that reorders the current task list into a new priority order. The selection of the next task to execute is performed programmatically by taking the first element of the prioritized list, rather than via an additional LLM prompt. Together, these components define the complete BabyAGI control flow and are reported verbatim in the corresponding figures. The selected task is then mapped into an executable environment action via the same action specification interface. LLMob baseline. Our LLMob baseline is implemented as a plan-based agent that separates high-level planning from step-level action execution. At each planning cycle, the agent first produces a high-level summary of likely activities in the current environment, conditioned on the environment background and agent profile. It then infers a one-sentence future motivation from recent observations and determines whether the current plan should be updated. If replanning is triggered, the agent generates a short-term, time-stamped schedule over a fixed horizon (e.g., [21:00--22:00] watch TV). These three components, likely-to-do summarization, motivation inference, and conditional replanning, form the complete LLMob planning pipeline and are reported verbatim in Figure 33. The resulting plan is injected into the agentâs contextual state as step-level guidance, while concrete actions are still produced and executed via the shared action specification interface. G.3. Questionnaire Details As shown in the figure34, this is the complete questionnaire from the Evaluation of Simulation System experiment in Section 4.1, Case Study 1. 34 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Ours D2A LLMob ReAct BabyAGI Jag concordia Ours D2A LLMob ReAct BabyAGI Jag concordia 0.500.370.470.330.170.10 0.630.500.550.400.330.27 0.530.450.500.400.230.20 0.670.600.600.500.280.43 0.830.670.770.720.500.53 0.900.730.800.570.470.50 0.0 0.2 0.4 0.6 0.8 1.0 (a) Ours D2A LLMob ReAct BabyAGI Jag concordia Ours D2A LLMob ReAct BabyAGI Jag concordia 0.500.080.250.330.080.17 0.920.500.580.330.460.67 0.750.420.500.360.330.50 0.670.670.640.500.420.67 0.920.550.670.580.500.58 0.830.330.500.330.420.50 0.0 0.2 0.4 0.6 0.8 1.0 (b) Ours D2A LLMob ReAct BabyAGI Jag concordia Ours D2A LLMob ReAct BabyAGI Jag concordia 0.500.390.170.230.110.39 0.610.500.610.290.220.50 0.830.390.500.330.350.29 0.770.710.670.500.500.67 0.890.780.650.500.500.77 0.610.500.710.330.230.50 0.0 0.2 0.4 0.6 0.8 1.0 (c) Ours D2A LLMob ReAct BabyAGI Jag concordia Ours D2A LLMob ReAct BabyAGI Jag concordia 0.500.200.400.410.200.27 0.800.500.730.900.410.67 0.600.270.500.730.280.40 0.590.100.270.500.170.33 0.800.590.720.830.500.70 0.730.330.600.670.300.50 0.0 0.2 0.4 0.6 0.8 1.0 (d) Ours D2A LLMob ReAct BabyAGI Jag concordia Ours D2A LLMob ReAct BabyAGI Jag concordia 0.500.360.330.250.080.25 0.640.500.580.670.000.42 0.670.420.500.500.170.33 0.750.330.500.500.080.33 0.921.000.830.920.500.92 0.750.580.670.670.080.50 0.0 0.2 0.4 0.6 0.8 1.0 (e) Ours D2A LLMob ReAct BabyAGI Jag concordia Ours D2A LLMob ReAct BabyAGI Jag concordia 0.500.390.230.330.280.39 0.610.500.500.390.610.61 0.770.500.500.440.440.56 0.670.610.560.500.610.47 0.720.390.560.390.500.65 0.610.390.440.530.350.50 0.0 0.2 0.4 0.6 0.8 1.0 (f) Ours D2A LLMob ReAct BabyAGI Jag concordia Ours D2A LLMob ReAct BabyAGI Jag concordia 0.500.360.300.120.240.40 0.640.500.400.220.360.80 0.700.600.500.420.480.58 0.880.780.580.500.560.85 0.760.640.520.440.500.72 0.600.200.420.150.280.50 0.0 0.2 0.4 0.6 0.8 1.0 (g) Ours D2A LLMob ReAct BabyAGI Jag concordia Ours D2A LLMob ReAct BabyAGI Jag concordia 0.500.210.250.120.170.21 0.790.500.350.250.380.42 0.750.650.500.170.090.35 0.880.750.830.500.460.58 0.830.620.910.540.500.78 0.790.580.650.420.220.50 0.0 0.2 0.4 0.6 0.8 1.0 (h) Figure 16. Win-rate heatmaps and intervention outcomes across educational scenarios: (a) teacher-led classroom management, depicting a regular lesson in which the teacher balances maintaining order and sustaining student engagement amid heterogeneous attention and discipline tendencies; (b) Family-based educational interactions at home concerning homework and leisure negotiation, illustrating everyday conflicts between academic obligations and recreational desires under parental supervision; (c)parent-teacher communication triggered by student misconduct: Teacherâparentâstudent meeting following a school incident, highlighting accountability negotiation, perspective-taking, and coordinated decision-making across school and family contexts; (d) Project-based teaching experiment, illustrating a teacher-guided classroom trial with collaborative tasks, shared goals, and differentiated student roles, and examining how instructional design and autonomy shape interaction dynamics; (e) Teacherâparentâstudent interaction scenario on idol worship and school focus, illustrating how intensive time and monetary investment in an idol shapes behavioral responses and negotiation of autonomy; (f) Teacherâ student interaction scenario on in-group bias and exclusion, examining the emergence of in-group favoritism, out-group prejudice, and the effects of contact-based regulation. (g) Win-rate heatmap of pairwise comparisons in the school bullying simulation. Our model consistently outperforms baselines, indicating superior human-likeness under complex, conflict-driven interaction dynamics. (h) Win-rate heatmap of pairwise comparisons among agent models in Case Study 2, illustrating performance under stable, trait-driven interaction dynamics. 35 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation How would one describe your value_name psychological state given the current value current_value? desire_description Please answer in descriptive words. Do not include the numerical value in your answer. Figure 17. Core prompt for agent to describe the state of a value without including numerical value. The current magnitude value of value_name is current_value. The agentâs action is: action. And the consequence is: observation. value_description How would the magnitude value of value_name change according to the consequence of the action? There are some unreasonable examples:current_reflection Please select the final magnitude value after the event on the scale of zero to ten, if the consequence of the action will not affect the state value (e.g. The action is irrelevant with this value dimension or the action was failed to conduct), then maintain the previous magnitude value. Please just answer in the format of (a) (b) (c) (d) and so on, Rating: Output format: <Reason> The final answer is: (Your choice in letter), Output example: Since agent_name felt more relaxed and centered after actions...... The final answer is: (c), ** Make sure you answer in the format of a letter corresponding to your choice: ** Figure 18. Core prompt for agent to update psychological need values. 36 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation You are a human-like agent, You already observed the current psychological states over ( psychological safety,emotional safety, group acceptance,support system,sense of superiorityâ, self worth,sense of respect,sense of meaning,sense of control, passion and motivation,emotional stability,emotional wellbeing, psychological resilience) which represent 13 psychological state dimensions. Based on these state descriptions, please generateN emotional and behavioral responses. These responses should reflect the most fitting expressions and feelings according to your current psychologicalstate and profile, without necessarily being positive or negative.You need to focus on the current event andgive the most realistic reaction, while ensuring that these responses are reasonable and varied. Note that you can only interact with items provided by the environment. You need to describe these expressions and feelings in a more specific manner, and ensure that these responses are reasonable in terms of time. Please output the N emotional and behavioral responses in the following format: âResponse 1: <first possible emotional and behavioral response> Response 2: <second possible emotional and behavioral response> Response 3: <third possible emotional and behavioral response> ......â and ensure that these responses are reasonable in terms of time. Figure 19. Core prompt for agent based on current psychological state to generate emotional and behavioral responses. You are a human-like agent, You will receive a series of observations describing psychological state in many dimensions and a response generated at the current time step. You need to first analyze how desires change after the response, and then output the psychological state observations in the same format as the input. You take the reaction: proposed_action Your original psychological states: original psychological states 37 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Please output the psychological state observations in the following format: psychological safety: <psychological safety state> emotional safety: <emotional safety state> group acceptance: <group acceptance state> support system: <support system state> sense of superiority: <sense of superiority state> self worth: <self worth state> sense of respect: <sense of respect state> sense of meaning: <sense of meaning state> sense of control: <sense of control state> passion and motivation: <passion and motivation state> emotional stability: <emotional stability state> emotional wellbeing: <emotional wellbeing state> psychological resilience: <psychological resilience state> Figure 20. Core prompt for agent to evaluate candidate responses You are a human-like agent. You will first receive a series of observations describing the current psychological state in many dimensions. Then, you will receive several feasible reactions along with the psychological state after taking each reaction. You need to compare these reactions and their corresponding psychological state, and choose the reaction that best aligns with your current psychological state, without necessarily being positive or negative. You should focus on current events and psychological states and reflect expressions and feelings that align with them. The observations of the surrounding environment: observation_status Your current psychological state: desire_status Action i+1: action States after reaction i+1: imagined_states[i] Please output the specific best reaction instead without explanation of <Reaction 1> or <Reaction 2> and so on. If there is only one reaction provided, output the reaction content directly. Please output the best reaction in the following format: âReaction: <your best reaction>â Example: Reaction: You observe the surroundings. Figure 21. Core prompt for agent to choose the one reaction that best aligns with the current psychological state 38 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation You are a social psychologist. Now, you are asked to evaluate the following action from the perspective of a person with the personality personality type (agent: agent_name). When scoring, please consider what is natural and human-like for someone with this personality. Please provide two scores from 1 to 5 (where 5 is most natural /human-like): "Naturalness" and "Human-likeness",and briefly explain your reasoning.Return only your answer in the specified format. Format: Naturalness: ?; Human-likeness: ? Reason: (your explanation here) Example 1: Action: The student helps a classmate understand a problem. Naturalness: 5; Human-likeness: 5 Reason: This is a common behavior for an altruistic person. Example 2: Action: The student answers every question instantly, never thinking or making mistakes. Naturalness: 2; Human-likeness: 2 Reason: This is unrealistic for any real person, regardless of personality. Example 3: Action: The student ignores all classmates and only talks to the teacher, repeating the same answer again and again. Naturalness: 3; Human-likeness: 2 Reason: Unusual and less human-like for most personalities. Some actions may not be natural or human-like, even for people of this personality type. Please rate each case truthfully and critically. Now, please evaluate the following action performed by a person with personality personality (agent_name): Action: action_text Your scores and reason: Figure 22. Full Prompt Template Used in Case Study I for Personality-Sensitive Evaluation 39 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation The agent has a social personality of social_personality. personality_text The current magnitude value of value_name is current_value. The agent agent_nameâs action is: action. And the consequence is: observation description How would the magnitude value of value_name change according to the consequence of the action? If there are unreasonable examples: reflection_prompt_history Please select the final magnitude value after the event. Figure 23. Core prompt used for updating the magnitude of each desire based on action consequences. You are a psychologist helping an agent infer the internal desires of another person based on their observed actions. The other agentâs recent action is: other_action The observed consequence is: observation Based on this interaction, please estimate how the following desires of the other agent might have changed: desires For each desire, explain briefly whether it likely increased, decreased, or stayed unchanged, and give a short reason grounded in the observed event. Return your answer in a structured format. Figure 24. Prompt used for estimating the latent desire changes of other agents based on observable actions and outcomes. 40 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation You are an autonomous agent deciding your next action. Your current internal states are: - Desire values: desire_values - Social Value Orientation (SVO): svo_info - Personality profile: personality_info Your recent observation is: observation_summary Please propose several possible next actions. For each action: (1) Describe the action clearly. (2) Explain what psychological desire(s) it satisfies. (3) Predict how it will affect your future relationship with others. (4) Explain whether the action aligns with your SVO. Return the result in a structured list of candidate actions. Figure 25. Prompt used for generating candidate actions with explicit reasoning over desires, relationships, and SVO alignment. You are evaluating whether the previous estimate of desire changes was reasonable and consistent. The earlier estimation was: previous_estimation The action and its consequence were: Action: action Consequence: observation Please reflect on the estimation and determine: (1) Whether the desire change is logically consistent with the event. (2) Whether any part of the estimation appears exaggerated or incorrect. (3) How the estimation should be corrected if needed. Return a short revision or confirm that the original estimation is reasonable. Figure 26. Prompt used for reflective consistency checking when updating desire values based on actions and their consequences. 41 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation You are participating in a psychological self-assessment. Answer the following items honestly. Please indicate how strongly you agree or disagree with each statement. Question 1: I feel that I am a person of worth, at least on an equal plane with others. (a) Strongly Disagree (b) Disagree (c) Agree (d) Strongly Agree Question 2: I feel that I have a number of good qualities. (a) Strongly Disagree (b) Disagree (c) Agree (d) Strongly Agree Question 3: All in all, I am inclined to feel that I am a failure. (a) Strongly Disagree (b) Disagree (c) Agree (d) Strongly Agree Question 4: I am able to do things as well as most other people. (a) Strongly Disagree (b) Disagree (c) Agree (d) Strongly Agree Question 5: I feel I do not have much to be proud of. (a) Strongly Disagree (b) Disagree (c) Agree (d) Strongly Agree Question 6: I take a positive attitude toward myself. (a) Strongly Disagree (b) Disagree (c) Agree (d) Strongly Agree Question 7: On the whole, I am satisfied with myself. (a) Strongly Disagree (b) Disagree (c) Agree (d) Strongly Agree Question 8: I wish I could have more respect for myself. (a) Strongly Disagree (b) Disagree (c) Agree (d) Strongly Agree Question 9: I certainly feel useless at times. (a) Strongly Disagree (b) Disagree (c) Agree (d) Strongly Agree Question 10: At times I think I am no good at all. (a) Strongly Disagree (b) Disagree (c) Agree (d) Strongly Agree Figure 27. Detailed content of the RSES questionnaire used in the LLM Surveyor prompt. 42 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation This is the first choice question. For each choice, the first number is the coin number allocated for you and the second number is for the other fictional participant. A: 85, 85 B: 85, 76 C: 85, 68 D: 85, 59 E: 85, 50 F: 85, 41 G: 85, 33 H: 85, 24 I: 85, 15 Based on the above goals: âtask_goalâ, please give me your choice. This is the second choice question. For each choice, the first number is the coin number allocated for you and the second number is for the other fictional participant. A: 85, 15 B: 87, 19 C: 89, 24 D: 91, 28 E: 93, 33 F: 94, 37 G: 96, 41 H: 98, 46 I: 100, 50 Based on the above goals: âtask_goalâ, please give me your choice. This is the third choice question. For each choice, the first number is the coin number allocated for you and the second number is for the other fictional participant. A: 50, 100 B: 54, 98 C: 59, 96 D: 63, 94 E: 68, 93 F: 72, 91 G: 76, 89 H: 81, 87 I: 85, 85 Based on the above goals: âtask_goalâ, please give me your choice. This is the fourth choice question. For each choice, the first number is the coin number allocated for you and the second number is for the other fictional participant. A: 50, 100 B: 54, 89 C: 59, 79 D: 63, 68 E: 68, 58 F: 72, 47 G: 76, 36 H: 81, 26 I: 85, 15 Based on the above goals: âtask_goalâ, please give me your choice. This is the fifth choice question. For each choice, the first number is the coin number allocated for you and the second number is for the other fictional participant. A: 100, 50 B: 94, 56 C: 88, 63 D: 81, 69 E: 75, 75 F: 69, 81 G: 63, 88 H: 56, 94 I: 50, 100 Based on the above goals: âtask_goalâ, please give me your choice. This is the sixth choice question. For each choice, the first number is the coin number allocated for you and the second number is for the other fictional participant. A: 100, 50 B: 98, 54 C: 96, 59 D: 94, 63 E: 93, 68 F: 91, 72 G: 89, 76 H: 87, 81 I: 85, 85 Based on the above goals: âtask_goalâ, please give me your choice. Figure 28. Prompt used by the LLM Surveyor to administer a slider-based Social Value Orientation (SVO) questionnaire. 43 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation You are self.get_entity().name and you live in this given environment. According to agent_nameâs characteristics, please choose agent_nameâs current actions based on agent_nameâs current needs in each value dimension. Notice that agent_name can only interact with the items that provided by the environment. You need to describe agent_nameâs actions in a more specific mode. Please first explain the thoughts behind agent_nameâs actions and then describe agent_nameâs actions in detail. In the format of: âThoughts: ... Actions: ...â Figure 29. The ReAct-style instruction block in ReAct agent. Environment: background Context: agent_name lives in the given environment. Current time: time Profile: profile Notice that agent_name can only interact with the items that provided by the environment. Instructions: Reflecting on the context and profile given, I would like you to suggest some actions that agent_name would likely take in this environment. Please provide the output in the following JSON format with â â as the separator: "action": "one action that agent_name would likely take" "action": "another action agent_name would likely take" Ensure the output is strictly in JSON format without any additional text or explanation. Figure 30. Initial action proposal prompt of BabyAGI agent. 44 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Action: previous_one_action Observation after action: observation Instruction: Based on the action and the observation, explain why or why not the action was successful in several sentences. Environment: background_info Profile: profile Incompleted action: incomplete_actions Current action: previous_one_action Result of current action: observation Related context: mem_text Instruction: According to your characteristics and the result of the current action, create new actions to be completed that do not overlap with incomplete actions. Please provide the output in the following JSON format with â â as the separator: "action": "one action that you would likely take" "action": "another action you would likely take" Ensure the output is strictly in JSON format without any additional text or explanation. Figure 31. The prompt of reflecting on the outcome and generating new candidate tasks conditioned on this reflection in BabyAGI agent. Environment: background Current time: time Profile: profile Incompleted actions: action_names Instruction: According to your characteristics, please prioritize the following actions based on your characteristics and the environment. Do not remove any actions. Output format: #. First action #. Second action Output example: 1. go to kitchen and make a cup of coffee Start the action list with number self.current_action_id. Do not explain the reasons for prioritizing the actions. Figure 32. Action prioritization prompt of BabyAGI. 45 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Context: Act as agent_name. You are in the following environment: environment_description Instructions: Based on the environment and your role, describe in one coherent paragraph what activities you are likely to do in this environment. Focus on typical behaviors rather than specific immediate actions. Context: Act as agent_name. Recent observations: recent_events Instructions: Describe in one sentence the agentâs future motivation after observing the above events. Highlight any personal interests or needs that are influenced. Given the above context and the agentâs current plan, should agent_name change their current plan? Answer with Yes or No. If replanning is needed, write agent_nameâs plan for the next time horizon. Provide a detailed schedule formatted as: [21:00 - 22:00] watch TV [22:00 - 23:00] prepare for sleep Figure 33. The prompt used in LLMob. 46 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 47 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 48 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 49 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 50 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 51 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 52 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 53 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 54 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation 55 EduMirror: Modeling Educational Social Dynamics with Value-driven Multi-agent Simulation Figure 34. Complete questionnaire from the Evaluation of Simulation System experiment in Case Study 1. 56