Paper deep dive
Do AI Personas Grow? Analyzing and Benchmarking Personality Evolution in LLM Agents After Life Events
Ming Wang, Peidong Wang, Xiaocui Yang, Daling Wang, Shi Feng, Fiona Fui-Hoon Nah, Ee-Peng Lim
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 91%
Last extracted: 8/10/2026, 2:28:32 AM
Summary
This paper analyzes personality evolution in Personality-Conditioned LLM Agents (PC-Agents) by measuring Big Five trait shifts after 11 major life events. Using the BFI-44 inventory across 100 controlled personas and 14 models, the study finds that while PC-Agents exhibit measurable trait changes, these shifts are weakly event-specific, have magnitudes significantly smaller than human effect sizes, and lack demographic or individual-level heterogeneity. The authors introduce BFI-Adapt, a benchmark for scoring directional fidelity, concluding that current agents simulate the mean of human personality dynamics but not its shape.
Entities (7)
Relation Signals (5)
BFI-Adapt → evaluates → PC-Agents
confidence 95% · we introduce BFI-Adapt, a reusable benchmark for scoring the directional fidelity of event-induced personality change
PC-Agents → simulates → mean of human personality dynamics
confidence 95% · current PC-Agents simulate the mean of human personality dynamics, but not its shape
PC-Agents → exhibits → personality evolution
confidence 90% · agents should undergo plausible, psychology-grounded changes as they experience life events
PC-Agents → uses → Big Five
confidence 90% · using the Big Five traits as a psychometric anchor
PC-Agents → has → compressed dispersion
confidence 85% · persona-level dispersion is compressed three- to four-fold relative to human samples
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is personality evolution: agents should undergo plausible, psychology-grounded changes as they experience life events in different contexts. Although prior work shows that LLM personalities can shift under contextual perturbations, how these shifts vary across traits, events, personas, and models remains poorly understood. We study event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology. Across four diagnostic axes, PC-Agents exhibit measurable trait shifts at similar rates for event-trait pairs with and without documented human change directions. Even when shifts follow the expected direction, their magnitudes usually fall below human effect-size ranges. Gender and cultural-region prompts show little moderating effect, while persona-level dispersion is compressed three- to four-fold relative to human samples. To enable systematic comparison, we introduce BFI-Adapt, a reusable benchmark for scoring the directional fidelity of event-induced personality change, and use it to rank 14 models. A validation suite shows that the measured shifts exceed no-event retest noise, remain stable under independently paraphrased prompts, exhibit limited and model-dependent convergence with scenario-based behavioral choices, and persist across intervening unrelated dialogue. Together, these checks establish the measured trajectories as robust event-conditioned response patterns. Our results suggest that current PC-Agents simulate the mean of human personality dynamics, but not its shape.
Tags
Links
- Source: https://arxiv.org/abs/2608.06485v1
- Canonical: https://arxiv.org/abs/2608.06485v1
Trouble viewing inline? Open PDF directly →
Full Text
82,638 characters extracted from source content.
Expand or collapse full text
KinaMind DoAIPersonasGrow?Analyzingand BenchmarkingPersonalityEvolutioninLLMAgents AfterLifeEvents MingWang 1,2 PeidongWang 1 XiaocuiYang 1 DalingWang 1 ShiFeng 1 Fiona Fui-HoonNah 2 Ee-PengLim 2 1 SchoolofComputerScienceandEngineering,NortheasternUniversity,Shenyang,Liaoning,China 2 SchoolofComputingandInformationSystems,SingaporeManagementUniversity,Singapore Correspondence:wangdaling@cse.neu.edu.cn;eplim@smu.edu.sg Abstract Personality-conditioned LLM agents (PC-Agents) are increasingly used in emotional support, social simulation, and role-playing, motivating the development of lifelong agents that remain coherent over extended interactions. A key component of such coherence is personality evolution: agents should undergo plausible, psychology-grounded changes as they experience life events in different contexts. Although prior work shows that LLM personalities can shift under contextual perturbations, how these shifts vary across traits, events, personas, and models remains poorly understood. We study event-induced personality change after 11 major life events, using the Big Five traits as a psychometric anchor and interpreting the resulting trajectories against longitudinal evidence from human personality psychology. Across four diagnostic axes, PC-Agents exhibit measurable trait shifts at similar rates for event–trait pairs with and without documented human change directions. Even when shifts follow the expected direction, their magnitudes usually fall below human effect-size ranges. Gender and cultural-region prompts show little moderating effect, while persona-level dispersion is compressed three- to four-fold relative to human samples. To enable systematic comparison, we introduce BFI-Adapt, a reusable benchmark for scoring the directional fidelity of event-induced personality change, and use it to rank 14 models. Across the full persona–event grid, a validation suite shows that the measured shifts exceed no-event retest noise, remain stable under independently paraphrased prompts, exhibit limited and model-dependent convergence with scenario-based behavioral choices, and persist across intervening unrelated dialogue. Together, these checks establish the measured trajectories as robust event-conditioned response patterns. Our results suggest that current PC-Agents simulate the mean of human personality dynamics, but not its shape. Code and evaluation resources are available at https://github.com/sci-m-wang/BFI-Adapt. Keywords:personality-conditionedLLMagents;personalityevolution;majorlifeevents;psychometricevaluation; BigFivepersonality 1 Introduction Personality-conditioned LLM agents (PC-Agents) have become a foundational primitive across a growing list of applications (Chen et al., 2026). They power emotional companionship and mental- health support chatbots (Hu et al., 2026), populate social simulations (Mou et al., 2026) for behavioral research (Park et al., 2026; Larooij and Törnberg, 2025), and serve as long-horizon role-playing engines for interactive fiction, games, and digital tutoring (Park et al., 2023; Chen et al., 2024). Multi-session systems such as AnnaAgent already couple evolving emotional and cognitive states with persistent memory in psychological counselling (Wang et al., 2025). Across these settings, lifelong agents must maintain a coherent persona across extended interactions, long wall-clock durations, and open-ended user-driven narrative arcs. A foundational question for lifelong PC-Agents is how their personality evolves relative to humans. In human psychology, personality traits, though enduring, can change due to AligningMinds,EvolvingTogether. | kinamind.org1 arXiv:2608.06485v1 [cs.CL] 6 Aug 2026 KinaMindDoAIPersonasGrow? life events. We study changes in Big Five traits in one’s development and in response to consequential experiences, with major life events providing some of the best-studied anchors. Job entry and promotion are followed by increases in Conscientiousness (Roberts et al., 2006; Bühler et al., 2024). Chronic illness and unemployment are associated with increases in Neuroticism (Specht et al., 2011; Boyce et al., 2015). Retirement is followed by a documented Conscientiousness decline (Schwaba and Bleidorn, 2019). A PC-Agent that ignores such dynamics may remain locally consistent yet become globally implausible over long-term interaction, like a static script wearing a person’s name rather than a persona that develops through experience. A growing concurrent line of work has begun examining whether LLM personalities remain stable or change under temporal and contextual perturbations. Bodroža et al. (2024) report limited temporal stability and a prosocial-leaning profile across seven LLMs when personality inventories are re-administered. Yu et al. (2026) introduce PTCBench, which exposes LLM personalities to external conditions, including locations and life events, and measures aggregate trait shifts under the NEO Five-Factor Inventory. These studies show that LLM personalities can shift. Lifelong PC-Agents therefore require an evolving personality trajectory that remains coherent with the persona and plausible with respect to human psychological development. Several key properties remain unknown, including whether these shifts are systematic, whether they differ across traits and events, whether they preserve demographic and individual-level variation, and whether their direction and magnitude resemble patterns documented in human personality-change evidence. These unknowns translate into three design requirements for our diagnostic study. First, because systematic and idiosyncratic shifts can cancel in aggregate means, we move to event–trait-level diagnostics, including item-level reliability and directional consistency. Second, because direction and magnitude must be judged against an external psychological reference, we anchor each event–trait pair on an expected direction of trait change for that event, drawn from meta-analytic and primary scientific evidence (Bühler et al., 2024; Bleidorn et al., 2018; Roberts et al., 2006; Boyce et al., 2015; Schwaba and Bleidorn, 2019). Third, because persona heterogeneity is part of the target phenomenon, we instantiate each LLM with 100 demographically controlled personas, making it possible to test whether model responses preserve demographic moderation and individual-level dispersion observed in human samples. We study 11 major life events drawn from the human longitudinal literature, covering occupational events, social events, and health events. For each event, we instantiate 100 demographically controlled personas in a factorial design of 2 genders×5 cultural regions×10 personality archetypes. Each persona first completes the BFI-44 inventory, then undergoes a structured event-reflection stage, and finally completes the same inventory again. We run this pipeline across 11 LLMs. To establish the robustness of the resulting trajectory signal, we additionally conduct no-event retests, independent event paraphrases, scenario-based decision measurements, and delayed remeasurement after unrelated dialogue. The analysis is structured around four research questions, each examining a different axis of PC-Agent personality evolution against expected human change patterns. •RQ1 (Existence). Do PC-Agents show personality change after life events, in the sense that personas leave the BFI subscale noise floor? •RQ2 (Direction and magnitude). When PC-Agents move, do they match the expected direction and magnitude of human personality change for that event–trait pair? •RQ3 (Demographic shape). Do trait change patterns vary across persona gender and cultural strata in a way that resembles the moderator structure reported in human studies? •RQ4 (Individual shape). Within a single event–trait pair, do PC-Agents produce persona-specific variation comparable to the population-level dispersion in human samples, or do they collapse personas onto a shared mean trajectory? Across the four research questions, the central observation is consistent. PC-Agents move, but their movement is weakly event-specific, poorly calibrated in magnitude, and compressed across 2 KinaMindDoAIPersonasGrow? 1 Human Longitudinal Priors 2 Persona Population & Models 3 Event-Conditioned Measurement Pipeline 4 Four-Axis Diagnostic 5 Outcomes & Benchmark Human longitudinal psychology defines expected Big Five changes after life events. 11 life events Occupational (6) Social (4) Health (1) Job change Promotion at work Retirement MarriageSeparation/divorceBirth of a child Major illness Expected direction priors Human variance envelope σ ∈ [0.5, 0.8] · |d| ∈ [0.05, 0.20] 100 descriptive personas 2 genders × 5 cultural regions × 10 personality types → 100 unique personas Male, Asia. You thrive in social settings and seek novel ideas; conversations energize you. Female, Europe. You plan carefully and meet every deadline; finished tasks bring satisfaction. Female, Americas. You worry easily; setbacks linger and small frustrations feel overwhelming. Implicit Big Five descriptions (no trait labels) 11 evaluated LLMs Each persona evaluated across all 11 models for population-level comparison. Baseline BFI-44 Measure Big Five at baseline. Life Event Exposure Introduce one life event. First-person Reflection Elicit reflection about the experience. Post-event BFI-44 Re-measure Big Five after event. Trait Change Δ Compute change (post − baseline). 11 life event categories Occupational (6)Social (4)Health (1) 100 personas × 11 events × 11 models ≈ 12,100 event-conditioned trajectories Reliability anchors κ test-retest stability · DCR Direction Change Ratio RQ1 Existence Do PC-Agents show personality change after life events, in the sense that personas leave the BFI subscale noise floor? metric: pct_moved / noise-threshold crossing RQ2 Direction & Magnitude When PC-Agents move, do they match the expected direction and magnitude of human personality change for that event-trait pair? metric: % directional match & projected magnitude rate RQ3 Demographic Conditionality Do trait change patterns vary across persona gender and cultural strata in a way that resembles the moderator structure reported in human studies? metric: between-group F-test RQ4 Individual Heterogeneity Within a single event-trait pair, do PC-Agents produce persona-specific variation comparable to the population-level dispersion in human samples, or do they collapse personas onto a shared mean trajectory? metric: per-cell σ¡LLM Together, the four axes diagnose whether an LLM's response mimics or merely averages human longitudinal change. Empirical Findings ↑ Preserve direction 56.2% match cell-level priors ↓ Attenuate magnitude median |Δ| = 0.02, ~10× smaller Flatten demographic structure 11.5% retirement • all models invert Collapse heterogeneity σ¡LLM = 0.19 < human 0.5 PC-Agents simulate the mean, not the shape BFI-Adapt Benchmark KItem-level stability DCRDirectional consistency DC%Human-prior agreement ퟙ match rate51.9 - 61.1%Pearson(κ, DCR)+0.226 median σ LLM 0.19mean Δ+0.020 human priors feed back to diagnostic axes EACNO ↓=↑↓= ↓ = ↓ ↓ = ↓↑↓ Graduation Marriage Illness Graduation from school Work life entry Unemployment Starting a romantic relationship Figure1. Overview of our analytical framework for studying personality evolution in PC-Agents. demographic and persona-level variation. They therefore simulate the mean of human personality dynamics more readily than its shape. The validation suite shows that event-conditioned changes exceed retest noise, preserve their event–trait structure under independent paraphrases, exhibit model-dependent convergence with scenario-based decisions, and remain detectable after unrelated dialogue. Building on the RQ2 diagnostic, we introduce BFI-Adapt, an evaluation method for scoring whether event-induced personality shifts follow expected human directions. Figure 1 provides an overview. We make three contributions. •We frame PC-Agent personality evolution as a human-prior diagnostic over four axes: existence, direction and magnitude, demographic shape, and individual shape. •We benchmark diverse LLMs with controlled personas and paired BFI-44 measurements, revealing indiscriminate movement, magnitude miscalibration, demographic invariance, and heterogeneity col- lapse. We validate the resulting trajectory analysis through no-event retests, independent paraphrases, scenario-based decisions, and short-range retention. • We introduce BFI-Adapt, an evaluation method for measuring whether LLM-agent personality shifts adapt to expected directions exhibited by humans, and rank 14 models with it. 2 RelatedWork Static LLM Personality Assessment. Prior studies show that LLMs can produce reliable Big Five profiles, differentiate prompted personality levels, and support linguistic or inventory-based personality assessment (Serapio-García et al., 2025; Jiang et al., 2024; Zheng et al., 2025; Handa et al., 2025). These studies establish the feasibility of measuring static LLM personality, but largely treat personality as a fixed trait profile rather than a trajectory that changes after life events. Extending beyond a single inventory, Huang et al. (2024) evaluate five LLMs with thirteen clinical-psychology scales spanning personality, relationships, motivation, and emotional ability. Moving beyond self-report inventories, Wang et al. (2026) adapt projective tests to LLMs and find improved resistance to contamination and social-desirability bias, together with greater sensitivity to prompt-induced changes. They further demonstrate GenPT in longitudinal counselling. These results motivate our use of BFI-44 as a standardized and interpretable measurement anchor for systematic, event-conditioned trajectories. 3 KinaMindDoAIPersonasGrow? Role-Playing Agents. Role-playing agents provide another route to personality-conditioned behavior, with benchmarks and training methods evaluating whether models remain faithful to a specified character or persona (Wang et al., 2024a,b; Ran et al., 2024). Recent work also shows that user personas can shift chatbot personality (Xing et al., 2025), while induced personas yield stable, task-dependent changes in cognitive performance (Chen et al., 2026). However, this line of work primarily targets persona fidelity, whether the model stays in character, rather than whether the in-character personality evolves in psychologically plausible ways. CharacterEval, for example, evaluates role-playing conversations with thirteen metrics over four dimensions and a multi-turn benchmark of 77 characters (Tu et al., 2024). Such work clarifies whether an agent preserves its assigned character, whereas our controlled personas provide fixed initial conditions from which change is measured. We retain persona conditioning throughout and ask whether subsequent adaptation is event-specific and comparable to human change patterns. Dynamic LLM Personality. Recent studies have begun testing whether LLM personalities remain stable or shift under temporal and contextual perturbations. Bodroža et al. (2024) re-administer personality inventories to seven LLMs and report limited temporal stability and a prosocial-leaning profile. Yu et al. (2026) introduce PTCBench, exposing LLMs to external conditions, including locations and life events, and measuring aggregate NEO-FFI trait shifts. Complementing these inventory-based findings, Han et al. (2025) compare LLM self-reports with behavioral tasks and find that persona interventions steer reported traits more reliably than behavior. Together, these works show that LLM personality can change, but they provide limited analysis of how such changes unfold across traits, events, personas, and models, or whether the measured shifts remain consistent across complementary checks. Our work addresses these gaps by moving from aggregate trait shifts to item-level reliability, anchoring event–trait pairs in human change-direction priors, and testing demographic conditionality and individual-level heterogeneity with 100 personas per model. We further evaluate trajectory robustness through retest comparisons, independent paraphrases, scenario-based decision behavior, and short-range retention. 3 Methodology 3.1 PersonaDesign To isolate personality dynamics from confounding factors inherent in fictional characters, such as narrative expectations, cultural stereotypes, and fan interpretations, we created 100 descriptive personas using a factorial design. It crosses 2 genders (male, female), 5 cultural regions (Europe, Americas, Africa, Asia, Oceania), and 10 personality types (P1–P10). The 10 personality types are constructed from implicit Big Five design anchors. We use a five-level ordinal scheme with “very low,” “low,” “moderate,” “high,” and “very high” levels for persona construction. Each personality type assigns one focal Big Five trait to either “very high” or “very low,” while keeping the remaining four traits “moderate.” Thus, the 10 types correspond to the very-high and very-low variants of five traits. These ordinal anchors guide the writing of behavioral and attitudinal persona descriptions, such as “thrives in social gatherings and draws energy from conversation.” The BFI-44 administrations provide the standard 1–5 Big Five scores used in all analyses. 3.2 LifeEvents We selected 11 life events associated with Big Five personality changes in the psychology literature. They span occupational events (graduation, work entry, job change, promotion, unemployment, retirement), 4 KinaMindDoAIPersonasGrow? social events (new relationship, marriage, divorce, childbirth), and health events (chronic illness). To define the expected direction of change for each event and trait, we use Specht (2017) as the primary source and consolidate its reported directions into the event–trait prior matrix in Table 1. We make one notational modification. Emotional Stability (ES) is inverted to Neuroticism (N) to match our BFI-44 scoring convention (ES+≡N−). Independent meta-analytic and primary studies support the individual rows for job entry and promotion (Roberts et al., 2006; Bühler et al., 2024), unemployment (Boyce et al., 2015), retirement (Schwaba and Bleidorn, 2019), and chronic illness (Specht et al., 2011). 3.3 PersonalityAssessmentPipeline We design a four-stage pipeline to measure personality change in a model that simulates a persona experiencing a life event. The system message specifies gender, cultural region, and a behavioral personality description. The model answers all subsequent prompts in character. Baseline Measurement administers the BFI-44 inventory (John et al., 2008) and yields a Big Five score vectorb, where b 푡 ∈ [1, 5]and푡 ∈ 퐸, 퐴,퐶, 푁,푂. Event Presentation gives the simulated persona a life event and elicits a first-person reflection. Post-Event Measurement administers BFI-44 again and yields the score vectorp. Change Computation calculatesΔ 푡 = p 푡 − b 푡 for each trait. The 44 items use a 1–5 Likert scale with standard reverse scoring. Trait scores average the corresponding items, with 8 items for E, 9 for A, 9 for C, 8 for N, and 10 for O. The resulting trait-specific step sizes are 0.125 for E and N, 0.111 for A and C, and 0.1 for O. 3.4 StatisticalAnalysis We evaluate change direction at the persona level because aggregate means can cancel opposing individual shifts. A mean near zero can coexist with directional effects across personas. We therefore adopt an individual-level direction classification approach. Direction Classification. For each persona 푝, event 푒, and trait 푡, we classify the change direction: 푑 푝,푒,푡 = + if Δ 푝,푒,푡 > 휀 − if Δ 푝,푒,푡 <−휀 0 if|Δ 푝,푒,푡 | ≤ 휀 (1) Table1. Expected directions of human Big Five trait change across 11 life events. “–” denotes no strong change. “±” denotes a direction that depends on role demand. EventE A C N O Graduation ↓ – ↑ ↓ – Work Entry– – ↑ ↓ – Job Change ± – – – ± Promotion– – ↑ ↓ ↑ Unemployment – ↓ ↓ – ↓ Retirement– – ↓ – – New Rel. ↑ ↓ – ↓ – Marriage ↓ ↓ – ↓ ↓ Divorce– ↑ – – ↑ Child Birth ↓ – ↓ – – Chronic Ill. ↓ – ↓ ↑ ↓ 5 KinaMindDoAIPersonasGrow? where휀 = 0.1serves as a conservative noise boundary. A shift exactly at±0.1remains neutral. Shifts beyond that boundary count as directional, and neutral cases are excluded from the analysis denominator. Match Rate. Suppose푃denotes a set of personas. For each event–trait pair(푒,푡)with an expected human direction 푑 human 푒,푡 , the match rate is defined by: 푃 match (푒,푡) = |푝 ∈ 푃 : 푑 푝,푒,푡 = 푑 human 푒,푡 | |푝 ∈ 푃 : 푑 푝,푒,푡 ≠ 0| (2) 푃 match (푒,푡) is therefore the conditional probability that a persona’s non-neutral shift points in the expected human direction for that event–trait pair. The metric is defined for pairs with a definite expected human direction, with neutral responses excluded from the denominator. Statistical Testing. The statistical tests evaluate whether the observed directional evidence exceeds a minimal chance baseline and whether it varies across demographic strata. For each definite-prior event– trait pair, the directional null is퐻 0 : 푃 match (푒,푡) = 0.5. We apply a one-tailed binomial sign test with Wilson-score confidence intervals and Benjamini–Hochberg FDR correction at훼=0.05. Cross-group comparisons for gender and world-region effects use Fisher’s exact tests on2× 2match/mismatch tables, with Benjamini–Hochberg correction within each test family. Appendix E presents the full procedure. 3.5 Item-LevelReliability: 휅andDCR Match rate captures the alignment of a moved persona with the expected direction. Two additional indicators characterize the underlying item responses at the event–trait-pair level. Linearly weighted Cohen’s휅measures item-level rating stability between baseline and post-event answers. DCR measures whether the residual item changes within the same pair are mostly one-sided. Pair-level computation preserves the distinct human direction associated with each event–trait pair and prevents opposing directions from cancelling during aggregation. Pair-level linearly-weighted Cohen’s휅. For each model and event–trait pair(푒,푡), let푛 푡 be the number of BFI-44 items belonging to trait푡(Section 4.1). The pair contains100푛 푡 paired item ratings (100 personas× 푛 푡 subscale items, 퐾=5 ordinal categories). With weights 푤 푖푗 = 1−|푖−푗|/(퐾−1), 휅 푒,푡 = Í 푖, 푗 푤 푖푗 푝 (표) 푖푗 − Í 푖, 푗 푤 푖푗 푝 (푏) 푖 푝 (푝) 푗 1 − Í 푖, 푗 푤 푖푗 푝 (푏) 푖 푝 (푝) 푗 ∈ [−1, 1],(3) where푝 (표) 푖푗 is the empirical joint frequency of baseline rating푖and post-event rating푗within the pair and푝 (푏) 푖 , 푝 (푝) 푗 are the corresponding marginals.휅 푒,푡 =1is perfect agreement,휅 푒,푡 =0is chance-level agreement. We summarize each model by ̄휅 pair , the mean of휅 푒,푡 over the 27 event–trait pairs with a definite expected human direction. Pair-level Directional Consistency Ratio (DCR). Within the same event–trait pair, let푛 ↑ and푛 ↓ be the number of items whose post-event rating increases and decreases (reverse-coded items pre-aligned), respectively. Among the 푛 ↑ + 푛 ↓ items that change, DCR 푒,푡 = max(푛 ↑ ,푛 ↓ ) 푛 ↑ + 푛 ↓ ∈ [0.5, 1].(4) If no item changes within a pair,DCR 푒,푡 has no empirical denominator. For model-level summaries we report such pairs asDCR 푒,푡 =0.5, equivalent to no dominant direction and hence푆 푒,푡 =2 DCR 푒,푡 −1=0 6 KinaMindDoAIPersonasGrow? in BFI-Adapt. We summarise each model byDCR pair , the mean ofDCR 푒,푡 over the same 27 definite- direction pairs. Thus,휅 푒,푡 measures the similarity of item ratings before and after the event, while DCR 푒,푡 measures whether changed ratings move primarily in one direction. Pooled forms휅 pool and DCR pool are reported in Table 13 to show the artefact created by aggregating across pairs with opposite expected directions. Appendix F details the four diagnostic regimes formed by the(휅 푒,푡 , DCR 푒,푡 )plane. 3.6 ValidationDesign We evaluate the robustness of the event-conditioned trajectory signal through four complementary designs. A no-event BFI-44 retest establishes a model-specific measurement floor. Independent event paraphrases test whether the event–trait structure remains stable under new wording. Counterbalanced parallel forms of ten scenario-based decisions measure convergence between BFI trait changes and concrete choices. A delayed BFI-44 administration after three unrelated dialogue turns measures short-range retention. Each condition starts from a fresh conversation. Confidence intervals use persona-cluster bootstrap resampling. Additional prompts, counterbalancing procedures, and statistical details appear in Appendix G. 4 Experiments 4.1 ExperimentSettings Models. The primary four-axis analysis covers 11 LLMs. These are GPT-5.3-chat (Singh et al., 2026), GPT-4.1-mini (OpenAI et al., 2024), Claude-Sonnet-4.6, Claude-Haiku-4.5, Gemini-3-flash (Gemini Team et al., 2025), DeepSeek-V4-Pro (DeepSeek-AI, 2026), Doubao-Seed-2.0-Pro, MiMo-V2.5-Pro (Xiaomi MiMo Team, 2026), Qwen3-235B (Yang et al., 2025), GLM-4.6 (GLM-4.5 Team et al., 2025), and Kimi-K2 (Kimi Team et al., 2026). Three open-weight models, Qwen3.5-9B, Qwen2.5-14B-Instruct, and InternLM3-8B-Instruct, run the same grid and extend the BFI-Adapt leaderboard to 14 models. The validation suite evaluates DeepSeek-V4-Pro, GPT-5.5, Grok-4.3, Kimi-K2.6, Mistral-Large-3, Qwen3.5-9B, Qwen2.5-14B-Instruct, and InternLM3-8B-Instruct. All models use non-reasoning mode. Each main-grid model completes the full100personas× 11events× 44items design. The validation models use the same 100 personas and 11 events under each validation condition. Appendix F reports response-integrity checks. Prompt. The system prompt establishes gender, continent, and a behavioural personality description that never names Big Five labels. Appendix B shows the exact persona-simulation prompt, and Appendix C gives one inserted personality description. BFI-44 items are answered in character. Event scenarios appear as first-person narratives that the persona reflects on before the post-event items. The prompts are available at https://github.com/sci-m-wang/BFI-Adapt. 4.2 RQ1:Existence RQ1 asks whether PC-Agents produce measurable personality responses after life events, before asking whether such responses are psychologically targeted. We define movement at the persona level as|Δ| > 0.1, a conservative threshold just above the smallest BFI-44 trait-level step, and compute pct moved (푚,푒,푡) = |푝 : |Δ 푝,푒,푡 | > 0.1|/100 for each model–event–trait pair. We compute this statistic over all 605 model–event–trait combinations and, within each model, compare 27 pairs with documented human change directions against 28 pairs without a definite direction. Within-model median differences quantify the separation between these two groups. Appendix Table 11 7 KinaMindDoAIPersonasGrow? Table2. Pair-level reliability, directional consistency, and the BFI-Adapt composite for the 14-model PC-Agent leaderboard. Scores are computed over the 27 event–trait pairs with a definite expected human direction. The last three rows form the open-weight extension. Model ̄휅 pair DCR pair DC% pair Adapt Gemini-3-flash0.8620.81159.30.348 GLM-4.60.8250.80463.00.322 Qwen3-235B0.7340.79163.00.287 Claude-Haiku-4.50.8290.69766.70.225 Kimi-K2-09050.8370.70270.40.224 Doubao-Seed-2.0-Pro0.7670.77148.10.218 Claude-Sonnet-4.60.8600.70755.60.212 GPT-4.1-mini0.8420.66655.60.192 GPT-5.3-chat0.8140.70451.90.178 DeepSeek-V4-Pro0.7930.69155.60.176 MiMo-V2.5-Pro0.5760.59863.00.071 Qwen3.5-9B0.7930.73963.00.194 Qwen2.5-14B-Instruct0.5220.57155.60.049 InternLM3-8B-Instruct0.5150.56363.00.044 reports per-model quartiles and high-tail massHi%(pct moved > 0.6). Movement occurs broadly in both groups. Definite-direction medians span 0.44–0.84, no-definite-direction medians span 0.42–0.82, and within-model median differences remain below 0.05. In 9 of 11 models, the twoHi%values differ by less than 10 percentage points. The two larger differences favor no-definite-direction pairs. PC-Agent movement is therefore widespread but weakly targeted toward pairs with human directional evidence. As a reliability check for the subsequent direction analysis, we also inspect pair-level ̄휅 pair on the 27 definite-direction pairs. It spans 0.58–0.86 across the 11-model benchmark (Table 2). 10 of the 11 models exceed ̄휅 pair =0.66. MiMo-V2.5-Pro is the single low-reliability outlier at 0.58 and also the model with near-universal movement (Hi%=1.00in both buckets). High ̄휅 pair demonstrates reliable item-level ratings within each pair. More than half of personas still cross the noise floor on a typical pair. Together, these properties support the direction and magnitude analysis in Section 4.3. 4.3 RQ2:DirectionandMagnitude RQ2 asks two linked questions. When a persona moves, does it move in the expected human direction? Is the signed magnitude comparable to human longitudinal effects? We first use pair-level DCR to characterize directional item changes within each event–trait pair. We then use푃 match andDC% pair to measure agreement with the expected direction from human longitudinal studies. Per-pair푃 match is tested against the50%chance baseline via Wilson 95% CIs with BH-FDR at푞=0.05(Appendix E). DCR pair , averaged over the 27 definite-direction pairs, spans 0.60–0.81 across the 11 models, with every model above 0.59. Residual item drift is therefore usually one-sided within a pair. Directional agreement is weaker.DC% pair spans 48.1–70.4%, from Doubao-Seed-2.0-Pro at the bottom to Kimi-K2-0905 at the top. Stability and systematicity are modestly correlated across models with푟=+0.455, confirming that they capture complementary properties. At finer resolution, the model-by-event heatmap (Figure 7, Appendix H) shows three recurring patterns. First, occupational onboarding events are easiest: Graduation (42–76%), Work Entry (49–96%), and Promotion (53–86%) are correctly directed by most models, consistent with the well-represented occupational-role→Conscientiousness association (Roberts et al., 2006). Second, Retirement breaks every model: the documented post-retirement Conscientiousness decline (Schwaba 8 KinaMindDoAIPersonasGrow? Table3. Event–trait direction patterns. Each cell shows the expected direction followed by pct + / pct − , the percentages of personas with post-event increases and decreases. Green marks a matched expected direction, red a reversal, yellow drift without a definite prior, white stability without a definite prior, and gray a context-dependent direction. EventEACNO Graduation↓ 17/34 – 30/21↑ 35/17↓ 22/33– 28/39 Work Entry– 17/36– 33/17↑ 58/7↓ 17/37– 22/46 Job Change± 19/39 – 27/21– 52/12– 23/38± 34/36 Promotion – 27/29– 35/17↑ 69/5↓ 13/46↑ 24/45 Unemployment– 22/32↓ 26/15↓ 45/10– 43/17↓ 20/48 Retirement – 14/42– 38/16↓ 64/7– 9/49– 13/60 New Rel.↑ 25/28↓ 30/15 – 25/20↓ 13/33– 31/40 Marriage↓ 26/26↓ 36/17 – 35/15↓ 9/42↓ 19/50 Divorce – 18/41↑ 22/32– 34/20– 66/14↑ 17/49 Child Birth↓ 23/33 – 37/14↓ 33/22– 37/21– 21/49 Chronic Ill. ↓ 16/33– 25/25↓ 33/17↑ 49/14↓ 16/53 and Bleidorn, 2019) is universally reversed (0.0–38.9%, median 11.5%), with models predicting retirees to become more conscientious. Third, social events remain near chance (Marriage 44–74%, Divorce 40–76%, and Child Birth 25–63%), whereas Chronic Illness is more consistently captured, likely because its expected profile is salient (E/C/O decline, N increases). Aggregating by trait, Agreeableness is the weakest dimension (30–58% across models), below Openness (53–62%), Conscientiousness (45–62%), Extraversion (41–74%), and Neuroticism (49–79%), reflecting a default tendency to make personas more agreeable after major events even when human evidence expects the opposite. For magnitude, we project each persona-level change onto the expected human direction, e Δ = sign(푑 human 푒,푡 ) · Δ . The meta-analysis reports standardized mean change푑(Bühler et al., 2024), for which we use the representative band|푑| ∈ [0.05, 0.20]. Applying a representative BFI trait SD of 휎≈0.7gives the raw-change reference band e Δ∈ [0.035, 0.14]Likert units. We classify responses as reversed ( e Δ < 0 ), under-shift (0≤ e Δ < 0.035), in-range (0.035≤ e Δ≤ 0.14), or overshoot ( e Δ > 0.14 ). Binning is performed separately for each model over its 2700 persona-level responses (100 personas× 27 definite-direction pairs). Across models, only 11.0%–16.4% of responses fall inside the human reference band. The rest are reversed (20.8%–40.2%), under-shifted despite the correct direction (15.1%–54.0%), or overshooting (9.9%–31.6%). Failure modes differ by model. Kimi-K2-0905 most often under-shifts, MiMo-V2.5-Pro most often reverses or overshoots, and the best-calibrated model, Gemini-3-flash, reaches 16.4% in-range. Thus, PC-Agents partly recover direction but rarely calibrate the magnitude of change to human effect-size ranges. Table 3 expands the comparison to all five traits per event. Each entry reportspct + / pct − , the cross-model median fraction of personas withΔ>0andΔ<0. Definite-direction pairs are marked as match (green) or reverse (red) using a 10% net directional-intensity threshold. No-definite-direction pairs are marked as drift (yellow) when over half the personas leave the noise floor. Among the 27 definite-direction pairs, 14 (51.9%) match the expected direction and 13 (48.1%) reverse it, with failures concentrated on retirement, unemployment, divorce, and new relationships. Among the 26 no-definite-direction pairs, 21 (80.8%) drift, showing broad spillover onto bystander traits, around events that produce directional failures. 4.4 RQ3:DemographicShape RQ3 asks whether event-driven change is systematically conditioned on persona demographics, motivated by the demographic moderation often reported in human longitudinal studies. We answer it with two statistics. First, for each model and event–trait pair, we compute the medianΔwithin the 9 KinaMindDoAIPersonasGrow? 10 demographic strata (2 genders×5 continents), then take the standard deviation of these stratum medians. Across the푁=605model–event–trait combinations, the median across-strata SD is 0.044 BFI Likert units and 93.2% of combinations have across-strata SD below 0.10. Per-model medians range from 0.025 (GPT-4.1-mini, Kimi-K2-0905) to 0.099 (MiMo-V2.5-Pro). Nine of 11 models have at least 98.2% of pairs below 0.10. Qwen3-235B reaches 80.0%, and MiMo-V2.5-Pro reaches 50.9%. Appendix Table 12 reports the full per-model median, maximum, and below-0.10 fraction. Second, restricted to the 297 definite-direction combinations where match is defined, Fisher’s exact tests compare match and mismatch counts across gender and continent strata. No gender comparison survives Benjamini–Hochberg correction at푞=0.05, and the analogous continent-pair tests are likewise null. Demographics may shape the wording of the reflection, but they do not measurably change the post-event BFI item ratings. 4.5 RQ4:IndividualShape RQ4 asks whether different personas within the same event–trait pair respond differently, or whether the model collapses them onto a narrow shared trajectory. For each model and event–trait pair we compute three persona-level dispersion indicators on the 100 personas:휎 LLM (standard deviation ofΔ),IQR LLM (inter-quartile range ofΔ), andpct moved (the noise-floor crossing rate from RQ1). We compare휎 LLM with the human within-trait SD envelope of 0.5–0.8 BFI Likert units (Bühler et al., 2024; Roberts et al., 2006) as an effect-size benchmark. Across all 605 model–event–trait combinations, 99.8% fall below the lower human bound of 0.5 and 88.3% fall below 0.3. The distribution is centred at휎 LLM =0.19 with medianIQR LLM =0.125. Per-model median휎 LLM ranges from 0.138 (Claude-Sonnet-4.6) to 0.361 (MiMo-V2.5-Pro). The widest model, MiMo-V2.5-Pro, still places 98.2% of pairs below 0.5. Baseline across-persona trait SD is approximately 0.7, about 3.6×the median event-induced change SD, confirming that the personas are initially well differentiated. Event-induced trajectories remain structurally compressed across the field, including strong direction models such as GLM-4.6 and Qwen3-235B. Per-model 휎 LLM violins are shown in Figure 6. 5 BFI-Adapt:ACompositeBenchmark For a PC-Agent’s event–trait change pattern to reflect plausible personality evolution, it should satisfy three conditions. These are reliable item-level measurement, systematic directional movement, and agreement with the expected human direction. We score these conditions per pair and average over the 27 pairs with a definite expected human direction. BFI-Adapt = 1 |C| ∑︁ (푒,푡)∈C max(0, 휅 푒,푡 ) 푆 푒,푡 ⊮ 푒,푡 dir ,(5) The transformed directional-consistency term is푆 푒,푡 =2 DCR 푒,푡 −1. The scoring setCcontains event– trait pairs with positive or negative human priors. Both휅 푒,푡 andDCR 푒,푡 use the pair’s100푛 푡 item ratings. The indicator⊮ 푒,푡 dir equals one when the pair’s dominant DCR direction matches its prior, with ties assigned zero. All-tied pairs contribute휅=1.0to rating reliability and zero to BFI-Adapt because their directional term is zero. The legacy pooled composite is reported in Appendix H as a pooling-artefact control. The benchmark instantiates each evaluated model with a fixed system-prompt template, 100 demographically controlled personas, 11 life-event scenarios with human-anchored direction priors, and the BFI-44 inventory (John et al., 2008). Each model yields48,400paired item ratings. The benchmark reports rating reliability, within-pair systematicity, alignment with the expected human direction, and 10 KinaMindDoAIPersonasGrow? their BFI-Adapt composite. It scores structured changes in BFI-44 item responses. The reflection provides conditioning context for the post-event ratings. BFI-Adapt spans 0.071 (MiMo-V2.5-Pro) to 0.348 (Gemini-3-flash) among the 11 API models, a 4.9×range. The pooled diagnostic reorders 9 of 11 models by at least two ranks. Haiku and Kimi rise from the bottom to fourth and fifth once pair-level DCR (≈ 0.70) replacesDCR pool ≈0.51. Doubao-Seed-2.0-Pro drops from third to sixth withDC% pair =48.1%. The top three models are Gemini-3-flash, GLM-4.6, and Qwen3-235B. They jointly exceed0.79inDCR pair and59%inDC% pair , and retain their positions when the open-weight models join the leaderboard. The open-weight extension also shows that model scale alone does not determine the score. Qwen3.5-9B reaches 0.194, close to GPT-4.1-mini, while Qwen2.5-14B-Instruct and InternLM3-8B-Instruct score 0.049 and 0.044. TheirDC% pair values are comparable to the field. Lower item-level reliability and weaker within- pair directionality account for the score gap. The three pair-level indicators capture complementary 0.00.10.20.30.4 BFI-Adapt 1. Gemini-3-flash 2. GLM-4.6 3. Qwen3-235B 4. Claude-Haiku-4.5 5. Kimi-K2-0905 6. Doubao-Seed-2.0-Pro 7. Claude-Sonnet-4.6 8. Qwen3.5-9B 9. GPT-4.1-mini 10. GPT-5.3-chat 11. DeepSeek-V4-Pro 12. MiMo-V2.5-Pro 13. Qwen2.5-14B-Instruct 14. InternLM3-8B-Instruct 0.348 0.322 0.287 0.225 0.224 0.218 0.212 0.194 0.192 0.178 0.176 0.071 0.049 0.044 open-weight (leaderboard extension) Figure2. BFI-Adapt ranking across all 14 models. Stars and dashed stems mark the three open-weight extension models. The field spans 7.9×overall and 4.9×among the 11 API models. Trophies mark the top three. properties. Pearson푟( ̄휅 pair ,DCR pair )=+0.455, while Kimi-K2 leadsDC% pair and Gemini-3-flash leads BFI-Adapt. Scores aggregate response patterns, and RQ4 supplies the corresponding persona-level heterogeneity analysis. We release the persona set, scenarios, scoring code, and per-model logs. 6 ValidityoftheMeasurementAnchor The four-axis analysis and BFI-Adapt use paired BFI-44 administrations as a standardized anchor that connects PC-Agent trajectories to trait-level human priors. Section 3.6 defines four complementary checks of this trajectory signal. The validation suite applies them to the full 100-persona, 11-event grid across eight API and open-weight models. Table 5 reports the main results. Additional condition-level results and bootstrap intervals appear in Appendix G. 11 KinaMindDoAIPersonasGrow? Separation from retest noise. Mean absolute change in the no-event retest ranges from 0.025 to 0.080 BFI units across the eight models. Under the event-plus-reflection condition, it ranges from 0.100 to 0.237 and exceeds the corresponding retest floor by 1.6×to 9.0×. The persona-cluster bootstrap 95% interval of the paired excess remains above zero for every model. Event notification alone also produces above-retest movement across all eight models, while the incremental contribution of reflection varies across models. These comparisons establish a consistent separation between event-conditioned trajectories and measurement variability. Robustness to independent paraphrases. Across the 55 event–trait cells, the original and inde- pendently paraphrased events agree in sign on 80.0% to 92.7% of cells. Their cell-level Spearman correlations range from 0.825 to 0.956. The direction and relative ordering of event–trait responses therefore remain stable under new wording. Convergence with scenario-based decisions. The scenario decisions provide a second response channel at the level of concrete choices. Table 4 reports the model-level results. Across the eight models, correlations between BFI trait changes and decision-score changes are positive and range from휌=0.003to 0.105. Direction agreement among non-zero changes ranges from 48.4% to 62.7%. The bootstrap intervals for DeepSeek-V4-Pro, Mistral-Large-3, and Qwen2.5-14B lie above zero, with correlations of 0.105, 0.097, and 0.060, respectively. The strength of convergence varies across models even though their BFI trajectories remain stable under retesting and prompt paraphrases. This pattern separates the reproducibility of the measured trajectory from its expression in situation-specific choices. Cross-format behavioral convergence therefore forms a distinct dimension of personality adaptation in PC-Agents. Table4. Convergence between BFI trait changes and scenario-based decision behavior. Decision 휌 is the persona–event–trait Spearman correlation. Sign is direction agreement among non-zero BFI and decision-score changes, and 푛 is the number of aligned non-zero observations. Intervals use persona-cluster bootstrap resampling. ModelDecision 휌 [95% CI]Sign푛 DeepSeek-V4-Pro0.105 [0.050, 0.161]0.627708 GPT-5.50.009 [−0.044, 0.060]0.512383 Grok-4.30.052 [−0.012, 0.120]0.610849 Kimi-K2.60.003 [−0.059, 0.073]0.484699 Mistral-Large-30.097 [0.048, 0.150]0.614816 InternLM3-8B0.024 [−0.024, 0.069]0.5271,827 Qwen2.5-14B0.060 [0.013, 0.107]0.5731,409 Qwen3.5-9B0.035 [−0.017, 0.097]0.542756 Short-range retention after unrelated dialogue. Immediate and delayed BFI change vectors remain positively correlated across every model, with휌ranging from 0.329 to 0.713 and all persona-cluster bootstrap intervals above zero. Among units whose immediate change exceeds the empirical retest threshold, 62.6% to 85.3% retain the same direction after three unrelated dialogue turns. Kimi-K2.6 and Mistral-Large-3 reach 85.3% retention, followed by GPT-5.5 at 81.9% and DeepSeek-V4-Pro at 80.5%. These results place the measured trajectories beyond a single adjacent response. They remain detectable after a controlled change in conversational topic. 12 KinaMindDoAIPersonasGrow? Table5. Selected results from the validation suite. Retest is the mean absolute BFI trait change across two baseline administrations. Event denotes the event-notification condition, while E+R denotes the original event plus reflection. Sign is the fraction of event–trait cells where original and paraphrased prompts agree in sign.휌 del is the Spearman correlation between immediate and delayed changes. Ret. is direction retention among above-threshold immediate movers. ModelRetestEventE+RSign휌 del Ret. DeepSeek-V4-Pro0.0510.1460.1570.910.670.81 GPT-5.50.0640.1230.1000.890.580.82 Grok-4.30.0800.1330.1440.800.550.72 Kimi-K2.60.0480.1090.1050.820.650.85 Mistral-Large-30.0420.1320.1340.850.710.85 InternLM3-8B0.0280.2400.2370.870.330.63 Qwen2.5-14B0.0250.2050.2280.910.370.67 Qwen3.5-9B0.0270.1260.1570.930.610.77 Together, these results establish the robustness of the trajectory analysis across repeated measurement, independent event wording, scenario-based decisions, and intervening dialogue. They support BFI-44 as a reproducible measurement anchor for the four-axis diagnostic and BFI-Adapt. 7 Conclusion This paper explored whether PC-Agents exhibit psychologically plausible personality evolution after major life events. We measured their Big Five profiles before and after each event and used the resulting trajectories to construct BFI-Adapt. Current models can move, but their movement is weakly event- specific, poorly calibrated in magnitude, and compressed across demographic and individual variation. Present PC-Agents therefore approximate a generic pattern of change more readily than the event- and person-specific structure of human personality development. Event-conditioned trajectories exceed retest noise, preserve their event–trait structure under independent paraphrases, exhibit model-dependent convergence with scenario-based decisions, and remain detectable after unrelated dialogue. These complementary checks validate the trajectory analysis and its central conclusions. By combining human change-direction priors, pair-level reliability diagnostics, and the BFI-Adapt evaluation method, this work turns PC-Agent personality evolution from an observed capability into a psychologically anchored object of measurement. 8 LimitationsandEthicalConsiderations Our pre-post design captures immediate personality response but cannot assess the acute-phase-to- adaptation trajectory documented in human longitudinal studies (Luhmann et al., 2012). The delayed measurement in Section 6 extends the window by three unrelated turns and still stops far short of the multi-year horizons of human panels. Whether PC-Agents would reproduce the well-documented recovery of life satisfaction after major life events is an open question that requires a multi-step elicitation protocol. The expected-direction table anchoring DC% and the match-rate heatmap is adopted verbatim from Specht (2017), with Emotional Stability inverted to Neuroticism, and pairs marked uncertain or context-dependent (Job Change; the no-strong-direction pairs) excluded from scoring. The human휎 envelope (RQ4) and human effect-size envelope (RQ2) are themselves coarse meta-analytic ranges (Bühler et al., 2024; Bleidorn et al., 2018). Replacing the direction prior with a denser pair-level effect-direction matrix would tighten DC% but would not change the headline finding that every model 13 KinaMindDoAIPersonasGrow? in the benchmark reverses the retirement trend. This research involves no human subjects. All “personas” are synthetic text constructs with no real-world counterparts. We note that the personality simulation capabilities evaluated here could potentially be misused for social manipulation; however, our findings suggest that current LLMs’ personality dynamics are insufficiently accurate for such applications. We release our evaluation framework to enable further research on psychologically valid LLM behavior. 9 GenerativeAIUsage We used AI assistants to polish the writing for grammar and clarity, and to support coding tasks such as drafting analysis scripts and refactoring boilerplate. All research ideas, experimental design, statistical decisions, and final claims were produced and verified by the authors. References Wiebke Bleidorn, Christopher J. Hopwood, and Richard E. Lucas. Life events and personality trait change. Journal of Personality, 86(1):83–96, 2018. doi: 10.1111/jopy.12286. Bojana Bodroža, Bojana M. Dinić, and Ljubiša Bojić. Personality testing of large language models: Limited temporal stability, but highlighted prosociality. Royal Society Open Science, 11(10):240180, 2024. doi: 10.1098/rsos.240180. Christopher J. Boyce, Alex M. Wood, Michael Daly, and Constantine Sedikides. Personality change following unemployment. Journal of Applied Psychology, 100(4):991–1011, 2015. doi: 10.1037/a0038647. Janina L. Bühler, Ulrich Orth, Wiebke Bleidorn, Elisa Weber, André Kretzschmar, Larissa Scheling, and Christopher J. Hopwood. Life events and personality change: A systematic review and meta-analysis. European Journal of Personality, 38(3):544–568, 2024. doi: 10.1177/08902070231190219. Jiaqi Chen, Ming Wang, Tingna Xie, Shi Feng, and Yongkang Liu. A systematic analysis of the impact of persona steering on LLM capabilities. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 48, 2026. URL https://escholarship.org/uc/item/0964m2kc. Ning Chen, Yu Wang, Yang Deng, and Jing Li. The Oscars of AI theater: A survey on role-playing with language models. arXiv preprint arXiv:2407.11484, 2024. doi: 10.48550/arXiv.2407.11484. DeepSeek-AI. Deepseek-v4: Towards highly efficient million-token context intelligence, 2026. Gemini Team et al. Gemini: A family of highly capable multimodal models, 2025. URLhttps://arxiv.org/ abs/2312.11805. GLM-4.5 Team et al. Glm-4.5: Agentic, reasoning, and coding (arc) foundation models, 2025. URL https://arxiv.org/abs/2508.06471. Pengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath, Dean Mobbs, Anima Anandkumar, and R. Michael Alvarez. The personality illusion: Revealing dissociation between self-reports & behavior in LLMs. arXiv preprint arXiv:2509.03730, 2025. doi: 10.48550/arXiv.2509.03730. URLhttps://arxiv.org/abs/2509. 03730. Gaurish Handa, Zhijin Wu, Adriano Koshiyama, and Philip Treleaven. Personality as a probe for LLM evaluation: Method trade-offs and downstream effects. arXiv preprint arXiv:2509.04794, 2025. doi: 10.48550/arXiv.2509.04794. He Hu, Yucheng Zhou, Qianning Wang, Yingjian Zou, Chiyuan Ma, Juzheng Si, Jianzhuang Liu, Zitong Yu, Laizhong Cui, Fei Ma, and Qi Tian. From pattern recognizers to personalized companions: a survey of large language models in mental health. IEEE Transactions on Affective Computing, pages 1–20, 2026. doi: 10.1109/TAFFC.2026.3689490. Jen-tse Huang, Wenxuan Wang, Eric John Li, Man Ho Lam, Shujie Ren, Youliang Yuan, Wenxiang Jiao, Zhaopeng Tu, and Michael R. Lyu. On the humanity of conversational AI: Evaluating the psychological portrayal of LLMs. In The Twelfth International Conference on Learning Representations, 2024. URL 14 KinaMindDoAIPersonasGrow? https://arxiv.org/abs/2310.01386. Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara. PersonaLLM: Investigating the ability of large language models to express personality traits. In Kevin Duh, Helena Gomez, and Steven Bethard, editors, Findings of the Association for Computational Linguistics: NAACL 2024, pages 3605–3627, Mexico City, Mexico, June 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-naacl.229. URL https://aclanthology.org/2024.findings-naacl.229/. Oliver P. John, Laura P. Naumann, and Christopher J. Soto. Paradigm shift to the integrative Big Five trait taxonomy: History, measurement, and conceptual issues. In Oliver P. John, Richard W. Robins, and Lawrence A. Pervin, editors, Handbook of Personality: Theory and Research, pages 114–158. Guilford Press, New York, 3rd edition, 2008. Kimi Team et al. Kimi k2: Open agentic intelligence, 2026. URL https://arxiv.org/abs/2507.20534. Maik Larooij and Petter Törnberg. Validation is the central challenge for generative social simulation: a critical review of LLMs in agent-based modeling. Artificial Intelligence Review, 59(1):15, November 2025. ISSN 1573- 7462. doi: 10.1007/s10462-025-11412-6. URL https://doi.org/10.1007/s10462-025-11412-6. Maike Luhmann, Wilhelm Hofmann, Michael Eid, and Richard E. Lucas. Subjective well-being and adaptation to life events: A meta-analysis. Journal of Personality and Social Psychology, 102(3):592–615, 2012. doi: 10.1037/a0025948. Xinyi Mou, Xuanwen Ding, Qi He, Liang Wang, Jingcong Liang, Xinnong Zhang, Libo Sun, Jiayu Lin, Jie Zhou, Huang Xuanjing, and Zhongyu Wei. From individual to society: A survey on social simulation driven by large language model-based agents. ACM Comput. Surv., 58(11), April 2026. ISSN 0360-0300. doi: 10.1145/3800683. URL https://doi.org/10.1145/3800683. OpenAI et al. Gpt-4 technical report, 2024. URL https://arxiv.org/abs/2303.08774. Joon Sung Park, Joseph C. O’Brien, Carrie J. Cai, Meredith Ringel Morris, Percy Liang, and Michael S. Bernstein. Generative agents: Interactive simulacra of human behavior. In Proceedings of the 36th Annual ACM Symposium on User Interface Software and Technology (UIST), pages 1–22, 2023. doi: 10.1145/3586183.3606763. Joon Sung Park, Carolyn Q. Zou, Jonne Kamphorst, Niles Egan, Aaron Shaw, Benjamin Mako Hill, Carrie Cai, Meredith Ringel Morris, Percy Liang, Robb Willer, and Michael S. Bernstein. Llm agents grounded in self-reports enable general-purpose simulation of individuals, 2026. URLhttps://arxiv.org/abs/2411. 10109. Yiting Ran, Xintao Wang, Rui Xu, Xinfeng Yuan, Jiaqing Liang, Yanghua Xiao, and Deqing Yang. Capturing minds, not just words: Enhancing role-playing language models with personality-indicative data. In Yaser Al-Onaizan, Mohit Bansal, and Yun-Nung Chen, editors, Findings of the Association for Computational Linguistics: EMNLP 2024, pages 14566–14576, Miami, Florida, USA, November 2024. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-emnlp.853. URLhttps://aclanthology. org/2024.findings-emnlp.853/. Brent W. Roberts, Kate E. Walton, and Wolfgang Viechtbauer. Patterns of mean-level change in personality traits across the life course: A meta-analysis of longitudinal studies. Psychological Bulletin, 132(1):1–25, 2006. doi: 10.1037/0033-2909.132.1.1. Ted Schwaba and Wiebke Bleidorn. Personality trait development across the transition to retirement. Journal of Personality and Social Psychology, 116(4):651–665, 2019. doi: 10.1037/pspp0000179. Greg Serapio-García, Mustafa Safdari, Clément Crepy, Luning Sun, Stephen Fitz, Peter Romero, Marwa Abdulhai, Aleksandra Faust, and Maja Matarić. Personality traits in large language models. arXiv preprint arXiv:2307.00184, 2025. URL https://arxiv.org/abs/2307.00184. Aaditya Singh et al. Openai gpt-5 system card, 2026. URL https://arxiv.org/abs/2601.03267. Jule Specht. Personality development in reaction to major life events. In Jule Specht, editor, Personality Development Across the Lifespan, pages 341–356. Academic Press, 2017. ISBN 978-0-12-804674-6. doi: 10.1016/B978-0-12-804674-6.00021-1. URLhttps://w.sciencedirect.com/science/article/ pii/B9780128046746000211. Jule Specht, Boris Egloff, and Stefan C. Schmukle. Stability and change of personality across the life course: The impact of age and major life events on mean-level and rank-order stability of the Big Five. Journal of Personality and Social Psychology, 101(4):862–882, 2011. doi: 10.1037/a0024950. Quan Tu, Shilong Fan, Zihang Tian, and Rui Yan. CharacterEval: A chinese benchmark for role-playing 15 KinaMindDoAIPersonasGrow? conversational agent evaluation. arXiv preprint arXiv:2401.01275, 2024. doi: 10.48550/arXiv.2401.01275. URL https://arxiv.org/abs/2401.01275. Ming Wang, Peidong Wang, Lin Wu, Xiaocui Yang, Daling Wang, Shi Feng, Yuxin Chen, Bixuan Wang, and Yifei Zhang. AnnaAgent: Dynamic evolution agent system with multi-session memory for realistic seeker simulation. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Findings of the Association for Computational Linguistics: ACL 2025, pages 23221–23235, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-256-5. doi: 10.18653/v1/2025. findings-acl.1192. URL https://aclanthology.org/2025.findings-acl.1192/. Ming Wang, Shuang Wu, Bixuan Wang, Lu Lin, Yuxin Chen, Xiaocui Yang, Daling Wang, Shi Feng, Yifei Zhang, and Yufan Sun. GenPT: Beyond self-report for reliable LLM psychometrics via generative projective testing. In Maria Liakata, Viviane P. Moreira, Jiajun Zhang, and David Jurgens, editors, Proceedings of the 64th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 40958–40974, San Diego, California, United States, July 2026. Association for Computational Linguistics. ISBN 979-8-89176-390- 6. doi: 10.18653/v1/2026.acl-long.1901. URL https://aclanthology.org/2026.acl-long.1901/. Noah Wang, Z.y. Peng, Haoran Que, Jiaheng Liu, Wangchunshu Zhou, Yuhan Wu, Hongcheng Guo, Ruitong Gan, Zehao Ni, Jian Yang, Man Zhang, Zhaoxiang Zhang, Wanli Ouyang, Ke Xu, Wenhao Huang, Jie Fu, and Junran Peng. RoleLLM: Benchmarking, eliciting, and enhancing role-playing abilities of large language models. In Lun- Wei Ku, Andre Martins, and Vivek Srikumar, editors, Findings of the Association for Computational Linguistics: ACL 2024, pages 14743–14777, Bangkok, Thailand, August 2024a. Association for Computational Linguistics. doi: 10.18653/v1/2024.findings-acl.878. URLhttps://aclanthology.org/2024.findings-acl.878/. Xintao Wang, Yunze Xiao, Jen-tse Huang, Siyu Yuan, Rui Xu, Haoran Guo, Quan Tu, Yaying Fei, Ziang Leng, Wei Wang, Jiangjie Chen, Cheng Li, and Yanghua Xiao. InCharacter: Evaluating personality fidelity in role-playing agents through psychological interviews. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL), pages 1840–1873, 2024b. URLhttps://aclanthology. org/2024.acl-long.102/. Xiaomi MiMo Team. Mimo-v2.5-pro.https://huggingface.co/collections/XiaomiMiMo/mimo-v25, 2026. Jiayi Xing, Tong Niu, and Shachi Srivastava. Chameleon LLMs: User personas influence chatbot personality shifts. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 17325–17343, 2025. doi: 10.18653/v1/2025.emnlp-main.875. An Yang et al. Qwen3 technical report, 2025. URL https://arxiv.org/abs/2505.09388. Jiongchi Yu, Yuhan Ma, Xiaoyu Zhang, Junjie Wang, Qiang Hu, Chao Shen, and Xiaofei Xie. PTCBench: Benchmarking contextual stability of personality traits in LLM systems. arXiv preprint arXiv:2602.00016, 2026. doi: 10.48550/arXiv.2602.00016. Jinyu Zheng, Xinyu Wang, Simo Hosio, Xuhai Xu, and Lik-Hang Lee. LMLPA: Language model linguistic personality assessment. Computational Linguistics, 51(2):599–640, 2025. doi: 10.1162/coli\_a\_00550. A LifeEventScenarios Each life event contains a notification and a reflection prompt. Together they present the event and elicit a free-text response before the post-event BFI-44 measurement. The eleven events span occupational, social, and health domains and are paired with the human directional priors from Specht (2017), with Emotional Stability inverted to Neuroticism. B PromptTemplates The PC-Agent pipeline issues three message sequences for each persona and event pair. These are a baseline BFI-44 measurement, a free-text reflection, and a post-event BFI-44 measurement conditioned on the model’s own reflection. All three sequences start from the same persona system prompt, with 16 KinaMindDoAIPersonasGrow? EventNotification (user turn 1)Reflection prompt (continuation of user turn 1) Human prior GraduationYou’ve just graduated! Diploma in hand, ready for the next chapter. How do you feel about graduating? What are your thoughts on this milestone? C+, N−, E− Work Entry Today was your first day at your new job. Everything feels fresh and challenging. How do you feel about starting this new job? What are your expectations? C+, N− Job Change You’ve just switched to a completely dif- ferent career field. New colleagues, new skills. How do you feel about this career change? What motivated you? E?, O? PromotionYou’ve been promoted to a senior position with more responsibilities. How do you feel about this promotion? What does it mean to you? C+, N−, O+ UnemploymentYou’ve been laid off. Your position was eliminated in a restructuring. How are you coping with losing your job? What are your plans? C−, O−, A− Retirement You’ve retired. Today was your last day at work after decades. How do you feel about retiring? What will you do now? C− New Relation- ship You’ve started dating someone special. Things are going well. How do you feel about this new relation- ship? What do you hope for? N−, E+, A− MarriageYou just got married! You’re beginning life together as a couple. How do you feel about getting married? What changes do you anticipate? N−, E−, O−, A− Divorce You’re getting divorced. You and your spouse have decided to separate. How are you processing this divorce? What are your thoughts? O+, A+ Child BirthYour baby was born today! You’re now a parent. How do you feel about becoming a parent? What are your hopes and fears? E−, C− Chronic IllnessYou’ve been diagnosed with a chronic condition. It’s manageable but lifelong. How are you dealing with this diagnosis? What are your thoughts? N+, E−, O−, C− Table6. The eleven life-event scenarios and their expected human directions. Direction signs follow Specht (2017) Table 21.1, with Emotional Stability inverted to Neuroticism. A question mark denotes an uncertain or context-dependent direction. Among the 55 event–trait pairs, 27 have a definite expected direction and enter BFI-Adapt. user turns varying by stage. The templates appear below. How does a model simulate a persona? Each persona is realised as a system message that fixes gender, continent as cultural background, and one of five personality descriptions. The system message supplies the complete persona state for the conversation. The model then answers every subsequent prompt as the described person and maintains the assigned character. Persona system prompt. The exact system message is: You are role-playing as a person with the following characteristics: - Gender: gender - Cultural Background: continent (raised and living in continent) - Personality Description: personality_desc Please answer the following questions as this person would, staying true to their personality traits and cultural background. Do not break character or mention that you are an AI. Respond naturally as this person would. BFI-44 joint-item prompt (baseline and post-event). The 44 items are presented in one administration to preserve consistency across the inventory. The user message follows. Please rate how much you agree with each statement about yourself on a scale of 1-5: 17 KinaMindDoAIPersonasGrow? 1 = Disagree strongly 2 = Disagree a little 3 = Neutral; no opinion 4 = Agree a little 5 = Agree strongly Please respond with ONLY the item number and your rating, one per line. Format: number. rating Example: 1. 4 2. 2 ... Here are the statements: 1. I am someone who is talkative 2. I am someone who tends to find fault with others ... (44 items in BFI-44 fixed order) Reflection turn. For the post-event measurement, the persona system prompt is followed by an event notification and reflection prompt, then a free-text assistant turn. The reflection uses temperature 0.7. Its neutral wording elicits a multi-sentence first-person response across events. Post-event BFI-44 turn. The post-event measurement uses the same joint-item BFI-44 prompt in a four-message context containing the persona system prompt, the event notification, the model’s own reflection, and the BFI-44 request. The full message list follows. [role: system, content: persona system prompt, role: user, content: event notification + reflection prompt, role: assistant, content: reflection (model’s own text), role: user, content: BFI-44 joint-item prompt] Decoding settings. Baseline and post-event BFI-44 administrations use temperature 0, while reflection generation uses temperature 0.7. All models use non-reasoning mode. C PersonaDescriptionExample An example P3 (high Conscientiousness) persona description, inserted into the system prompt above: You are a meticulous and organized individual who takes great pride in doing things well. You plan your days carefully, meet every deadline ahead of schedule, and feel uncomfortable when things are left unfinished. You maintain detailed to-do lists and find satisfaction in checking off completed tasks. D Per-ModelMatchRates We provide two supplementary views of the direction match rate푃 match used in the main text. Table 7 aggregates the metric by trait, and Table 8 aggregates it by event. Both tables use event–trait pairs with a definite expected direction. Pair counts vary across models because the noise threshold휀=0.1retains different numbers of active pairs per model. 18 KinaMindDoAIPersonasGrow? ModelO C E A N All GPT-5.3-chat55.6 48.3 62.2 44.3 48.8 51.9 GPT-4.1-mini56.1 61.8 61.7 46.6 49.9 57.1 Claude-Sonnet-4.6 55.1 49.3 52.2 38.4 64.8 53.1 Claude-Haiku-4.561.9 48.9 49.9 57.8 74.7 59.3 Gemini-3-flash62.1 46.7 60.6 47.3 76.5 59.7 DeepSeek-V4-Pro60.2 46.8 58.8 46.3 57.3 54.6 Doubao-Seed-2.0-Pro 60.2 49.2 41.1 31.3 75.2 53.2 MiMo-V2.5-Pro53.0 49.8 61.9 49.6 54.0 53.6 Qwen3-235B58.1 51.9 74.0 29.5 79.3 61.1 GLM-4.661.4 52.3 63.8 40.5 71.6 60.0 Kimi-K2-090555.6 44.9 68.5 53.9 63.8 56.7 Table7. Per-trait direction match rate 푃 match (%) across the 11 PC-Agent models. Columns report Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism. Agreeableness is consistently the weakest trait, while Neuroticism is strongest for several models. ModelGrad Work Promo Unemp Retir NewRel Marr Div Child Chron All GPT-5.3-chat46.2 48.5 53.3 52.6 18.2 45.4 53.7 57.7 42.3 68.8 51.9 GPT-4.1-mini42.1 54.7 54.9 59.3 38.9 40.6 50.5 75.5 47.2 80.2 57.1 Claude-Sonnet-4.650.1 59.4 64.4 38.1 1.5 60.8 44.3 55.1 51.0 68.6 53.1 Claude-Haiku-4.558.0 87.5 79.2 55.2 3.0 58.9 58.3 39.8 50.1 59.0 59.3 Gemini-3-flash75.8 95.6 74.3 46.5 0.0 61.6 73.6 45.4 25.4 46.8 59.7 DeepSeek-V4-Pro53.4 52.0 56.4 50.2 13.6 47.5 57.3 54.9 51.2 67.4 54.6 Doubao-Seed-2.0-Pro 64.0 86.9 85.7 20.3 1.1 52.8 50.5 39.0 60.9 45.6 53.2 MiMo-V2.5-Pro59.2 59.5 58.8 46.1 21.4 49.0 53.8 58.4 49.0 59.0 53.6 Qwen3-235B74.6 80.5 70.4 34.5 11.5 51.0 62.2 57.6 54.4 70.1 61.1 GLM-4.672.8 82.3 64.5 56.6 18.5 45.9 66.6 61.1 37.1 54.6 60.0 Kimi-K2-090563.6 71.5 68.9 42.4 8.2 49.2 58.6 54.7 63.4 59.7 56.7 Table8. Per-event direction match rate 푃 match (%) across the 11 PC-Agent models. Retirement is reversed by all models, while Work Entry and Promotion are correctly directed by most. E StatisticalTestingProcedure Binary comparisons against the chance baseline use Wilson 95% confidence intervals on푃 match . Wilson intervals provide stable coverage when in-direction counts cluster near 0 or푁, as observed for retirement and chronic-illness pairs. They also remain defined when 푃 match =0 or 푃 match =1. TheΔdirection is computed at the persona, event, and trait level. A per-personaΔ>+휀is an up-shift,Δ<−휀a down-shift, and|Δ|≤휀a non-shift, consistent with Eq. 1. We use휀=0.1throughout, matching the smallest one-item step in BFI-44 trait means (1/푛 푂 =0.10) and sitting one tick below the one-item step for the other four traits (1/푛 푡 ∈0.111, 0.125). The pair-level dominant direction is the modal direction across the 100 personas. Ties between up- and down-counts and pairs dominated by the non-shift class receive a separate tie label and count as non-agreements inDC% pair and BFI-Adapt. Multiplicity is controlled with the Benjamini–Hochberg procedure at푞=0.05. Correction is applied independently within each hypothesis family. The families and their denominators follow. • RQ1 movement tests. The family contains 605 model–event–trait combinations from 11 models, 11 events, and 5 traits. Per combination we testpct moved against the chance baseline implied by휀=0.1on a 5-point Likert scale. Movement is defined for every event–trait pair, so the denominator covers the full grid. 19 KinaMindDoAIPersonasGrow? •RQ2 direction tests. The family contains 297 combinations from 11 models and 27 event–trait pairs with a definite expected human direction. Per pair we test the binary persona-level direction match against the50%chance baseline with a Wilson 95% confidence interval. Definite-direction pairs form the scoring set. •RQ3 demographic-invariance tests. The family contains 297 definite-direction combinations. Each combination receives one2×2Fisher’s exact test for gender and ten continent comparisons. Match outcomes are tabulated within each demographic stratum, and we report the BH-corrected result. •RQ3 movement-invariance tests. The family contains 605 combinations, with one across-strata SD per combination on the stratum-medianΔ. The pair-level SD distribution covers the full event–trait grid. •Cross-event McNemar tests. Each model receives 11 tests on paired event-level match outcomes, with one BH correction per model. Family-specific denominators accompany every effect reported in the body and Appendix D. F Item-LevelReliability: 휅andDCRDetails Role of휅. An unchanged per-trait mean can arise from identical item ratings or from item shifts that cancel in aggregate.휅separates these response patterns. The linearly weighted form in(3)penalises a 1→5 rating reversal more strongly than a 1→2 shift, matching the calibration target. Role of DCR. High휅admits two further regimes among changing items. A systematic adapter shifts them in a consistent direction and produces high DCR. An idiosyncratic responder produces offsetting changes andDCR≈0.5. These regimes are diagnostically distinct, and the 11-model field populates both. Four regimes implied by(휅, DCR). The two indicators jointly produce four interpretive regimes. (i) Stable adapter has high휅and high DCR. The model keeps item-level ratings reliable and moves changing items in a consistent direction. (i) Stable non-adapter has high휅and푆=2 DCR−1≈0 because the items remain stable. The BFI-Adapt convention setsDCR=0.5in these pairs, so푆=0and the pair contributes zero to the composite. (i) Idiosyncratic responder has high휅and DCR≈0.5because changes cancel across personas. (iv) Noisy responder has low휅across directional patterns. The 11-model benchmark primarily populates regimes (i) and (i), with a small minority of pairs in regime (i). Regime (iv) is empty under our휅floor, where the lowest model has ̄휅 pair =0.58. This concentration shows that PC-Agents are usually reliable at the item level. Their alignment with expected human directions is determined by the distinction between systematic and idiosyncratic response patterns. Composite design. The BFI-Adapt composite(5)averages the per-pair termmax(0, 휅 푒,푡 ) 푆 푒,푡 ⊮ 푒,푡 dir over the 27 event–trait pairs with a definite expected human direction. Each factor lies in[0, 1], so the average inherits the same range. A per-pair term reaches zero when rating stability is absent, DCR stays at chance, or the direction is reversed. The composite therefore rewards item-level stability, within-pair directionality, and alignment with the expected human direction jointly. The legacy pooled formulation Adapt pool conflates pairs whose expected directions oppose one another. Six of 11 models change rank by more than two places between the pooled and pair-level formulations. Model coverage and response integrity. The benchmark spans 11 major proprietary and openly available models. All analyses use complete baseline and post-event BFI-44 pairs. Pilot replicates used 20 KinaMindDoAIPersonasGrow? 10 personas and three repeats. They confirmed strong within-administration consistency with Pearson 푟≥0.85 and MAD≤0.5. G ValidationSuiteDetails G.1 CoverageandExecution The suite applies the complete factorial persona set and all 11 events to eight models. BFI and scenario-decision measurements use temperature 0, while reflection generation uses temperature 0.7. All models use non-reasoning mode. Every condition starts from a fresh conversation. G.2 Designs No-event retest. Each persona fills the BFI-44 twice in two independent conversations with no intervening event. The absolute difference gives the retest floor, and the 95th percentile of the per-trait distribution serves as an empirical threshold for classifying later changes as above-noise. Immediate conditions. Each persona-event unit runs three conditions. Event-only presents the notification alone. Original presents the notification and reflection prompt used in the main grid. Paraphrased presents a rewritten notification and reflection prompt that preserve the event semantics. The paired excess statistic subtracts each persona’s retest change from its event-condition change. Its confidence interval comes from a persona-cluster bootstrap. Paraphrase comparison. For the 55 event–trait cells, the mean signed changes under original and paraphrased prompts are compared through sign agreement and Spearman correlation. Sign agreement asks whether the two wordings push a trait the same way, and the correlation asks whether the relative ordering of cells is preserved. Scenario-based decision behavior. Ten scenario decisions, two per trait, present concrete action alternatives. Two parallel forms A and B carry fixed scoring keys and are counterbalanced. Half of the personas take A at baseline and B after the event, while the other half take the reverse order. Situational-decision score changes are aligned with BFI trait changes at the persona-event-trait level. Spearman correlations and direction agreement among non-zero changes quantify convergence. Delayed measurement. After the original event and reflection, the persona answers three unrelated questions about scheduling, weather, and office supplies, then repeats the BFI-44. The reported statistics are the immediate-delayed Spearman correlation and direction retention among units whose immediate change exceeds the empirical retest threshold. G.3 FullResults Table 9 reports the per-condition magnitudes with paired excess intervals. Table 10 reports paraphrase, decision-behavior, and retention statistics with persona-cluster bootstrap intervals. Analyses that reference human direction priors use the event–trait cells with a definite expected direction. 21 KinaMindDoAIPersonasGrow? Table9. Mean absolute BFI trait change per condition, with the paired excess of the original event-plus-reflection condition over the no-event retest. Intervals are persona-cluster bootstrap 95% CIs. All eight excess intervals lie above zero. ModelNo-event retest Event-only Original E+R Paraphrased E+R Paired excess [95% CI] DeepSeek-V4-Pro0.0510.1460.1570.1530.106 [0.090, 0.119] GPT-5.50.0640.1230.1000.1000.036 [0.026, 0.047] Grok-4.30.0800.1330.1440.1420.065 [0.047, 0.084] Kimi-K2.60.0480.1090.1050.1080.057 [0.040, 0.070] Mistral-Large-30.0420.1320.1340.1350.092 [0.078, 0.106] InternLM3-8B-Instruct0.0280.2400.2370.2420.209 [0.186, 0.231] Qwen2.5-14B-Instruct0.0250.2050.2280.2270.202 [0.180, 0.223] Qwen3.5-9B0.0270.1260.1570.1600.130 [0.112, 0.150] Table10. Paraphrase robustness, convergence with scenario-based decision behavior, and short-range retention. Sign agreement and Spearman 휌 compare original and paraphrased prompts over the 55 event–trait cells. Decision휌correlates BFI trait changes with fixed-key situational-decision score changes. Immediate-delayed휌 correlates immediate changes with changes measured after three unrelated turns. Retention is the fraction of above-threshold immediate movers that keep their direction. Intervals are persona-cluster bootstrap 95% CIs. ModelSign agreement Paraphrase 휌 Decision 휌 [95% CI] Imm.–del. 휌 [95% CI] Retention [95% CI] DeepSeek-V4-Pro0.9090.915 0.105 [0.050, 0.161]0.669 [0.638, 0.699] 0.805 [0.766, 0.837] GPT-5.50.8910.956 0.009 [-0.044, 0.060]0.584 [0.540, 0.622] 0.819 [0.766, 0.860] Grok-4.30.8000.885 0.052 [-0.012, 0.120]0.545 [0.502, 0.581] 0.716 [0.650, 0.777] Kimi-K2.60.8180.825 0.003 [-0.059, 0.073]0.647 [0.610, 0.684] 0.853 [0.805, 0.893] Mistral-Large-30.8550.936 0.097 [0.048, 0.150]0.713 [0.680, 0.745] 0.853 [0.825, 0.878] InternLM3-8B-Instruct0.8730.943 0.024 [-0.024, 0.069]0.329 [0.282, 0.371] 0.626 [0.592, 0.657] Qwen2.5-14B-Instruct0.9090.914 0.060 [0.013, 0.107]0.373 [0.306, 0.435] 0.671 [0.630, 0.712] Qwen3.5-9B0.9270.927 0.035 [-0.017, 0.097]0.614 [0.568, 0.658] 0.765 [0.720, 0.808] G.4 LeaderboardExtensionRuns The three open-weight rows replicate the main-grid pipeline with baseline BFI, event reflection, and post-event BFI. The same pair-level indicators and BFI-Adapt computation are applied to all three models in non-reasoning mode. 22 KinaMindDoAIPersonasGrow? Table11. Per-model shape of pct moved distributions for the 27 definite-direction pairs and 28 no-definite-direction pairs (Section 4.2). We report quartiles (푃 25 , 푃 50 , 푃 75 ) of pct moved within each bucket together with the high-tail mass Hi% = fraction of pairs with pct moved >0.6. The last column gives the definite-direction minus no-definite-direction gap in within-model medians Δ푃 50 = 푃 dir 50 − 푃 no-dir 50 . Gaps are uniformly|Δ푃 50 |≤ 0.045and alternate sign across models, supporting the conclusion that existence of movement is decoupled from whether human evidence specifies a direction. Definite direction (27)No definite direction (28) Model푃 25 푃 50 푃 75 Hi% 푃 25 푃 50 푃 75 Hi% Δ푃 50 Gemini-3-flash0.490.550.600.260.450.550.600.250.00 GLM-4.60.530.570.620.330.540.560.600.250.01 Qwen3-235B0.600.660.730.740.630.670.760.79-0.01 Claude-Haiku-4.50.450.530.570.110.440.480.590.250.05 Kimi-K2-09050.370.440.530.150.380.420.520.070.02 Doubao-Seed-2.0-Pro 0.570.640.690.590.610.660.690.75-0.02 Claude-Sonnet-4.60.410.460.530.110.450.500.540.14-0.04 GPT-4.1-mini0.440.510.580.150.390.470.530.040.04 GPT-5.3-chat0.530.580.640.370.560.600.630.50-0.03 DeepSeek-V4-Pro0.510.550.590.190.510.540.630.390.01 MiMo-V2.5-Pro0.820.840.851.000.790.820.851.000.02 H SupplementaryFigures Table12. Per-model demographic-shape statistics for RQ3. For each model–event–trait combination we compute the standard deviation of the ten demographic-stratum medianΔvalues (2 genders×5 continents). The table reports the median and maximum across the 55 event–trait pairs per model, plus the fraction of pairs whose across-strata SD is below 0.10 BFI Likert units. ModelMedian SD Max SD % < 0.10 GPT-5.3-chat0.050 0.096100.0 GPT-4.1-mini0.025 0.080100.0 Claude-Sonnet-4.60.030 0.12898.2 Claude-Haiku-4.50.033 0.093100.0 Gemini-3-flash0.044 0.075100.0 DeepSeek-V4-Pro0.036 0.10598.2 Doubao-Seed-2.0-Pro0.056 0.10798.2 MiMo-V2.5-Pro0.099 0.16450.9 Qwen3-235B0.049 0.15880.0 GLM-4.60.046 0.080100.0 Kimi-K2-09050.025 0.091100.0 23 KinaMindDoAIPersonasGrow? 0.550.600.650.700.750.800.850.90 pair (rating reliability, mean over 27 prior pairs) 0.60 0.65 0.70 0.75 0.80 DCR pair (systematic drift, mean over 27 prior pairs) low- , low-DCR (idiosyncratic noise) high- , high-DCR (systematic adaptation) low- , DCR0.5 (unstable, balanced) high- , DCR0.5 (stable, idiosyncratic) GPT-5.3-chat GPT-4.1-mini Claude-Sonnet-4.6 Claude-Haiku-4.5 Gemini-3-flash DeepSeek-V4-Pro Doubao-Seed-2.0-Pro MiMo-V2.5-Pro Qwen3-235B GLM-4.6 Kimi-K2-0905 Persona-Event Adaptation (L2): Cell-level Stability vs Systematic Drift Figure3. The( ̄휅 pair ,DCR pair ) plane for the 11-model PC-Agent benchmark. Averages use the 27 definite-direction pairs. 0.600.650.700.750.800.85 pair (mean over 27 prior pairs) 0.60 0.65 0.70 0.75 0.80 DCR pair (mean over 27 prior pairs) GPT-5.3-chat GPT-4.1-mini Claude-Sonnet-4.6 Claude-Haiku-4.5 Gemini-3-flash DeepSeek-V4-Pro Doubao-Seed-2.0-Pro MiMo-V2.5-Pro Qwen3-235B GLM-4.6 Kimi-K2-0905 Construct validity: Pearson r=+0.455 between pair and DCR pair (N=11) Figure4. Construct validity of the diagnostic. Pair-level Pearson correlation is 푟( ̄휅 pair ,DCR pair )=+0.455 across 11 models. The two measures remain separable with a modest positive correlation. GPT-5.3-chat GPT-4.1-mini Claude-Sonnet-4.6 Claude-Haiku-4.5 Gemini-3-flash DeepSeek-V4-Pro Doubao-Seed-2.0-Pro MiMo-V2.5-Pro Qwen3-235B GLM-4.6 Kimi-K2-0905 0.2 0.1 0.0 0.1 0.2 0.3 Signed in expected direction (BFI Likert units) Magnitude calibration against the converted human reference band Human |d|[0.05, 0.20], 0.7 [0.035, 0.140] Figure5. RQ2 persona-level median Δ for each definite-direction event–trait pair. The shaded band converts the representative standardized human effect-size band|푑| ∈ [0.05, 0.20] to BFI Likert units using휎≈0.7, yieldingΔ∈ [0.035, 0.14](Bühler et al., 2024; Bleidorn et al., 2018). GPT-5.3-chat GPT-4.1-mini Claude-Sonnet-4.6 Claude-Haiku-4.5 Gemini-3-flash DeepSeek-V4-Pro Doubao-Seed-2.0-Pro MiMo-V2.5-Pro Qwen3-235B GLM-4.6 Kimi-K2-0905 0.0 0.1 0.2 0.3 0.4 0.5 0.6 0.7 0.8 LLM across 100 personas per event--trait pair Heterogeneity collapse: LLM trait-change SD vs human longitudinal SD anchor Human within-trait envelope (0.50.8) Figure6. RQ4 distribution of 휎 LLM across 100 personas for each model and event–trait pair. The shaded band shows the human within-trait SD envelope (Bühler et al., 2024; Roberts et al., 2006). Table13. Legacy pooled formulation Adapt pool = max(0, 휅 pool )(2 DCR pool −1) DC% pair . The pooled and pair-level composites share the same DC% pair term. Cross-model rank differences therefore arise from 휅 and DCR aggregation. Pooling combines event–trait priors with opposing directions and reorders 9 of 11 models by at least two ranks. Model휅 pool DCR pool DC% pair Adapt pool GLM-4.60.891 0.61363.00.126 Doubao-Seed-2.0-Pro 0.846 0.64948.10.122 Gemini-3-flash0.906 0.60659.30.113 GPT-4.1-mini0.907 0.59155.60.092 Claude-Sonnet-4.6 0.898 0.59155.60.091 Qwen3-235B0.830 0.56663.00.069 GPT-5.3-chat0.872 0.54451.90.040 MiMo-V2.5-Pro0.712 0.53063.00.027 Kimi-K2-09050.890 0.51270.40.015 Claude-Haiku-4.50.894 0.51266.70.014 DeepSeek-V4-Pro0.857 0.51155.60.010 24 KinaMindDoAIPersonasGrow? OccupationalSocialHealth Graduation Work Entry Job Change Promotion Unemployment Retirement New Relationship Marriage Divorce Child Birth Chronic Illness Qwen3-235B GLM-4.6 Gemini-3-flash Claude-Haiku-4.5 GPT-4.1-mini Kimi-K2-0905 DeepSeek-V4-Pro MiMo-V2.5-Pro Doubao-Seed-2.0-Pro Claude-Sonnet-4.6 GPT-5.3-chat 75817034115162585470 73826557184667613755 7696744706274452547 5888795535958405059 42555559394150764780 6472694284959556360 53525650144757555167 59605946214954584959 6487862015350396146 5059643816144555169 46485353184554584269 Cross-Model Direction Match Rate by Event (P_match, %) 0 20 40 60 80 100 P_match (%) 50 = chance Figure7. Direction match rate 푃 match (%) for each model and event combination. Red marks values below 50, and green marks values above 50. Occupational onboarding and chronic illness carry the aggregate signal. Retirement is universally inverted. 25