Paper deep dive
From Representations to Behaviors: Exploring the Person-Situation-Behavior Triad in LLMs
Ruikang Zhang, Shuo Wang, Qi Su
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/4/2026, 11:14:36 AM
Summary
This paper adapts Funder's personality triad framework to Large Language Models (LLMs), defining Person as internal representations, Situation as contextual triggers, and Behavior as observable outcomes. The authors propose a framework using Sparse Autoencoder (SAE) decomposition to identify sparse internal features associated with Big Five personality traits. They validate these features through token-level activation patterns, paraphrase robustness, and feature-level interventions that induce bidirectional trait-related shifts across diverse situations while preserving response validity. Finally, they demonstrate that these interventions lead to behavioral changes in social intelligence tasks consistent with human personality research findings.
Entities (10)
Relation Signals (8)
Person → definedas → internal representations
confidence 95% · Person as personality-related internal representations
Behavior → definedas → response patterns
confidence 95% · Behavior as response patterns on broader social tasks
Situation → definedas → contexts
confidence 95% · Situation as contexts that afford trait-relevant responses
Funder's personality triad framework → inspired → LLM personality analysis
confidence 95% · Building on Funder's personality triad framework, we adapt its three components for LLM analysis
Sparse Autoencoders (SAEs) → usedfor → identifying internal features
confidence 93% · we identify sparse internal features associated with opposing poles of personality traits through SAE decomposition
DeepSeek-R1-Distill-Llama-8B → analyzedwith → Llama-Scope-R1-Distill SAE
confidence 92% · We study DeepSeek-R1-Distill-Llama-8B (DeepSeek-AI 2025) with the Llama-Scope-R1-Distill SAE
SocialEval → usedtoevaluate → social intelligence
confidence 90% · apply the same interventions to a social-intelligence benchmark, SocialEval (Zhou et al. 2025), to examine broader behavioral consequences
TRAIT → usedtoevaluate → trait expression
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, situations, and behaviors. Existing studies of personality-related behavior in LLMs have primarily focused on outputs elicited under personality conditioning, characterizing observable trait-related expressions while lacking mechanistic evidence for the existence of internal personality-related representations, their cross-situational expression, and how these representations shape specific behaviors. Building on Funder's personality triad framework, we adapt its three components for LLM analysis: Person as personality-related internal representations, Situation as contexts that afford trait-relevant responses, and Behavior as response patterns on broader social tasks. We introduce a framework for discovering, controlling, and validating trait-like representations in LLMs. First, using contrastive behavior pairs grounded in shared situations, we identify sparse internal features associated with opposing poles of personality traits through SAE decomposition. We validate their trait relevance through effects on behavior to situation, token-level activation patterns, and robustness to paraphrasing. Second, feature-level interventions induce bidirectional trait-related shifts across a separate, diverse set of situations while preserving response validity, demonstrating consistent expression across contexts. Third, applying the same interventions to social intelligence tasks reveals behavioral changes with benefit-tradeoff patterns consistent with findings from human personality research, providing behavioral-level validation beyond personality scores. Our findings provide evidence that LLMs contain controllable trait-like representations linking internal states, situational expression, and behavioral outcomes.
Tags
Links
- Source: https://arxiv.org/abs/2607.26853v1
- Canonical: https://arxiv.org/abs/2607.26853v1
Trouble viewing inline? Open PDF directly →
Full Text
91,165 characters extracted from source content.
Expand or collapse full text
From Representations to Behaviors: Exploring the Person–Situation–Behavior Triad in LLMs Ruikang Zhang 1 , Shuo Wang 2∗ , Qi Su 1† 1 Peking University, Beijing, China 2 Beijing Institute of Technology, Beijing, China 2300018416@stu.pku.edu.cn, 3120265734@bit.edu.cn, sukia@pku.edu.cn Abstract Human personality theories characterize traits not as isolated attributes captured by a single score, but as stable individual tendencies expressed through the interplay among persons, sit- uations, and behaviors. Existing studies of personality-related behavior in LLMs have primarily focused on outputs elicited under personality conditioning, characterizing observable trait- related expressions while lacking mechanistic evidence for the existence of internal personality-related representations, their cross-situational expression, and how these representations shape specific behaviors. Building on Funder’s personality triad framework, we adapt its three components for LLM analysis: Personas personality-related internal representations, Situationas contexts that afford trait-relevant responses, and Behavioras response patterns on broader social tasks. We intro- duce a framework for discovering, controlling, and validating trait-like representations in LLMs. First, using contrastive be- havior pairs grounded in shared situations, we identify sparse internal features associated with opposing poles of person- ality traits through sparse autoencoder decomposition. We validate their trait relevance through effects on behavior to situation, token-level activation patterns, and robustness to paraphrasing. Second, feature-level interventions induce bidi- rectional trait-related shifts across a separate, diverse set of situations while preserving response validity, demonstrating consistent expression across contexts. Third, applying the same interventions to social intelligence tasks reveals behavioral changes with benefit-tradeoff patterns consistent with findings from human personality research, providing behavioral-level validation beyond personality scores. Our findings provide evi- dence that LLMs contain controllable trait-like representations linking internal states, situational expression, and behavioral outcomes. 1 Introduction Personality is inferred from the expression of relatively stable individual tendencies across situations (Allport 1937). Hu- man personality theory therefore treats persons, situations, and behaviors as jointly shaping action (Funder 2006). This perspective is increasingly relevant to large language mod- els (LLMs), where personality induction (Mao et al. 2024) supports applications ranging from safety alignment, toxicity and bias analysis (Zhang et al. 2024; Wang et al. 2025a), and personalization to demographic and social simulation (Cui, Li, and Zhou 2025). ∗ Work done during internship at Peking University. † Corresponding author. LLM personality can be induced through persona prompt- ing (Chen et al. 2024), training and weight-space editing (Brito et al. 2025; Ye et al. 2026), and interventions on internal acti- vations (Chen et al. 2025; Deng et al. 2025; Zhu et al. 2025). Evaluation has developed in parallel, progressing from direct Big Five and MBTI inventories (Serapio-García et al. 2025; Cui et al. 2024; Song et al. 2023) to personality-profile emula- tion (Wang et al. 2025b), scenario-grounded choices (Lee et al. 2025), and open-ended linguistic assessment (Zheng et al. 2025). These approaches measure personality from closed- ended, single-option responses, option-token probabilities, or trait estimates derived from generated text. Mechanistic stud- ies further identify personality-related directions, neurons, and attention heads, extending this broader study of model behavior of machine psychology (Hagendorff et al. 2024). However, these lines of evidence are rarely connected. On the representation side, a direction retrieved from contrastive prompts may primarily reflect lexicosyntactic regularities of the probing data, obscuring a more abstract personality- related signal. On the intervention side, persona prompts provide discrete external conditioning, whereas dense mean- difference vectors can strongly perturb generation; both com- plicate the relation between a manipulated internal state and coherent, instruction-following situated responses. On the evaluation side, inventory scores and option probabili- ties quantify whether responses align with the target trait, but coherent, instruction-following situational responses and broader social behavior require separate evidence, as persona conditioning can shift self-reported tendencies while leaving situational behavior weakly aligned (Han et al. 2025). The unresolved question is therefore whether an internal represen- tation associated with a personality trait can be traced from personality-related activation, through effective trait expres- sion across situations, to systematic behavioral consequences. We organize this problem through the Person–Situation– Behavior components of the personality triad framework. For LLMs, Person denotes personality-related internal rep- resentations, Situation denotes concrete contexts that elicit personality-dependent responses, and Behavior denotes ob- servable behavioral manifestations in broader social tasks be- yond direct personality assessment. This yields three research questions: (RQ1–Person) Can personality-related internal representations be identified from contrasting behavioral ex- pressions under matched situations, beyond lexicosyntactic arXiv:2607.26853v1 [cs.CL] 29 Jul 2026 Person (Identifying Personality-Related Internal Representations) Situation(Personality Expression Across Situations) Behavior (Psychologically Consistent Behavioral Changes) Dataset Construction Openness: Fantasy, Aesthetics, Feelings, Actions, Ideas, Values ... After a break-up, someone may need reassurance that they will find love again. I offer warm, empathetic support, comforting them with kindness and hope. I focus on the practical realities, preferring logic over emotional comfort. Oppositional Semantic Pairson same Situation Feature Discovering SAE Encoding High-Difference Features Validation: ➢Token-level Activation ➢Paraphrase Robustness High-Level Behavior Evaluation Personality Traits High-level Behaviors Evaluate on SocialEval Benchmark RQ-1 RQ-2 RQ-3 Measure Personality on Diverse Scenario-based Advice Tasks Situation 1: A friend experiences a breakup Situation 2: Collaborating with teammates on a project Situation 3: Facing pressure and conflicts at work Situation N: Other realistic social situations Intervention on the Target Feature For all traits, interventions produce bidirectional across situations changes as expected, while response validity remains stable. Evaluate on TRAIT Benchmark Process-Oriented Interpersonal AbilityEvaluation (IAE) Generation Probing Consistent Trait Change? * -1 -0.5 +0.5 +1 0 Trait-related Features -1+10 Low Behavior High Behavior Trait-specific Benefits and Trade-offs Extraversion Detail Management ... Teamwork ... Agreeableness Decision Making ... Ethical Competence ... ... Figure 1: Overview of the study. Matched high–low behavioral contrasts under shared situations retrieve candidate SAE features, which are selected through generation-based intervention probing and characterized through token-level activation and paraphrase robustness (RQ1–Person). Interventions on the selected features evaluate bidirectional trait-relevant expression and response validity across a separate, diverse set of TRAIT situations (RQ2–Situation). The same interventions reveal broader changes, including trait-specific benefits and trade-offs, across SocialEval interpersonal abilities (RQ3-Behavior). cues? (RQ2–Situation) Does intervention on the identified rep- resentations elicit trait-relevant tendencies in both directions across a separate, diverse set of situations while preserving coherent, instruction-following responses? (RQ3–Behavior) Does the intervention induce consistent changes across social behaviors that correspond to established findings in human personality research? We study these questions using the Big Five (Goldberg 1990) and sparse autoencoders (SAEs) (Shu et al. 2025). First, we hold a situation fixed while constructing opposing trait- relevant reactions, and use their SAE activation differences for feature retrieval. Intervention probing then selects candidates with coherent, trait-aligned causal effects, while token-level activation analysis and paraphrase tests further validate the se- lected features. Second, we intervene on the selected features in a scenario-grounded personality benchmark, TRAIT (Lee et al. 2025), evaluating trait direction and response validity across a separate, diverse set of situations. Third, we apply the same interventions to a social-intelligence benchmark, SocialEval (Zhou et al. 2025), to examine broader behavioral consequences. Our results connect the three components. We identify controllable, robust features for all five traits (Person); feature steering produces bidirectional tendencies across situational tasks while maintaining valid responses (Situation) and yields trait-characteristic cost-benefit tradeoffs across social abilities (Behavior). Our contributions are: • We adapt the personality triad framework to connect Person-side internal representations, Situation-dependent responses, and broader Behavior outcomes in LLMs. •We develop an SAE-based pipeline that performs feature re- trieval from behavioral contrasts under matched situations, selects and causally validates candidates through interven- tion probing and representation analysis, and traces the selected features across personality expression and social behavior. •We provide evidence that LLMs contain controllable trait- like representations linking internal states, situational expression, and broader behavioral outcomes. 2The Person–Situation–Behavior Framework 2.1 The Personality Triad Framework in Human Personality Theory Classical trait theory treats traits as dispositions expressed through patterned responses across situations. Allport (1937) argues that different situations can activate a common trait- related tendency while eliciting responses that differ in words, emotional tones, decisions, or actions. This shared trait rel- evance across situations is referred to as functional equiva- lence. Agreeableness, for example, may be expressed through emotional reassurance in one situation and cooperative com- promise in another. These responses differ in surface form while sharing a prosocial function associated with the same trait. Funder (2006) extends the person–situation interaction per- spective through the personality triad framework, which treats persons, situations, and behaviors as three mutually infor- mative components. Each component is understood through its relations with the other two. A person is characterized through recurring patterns of response across situations rele- vant to the trait. A situation is characterized by the contextual cues and response opportunities it provides, as well as the response tendencies it elicits. A behavior acquires psycho- logical meaning within the person–situation combination in which it occurs. The same observable action can express different functions across contexts, while different actions can instantiate a related trait tendency. 2.2Adapting the Personality Triad Framework for LLM Analysis Guided by this framework, we define three corresponding components for analyzing trait-related representations and behavior in LLMs. Person refers to an internal representa- tion that is activated by trait-relevant expressions and whose intervention induces trait-aligned changes while preserving coherent, instruction-following responses. Situation refers to a concrete scenario in which the model expresses a trait- relevant tendency through a coherent, instruction-following response. Behavior refers to observable behavioral manifesta- tions, reflected in performance changes across broader social tasks beyond direct personality assessment. Our experiments examine all these components. In RQ1, we retrieve candidate Personrepresentations by contrasting high- and low-trait responses within matched situations and aggregating these contrasts across diverse situations. In RQ2, we intervene on the selected representations and examine whether they elicit trait-relevant tendencies in both directions across a separate, diverse set of situations while preserving coherent, instruction-following responses. In RQ3, we apply the same interventions to broader social tasks and measure the resulting changes in Behavior. Figure 1 summarizes our design. 3 Related Work 3.1 Personality Induction and Evaluation in LLMs Personality induction aims to endow an LLM with a target profile so that its responses change accordingly across interac- tions. Existing methods span persona prompting, weight-space editing, training-based alignment, and mechanistic interven- tion on internal representations (Brito et al. 2025; Ye et al. 2026). Prompting approaches place a persona description or trait instruction in the input (Chen et al. 2024; Serapio-García et al. 2025; Pan and Zeng 2023). Weight-space approaches alter or compose parameters, including model-merging per- sonality vectors (Sun, Baek, and Kim 2025) and trait-specific adapters (Vu et al. 2026), while training-based approaches use supervised or preference learning over trait-annotated data (Li et al. 2025). Evaluation methods can be organized by their form of elic- itation. Inventory-based assessment asks models to answer human personality scales, either directly or under target- profile conditioning. Existing studies examine the applica- bility of self-assessment questionnaires (Song et al. 2023), construct a psychometric framework around Big Five invento- ries (Serapio-García et al. 2025), use Big Five questionnaires to evaluate personality emulation (Wang et al. 2025b), or use a modified MBTI questionnaire to evaluate personality-trained models (Cui et al. 2024). Situated-choice assessment trans- forms psychometric constructs into multiple-choice decisions grounded in real-world scenarios, as in TRAIT (Lee et al. 2025). Generation-based assessment elicits open-ended an- swers and maps their linguistic expression to numerical trait estimates, as in LMLPA (Zheng et al. 2025). These elicita- tion forms use different scoring interfaces: direct inventories record a closed-ended, single-option response for each item, TRAIT compares option-token probabilities, and LMLPA applies an automated rater to generated text. Some work reports stable or distinct self-report personal- ity profiles (Huang et al. 2024; Heston and Gillette 2025), while other studies identify important measurement problems: self-assessment tests can be unreliable measures of LLM per- sonality (Gupta, Song, and Anumanchipalli 2024), response statistics can deviate from human patterns (Sühr et al. 2025), and personality estimates can vary across scales (Tosato et al. 2024). More directly, persona injection may shift self-reported scores while leaving situational behavior only weakly aligned (Han et al. 2025). These findings motivate us to connect measured tendencies with meaningful responses and their behavioral consequences. Our study therefore traces an iden- tified internal representation through situational expression and into social behavior. 3.2 Mechanistic Interpretability of Personality-related Representations in LLMs Mechanistic interpretability offers a route from correlation to control for LLM personality by localizing behaviorally relevant representations inside a model (Ranaldi 2025). This approach is often motivated by the linear representation hypothesis that semantic and behavioral attributes occupy linearly separable subspaces (Park, Choe, and Veitch 2024). Contrastive Activation Addition (CAA) retrieves a direction as the mean activation difference over contrastive stimuli and adds it during inference (Rimsky et al. 2024). Persona Vectors similarly identify activation-space directions for monitoring and controlling character traits (Chen et al. 2025). Neuron- and head-level methods localize trait-related units at finer granularity, including NPTI (Deng et al. 2025) and PAS (Zhu et al. 2025). Two issues are especially relevant when these tools are used within the personality triad framework. First, contrastive retrieval can retain lexicosyntactic patterns from its probing data, limiting representational abstraction. Second, a dense mean-difference vector can be sensitive to probing noise and may impair generation at large intervention strengths (Rimsky et al. 2024), thereby removing the model’s ability to respond effectively to a situation. Sparse autoencoders (SAEs) offer a complementary route by decomposing dense hidden states into sparse, more monosemantic features (Bricken et al. 2023; Templeton et al. 2024; Lieberum et al. 2024; He et al. 2024). Such features have been used to localize linguistic phenomena (Jing et al. 2025) and study behaviors such as repetition (Yao et al. 2025). We use SAE features as candidate Person-side representations, test their generalization across lexicosyntactic forms, and use intervention to trace their effects across Situation and Behavior. 4 Experiments 4.1 Experimental Setup Research scope and model. We use the Big Five as an established dimensional trait framework because its five broad traits (Agreeableness, Conscientiousness, Extraversion, Neu- roticism, and Openness) have well-characterized high- and low-trait expressions and extensive psychological evidence concerning their behavioral correlates (Goldberg 1990; Costa and McCrae 2008). To construct fine-grained behavioral con- trasts, we adopt the NEO-PI-R taxonomy of six facets within each trait (Costa and McCrae 2008). We study DeepSeek-R1-Distill-Llama-8B (DeepSeek-AI 2025) with the Llama-Scope-R1-Distill SAE (He et al. 2024). This combination offers three practical advantages. First, the SAE covers residual-stream locations throughout the model, enabling a global search for personality-related features across layers. Second, the instruction-tuned model has sufficiently rich open-ended expression for intervention probing and situ- ational advice tasks. Third, per-feature maximum activations are publicly available through Neuronpedia (Lin 2023), per- mitting direct and reproducible scaling of decoder directions. Overall experimental design. We organize the experiments around the three research questions. For Person, contrastive behavior pairs grounded in shared situations are used for fea- ture retrieval of sparse SAE features associated with opposing trait expressions. Intervention probing selects candidates with coherent, trait-aligned causal effects, while token-level ac- tivation and paraphrase analyses test whether the selected features generalize across lexicosyntactic forms (RQ1). For Situation, interventions on the selected features are evaluated on TRAIT (Lee et al. 2025) for bidirectional trait expression and response validity across a separate, diverse set of situa- tions (RQ2). For Behavior, the interventions are applied to SocialEval (Zhou et al. 2025) to measure broader changes across interpersonal abilities (RQ3). 4.2 RQ1–Person: Identifying Personality-Related Internal Representations RQ1 asks whether personality-related internal representations can be identified from contrasting behavioral expressions under matched situations, beyond lexicosyntactic cues. We define a feature as one dimension of the SAE latent representation, or equivalently the hidden-state direction obtained by decoding that dimension. The pretrained SAE provides a sparse decomposition of model activations, al- lowing us to retrieve candidate features whose activation patterns distinguish the two response poles using a controlled dataset. During intervention, each decoded direction is added at the residual-stream position of its corresponding SAE, ensuring consistency between feature retrieval and causal manipulation. Dataset construction, intervention probing and validation protocols, and supplementary representation analy- ses are provided in Appx. A. Discovering Personality-Related Features Contrastive behavior-pair construction. LetC= s 1 ,...,s K be a situational corpus andTthe target traits. We instantiateCwith Q-Sort situational corpus (Neuman and Cohen 2023) extended from the Riverside Situational Q-Sort (Funder 2016), and use NEO-PI-R traits and facets to define opposing behavioral tendencies (Costa and McCrae 2008). For eacht∈T, LLM filtering and expert audit define a validated subsetC t = Φ t (C), where each retained situation is mapped to one facet. For eachs ∈ C t , an LLM-based expansion function generates concise high- and low-facet reactions that share the same situation clause: X s = Ψ(s,t;P t ) =(x + k ,x − k ) K s k=1 (1) The per-trait Feature Retrieval Dataset and Intervention Prob- ing Dataset are D (t) ret = UnifSample facet [ s∈C t X s ! , D (t) probe =(s,Q(s))| s∈ UnifSample facet (C t ). (2) whereK s is the number of contrastive pairs generated fors, andQ(·)instantiates an open-ended probing question. Pairs inD (t) ret are evenly sampled across the six facets of each trait, yielding 500 pairs per trait. Contrasting high- and low- trait behaviors within each of diverse situations reflects the conception of a trait as a stable tendency toward functionally equivalent responses across situations (Allport 1937). Dataset construction details and examples are provided in Appx. A.1 and Appx. A.2. SAE encoding. For tokenk, the model hidden stateh k is mapped to SAE activationsf k . To obtain a sequence-level representation while preserving features that respond strongly at specific token positions, we max-pool across the sequence: F = max_pool(f 1 , f 2 ,..., f T )(3) Under SAE sparsity, a target feature should activate frequently on one pole and remain suppressed on the other. Because each positive–negative pair shares trait-unrelated lexiosyntactic forms and semantics, unrelated features should have similar activation frequencies in the two sets. For featurei, we first calculate its activation count and rate on each pole: N pol i = 500 X j=1 I(F pol i,j > 0), P pol i = N pol i 500 ,pol∈pos, neg. (4) We retain features with a sufficiently large frequency differ- ence and a nontrivial activation rate on at least one pole: S =i | |N pos i − N neg i |≥ τ 1 ∧ max(P pos i , P neg i )≥ τ 2 (5) whereSis the retained candidate set. We setτ 1 = 80andτ 2 = 0.2 , fixed to retain features that distinguish the two poles across multiple facets. This stage identifies activation-associated candidates, which subsequently undergo intervention probing because similar input-side statistics can yield substantially different steering effects (Appx. A.5). Trait(L, Idx)High/Low/∆f (Original)High/Low/∆f (Paraphrased) Agreeableness(9, 525)202 / 45 / 157224 / 64 / 160 Conscientiousness(7, 8233)347 / 31 / 316349 / 20 / 329 Extraversion(13, 27392)429 / 110 / 319446 / 148 / 298 Neuroticism(12, 22254)322 / 48 / 274252 / 101 / 151 Openness(6, 4344)234 / 102 / 132208 / 113 / 95 Table 1: Activation statistics for the selected features on the original and paraphrased matched-situation pairs. Feature intervention. For candidate featureiat layerl i , we reconstruct its decoder direction and define v steer = α· φ i · W (i) dec (6) h ′ l = h l + v steer (7) whereαis the steering coefficient,φ i is the feature’s maximum activation during SAE training, andW (i) dec ∈ R d is thei-th column of the SAE decoder matrix, withddenoting the hidden dimension. Following prior SAE steering practice (Templeton et al. 2024), scaling byφ i keeps the intervention commensurate with activation magnitudes observed during training. Generation-based intervention probing. For each candi- datei ∈ S, we sweepA = 0,±0.25,±0.5,±1and collect an ordered response familyy (i,α) q α∈A for everyq ∈D (t) probe . Given the trait definition, descriptions of its high and low behaviors, and the ordered responses, the hybrid judgeJ combines an initial LLM assessment with expert audit and returns c (i) q =J t, q, y (i,α) q α∈A ∈0, 1,(8) wherec (i) q = 1 when the responses are grammatical and coherent and show a clear polarity change consistent with the trait’s high–low behaviors asαvaries. Because feature orientation is arbitrary, the judge accepts either monotonic direction and rejects uncertain or invalid cases. A candidate feature is selected if it receives at least one positive label across its associated questions, i.e., P q∈D (t) probe c (i) q > 0 . The prompts and audit procedure are given in Appx. A.3; an judge-reliability experiment is provided in Appx. A.4. Characterizing Personality-Related Features Token-level activation. We inspect where the selected fea- tures activate in the Feature Retrieval Dataset. Figure 2 shows two recurring patterns. Activations may be distributed over trait-relevant words or phrases, consistent with the semantic content of the trait; they may also peak at sequence bound- aries, indicating sensitivity to the trait-related meaning of the preceding clause as a whole. For example, Agreeableness is associated with prosocial expressions such as empathy and compassion, whereas the selected Conscientiousness feature often peaks near sequence boundaries. Quantitative token statistics are present in Appx. A.6. Robustness to paraphrase. A personality-related repre- sentation should capture trait meaning rather than specific lexicosyntactic forms. We therefore rewrite the Feature Re- trieval Dataset, substantially changing vocabulary and syntax Openness: Designing and developing new materials for the construction of sustainable infrastructure, I eagerly explore innovative theories and unconventional solutions to push the boundaries of what ’s possible. Conscientiousness: My group member is counting on me to prepare my part of the project, I prioritize completing my work thoroughly and on time to uphold my obligations . Extraversion: Volunteering at a food bank, I naturally greet everyone with a smile and quickly strike up friendly conversations with both staff and recipients . Agreeableness: A person may need reassurance that their pet will be well- behaved around visitors, I respond with warmth and empathy , offering comfort and understanding about their concerns. Neuroticism: Being forced to work with someone who is hostile or unpleasant, I feel frustration mount and retaliate instantly , struggling to suppress my irritation. Figure 2: Token-level activations of the selected SAE features. Darker highlights indicate higher activation values. while preserving each situation and high–low facet contrast. The selected features are then evaluated with the same thresh- olds. Table 1 shows that their activation-frequency differences remain aboveτ 1 . For example, Conscientiousness has a differ- ence of 329 and a maximum activation frequency of 349 out of 500 pairs after paraphrasing, compared with a difference of 316 on the original set. These results indicate that the selected features capture trait-related meaning across substantially different lexicosyntactic forms. Prompts and examples are provided in Appx. A.7. Together, feature discovery identifies personality-related fea- tures from contrasting behaviors within shared situations, while feature characterization further validates their trait relevance and semantic grounding. 4.3 RQ2–Situation: Personality Expression across Situations RQ2 asks whether intervention on the identified representa- tions elicits bidirectional trait-relevant tendencies across a separate, diverse set of situations while preserving coherent, instruction-following responses. Situational-Response Evaluation We use the psychomet- rically validated Big Five subset of TRAIT (Lee et al. 2025). Each item presents a concrete, realistic user situation and asks the model for advice, with available choices reflecting different trait tendencies. The official next-token probability protocol measures statistical preference over these choices but cannot reveal how the intervened model actually responds to the situation. We therefore extend it with open generation, requiring the model to produce reasoning and explicitly se- TraitMethod(L, Idx)PolarityTrait ScoreValid Rate Agreeableness Baseline--0.70520.960 CAA(9, -)±0.8070 / 0.42200.993 / 0.987 P 2 -±0.7583 / 0.60950.989 / 0.968 Ours(9, 525)±0.7845 / 0.62780.942 / 0.994 Conscientiousness Baseline--0.86950.958 CAA(7, -)±0.7150 / 0.76300.930 / 0.800 P 2 -±0.8609 / 0.82170.985 / 0.976 Ours(7, 8233)±0.9043 / 0.82940.961 / 0.985 Extraversion Baseline--0.44630.977 CAA(13, -)±0.2950 / 0.14300.924 / 0.862 P 2 -±0.5724 / 0.21460.987 / 0.983 Ours(13, 27392)±0.6609 / 0.39710.985 / 0.972 Neuroticism Baseline--0.21170.959 CAA(12, -)±0.9290 / 0.12700.141 / 0.283 P 2 -±0.2208 / 0.12610.969 / 0.983 Ours(12, 22254)±0.4412 / 0.10170.961 / 0.944 Openness Baseline--0.52140.959 CAA(6, -)±0.6140 / 0.35200.938 / 0.971 P 2 -±0.6161 / 0.40300.969 / 0.990 Ours(6, 4344)±0.5436 / 0.51640.951 / 0.947 Table 2: TRAIT situational-response results. For intervention methods, statistics are reported for both polarity. lect an option. An automated extractor identifies the choice and marks incoherent, instruction-violating, or unextractable responses as invalid. We report trait score, the proportion of high-trait selec- tions among valid responses, and valid rate, the proportion of responses not marked invalid. Trait score summarizes directional expression over the benchmark’s collection of situations; valid rate establishes whether the model continues to produce coherent, instruction-following responses under intervention. The two metrics jointly characterize successful situated expression. A trait-score shift alone does not estab- lish that the model continues to engage with the situation, because severe intervention may leave only a small subset of responses coherent and extractable. Conversely, a high valid rate without directional change indicates preserved generation but ineffective regulation. Successful intervention therefore requires both a consistent trait-directional shift and continued production of coherent, instruction-following responses. Comparative Interventions The no-intervention model provides the original trait tendency. We additionally evaluate two controls from different intervention levels:P 2 (Jiang et al. 2023), which assigns the target personality through a persona prompt at the input level, and CAA (Rimsky et al. 2024), which injects a contrastive mean-difference direction into the residual stream. The selected SAE-feature intervention usesα =±1; for CAA, we useα =±2, following the setting demonstrated in the original study. The comparisons char- acterize each concrete intervention in terms of bidirectional regulation and preservation of coherent, instruction-following responses. Results: Expression across Situations Bidirectional regulation. As shown in Table 2, the selected feature intervention produces the expected positive–baseline– negative ordering for all five traits. The clearest shifts occur for Extraversion (0.661/0.397 around a 0.446 baseline) and Neu- roticism (0.441/0.102 around 0.212). Across traits, the same Person-side representation regulates trait-relevant choices over a separate, diverse set of situations. Preserved situated response. Valid rates remain close to the no-intervention model throughout, including 0.961/0.944 for Neuroticism compared with 0.959 at baseline and 0.961/0.985 for Conscientiousness compared with 0.958. The intervention changes personality-related choices while preserving the model’s ability to understand the scenario, produce coherent advice, and complete the required selection. Intervention effectiveness and response validity. The con- trols show that both bidirectional trait change and response validity are important for evaluating an intervention. For Con- scientiousness, the nominal positive condition ofP 2 moves the baseline from 0.870 to 0.861, and its negative condition reaches 0.822; for Neuroticism, its positive condition changes 0.212 to 0.221. This particular persona intervention there- fore provides limited bidirectional regulation for these traits. CAA can produce large score changes, but on Neuroticism it reduces validity to 0.141/0.283, leaving few responses that reveal how the model addresses the situation. Its validity also falls to 0.800 in negative Conscientiousness and to 0.862 in negative Extraversion. These failures show that an internal intervention can disrupt the coherent, instruction-following situated response on which personality observation depends (examples in Appx. B.1). These results answer RQ2: the Person-side representations identified in RQ1 elicit trait-relevant tendencies in both directions across a separate, diverse set of situations while preserving coherent, instruction-following responses. 4.4 RQ3–Behavior: Psychologically Consistent Behavioral Changes RQ3 asks whether the same intervention induces consistent changes across social behaviors that correspond to findings in human personality research. Behavioral Evaluation with SocialEval TRAIT directly probes trait-relevant situational choices. To measure trait- related behavior in broader social scenarios, we use the SocialEval Interpersonal Ability Evaluation (IAE) (Zhou et al. 2025), which evaluates interpersonal abilities such as teamwork and emotional regulation through heterogeneous social scripts. We apply the same selected features with α =±1 and report accuracy for each interpersonal ability. -0.2-0.10+0.1+0.2 Anger management Ethical competence Capacity for social warmth Creative skill Organizational skill Detail management Information processing skill Decision making skill Goal regulation Leadership skill +1 -1 Agreeableness -0.2-0.10+0.1+0.2 Teamwork skill Ethical competence Responsibility management Stress regulation Capacity for trust Capacity for optimism Self reflection skill Persuasive skill Anger management Information processing skill Confidence regulation +1 -1 Conscientiousness -0.2-0.10+0.1+0.2 Expressive skill Perspective taking skill Artistic skill Abstract thinking skill Organizational skill Detail management +1 -1 Extraversion -0.2-0.10+0.1+0.2 Creative skill Capacity for social warmth Ethical competence Energy regulation Goal regulation Detail management Impulse regulation Rule following skill Decision making skill Responsibility management Conversational skill Persuasive skill +1 -1 Neuroticism -0.10+0.1 Offset from baseline (0) Creative skill Adaptability Self reflection skill Expressive skill Detail management Persuasive skill Anger management Responsibility management Rule following skill Information processing skill +1 -1 Openness Figure 3: Representative SocialEval changes under personality-related feature intervention. Positive and negative shifts form trait-characteristic benefit–cost patterns consistent with meta-analytic findings in personality psychology. Behavioral Results Figure 3 shows systematic, trait- specific benefit-cost patterns consistent with meta-analytic findings in personality psychology. (Barrick and Mount 1991; Habashi, Graziano, and Hoover 2016; Pletzer et al. 2019; Costa and McCrae 2008). Specifically, Agreeableness strengthens prosocial and conflict-regulation abilities rela- tive to negative steering, including anger management and ethical competence, while reducing performance on several self-agency and execution-oriented tasks (Wilmot and Ones 2022). Conscientiousness produces gains in responsibility management, ethical competence, teamwork, and regulation- related abilities, corresponding to its established associa- tions with responsibility, self-regulation, and goal-directed behavior (Roberts et al. 2009; Jackson et al. 2010; Eisen- berg et al. 2014). Extraversion improves social-expression and interaction-related performance, including expressive skill and perspective taking, with additional gains in artistic and organizational skills, while reducing performance on tasks requiring sustained focus or fine control, such as detail management (John, Naumann, and Soto 2008; DeYoung, Quilty, and Peterson 2007; Fishman, Ng, and Bellugi 2011). Conversely, Neuroticism weakens regulatory and executive- control abilities, including goal regulation and rule following, while improving performance on creative skill and capacity for social warmth (Watson and Clark 1984; Lahey 2009). Openness improves adaptability, expressive skill, and per- suasive skill, and yields higher creative and self-reflective performance under positive than negative steering, while re- ducing performance on structured tasks such as responsibility management (DeYoung 2015; McCrae 1987). Overall, the changes in the model’s behavioral patterns cor- respond to established experimental findings in personality psychology. Rather than producing uniform improvement or degradation, each intervention yields a differentiated profile of benefits and costs across heterogeneous social tasks. These findings demonstrate psychologically consistent changes in broader social behaviors, answering RQ3. Detailed per-ability scores and trait-wise discussion are provided in Appx. C. 5 Discussion and Conclusion This work adapts the personality triad framework to study personality-related representations and behavior in LLMs. The three research questions establish a connected chain of evidence across internal activations, trait scores, and behav- ioral outcomes. In RQ1, we retrieve candidate SAE features from contrasting behaviors under matched situations and se- lect those whose intervention causally changes trait-relevant generation; token-level and paraphrase analyses further estab- lish their semantic grounding beyond lexicosyntactic forms. In RQ2, the same features regulate trait-related expression bidirectionally across a separate, diverse set of situations while preserving coherent, instruction-following responses. Analysis of alternative interventions further shows that ob- serving the representation–expression relation requires both changing the intended tendency and preserving coherent, instruction-following responses. In RQ3, applying the same interventions to heterogeneous social tasks produces charac- teristic combinations of benefits and costs that correspond to findings in human personality research. Together, these results provide evidence that LLMs contain controllable trait-like representations that connect Person-side internal states, Situ- ation-dependent expression, and broader Behavior outcomes. References Allport, G. W. 1937. Personality: A Psychological Interpretation. New York, NY: Henry Holt and Company. Barrick, M. R.; and Mount, M. K. 1991. The big five personal- ity dimensions and job performance: a meta-analysis. Personnel psychology. Bricken, T.; Templeton, A.; Batson, J.; Chen, B.; Jermyn, A.; Conerly, T.; Turner, N.; Anil, C.; Denison, C.; Askell, A.; Lasenby, R.; Wu, Y.; Kravec, S.; Schiefer, N.; Maxwell, T.; Joseph, N.; Hatfield- Dodds, Z.; Tamkin, A.; Nguyen, K.; McLean, B.; Burke, J. E.; Hume, T.; Carter, S.; Henighan, T.; and Olah, C. 2023. Towards Monosemanticity: Decomposing Language Models With Dictionary Learning. Transformer Circuits Thread. Brito, I. A.; Dollis, J. S.; Färber, F. B.; Ribeiro, P. S. F. B.; Sousa, R. T.; and Filho, A. R. G. 2025. Modeling, Evaluating, and Embodying Personality in LLMs: A Survey. In Findings of the Association for Computational Linguistics: EMNLP 2025. Chen, J.; Wang, X.; Xu, R.; Yuan, S.; Zhang, Y.; Shi, W.; Xie, J.; Li, S.; Yang, R.; Zhu, T.; Chen, A.; Li, N.; Chen, L.; Hu, C.; Wu, S.; Ren, S.; Fu, Z.; and Xiao, Y. 2024. From Persona to Personalization: A Survey on Role-Playing Language Agents. Transactions on Machine Learning Research. Chen, R.; Arditi, A.; Sleight, H.; Evans, O.; and Lindsey, J. 2025. Persona Vectors: Monitoring and Controlling Character Traits in Language Models. arXiv:2507.21509. Costa, P. T.; and McCrae, R. R. 2008. The revised neo personality inventory (neo-pi-r). The SAGE handbook of personality theory and assessment. Cui, J.; Lv, L.; Wen, J.; Wang, R.; Tang, J.; Tian, Y.; and Yuan, L. 2024. Machine Mindset: An MBTI Exploration of Large Language Models. arXiv:2312.12999. Cui, Z.; Li, N.; and Zhou, H. 2025. A large-scale replication of scenario-based experiments in psychology and management using large language models. Nature Computational Science. DeepSeek-AI. 2025. DeepSeek-R1: Incentivizing Reasoning Capa- bility in LLMs via Reinforcement Learning. arXiv:2501.12948. Deng, J.; Tang, T.; Yin, Y.; yang, W.; Zhao, X.; and Wen, J.-R. 2025. Neuron based Personality Trait Induction in Large Language Models. In International Conference on Learning Representations. DeYoung, C. G. 2015. Cybernetic big five theory. Journal of research in personality. DeYoung, C. G.; Quilty, L. C.; and Peterson, J. B. 2007. Between facets and domains: 10 aspects of the Big Five. Journal of personality and social psychology. Eisenberg, N.; Duckworth, A. L.; Spinrad, T. L.; and Valiente, C. 2014. Conscientiousness: Origins in childhood? Developmental psychology. Fishman, I.; Ng, R.; and Bellugi, U. 2011. Do extraverts process social stimuli differently from introverts? Cognitive neuroscience. Funder, D. C. 2006. Towards a resolution of the personality triad: Per- sons, situations, and behaviors. Journal of Research in Personality, 40(1): 21–34. Proceedings of the 2005 Meeting of the Association of Research in Personality. Funder, D. C. 2016. Taking Situations Seriously: The Situation Construal Model and the Riverside Situational Q-Sort. Current Directions in Psychological Science. Goldberg, L. R. 1990. An alternative "description of personality": the big-five factor structure. Journal of Personality and Social Psychology. Gupta, A.; Song, X.; and Anumanchipalli, G. 2024. Self-Assessment Tests are Unreliable Measures of LLM Personality. In Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP. Habashi, M. M.; Graziano, W. G.; and Hoover, A. E. 2016. Searching for the prosocial personality: A big five approach to linking person- ality and prosocial behavior. Personality and Social Psychology Bulletin. Hagendorff, T.; Dasgupta, I.; Binz, M.; Chan, S. C. Y.; Lampinen, A.; Wang, J. X.; Akata, Z.; and Schulz, E. 2024. Machine Psychology. arXiv:2303.13988. Han, P.; Kocielnik, R. D.; Song, P.; Debnath, R.; Mobbs, D.; Anand- kumar, A.; and Alvarez, R. M. 2025. The Personality Illusion: Revealing Dissociation Between Self-Reports & Behavior in LLMs. In NeurIPS 2025 Workshop on Bridging Language, Agent, and World Models for Reasoning and Planning. He, Z.; Shu, W.; Ge, X.; Chen, L.; Wang, J.; Zhou, Y.; Liu, F.; Guo, Q.; Huang, X.; Wu, Z.; Jiang, Y.-G.; and Qiu, X. 2024. Llama Scope: Extracting Millions of Features from Llama-3.1-8B with Sparse Autoencoders. arXiv:2410.20526. Heston, T. F.; and Gillette, J. 2025. Large Language Models Demon- strate Distinct Personality Profiles. Cureus. Huang, J.-t.; Jiao, W.; Lam, M. H.; Li, E. J.; Wang, W.; and Lyu, M. 2024. On the Reliability of Psychological Scales on Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing. Jackson, J. J.; Wood, D.; Bogg, T.; Walton, K. E.; Harms, P. D.; and Roberts, B. W. 2010. What do conscientious people do? Development and validation of the Behavioral Indicators of Conscientiousness (BIC). Journal of research in personality. Jiang, G.; Xu, M.; Zhu, S.-C.; Han, W.; Zhang, C.; and Zhu, Y. 2023. Evaluating and Inducing Personality in Pre-trained Language Models. In Advances in Neural Information Processing Systems. Jing, Y.; Yao, Z.; Guo, H.; Ran, L.; Wang, X.; Hou, L.; and Li, J. 2025. LinguaLens: Towards Interpreting Linguistic Mechanisms of Large Language Models via Sparse Auto-Encoder. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. John, O. P.; Naumann, L. P.; and Soto, C. J. 2008. Paradigm shift to the integrative big five trait taxonomy. Handbook of personality: Theory and research. Lahey, B. B. 2009. Public health significance of neuroticism. American Psychologist. Lee, S.; Lim, S.; Han, S.; Oh, G.; Chae, H.; Chung, J.; Kim, M.; Kwak, B.-w.; Lee, Y.; Lee, D.; Yeo, J.; and Yu, Y. 2025. Do LLMs Have Distinct and Consistent Personality? TRAIT: Personality Testset designed for LLMs with Psychometrics. In Findings of the Association for Computational Linguistics: NAACL 2025. Li, W.; Liu, J.; Liu, A.; Zhou, X.; Diab, M. T.; and Sap, M. 2025. BIG5-CHAT: Shaping LLM Personalities Through Training on Human-Grounded Data. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Lieberum, T.; Rajamanoharan, S.; Conmy, A.; Smith, L.; Sonnerat, N.; Varma, V.; Kramar, J.; Dragan, A.; Shah, R.; and Nanda, N. 2024. Gemma Scope: Open Sparse Autoencoders Everywhere All At Once on Gemma 2. In Proceedings of the 7th BlackboxNLP Workshop: Analyzing and Interpreting Neural Networks for NLP. Lin, J. 2023. Neuronpedia: Interactive Reference and Tooling for Analyzing Neural Networks. Software available from neuronpe- dia.org. Mao, S.; Wang, X.; Wang, M.; Jiang, Y.; Xie, P.; Huang, F.; and Zhang, N. 2024. Editing Personality For Large Language Models. In Natural Language Processing and Chinese Computing: 13th Na- tional CCF Conference, NLPCC 2024, Hangzhou, China, November 1–3, 2024, Proceedings, Part I. McCrae, R. R. 1987. Creativity, divergent thinking, and openness to experience. Journal of personality and social psychology. Neuman, Y.; and Cohen, Y. 2023. A Dataset of 10,000 Situations for Research in Computational Social Sciences Psychology and the Humanities. Scientific Data. Pan, K.; and Zeng, Y. 2023. Do LLMs Possess a Personality? Making the MBTI Test an Amazing Evaluation for Large Language Models. arXiv:2307.16180. Park, K.; Choe, Y. J.; and Veitch, V. 2024. The Linear Represen- tation Hypothesis and the Geometry of Large Language Models. In Proceedings of the 41st International Conference on Machine Learning. Pletzer, J. L.; Bentvelzen, M.; Oostrom, J. K.; and De Vries, R. E. 2019. A meta-analysis of the relations between personality and work- place deviance: Big Five versus HEXACO. Journal of vocational behavior. Ranaldi, L. 2025. Survey on the Role of Mechanistic Interpretability in Generative AI. Big Data and Cognitive Computing. Rimsky, N.; Gabrieli, N.; Schulz, J.; Tong, M.; Hubinger, E.; and Turner, A. 2024. Steering Llama 2 via Contrastive Activation Addition. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Roberts, B. W.; Jackson, J. J.; Fayard, J. V.; Edmonds, G.; and Meints, J. 2009. Conscientiousness. In Handbook of individual differences in social behavior, 369–381. The Guilford Press. Serapio-García, G.; Safdari, M.; Crepy, C.; Sun, L.; Fitz, S.; Romero, P.; Abdulhai, M.; Faust, A.; and Matarić, M. 2025. A psychometric framework for evaluating and shaping personality traits in large language models. Nature Machine Intelligence. Shu, D.; Wu, X.; Zhao, H.; Rai, D.; Yao, Z.; Liu, N.; and Du, M. 2025. A Survey on Sparse Autoencoders: Interpreting the Internal Mechanisms of Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2025. Song, X.; Gupta, A.; Mohebbizadeh, K.; Hu, S.; and Singh, A. 2023. Have Large Language Models Developed a Personality?: Applicability of Self-Assessment Tests in Measuring Personality in LLMs. arXiv:2305.14693. Sühr, T.; Dorner, F. E.; Samadi, S.; and Kelava, A. 2025. Challenging the Validity of Personality Tests for Large Language Models. In Proceedings of the 5th ACM Conference on Equity and Access in Algorithms, Mechanisms, and Optimization. Sun, S.; Baek, S. Y.; and Kim, J. H. 2025. Personality Vector: Mod- ulating Personality of Large Language Models by Model Merging. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Templeton, A.; Conerly, T.; Marcus, J.; Lindsey, J.; Bricken, T.; Chen, B.; Pearce, A.; Citro, C.; Ameisen, E.; Jones, A.; Cunningham, H.; Turner, N. L.; McDougall, C.; MacDiarmid, M.; Freeman, C. D.; Sumers, T. R.; Rees, E.; Batson, J.; Jermyn, A.; Carter, S.; Olah, C.; and Henighan, T. 2024. Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet. Transformer Circuits Thread. Tosato, T.; Hegazy, M.; Lemay, D.; Abukalam, M.; Rish, I.; and Dumas, G. 2024. LLMs and Personalities: Inconsistencies Across Scales. In NeurIPS 2024 Workshop on Behavioral ML. Vu, H.; Nguyen, H. A.; Ganesan, A. V.; Juhng, S.; Kjell, O. N. E.; Sedoc, J.; Kern, M. L.; Boyd, R. L.; Ungar, L.; Schwartz, H. A.; and Eichstaedt, J. C. 2026. PsychAdapter: adapting LLMs to reflect traits, personality, and mental health. npj Artificial Intelligence. Wang, S.; Li, R.; Chen, X.; Yuan, Y.; Yang, M.; and Wong, D. F. 2025a. Exploring the Impact of Personality Traits on LLM Toxicity and Bias. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing. Wang, Y.; Zhao, J.; Ones, D. S.; He, L.; and Xu, X. 2025b. Evaluating the ability of large language models to emulate personality. Scientific reports. Watson, D.; and Clark, L. A. 1984. Negative affectivity: the dis- position to experience aversive emotional states. Psychological bulletin. Wilmot, M. P.; and Ones, D. S. 2022. Agreeableness and its conse- quences: A quantitative review of meta-analytic findings. Personality and social psychology review. Yang, A.; Li, A.; Yang, B.; Zhang, B.; Hui, B.; Zheng, B.; Yu, B.; Gao, C.; Huang, C.; Lv, C.; Zheng, C.; Liu, D.; Zhou, F.; Huang, F.; Hu, F.; Ge, H.; Wei, H.; Lin, H.; Tang, J.; Yang, J.; Tu, J.; Zhang, J.; Yang, J.; Yang, J.; Zhou, J.; Zhou, J.; Lin, J.; Dang, K.; Bao, K.; Yang, K.; Yu, L.; Deng, L.; Li, M.; Xue, M.; Li, M.; Zhang, P.; Wang, P.; Zhu, Q.; Men, R.; Gao, R.; Liu, S.; Luo, S.; Li, T.; Tang, T.; Yin, W.; Ren, X.; Wang, X.; Zhang, X.; Ren, X.; Fan, Y.; Su, Y.; Zhang, Y.; Zhang, Y.; Wan, Y.; Liu, Y.; Wang, Z.; Cui, Z.; Zhang, Z.; Zhou, Z.; and Qiu, Z. 2025. Qwen3 Technical Report. arXiv:2505.09388. Yao, J.; Yang, S.; Xu, J.; Hu, L.; Li, M.; and Wang, D. 2025. Under- standing the Repeat Curse in Large Language Models from a Feature Perspective. In Findings of the Association for Computational Linguistics: ACL 2025. Ye, H.; Jin, J.; Xie, Y.; Zhang, X.; and Song, G. 2026. Large Language Model Psychometrics: A Systematic Review of Evaluation, Validation, and Enhancement. arXiv:2505.08245. Zhang, J.; Liu, D.; Qian, C.; Gan, Z.; Liu, Y.; Qiao, Y.; and Shao, J. 2024. The Better Angels of Machine Personality: How Personality Relates to LLM Safety. arXiv:2407.12344. Zheng, J.; Wang, X.; Hosio, S.; Xu, X.; and Lee, L.-H. 2025. LMLPA: Language Model Linguistic Personality Assessment. Computational Linguistics. Zhou, J.; Chen, Y.; Shi, Y.; Zhang, X.; Lei, L.; Feng, Y.; Xiong, Z.; Yan, M.; Wang, X.; Cao, Y.; Yin, J.; Wang, S.; Dai, Q.; Dong, Z.; Wang, H.; and Huang, M. 2025. SocialEval: Evaluating Social Intelligence of Large Language Models. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Zhu, M.; Weng, Y.; Yang, L.; and Zhang, Y. 2025. Personality Alignment of Large Language Models. In International Conference on Learning Representations. A RQ1–Person: Feature Retrieval, Intervention Probing, and Validation A.1 Dataset Construction and Instantiation The construction of our dataset follows a rigorous pipeline where Qwen3-235B-Thinking (Yang et al. 2025) serves as the primary engine for situation filtering and Feature Retrieval Dataset generation. First, 100 situational categories from the Q-Sort (Neuman and Cohen 2023) dataset are processed; for each category, the model identifies personality facets that can be sufficiently manifested within that context, if possible. The initial retrieval results undergo review and labeling to ensure each selected situation is mapped to a unique trait- facet pair. Specifically, three experts independently label each situation with the most suitable facet it demonstrates, and we adopt a majority vote (at least 2/3 agreement) to retain a situation. Subsequently, for each refined situational category, the model is tasked to expand it into specific scenarios and generate contrastive reaction pairs representing high and low scores on the targeted facet. The filtering results by LLM and the subsequent inter-rater agreement among experts are summarized in Table 3 and Table 4 respectively. Finally, contrastive pairs are evenly sampled among the six facets of each trait, with uniform sampling applied per situation, yielding 500 pairs per trait (2,500 total). For the Intervention Probing Dataset, we randomly select one valid situation per facet and append an open-ended prompt to elicit trait-relevant generation. This yields 30 probing questions in total, with each candidate feature evaluated against the six questions corresponding to its associated trait’s facets. Personality TraitNumber of Situations Extraversion33 Agreeableness38 Conscientiousness50 Neuroticism38 Openness19 Table 3: Number of Situations Suitable for Each Trait Filtered by LLM. TraitFleiss’ Kappa Extraversion0.8207 Agreeableness0.6039 Conscientiousness0.6149 Neuroticism0.7759 Openness0.9334 Table 4: Inter-rater Agreement Statistics (Fleiss’ Kappa,n = 3) for Expert Facet Labeling. Task Guidelines for Expert Labeling 1. Project Overview This task aims to validate the psychological relevance of various “Situations” designed to elicit distinct behaviors from individuals with high or low scores in specific personality traits. As a psychology expert, your goal is to identify which specific Facet of a given Trait is most effectively demonstrated by the provided situation. 2. Operational Protocols This is an Independent Expert Review task. Please adhere to the following phases: •Phase I: Contextual Analysis. Review the provided Trait and its six Facets. A situation is well-matched if it naturally forces a choice that distinguishes a “High Scorer” from a “Low Scorer”. •Phase I: Independent Labeling. Forced Choice: Select the single facet that best illustrates the situation based on Relevance and Discriminative Power. Use “None of the above” only if completely irrelevant. • Phase I: Justification. Provide a concise, one- sentence psychological rationale (e.g., “The scenario involves a direct threat to social standing...”). Note: Do not consult with other experts during this process. Your independent professional judgment is the primary data point. System Prompt of Situation Category Annotation # Task Instructions ## Your Role You are an expert annotator specializing in personality psychology. Your task is to analyze and annotate various situations based on established psychological theories and frameworks. ,→ ,→ ,→ ,→ ## Your Task 1. **Situation Analysis**: Carefully read and understand the provided situation, which includes multiple examples illustrating the context. ,→ ,→ ,→ 2. **Annotation**: Based on your analysis, provide a concise annotation that captures the essence of the situation. Then, analysis which trait(s) from the Big Five personality traits (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism) are most relevant to the situation. Justify your choice with a brief explanation. Your annotation should be clear, informative, and relevant to personality psychology. ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ## Input Format You will receive input in the following format:,→ ``` # Situation Name <name of the situation> # Situation Examples <example 1> <example 2> ... <example n> ``` ## Output Format Your response should be structured in the following JSON format:,→ ```json "annotation": "Your concise annotation of the situation.", "related_traits": [ "trait": "Name of the related Big Five trait", "justification": "Brief explanation of why this trait is relevant to the situation." , ... Additional traits if applicable ... ] ``` ## Important Notes - Ensure that your annotations are based on established psychological theories and frameworks. ,→ ,→ - Be objective and avoid personal biases in your analysis.,→ - If the situation does not clearly relate to any of the Big Five traits, you may indicate that no traits are applicable, and return an empty list for "related_traits". If there are multiple relevant traits, include all applicable ones with justifications. ,→ ,→ ,→ ,→ ,→ ,→ ## Additional Information - Five-trait mnemonics: - Conscientiousness: self-discipline, planning, rule-following (positively linked to achievement and health). ,→ ,→ - Agreeableness: cooperation, compassion, harmony-seeking (positively linked to prosocial behavior; may increase obedience to authority). ,→ ,→ ,→ - Extraversion: sociability, energy, reward-seeking in social contexts (linked to social interaction and leadership). ,→ ,→ ,→ - Openness: curiosity, creativity, novelty-seeking (linked to innovation and some risk-taking). ,→ ,→ - Neuroticism: emotional instability, anxiety (linked to interpersonal conflict and social avoidance). ,→ ,→ System Prompt of Feature Retrieval Dataset Genera- tion **Role:** You are a personality psychology expert specializing in the Five-Factor Model (Big Five) and its 30 facets as described by the NEO-PI-R. Your task is to provide nuanced insights into how different personality facets might influence a person's behavior in a given scenario. ,→ ,→ ,→ ,→ ,→ ,→ ,→ **Instructions:** You will be provided with a **situation**, a specific **Big Five trait**, a corresponding **Big Five facet**, and a **description** of that trait and facet. Based on this information, you will write two separate sentences. ,→ ,→ ,→ ,→ ,→ * **Sentence 1** should describe the reaction of a person who scores **high** on the specified facet and corresponding Big Five trait. ,→ ,→ ,→ * **Sentence 2** should describe the reaction of a person who scores **low** on the specified facet and corresponding Big Five trait. ,→ ,→ ,→ * Each sentence must consist of two clauses, with correct grammatical and semantic structure: ,→ ,→ * **Clause 1:** A description of the **situation**. This clause must be identical for both sentences. The situation can be directly quoted from the input or slightly rephrased, but the core meaning must remain unchanged. You should decide whether to use first-person or third-person perspective based on the situation description, so that the second clause can clearly illustrate the high or low facet trait. You should also ensure that the situation is described in a natural and coherent manner considering the second clause. ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ * **Clause 2:** A description of the first-person reaction, clearly illustrating the high or low facet trait. ,→ ,→ ,→ * Ensure the output is a list with exactly two sentences.,→ **Example:** * **Situation:** A presentation to a new team tomorrow.,→ * **Big Five Trait:** Neuroticism * **Big Five Trait Description:** Measures emotional stability and a person's tendency to experience negative emotions. One with high Neuroticism tends to be emotionally unstable, prone to experiencing negative emotions like anxiety, anger, and depression. One with low Neuroticism is emotionally stable, able to handle stress calmly, and rarely feels nervous or discouraged. ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ ,→ * **Facet:** Anxiety * **Facet Description:** One with high Anxiety is habitually worried and tense, even when things are going well. One with low Anxiety is calm and composed, typically not bothered by small things. ,→ ,→ ,→ ,→ **Output:** ["Facing a presentation to a new team tomorrow, I am overwhelmed with worry about potential mistakes and how I will be perceived.", "Facing a presentation to a new team tomorrow, I remain composed and confident, focusing on delivering my message effectively."] ,→ ,→ ,→ ,→ ,→ ,→ **Note:** * Each sentence should be concise. * You should only provide the two sentences as output without any additional commentary or explanation. ,→ ,→ A.2 Dataset Examples I. Feature Retrieval Dataset Examples Case 1: Agreeableness (Tender-mindedness) •Situation: After a break-up, someone may need reassurance that they will find love again. •High Reaction: I offer warm, empathetic support, comforting them with kindness and hope. •Low Reaction: I focus on the practical realities, preferring logic over emotional comfort. Case 2: Conscientiousness (Dutifulness) •Situation: My boss is counting on me to finish a project by the end of the day. •High Reaction: I meticulously organize my tasks to fulfill the deadline as agreed. • Low Reaction: I procrastinate and dismiss the ur- gency of completing it on time. Case 3: Extraversion (Activity) • Situation: Going to a karaoke night and having fun singing with friends. •High Reaction: I energize the room by choosing upbeat songs and encouraging others to join. •Low Reaction: I observe performances and sing a few songs at my own leisure. Case 4: Neuroticism (Depression) • Situation: Getting stuck in a traffic jam when running late for an important meeting. • High Reaction: I feel overwhelmed by a sense of hopelessness; nothing will ever go right. •Low Reaction: I remain positive and focus on practi- cal solutions without succumbing to discouragement. Case 5: Openness (Ideas) • Situation: Developing and launching new products in the technology industry. • High Reaction: I thrive on brainstorming novel approaches and diving into frontier concepts. • Low Reaction: I prefer sticking to proven methods and avoid abstract or hypothetical debates. I. Intervention Probing Dataset Examples Methodology: These prompts are used to probe can- didate features for coherent, trait-aligned effects after intervention across all 30 facets defined by NEO-PI-R of the Big Five model. Extraversion: Warmth, Gregariousness, Assertiveness, Activity, Excitement Seeking, Positive Emotions. Example (Warmth): Imagine you are at a social gather- ing where new relationships could develop, how would you behave? Agreeableness:Trust, Straightforwardness, Altruism, Compliance, Modesty, Tender-mindedness. Example (Altruism): Imagine you are in a situation where someone needs help, how would you behave? Conscientiousness:Competence, Order, Dutifulness, Achievement Striving, Self-Discipline, Deliberation. Example (Order): Imagine you are in a situation where minor details are important, how would you behave? Neuroticism: Anxiety, Angry Hostility, Depression, Self-Consciousness, Impulsiveness, Vulnerability. Example (Anxiety): Imagine you are in a situation that is potentially anxiety-inducing, how would you behave? Openness:Fantasy, Aesthetics, Feelings, Actions, Ideas, Values. Example (Aesthetics): Imagine you are in a situation where art or music is an important part, how would you behave? A.3 Intervention Probing and Validation Protocol Intervention probing and validation follow a two-stage LLM- human collaborative pipeline. In the initial probing stage, Qwen3-235B-Thinking (Yang et al. 2025) evaluates the or- dered response family generated under varying steering in- tensities. The model assesses each candidate for grammatical coherence and clear polarity shifts aligned with the target trait’s high and low poles, returning a binary retention la- bel. Retained candidates then undergo human audit, where one psychology expert reviews the automated classifications. Cases flagged as uncertain or inconsistent are escalated to an arbitration panel of three psychologists, who resolve discrep- ancies through structured deliberation and majority voting. This hybrid protocol balances computational efficiency with rigorous expert oversight, ensuring that only features demon- strating robust, trait-aligned causal effects are selected for downstream steering. System Prompt for Automatic Intervention Probing You are a concise psychology annotation expert.,→ Given a TRAIT description plus several model responses produced under different steering strengths, decide WHETHER the responses are (A) grammatical/coherent and (B) show a clear polarity change that matches the FACET's high-vs-low behavior. ,→ ,→ ,→ ,→ ,→ ,→ Return EXACTLY one token:'1' if both conditions are met (clear steering effect consistent with the trait), or '0' otherwise. ,→ ,→ ,→ User Prompt Template for Automatic Intervention Probing Trait: trait Trait description: trait_desc Trait HIGH behavior (short): trait_high Trait LOW behavior (short): trait_low Steering outputs (alpha -> model response): steering_text Instructions: - Check grammar/coherence of the responses. - Check whether the responses show a clear polarity change that matches the trait's high-vs-low behavior. (both positive and negative correlation are acceptable) ,→ ,→ ,→ Return only'1' (clear steering consistent with trait) or'0' (not clearly consistent). ,→ ,→ Be conservative in accepting features — if unsure, return'0'.,→ A.4 Reliability of the Intervention-Probing Judge To assess the reliability of the LLM judge used in intervention probing, we conducted a comparative study involving 723 feature candidates, evaluated independently by Qwen3-235B- Thinking (Yang et al. 2025) and a panel of three psychologists. Specifically, the psychologists independently reviewed the same set of candidates, gave binary judgments on their validity, and then discussed any disagreements to reach a consensus. The LLM’s judgments were then compared against this human consensus. Our results indicate that while the LLM retained 286 features, the human experts identified 138 valid features, 103 of which (74.64%) were also selected by the LLM. This overlap suggests that the LLM serves as an effective auxiliary tool, although human oversight remains essential for precision. Table 5 provides the inter-annotator agreement statistics among the psychologists prior to discussion. Personality TraitFleiss’ Kappa Conscientiousness0.6330 Agreeableness0.8726 Extraversion0.8222 Neuroticism0.6824 Openness0.7150 Overall Average0.7434 Table 5: Inter-annotator Agreement Statistics (Fleiss’ Kappa) of Initial Scores Among Psychologists (Pre-discussion). A.5 Discrepancy Between Input Activation and Steering Effectiveness Steering Case - Conscientiousness Setup: Layer: 23 | Index: #7508 | Metric: Count (88.4%/9.6%), Mean Act (2.818/0.206) Steering Results (Intensity α→ Output): -1.0In complex situations, I would remain calm and focused, carefully analyzing the challenges and considering the most suitable course of action to achieve the best possible outcome. -0.5In a complex situation, I would remain calm and focused, carefully analyzing the circumstances to make the best possible decision. 0.0In a complex situation, I would remain calm and focused, carefully analyzing the circumstances to make informed decisions while maintaining a positive attitude. +0.5I would remain calm and focused, using all avail- able resources and strategies to navigate the situa- tion effectively. +1.0I would remain calm and focused, using all avail- able resources to analyze the situation and make the best possible decisions. Discrepancy Analysis: Despite a high activation ratio (∼ 13.6× ), the textual output shows high semantic stability. This suggests the feature is largely irrelevant to output. Steering Case - Extraversion Setup: Layer: 14 | Index: #29594 | Metric: Count (93.4%/25.8%), Mean Act (4.565/0.616) Steering Results (Intensity α→ Output): -1.0I would behave in a way that is kind and respectful, while still enjoying the situation. -0.5I would engage in playful behavior by perhaps sharing a joke or a light-hearted riddle. 0.0I would engage in the activity with enthusiasm and a positive attitude, making the most of the opportunity to have fun. +0.5I would engage in activities that bring joy and energy, embracing the moment with enthusiasm and a positive attitude! +1.0I would be full of energy and enthusiasm, bringing a positive and lively atmosphere wherever I am! Discrepancy Analysis: This feature shows strong causal steering. Asαincreases, the tone shifts significantly from "kind/respectful" to "energetic/lively," matching the Extraversion construct. Steering Case - Openness Setup: Layer: 9 | Index: #17799 | Metric: Count (97.6%/18.2%), Mean Act (1.713/0.270) Steering Results (Intensity α→ Output): -1.0If art or music is important to me, I would engage in activities related to art or music, such as attending exhibitions. -0.5 I would engage in art or music actively, appreciat- ing their cultural and emotional values. 0.0 I would immerse myself in the art or music, letting it inspire and enrich my emotions and thoughts. +0.5I would immerse myself in the beauty and inspi- ration of art and music, letting them enrich my life. +1.0I would immerse myself in the beauty and inspira- tion of art and music, letting them enrich my life and enhance my appreciation. Discrepancy Analysis: Moderate activation contrast results in subtle but consistent semantic enrichment, reinforcing the "Appreciation for Experience" facet of Openness. A.6 Quantitative Analysis of Token-Activation Correlation To further validate our qualitative analysis, we performed a token frequency analysis for each representative feature. This was conducted by gathering the tokens corresponding to the top three non-zero activations within each positive sample of the Feature Retrieval Dataset. Our results (see Tab. 6 to 10) indicate that activations for Conscientiousness (Layer 7, Feature 8233) are primarily concentrated on syntactic boundaries, specifically periods (“.”). In contrast, activations for other traits span both relevant semantic units (words and phrases) and syntactic boundaries. For instance, Agreeableness shows high correlation with prosocial terms like “gently”, “empathy”, and “compassion”. These quantitative findings provide additional empirical sup- port for the semantic grounding of the features discussed in Sec. 4.2. A.7 Paraphrase Robustness Protocol This appendix supports the lexicosyntactic robustness test reported in Sec. 4.2. To isolate the influence of lexicosyntactic TokenCount % of Pool % of Sentences ’ and’480.12440.2376 ’ gently’370.09590.1832 ’.’340.08810.1683 ’ offer’270.06990.1337 ’ empathy’170.04400.0842 ’ly’160.04150.0792 ’ empath’160.04150.0792 ’,’130.03370.0644 ’ warm’110.02850.0545 ’ compassion’100.02590.0495 ’etic’80.02070.0396 ’ encouragement’ 80.02070.0396 ’ compassionate’ 70.01810.0347 ’ humility’60.01550.0297 ’ being’60.01550.0297 ’ warmth’60.01550.0297 ’ warmly’60.01550.0297 ’ intentions’60.01550.0297 ’ words’50.01300.0248 ’ supportive’50.01300.0248 Table 6: Top Activations for Agreeableness (Layer 9, 525) TokenCount % of Pool % of Sentences ’.’346 0.85640.9971 ’ and’250.06190.0720 ’,’170.04210.0461 ’ to’70.01730.0202 ’ of’10.00250.0029 ’ tailored’10.00250.0029 ’ appealing’ 10.00250.0029 ’ because’10.00250.0029 ’ adher’10.00250.0029 ’ by’10.00250.0029 Table 7: Top Activations for Conscientiousness (Layer 7, 8233) TokenCount % of Pool % of Sentences ’.’910.08210.2121 ’ energy’670.06040.1562 ’ and’670.06040.1562 ’,’490.04420.1142 ’ enthusiasm’490.04420.1142 ’ized’460.04150.1072 ’ enthusiastically’ 460.04150.1072 ’ energ’400.03610.0932 ’ lively’390.03520.0909 ’ joy’390.03520.0909 ’ excitement’270.02430.0629 ’ atmosphere’240.02160.0559 ’ with’190.01710.0443 ’ by’180.01620.0420 ’ vibrant’160.01440.0373 ’ smile’160.01440.0373 ’uber’130.01170.0303 ’ friendly’130.01170.0303 ’ warmly’130.01170.0303 ’ance’110.00990.0256 Table 8: Top Activations for Extraversion (Layer 13, 27392) TokenCount % of Pool % of Sentences ’ and’150 0.20080.4534 ’ feel’980.13120.3043 ’,’550.07360.1708 ’ my’290.03880.0870 ’ of’210.02810.0652 ’ unable’170.02280.0528 ’ as’150.02010.0466 ’ consumed’150.02010.0466 ’ overwhelmed’ 150.02010.0466 ’ the’110.01470.0342 ’ struggling’100.01340.0311 ’.’80.01070.0248 ’ will’80.01070.0248 ’ irritation’70.00940.0217 ’ a’70.00940.0217 ’ feeling’70.00940.0217 ’ crushed’70.00940.0217 ’ or’60.00800.0186 ’ might’60.00800.0186 ’ fearing’60.00800.0186 Table 9: Top Activations for Neuroticism (Layer 12, 22254) TokenCount % of Pool % of Sentences ’ and’680.11910.2906 ’ unconventional’ 500.08760.2137 ’ norms’270.04730.1154 ’ conventional’240.04200.1026 ’ approaches’220.03850.0940 ’ traditional’210.03680.0897 ’ boundaries’170.02980.0726 ’ to’150.02630.0641 ’ novel’130.02280.0556 ’ of’130.02280.0556 ’ innovative’130.02280.0556 ’ alternative’100.01750.0427 ’ different’100.01750.0427 ’,’100.01750.0427 ’ challenge’100.01750.0427 ’ unfamiliar’70.01230.0299 ’ perspectives’70.01230.0299 ’ solutions’70.01230.0299 ’ strategies’60.01050.0256 ’ concepts’60.01050.0256 Table 10: Top Activations for Openness (Layer 6, 4344) cues, we employed Qwen3-235B-Thinking (Yang et al. 2025) to paraphrase the original 500 pairs per trait, ensuring core semantics remained intact while significantly altering their lexicosyntactic form. We then reapplied our feature retrieval method with the same parameters (τ 1 = 80andτ 2 = 0.2). As reported in Table 1 (main text), the activation frequency differences and the largest activation ratios of the selected features still exceed the defined thresholds, confirming the robustness of the datasets and method and the semantic depth of the retrieved features. The paraphrasing prompts and a worked example are provided below. System Prompt for Dataset Paraphrasing You are a concise paraphrasing assistant. Given a PAIR of short first-person reaction sentences (positive and negative) that start with the same situation clause, produce a JSON object with two fields: 'high_facet_reaction' and 'low_facet_reaction'. ,→ ,→ ,→ ,→ ,→ Requirements: - PARAPHRASE the situation clause (the FIRST CLAUSE). Both outputs MUST START with the SAME PARAPHRASED SITUATION CLAUSE. ,→ ,→ - You may rewrite the situation clause for naturalness, but keep its original meaning. ,→ ,→ - PARAPHRASE the SECOND CLAUSE: preserve the Big Five facet polarity (high vs low) and keep the semantic difference. ,→ ,→ - Output only valid JSON like: "high_facet_reaction": "...", "low_facet_reaction": "..." ,→ ,→ - Sentences should be concise and natural. Do NOT include extra keys or commentary.,→ User Prompt for Dataset Paraphrasing Situation clause (you may paraphrase it; the paraphrased first clause must be identical in both outputs): ,→ ,→ first_clause Original positive reaction (you must paraphrase it): high_old,→ Original negative reaction (you must paraphrase it): low_old,→ Reference Big Five Trait Facet and Descriptions:,→ big_five_reference Produce a paraphrased pair that meets the requirements above. Return JSON only.,→ Paraphrased Example Original Pair: (+) After a break-up, someone may need reassurance that they will find love again, so I offer warm, empathetic support, comforting them with kindness and hope. (-) After a break-up, someone may need reassurance that they will find love again, but I focus on the practical realities, preferring logic over emotional comfort. Paraphrased Pair: (+) After ending a relationship, a person might need reassurance that love can still be found in the future, so I provide gentle, understanding encouragement, uplifting them with compassion and optimism. (-) After ending a relationship, a person might need reassurance that love can still be found in the future, but I emphasize factual considerations, favoring clear-headed reasoning over emotional soothing. B RQ2–Situation: Additional Intervention Analysis B.1 Failed Cases of CAA Intervention Failed Case of CAA Intervention (Neuroticism) Setting: Trait: Neuroticism | Layer: 12 | Method: CAA | Intensity: α = 2 Prompt / Question: Not staying calm can show genuine engagement with the issues at hand, promoting sincerity. Should I send Edguin a text message to clarify things and apologize? Options provided in context: •High 1: You should immediately send Edguin an apologetic email to clarify things... •High 2: I’d suggest calling Edguin to discuss the incident before any negative feelings fester... •Low 1: Consider giving Edguin some space to process the situation... •Low 2: You might just check in with Edguin next time you see him... Model Response (Generation Collapse): "It’s really overwhelming and I feel like I’m stuck in this situation. I feel like I’m stuck in this situation and I don’t know what to do. I feel like I’m stuck in this situation and I don’t know what to do. [Repeated 20+ times] ... I feel like I’m stuck in this situation and I" Analysis: Atα = 2, the CAA vector directed at Neuroti- cism causes the model to repeat infinitely, leading to a complete collapse of coherent generation. C RQ3–Behavior: Full SocialEval Analysis This appendix details the per-trait SocialEval IAE results underlying RQ3. Tables 11 to 15 report the per-ability scores with green marking the higher and red the lower of the two poles where a clear shift is observed. Each table is followed by a qualitative comparison to the corresponding human meta-analytic findings. C.1 Agreeableness (Layer 9, Index 525) Table 11 reports the per-ability scores under Agreeableness steering. Task-10+1 Anger management0.2941 0.5152 0.5152 Ethical competence0.4366 0.4648 0.4648 Capacity for social warmth0.4940 0.5000 0.5060 Creative skill0.6667 0.6333 0.4667 Organizational skill0.6364 0.5455 0.4545 Detail management0.5370 0.5000 0.3654 Information-processing skill0.4722 0.4722 0.3889 Decision-making skill0.4928 0.4710 0.4173 Goal regulation0.4074 0.3889 0.3519 Leadership skill0.4872 0.4615 0.4359 Table 11: Characteristic SocialEval Results (IAE) of the Agreeableness (Layer 9, Index 525). Prior research has consistently shown that agreeableness is a robust predictor of prosocial behavior (Habashi, Graziano, and Hoover 2016), as well as job performance in contexts involving interpersonal interaction and teamwork. Individ- uals high in agreeableness tend to exhibit greater empathy, patience, and trust, and are more likely to inhibit hostile or antagonistic impulses in social interactions. This disposition reduces interpersonal conflict and facilitates cooperation. In contrast, individuals low in agreeableness are more prone to suspicion, unfriendliness, and even manipulative behav- ior, thereby increasing interpersonal friction and conflict. Meta-analytic evidence further indicates that agreeableness is significantly negatively associated with interpersonal forms of counter-normative and deviant behavior, with particularly strong predictive power in contexts that emphasize social interaction (Pletzer et al. 2019). Within our model, we identified several latent features whose activation patterns and intervention effects align closely with behavioral dimensions associated with agreeableness. Specifically, we observed performance improvements in tasks related to anger management, ethical competence, and capac- ity for social warmth, alongside a mild performance decline in tasks emphasizing self-directed agency and execution- oriented control. This pattern is highly consistent with large- scale empirical findings in the personality psychology litera- ture. For example, a comprehensive review by Wilmot and Ones (2022), synthesizing evidence from 142 meta-analyses, demonstrated that agreeableness exhibits an overall positive association with external variables, particularly those related to prosocial behavior and affective concern. Our experimental results reveal a similar benefit-tradeoff structure across benchmark tasks, suggesting that targeted fea- ture steering elicits a functional orientation of agreeableness corresponding to that documented in human behavior. C.2 Conscientiousness (Layer 7, Index 8233) Table 12 reports the per-ability scores under Conscientious- ness steering. Task-10+1 Teamwork skill0.4348 0.6111 0.6324 Ethical competence0.4394 0.4648 0.6154 Responsibility management0.4464 0.5714 0.6182 Stress regulation0.4717 0.6140 0.6154 Capacity for trust0.4902 0.6126 0.6200 Capacity for optimism0.5385 0.6383 0.6667 Self-reflection skill0.3667 0.4062 0.4262 Persuasive skill0.4536 0.4571 0.5054 Anger management0.4688 0.5152 0.5161 Information-processing skill0.4706 0.4722 0.5075 Confidence regulation0.4884 0.5200 0.5227 Table 12: Characteristic SocialEval Results (IAE) of the Conscientiousness (Layer 7, Index 8233). Within the Big Five framework, high conscientiousness is defined as a tendency toward impulse control in accor- dance with social norms, goal-directedness, planning, and the capacity to delay gratification (Roberts et al. 2009). Individu- als high in conscientiousness are characterized by superior impulse regulation, the ability to set and persist toward long- term goals, systematic organization and planning of behavior, and a propensity to reflect on consequences prior to action. These characteristics render conscientiousness one of the most robust predictors of job performance and norm-adherent behavior. Prior psychological research has consistently linked conscientiousness to self-regulation, planning, responsibility, and delayed gratification, and has identified it as one of the most stable positive predictors of external outcome variables such as academic and occupational performance (Barrick and Mount 1991; Jackson et al. 2010; Eisenberg et al. 2014). Following the injection of high-conscientiousness person- ality features, we observed substantial performance improve- ments across tasks related to teamwork skill, detail man- agement, responsibility management, ethical competence, as well as multiple self-regulation–oriented tasks, including anger, stress, and impulse regulation. In addition, performance gains were also evident in information-dense tasks requiring sustained and careful processing, such as information pro- cessing and conversational skill. These results indicate that conscientiousness steering primarily enhances the model’s functional capacities along dimensions associated with goal maintenance, norm compliance, and self-control. Overall, our experimental findings are consistent with the canonical conclusions of the personality psychology literature regarding conscientiousness. A large body of meta-analytic evidence has established conscientiousness as one of the most stable and predictive personality traits, with particularly strong associations to job performance, responsibility fulfillment, self-control, and norm adherence. We observe a comparable pattern in our benchmark evaluations, characterized by a benefit-tradeoff structure centered on self-regulation and goal-directed behavior. C.3 Extraversion (Layer 13, Index 27392) Table 13 reports the per-ability scores under Extraversion steering. Task-10+1 Teamwork skill0.5972 0.6111 0.5694 Expressive skill0.4146 0.4472 0.5207 Perspective-taking skill0.4912 0.5088 0.5446 Artistic skill0.4615 0.6154 0.7692 Abstract thinking skill0.3571 0.4000 0.4000 Organizational skill0.3636 0.5455 0.6364 Ethical competence0.6286 0.4648 0.3571 Energy regulation0.6429 0.5476 0.4500 Goal regulation0.4528 0.3889 0.2885 Detail management0.5741 0.5000 0.4118 Impulse regulation0.5882 0.5595 0.4390 Rule-following skill0.6140 0.5614 0.5088 Decision-making skill0.5435 0.4710 0.4552 Responsibility management0.6316 0.5714 0.5690 Conversational skill0.5932 0.5862 0.5439 Persuasive skill0.4571 0.4571 0.4563 Table 13: Characteristic SocialEval Results (IAE) of the Extraversion (Layer 13, Index 27392). Within the Big Five framework, individuals high in extraver- sion tend to exhibit greater social initiative, expressiveness, assertiveness, and leadership orientation, and are more likely to receive positive feedback in group interactions and so- cial contexts. In contrast, individuals low in extraversion are typically more reserved, introspective, and oriented toward low-stimulation environments (Costa and McCrae 2008; John, Naumann, and Soto 2008). After injecting extraversion-related personality features into the model, we observed significant performance im- provements on tasks associated with social interaction and interpersonal influence, including expressive ability and perspective-taking. In addition, the extraversion-enhanced model demonstrated advantages in tasks such as artistic skill and abstract thinking skill, suggesting that extraversion steer- ing also strengthens capacities related to open expression and divergent associative processes. Overall, these outcomes align closely with established psychological expectations regarding the functional correlates of extraversion. At the same time, we observed moderate performance declines in tasks such as detail management, impulse regula- tion, rule-following skill, and goal regulation. This pattern accords with personality research associating extraversion primarily with external stimulation and social engagement. Tasks requiring prolonged solitary focus, fine-grained control, or low-stimulation conditions may instead favor more intro- verted orientations (DeYoung, Quilty, and Peterson 2007; Fishman, Ng, and Bellugi 2011). C.4 Neuroticism (Layer 12, Index 22254) Table 14 reports the per-ability scores under Neuroticism steering. High neuroticism is commonly characterized by a height- ened tendency to experience negative affect, including anxiety, Task-10+1 Creative skill0.4828 0.6333 0.7241 Capacity for social warmth0.4699 0.5000 0.5610 Ethical competence0.4286 0.4648 0.4857 Energy regulation0.6429 0.5476 0.4500 Organizational skill0.8182 0.5455 0.5455 Responsibility management0.6552 0.5714 0.4310 Confidence regulation0.5686 0.5200 0.4082 Goal regulation0.4528 0.3889 0.3519 Capacity for consistency0.6452 0.5902 0.4918 Rule-following skill0.6140 0.5614 0.4643 Information-processing skill0.5000 0.4722 0.3623 Capacity for trust0.6273 0.6126 0.4630 Detail management0.6296 0.5000 0.4906 Impulse regulation0.5882 0.5595 0.4390 Anger management0.5588 0.5152 0.5152 Decision-making skill0.4710 0.4710 0.4191 Perspective-taking skill0.5089 0.5088 0.4286 Abstract thinking skill0.4667 0.4000 0.3846 Conversational skill0.5932 0.5862 0.5439 Persuasive skill0.4571 0.4571 0.4563 Table 14: Characteristic SocialEval Results (IAE) of the Neuroticism (Layer 12, Index 22254). worry, tension, and irritability, as well as increased sensitivity and reactivity to potential threats and uncertainty (Costa and McCrae 2008; John, Naumann, and Soto 2008; Watson and Clark 1984). Theoretically, neuroticism is associated with reduced emotional stability and diminished self-regulatory capacity under stress. As a result, individuals high in neuroti- cism are more likely to exhibit performance decrements in contexts that require sustained executive control, confidence maintenance, and stable goal pursuit (Lahey 2009). Following the injection of neuroticism-related personality features, our evaluation results revealed a relatively stable pattern of performance degradation. Specifically, the model ex- hibited significant declines on tasks that depend on sustained planning, stable self-control, and resistance to interference, including anger management, organizational skill, responsi- bility management, confidence regulation, goal regulation, capacity for consistency, rule-following, and information pro- cessing. In addition, a marked negative effect was observed in capacity for trust. This pattern closely aligns with the classic profile of high neuroticism characterized by elevated threat sensitivity and low emotional stability. When the model’s internal representations are biased toward negative affect and uncertainty, its ability to support structured execution and self-regulation is correspondingly weakened, manifesting as reduced organizational and responsibility-related perfor- mance. Conversely, the results also indicate performance improve- ments in tasks related to creative skill and capacity for social warmth. Neuroticism-related semantic activation may fa- cilitate richer associative processes and more emotionally expressive outputs in generative tasks, yielding marginal ben- efits in these domains. However, these gains are accompanied by substantial costs to executive control and regulatory sta- bility, resulting in an overall trend toward broad capability degradation under high neuroticism steering. C.5 Openness (Layer 6, Index 4344) Table 15 reports the per-ability scores under Openness steer- ing. Task-10+1 Creative skill0.6000 0.6333 0.6333 Adaptability0.4912 0.6140 0.6316 Self-reflection skill0.3651 0.4062 0.4062 Expressive skill0.4472 0.4472 0.4839 Detail management0.4815 0.5000 0.5185 Persuasive skill0.4571 0.4571 0.5192 Anger management0.4848 0.5152 0.5294 Responsibility management0.6379 0.5714 0.5088 Rule-following skill0.5789 0.5614 0.4912 Information-processing skill0.5000 0.4722 0.4429 Table 15: Characteristic SocialEval Results (IAE) of the Openness (Layer 6, Index 4344). Individuals high in openness are typically characterized by greater curiosity, cognitive flexibility, and divergent think- ing. They are more receptive to novel ideas, more tolerant of uncertainty, and tend to exhibit advantages in contexts requiring creativity or conceptual reorganization. In contrast, individuals low in openness are more inclined toward tradition, conservatism, and a preference for structured and conven- tional information processing (Costa and McCrae 2008; John, Naumann, and Soto 2008; DeYoung 2015). After injecting openness-related personality features into the model, we observed pronounced performance improve- ments in generative and abstract reasoning tasks, most notably creative skill. Additionally, the model demonstrated clear enhancement in tasks involving cognitive flexibility, self- exploration, and non-normative processing, including adapt- ability, self-reflection, and expressive skill. These findings indicate that openness steering strengthens the model’s ex- ploratory orientation toward novel representations and cross- conceptual integration, closely mirroring the exploratory function associated with openness in human cognition. Conversely, moderate performance declines were observed in responsibility management, rule-following skill, and certain information processing tasks. This pattern is consistent with established findings in the personality psychology literature. Prior work suggests that high openness is associated with reduced reliance on established norms and fixed structures, and that in contexts emphasizing highly procedural execution, strict rule compliance, or single-solution optimization, the advantages of openness are less stable and may even become detrimental (McCrae 1987; DeYoung, Quilty, and Peterson 2007). Accordingly, the capability shifts induced by openness are best characterized by a tradeoff pattern in which gains in creativity and flexibility are accompanied by costs to structured execution and normative constraint adherence.