Paper deep dive
Persona-E$^2$: A Human-Grounded Dataset for Personality-Shaped Emotional Responses to Textual Events
Yuqin Yang, Haowu Zhou, Haoran Tu, Zhiwen Hui, Shiqi Yan, HaoYang Li, Dong She, Xianrong Yao, Yang Gao, Zhanpeng Jin
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/14/2026, 1:51:03 AM
Summary
Persona-E^2 is a large-scale, human-grounded dataset designed to study how individual personality traits (MBTI and Big Five) influence emotional appraisals of textual events across news, social media, and life narratives. The study addresses 'personality illusion' in LLMs, demonstrating that personality-conditioned information improves the accuracy of simulated emotional responses and that affective disagreement is a structured signal of personality rather than mere noise.
Entities (5)
Relation Signals (4)
Persona-E^2 â incorporates â MBTI
confidence 100% ¡ Persona-E^2 (Persona-Event2Emotion), a large-scale dataset grounded in annotated MBTI and Big Five traits
Persona-E^2 â incorporates â Big Five Inventory
confidence 100% ¡ incorporates the popular Myers-Briggs Type Indicator (MBTI) and the robust Big Five Inventory (BFI)
LLMs â sufferfrom â Personality Illusion
confidence 95% ¡ role-playing Large Language Models (LLMs) attempt to simulate such nuanced reactions, they often suffer from 'personality illusion'
Big Five Inventory â alleviates â Personality Illusion
confidence 90% ¡ BFI out-performs MBTI in mitigating 'personality illusion.'
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Most affective computing research treats emotion as a static property of text, focusing on the writer's sentiment while overlooking the reader's perspective. This approach ignores how individual personalities lead to diverse emotional appraisals of the same event. Although role-playing Large Language Models (LLMs) attempt to simulate such nuanced reactions, they often suffer from "personality illusion'' -- relying on surface-level stereotypes rather than authentic cognitive logic. A critical bottleneck is the absence of ground-truth human data to link personality traits to emotional shifts. To bridge the gap, we introduce Persona-E$^2$ (Persona-Event2Emotion), a large-scale dataset grounded in annotated MBTI and Big Five traits to capture reader-based emotional variations across news, social media, and life narratives. Extensive experiments reveal that state-of-the-art LLMs struggle to capture precise appraisal shifts, particularly in social media domains. Crucially, we find that personality information significantly improves comprehension, with the Big Five traits alleviating "personality illusion.'
Tags
Links
- Source: https://arxiv.org/abs/2604.09162v1
- Canonical: https://arxiv.org/abs/2604.09162v1
Trouble viewing inline? Open PDF directly â
Full Text
106,806 characters extracted from source content.
Expand or collapse full text
Persona-E 2 : A Human-Grounded Dataset for Personality-Shaped Emotional Responses to Textual Events Yuqin YangHaowu ZhouHaoran TuZhiwen HuiShiqi Yan HaoYang LiDong SheXianrong YaoYang GaoZhanpeng Jin â School of Future Technology South China University of Technology, Guangzhou, China ftyuqin_yang, 202364870491, 202330691461, 202364870731@mail.scut.edu.cn 202364871202, ftlhy, ftdshe, ftxryao@mail.scut.edu.cn gaoyang2025, zjin â @scut.edu.cn Abstract Most affective computing research treats emo- tion as a static property of text, focusing on the writerâs sentiment while overlooking the readerâs perspective. This approach ignores how individual personalities lead to diverse emotional appraisals of the same event. Al- though role-playing Large Language Models (LLMs) attempt to simulate such nuanced re- actions, they often suffer from âpersonality illusionâârelying on surface-level stereotypes rather than authentic cognitive logic. A criti- cal bottleneck is the absence of ground-truth human data to link personality traits to emo- tional shifts. To bridge the gap, we introduce Persona-E 2 (Persona-Event2Emotion), a large- scale dataset grounded in annotated MBTI and Big Five traits to capture reader-based emo- tional variations across news, social media, and life narratives. Extensive experiments reveal that state-of-the-art LLMs struggle to capture precise appraisal shifts, particularly in social media domains. Crucially, we find that person- ality information significantly improves com- prehension, with the Big Five traits alleviating âpersonality illusion.â 1 Introduction âTwo individuals can construe their situations quite similarly (agree on all the facts), and yet react with very different emotions, because they have appraised the adaptational significance of those facts differently.â (Lazarus, 1991) The study of affective appraisal of events has long been central to affective computing and cog- nitive psychology (Plaza-del Arco et al., 2024). While appraisal theories suggest that emotions emerge through individualized appraisals shaped by goals and dispositions (Lazarus, 1991; Scherer and Wallbott, 1994), NLP research has largely fo- cused on writer-expressed sentiments and reader- based unified emotional labels (Plaza-del Arco et al., 2024). This focus overlooks reader-based nuanced perception (Buechel and Hahn, 2017b), which is critical for applications, including empa- thetic agents, mental health support, and personal- ized AI assistants, that must not only process the texts but also reason about how different individu- als appraise the same event diversely. Recent interest in role-playing LLMs aims to simulate individualized reactions by injecting rich personality profiles into prompts (Tseng et al., 2024; Chen et al., 2024; Hu and Collier, 2024; Mao et al., 2024). Despite this promise, these methods often exhibit âpersonality illusionâ (Han et al., 2025): models tend to imitate stereotypical behaviors rather than adopting the cognitive ap- praisal patterns based on personality. Crucially, LLM-generated labels lack grounding in authentic feedback (Li et al., 2025a), making them insuffi- cient for evaluating whether models truly capture emotional diversity (Samuel et al., 2025). Thus, the field still lacks a human-grounded dataset to vali- date and enhance personality-conditioned emotion elicitation. To address the gap, we introduce a novel dataset, Persona-E 2 (Persona-Event2Emotion), which in- corporates the popular Myers-Briggs Type Indica- tor (MBTI) (Myers et al., 1962; John et al., 1991) and the robust Big Five Inventory (BFI) (John et al., 2010) traits into reader-based emotion labeling. As shown in Fig. 1 by engaging annotators with as- sessed personality profiles to label events across diverse domains (News, Social Media, Life Experi- ence narratives), Persona-E 2 enables a controlled analysis of the personality effect on the appraisals of identical textual events (Troiano et al., 2023). Notably, unlike previous corpora, Persona-E 2 pri- oritizes annotation density (36 labels per event) to capture diverse, trait-shaped responses (Tab. 1). To evaluate the utility of Persona-E 2 , we address three key research questions, through the experi- mental design in Sec. 5: arXiv:2604.09162v1 [cs.CL] 10 Apr 2026 News Media Event2Emo Data Collection Data process Life Evaluation Multi- Dimensional LLM Scoring Content Safety Filtering Expert Verification Persona-E 2 RQ1: Affective Divergence RQ2: LLM Simulation RQ3: Cognitive Soundness Persona Group Figure 1: Overview of the Persona-E 2 framework. Events from three domains undergo multi-stage data processing. High-quality stimuli are then annotated by a Persona Group, serving to evaluate three research questions. â˘RQ1. Affective Divergence: How do emo- tional responses diverge across the General Writer, General Reader, and Persona Reader, and how is this variance modulated by source domain and personality traits? â˘RQ2. LLM Simulation: Can LLMs effec- tively simulate Persona Reader responses, par- ticularly when faced with elicitation conflicts? â˘RQ3. Cognitive Soundness: Do LLMs gen- erate psychologically grounded rationales for their predictions, and what methods can en- hance their cognitive validity? Our analysis reveals that affective appraisal is a domain-sensitive process, with disagreement serv- ing as a structured personality signal. While LLMs struggle to predict precise appraisal shifts, partic- ularly in social media domains, personality traits improve LLMsâ comprehension, and BFI outper- forms MBTI in mitigating âpersonality illusion.â Finally, we release the dataset to support commu- nity development. 2 Related Work Extended discussions are provided in Appendix A. 2.1 Event-Elicited Emotion Analysis Early research established the baseline for un- derstanding emotions elicited by events. Classic works like ISEAR (Scherer and Wallbott, 1994), SocialIQA (Sap et al., 2019b) and others (Rashkin et al., 2018; Troiano et al., 2019; Forbes et al., 2020) analyzed first-person narratives and social commonsense, treating events as primitive stimuli for affective responses. Subsequent studies intro- duced appraisal theory to interpret these cognitive layers in depth (Troiano et al., 2022, 2023). Cru- cially, the field is shifting from writer-expressed sentiment to reader-based perception (Buechel and Hahn, 2017b). Benchmarks such as GoodNew- sEveryone (Bostan et al., 2020), iNews (Hu and Collier, 2025) and RESEMO (Hu et al., 2024a) focus on how audiences react to news and social media. However, most existing resources rely on aggregating annotations into a single ground truth, which obscures the inter-individual variability es- sential for understanding diverse emotional elicita- tion (Plank, 2022; Soni et al., 2024). 2.2 Personality-Conditioned Affective Computing Research on personalityâemotion interaction typ- ically utilizes the MBTI (Myers et al., 1962) and the BFI (John et al., 2010) via three paradigms. Explicit methods link self-reported traits to text or dialogue, as seen in datasets like PAN- DORA (Gjurkovi Ě c et al., 2021), and Person- aTAB (Inoue et al., 2025), though they primar- ily capture writer expression rather than reader elicitation. Implicit methods infer traits from be- havioral data but often lack ground truth (Gao et al., 2013; Wang et al., 2024; Hu et al., 2024b; Shen et al., 2025). Recently, LLM-based simula- tion has emerged to generate persona-specific re- sponses (Tseng et al., 2024), such as Big5-Chat (Li et al., 2025a), PersonaGym (Samuel et al., 2025), and PersonalityEdit (Mao et al., 2024). Stud- ies show that as richer prompts with profiles are introduced, the behavioral fidelity of simulated DatasetYear#Events#AnnotationsPerspective#EmotionsPersonality News-based Domain GoodNewsEveryone (Bostan et al., 2020)20205,00015,000Writer + Reader15â NewsMTSC (Hamborg and Donnay, 2021)202111,02956,000Writer7â iNews (Hu and Collier, 2025)20252,89914,550Reader6â Social Media Domain SemEval-2018 Task 1 (Mohammad et al., 2018)201822,000700,000Writer4â GoEmotions (Demszky et al., 2020)202058,000118,000Reader27â SMP2020-EWECT (BrownSweater, 2020)202034,76836,374Writer6â SenWave (Yang et al., 2025b)202510,00020,000Writer10â Life Experience Domain ISEAR (Scherer and Wallbott, 1994)19947,6667,666Writer7â Event2Mind (Rashkin et al., 2018)201824,71657,000Writer + ReaderOVâ ATOMIC (Sap et al., 2019a)201924,00072,000ExperiencerOVâ EmpatheticDialogues (Rashkin et al., 2019)201924,85024,850Writer32â Social IQA (Sap et al., 2019b)201937,58837,588ReaderOVâ Crowd-enVENT (Troiano et al., 2023)20236,59111,091Writer + Reader13â Cross-domain Integration (Ours) Persona-E 2 20263,111111,996Reader7â Table 1: A unified comparison of emotion-annotated datasets across three sources: News, Social Media, and Life Experience. Note: OV: Open vocabulary. agents improves accordingly (Bai et al., 2025; Hu and Collier, 2024). Despite their promise, recent works indicate a âpersonality illusionâ (Han et al., 2025) where models mimic linguistic styles with- out adopting the underlying appraisal mechanisms. This highlights a critical gap: the lack of a human- grounded dataset to rigorously evaluate whether LLMs truly capture trait-driven emotional diver- sity. 3 Persona-E 2 Dataset Construction To construct a rigorously controlled dataset for reader-based emotion elicitation, we designed a pipeline integrating heterogeneous event sourcing and a multi-stage filtering process. 3.1 Event Sources To ensure affective variety and broad coverage, we gather events from three complementary do- mainsânews, social media, and life experience narrativesâcovering both digital-world and real- world contexts (Appendix B.1). These sources in- clude two distinct elicitation modes: a) First-person projection, where personal experience drives af- fective memory, and b) Third-person observation, involving detached, personality-shaped appraisals. This constitutes a large-scale reader-centered emo- tion dataset that integrates the diversity of source domains and elicitation modes. News We crawled factual reports from main- stream news websites, trending topics and verified institutional accounts. These well-structured texts provide socially significant events that elicit emo- tions from third-person perspective observations. Social Media We collected posts from several public channels on Reddit. Social content brings greater topical breadth and more ambiguous con- text, which may elicit empathy or judgment from annotators. To capture more social norms and in- terpersonal dynamics, we also introduced a subset of events from Social Chemistry 101 (Forbes et al., 2020) to ensure sufficient emotion-eliciting stimuli. Life ExperienceLife experience narratives were obtained from specific channels dedicated to expe- rience sharing. These events focus on the quotidian experiences of everyday life, ranging from minor frustrations to moments of gratitude. Such narra- tives are designed to invite first-person projection, serving as a counterpart to the detached perspective of news. 3.2 Event Filtering Pipeline Raw collections inevitably contain noise, safety risks, and large quantities of content that lack emo- tional significance. To ensure high-quality stimuli, we implemented a 3-stage filtering procedure. Stage 1: Content Safety Filtering. For life experience and social media domains, we first prune toxic or sensitive content using NSFW clas- sifiers (Albouzidi, 2023; TostAI, 2023). More de- tailed information is illustrated in Appendix B.2.1. N e w s S o c i a l M e d i a L i f e E x p e r i e n c e Persona-E² 112k News International (98) Environment (56) Health (193) Education (75) Economy (84) Technology (42) Social (905) Entertainment (42) Sports (21) Social Media Emotional (369) Info Sharing (7) Discussion (71) Humor (30) Self-Record (379) Life Experience Interpersonal (309) Norm Trans (79) Reputation (23) Routine Daily (195) Consequences (133) ShortMedium Long 0 250 500 750 1000 1250 Count Text Length 1206 143 167 News ShortMedium Long 0 100 200 300 400 500 462 84 310 Social Media ShortMedium Long 0 200 400 600 84 511 144 Life Experience Figure 2: Hierarchical composition of the Persona-E 2 dataset. The sunburst chart illustrates three levels from inner to outer: original data sources, three primary domains, and fine-grained semantic subcategories. Note: Text Length: Short (0â30), Medium (30â100), Long (100â300). Stage 2: Multi-Dimensional LLM Scoring.We utilize QWEN3-MAX (Team, 2025) to translate non-English materials into English and filter out events that lack emotional significance. The final weighted score is computed as: Score = 0.35V + 0.30A + 0.20R + 0.15I (1) Here,V,A,I, andRrepresent personality vari- ability, emotional arousal, emotional implicit- ness, and source relevance, respectively (see Ap- pendix B.2.2). Applying source-specific thresholds (Appendix B.2.3), we selected 6,348 candidates from an initial pool of 76,773 events. Stage 3: Expert Verification. Finally, a 5- member expert panel conducted a rigorous au- ditâremoving factual errors, translation bias, and hate speechâyielding a set of 3,111 events. En- glish examples are provided in Appendix B.4. 3.3 Annotation Protocol As shown in Tab. 1, we adopted Ekmanâs six emotions (disgust, fear, anger, sadness, surprise, joy) plus neutral (Ekman, 1992). Following the ISEAR (Scherer and Wallbott, 1994), we used a reader-centric question: âHow would you feel when reading this event?", to capture elicitation rather than semantics. During this process, no role- playing was involved for human annotators so that each data point is anchored in an authentic persona. Annotation Unit We define an annotation unit as a reader-centric emotional response to a single textual event conditioned on a specific personality profile. Each unit is a tuple integrating the event context, a unique annotator ID, the annotatorâs mea- sured trait scores, and the corresponding emotion label. Labeling Process As shown in Tab. 1, we adopted Ekmanâs six emotions (disgust, fear, anger, sadness, surprise, joy) plus neutral (Ekman, 1992). Following the ISEAR (Scherer and Wallbott, 1994), we adopted the question style: âHow would you feel when reading this event?", to capture elicita- tion rather than semantics. Annotator Recruitment.We recruited 36 anno- tators, profiling their personalities via MBTI (My- ers et al., 1962) and BFI (John et al., 2010) ques- tionnaires (Appendix C.3). Crucially, the anno- tation process involved no role-playing. Conse- quently, annotators were strictly instructed to re- port their genuine emotional reactions for each of the 3111 events. Quality Control. To ensure high-fidelity re- sponses, we implemented reader-centric training, mandatory guideline review and behavior moni- toring (Appendix C.2). The annotation task was distributed over several weeks to maintain annota- tor attention and label quality. 4 Dataset Analysis 4.1 Descriptive Analysis Persona-E 2 comprises 112k high-quality annota- tions across three domains: news (49%), social me- dia (27%), and life experiences (24%). As shown 0100200 45 (10.2%) 14 (3.4%) 2 (7.7%) 110 (21.7%) 6 (9.1%) 9 (6.4%) 54 (19.6%) 133 (27.8%) 5 (13.5%) 49 (15.6%) 8 (22.2%) 21 (8.5%) 12 (9.2%) 43 (9.8%) 75 (18.0%) 0 (0.0%) 28 (5.5%) 0 (0.0%) 1 (0.7%) 18 (6.6%) 36 (7.5%) 2 (5.4%) 9 (2.9%) 2 (5.6%) 9 (3.6%) 2 (1.5%) Reddit SocialChemistry IUTB FMylife RAK BenignExistence The Paper Weibo Trending WeChat Article Today's Headlines Independent News BBC News ABC News Count Disgust Writer Reader 050100150 122 (27.7%) 9 (2.2%) 5 (19.2%) 47 (9.3%) 11 (16.7%) 15 (10.7%) 33 (12.0%) 36 (7.5%) 1 (2.7%) 29 (9.2%) 6 (16.7%) 76 (30.6%) 32 (24.6%) 13 (3.0%) 0 (0.0%) 1 (3.8%) 23 (4.5%) 0 (0.0%) 1 (0.7%) 19 (7.0%) 29 (6.1%) 1 (2.7%) 37 (11.8%) 4 (11.1%) 60 (24.2%) 31 (23.8%) Fear 0100200 61 (13.9%) 88 (21.2%) 1 (3.8%) 81 (16.0%) 10 (15.2%) 4 (2.9%) 30 (10.9%) 53 (11.1%) 5 (13.5%) 18 (5.7%) 4 (11.1%) 34 (13.7%) 15 (11.5%) 57 (13.0%) 140 (33.7%) 0 (0.0%) 110 (21.7%) 1 (1.5%) 0 (0.0%) 81 (29.7%) 105 (22.0%) 3 (8.1%) 17 (5.4%) 2 (5.6%) 4 (1.6%) 6 (4.6%) Anger 0100200 98 (22.3%) 53 (12.7%) 4 (15.4%) 82 (16.2%) 11 (16.7%) 14 (10.0%) 67 (24.4%) 137 (28.7%) 8 (21.6%) 56 (17.8%) 5 (13.9%) 43 (17.3%) 27 (20.8%) 157 (35.7%) 170 (40.9%) 1 (3.8%) 140 (27.6%) 3 (4.5%) 5 (3.6%) 41 (15.0%) 99 (20.7%) 5 (13.5%) 42 (13.4%) 6 (16.7%) 31 (12.5%) 19 (14.6%) Sadness 0100200300 51 (11.6%) 198 (47.6%) 4 (15.4%) 85 (16.8%) 13 (19.7%) 44 (31.4%) 79 (28.7%) 101 (21.1%) 7 (18.9%) 142 (45.2%) 11 (30.6%) 58 (23.4%) 41 (31.5%) 90 (20.5%) 29 (7.0%) 6 (23.1%) 41 (8.1%) 1 (1.5%) 9 (6.4%) 61 (22.3%) 77 (16.1%) 9 (24.3%) 101 (32.2%) 13 (36.1%) 107 (43.1%) 52 (40.0%) Neutral 050100150 50 (11.4%) 24 (5.8%) 8 (30.8%) 47 (9.3%) 7 (10.6%) 28 (20.0%) 3 (1.1%) 7 (1.5%) 6 (16.2%) 7 (2.2%) 0 (0.0%) 8 (3.2%) 0 (0.0%) 62 (14.1%) 1 (0.2%) 0 (0.0%) 117 (23.1%) 1 (1.5%) 5 (3.6%) 44 (16.1%) 107 (22.4%) 15 (40.5%) 75 (23.9%) 6 (16.7%) 24 (9.7%) 14 (10.8%) Surprise 050100150 13 (3.0%) 30 (7.2%) 2 (7.7%) 55 (10.8%) 8 (12.1%) 26 (18.6%) 9 (3.3%) 11 (2.3%) 5 (13.5%) 13 (4.1%) 2 (5.6%) 8 (3.2%) 3 (2.3%) 18 (4.1%) 1 (0.2%) 18 (69.2%) 48 (9.5%) 60 (90.9%) 119 (85.0%) 9 (3.3%) 25 (5.2%) 2 (5.4%) 33 (10.5%) 3 (8.3%) 13 (5.2%) 6 (4.6%) Joy Figure 3: Comparison of emotion distributions between Writers and Readers showing the percentage distribution of seven emotions across sources . Note: IUTB, ASK stand for r/iusedtobelieve and RandomActsofKindness. in Fig. 3, a significant emotional divergence exists between the writersâ sentiment (Hartmann, 2022) and the readersâ actual emotions. This divergence is domain-dependent: slight in factual news but pro- nounced in life experiences, where interpersonal narratives trigger first-person projection and emo- tional transmission. Moreover, the observed shift from writer to reader emotion leads us to a deeper discussion of RQ1 (Sec. 5.2). 4.2 Annotation Reliability In affective computing, inter-annotator disagree- ment is often dismissed as label noise (Plank, 2022). Moving beyond this, we report the annotation reli- ability by testing the personality-aware agreement gap. This shifts the focus from universal consen- sus to trait-conditioned alignment, grounded in the premise that subjective disagreement is structured by latent personality profiles. Thus, annotators with similar personalities should exhibit higher consen- sus than random groupings. The validation of this hypothesis via BFI and MBTI clustering ensures the datasetâs reliability. In-Group Agreement To validate the hypothe- sis, we applied K-means clustering on BFI vectors. Unlike MBTIâs discrete categories, this method adapts to continuous traits and ensures balanced clusters, addressing the statistical instability that arises from MBTIâs sparse subgroups. The choice ofk = 6was empirical, aimed at balancing the number of annotators per cluster with the captured personality diversity. To ensure the robustness of our groupings, we conducted a sensitivity analysis across various algorithms (K-means, GMM, Hier- archical) and cluster counts (k â [3, 9]). As de- tailed in Appendix E.1, the Personality Agreement Gap (PAG) consistently remains positive across all settings, confirming that trait-aligned grouping captures shared interpretative logic rather than clus- tering artifacts. As shown in Fig. 4, in-group Top-1 agreement consistently outperforms the global average. More specifically, Cluster 0 achieves a +11.5% gain on Top-1 agreement over the baseline. This confirms that grouping by traits uncovers shared interpreta- tive logic. Top1Top2Top3Top4Top5Top6Top7 Top-K Emotions Cluster 0 Cluster 1 Cluster 2 Cluster 3 Cluster 4 Cluster 5 Overall Groups 58.929.411.70.00.00.00.0 56.026.314.92.80.00.00.0 56.824.712.25.01.20.10.0 49.024.614.27.83.50.80.1 55.125.512.85.61.00.00.0 54.525.913.95.10.60.00.0 47.423.514.08.34.41.90.5 0 20 40 60 80 100 Figure 4: Top-Kemotion distribution across BFI clus- ters. The heatmap shows the average vote share (%) of the k-th most frequent emotion per event. Personality Grouping Analysis As shown in Tab. 2, we quantify the personality effect by calcu- lating the Personality Agreement Gap (PAG = Agr in âAgr out ), representing the Top-1 agreement delta between In-Group and Out-Group pairs with equalized sample sizes. ClusterAgr in Agr out PAG Cluster 0 61.07 35.11 +25.96 Cluster 1 57.63 37.59 +20.03 Cluster 2 57.26 43.39 +13.87 Cluster 3 49.39 41.12 +8.27 Cluster 4 54.73 40.38 +14.36 Cluster 5 54.76 37.11 +17.65 (a) BFI K-means (K = 6) ClusterAgr in Agr out PAG Cluster 0 58.05 37.63 +20.42 Cluster 1 56.32 38.13 +18.19 Cluster 2 51.05 45.15+5.9 Cluster 3 49.39 41.37 +8.02 (b) BFI K-means (K = 4) TypeAgr in Agr out PAG ESTP 64.33 37.35 +26.98 INTP 60.03 36.83 +23.20 INTJ 61.53 37.58 +23.95 ESTJ 56.02 38.39 +17.63 ENTJ 55.30 35.84 +19.46 ISTJ 50.83 41.16 +9.68 (c) MBTI Grouping (N ⼠3) GroupNOpen.Cons.Extra.Agree.Neuro. Cluster 03 0.8230.8680.6910.622 0.358 Cluster 14 0.718 0.536 0.482 0.797 0.638 Cluster 27 0.553 0.647 0.629 0.673 0.435 Cluster 311 0.695 0.828 0.604 0.8010.206 Cluster 46 0.556 0.569 0.455 0.580 0.592 Cluster 55 0.612 0.763 0.494 0.673 0.335 (d) BFI Mean Features (K = 6) GroupNOpen.Cons.Extra.Agree.Neuro. Cluster 04 0.8150.8360.6280.636 0.336 Cluster 15 0.679 0.517 0.429 0.629 0.729 Cluster 216 0.560 0.661 0.559 0.682 0.432 Cluster 311 0.695 0.828 0.604 0.8010.206 (e) BFI Mean Features (K = 4) Table 2: Personality-based Top-1 agreement analysis. Top: Comparison of In-Group (Agr in ) and Out-Group (Agr out ) agreement levels (Ncontrolled). Bottom: Mean BFI traits for clusters. Note: Open., Cons., Extra., Agree., and Neuro. represent the Big Five traits; N denotes the number of annotators. â˘BFI Grouping: Cluster 0 (High Conscien- tiousness/Openness) shows a massive PAG of +25.96% compared to the out-group (Tab. 2a), indicating a convergent appraisal pattern. Con- versely, Cluster 3 (Low Neuroticism) exhibits the smallest PAG (+8.3%). This may sug- gest that traits like high Neuroticism act as strict âperceptual filtersâ that funnel reactions into specific categories. Consistent patterns are also revealed across differentKsettings (Tab. 2b). ⢠MBTI Grouping: Similar patterns appear in MBTI types (Tab. 2c). ESTP types achieve peak PAG (+26.98%) by prioritizing social cues (Pickett et al., 2004), while ISTJ yields lower consensus. Positive PAG values reveal that in-group agreement consistently exceeds out-group agreement, provid- ing empirical evidence that affective disagreement is structured by trait-aligned patterns. 5 Experiment 5.1 Experimental Setup We designed experiments to assess the necessity of personality modeling in emotion analysis and the capability of LLMs to simulate these affective shifts. To evaluate performance for RQ2 and RQ3, we employed leading open and closed-source mod- els. Model List:GPT-5.1 (OpenAI, 2025), LLAMA- 3-8B (Grattafiori et al., 2024), QWEN3-8B (Yang et al., 2025a), GEMMA-3-12B (Team et al., 2025), and MINISTRAL-3-8B (AI, 2025). Computing Environment:Open-source models are evaluated on a cloud computing platform using NVIDIA A100 GPUs. Closed-source models are accessed through their APIs. Appendix D presents details on model version and hyper-parameters. 5.2 RQ1. Dataset Affective Divergence How do emotional responses diverge across the General Writer, General Reader, and Persona Reader? Fig. 3 illustrates a gap between writer and reader-based emotions, which we hypothesize is modulated by domain and personality-driven ap- praisal patterns (Buechel and Hahn, 2017b). To validate the assumption, we differentiated three af- fective layers: (1) General Writer (GW): semantic sentiment predicted by a pre-trained emotion clas- sifier (Hartmann, 2022) that provides an identical label space to our human annotations; (2) Gen- eral Reader (GR): majority-vote elicitation; and (3) Persona Reader (PR): trait-conditioned elicitation. To quantify this divergence, we categorized seven emotions into positive (surprise, joy) and negative (disgust, fear, anger, sadness) polarities. We then computed affective transition matrices, where ele- ment(i,j)represents the probability of emotioni shifting to emotionjacross the three perspectives. Domain-Driven Divergence.Comparing seman- tic sentiment (GW) with majority-vote elicitation (GR) reveals that emotional appraisal is a domain- sensitive reconstruction rather than a direct transfer (Fig. 5). Appendix E.2.1 presents the detailed data analysis related to this finding. First, news acts as a rational buffer: it maintains a high neutral- ity transfer and emotional resonance rate (Tab. 8). Disgust Fear Anger SadnessNeutral Surprise Joy Reader Emotion Disgust Fear Anger Sadness Neutral Surprise Joy Writer Emotion 12.0%8.7%29.7%9.1%14.5%21.7%4.3% 1.9%29.7%8.5%11.8%31.1%13.2%3.8% 6.5%8.4%26.6%5.8%27.3%24.0%1.3% 4.5%8.7%12.2%42.4%16.4%13.4%2.4% 2.6%10.9%7.4%6.7%43.6%19.0%9.7% 0.0%16.0%4.0%16.0%20.0%40.0%4.0% 6.5%0.0%0.0%8.7%32.6%17.4%34.8% (a) News Disgust Fear Anger SadnessNeutral Surprise Joy Reader Emotion Disgust Fear Anger Sadness Neutral Surprise Joy Writer Emotion 23.4%0.0%18.8%20.3%25.0%12.5%0.0% 5.3%6.8%6.8%47.7%15.9%15.2%2.3% 16.9%0.0%43.5%21.4%12.3%5.8%0.0% 4.4%1.3%14.5%60.4%8.2%6.9%4.4% 21.5%0.4%27.3%32.4%13.7%3.9%0.8% 8.8%2.5%18.8%26.2%20.0%17.5%6.2% 6.2%0.0%8.3%47.9%16.7%12.5%8.3% (b) Social Media Disgust Fear Anger SadnessNeutral Surprise Joy Reader Emotion Disgust Fear Anger Sadness Neutral Surprise Joy Writer Emotion 5.5%3.9%21.3%21.3%8.7%19.7%19.7% 2.6%7.7%11.5%11.5%6.4%12.8%47.4% 6.2%3.1%30.2%18.8%4.2%22.9%14.6% 3.6%2.7%15.3%35.1%7.2%11.7%24.3% 2.1%2.7%9.6%16.4%13.0%14.4%41.8% 4.4%3.3%8.9%18.9%8.9%13.3%42.2% 3.3%1.1%7.7%16.5%2.2%22.0%47.3% 0.0 0.2 0.4 0.6 0.8 1.0 (c) Life Experience Figure 5: Affective transition matrices between General Writer and Reader. Red boxes highlight: (a) high resonance in News, (b) negative bias in Social Media, and (c) positive shift in Life Experience. Note: Negative: disgust, fear, anger, sadness; Positive: surprise, joy. This aligns with professional journalismâs role in minimizing cognitive appraisal variance through factual grounding (Lazarus, 1991; Alm, 2008). Sec- ond, social media functions as an âEmotional Black Holeâ (Baumeister et al., 2001). We observed a se- vere negativity bias, where significant portions of neutral and positive GW sentiments shift to nega- tive GR reactions (81.6% of neutral and 59.35% of positive sentiments transfer into negative in Tab. 9). This reflects a âForced Sidingâ mechanism, where ambiguity is treated as a vacuum filled by reader- side hostility (Tajfel, 1974; Baumeister et al., 2001). Conversely, life narratives trigger a âPsychological Immune Systemâ (Gilbert et al., 1998). Readers exhibit an optimism bias, filtering distress to priori- tize empathetic resonance (43.3% of negative and 56.2% of neutral sentiments transfer to positive in Tab. 10), a cognitive nuance often overlooked by traditional sentiment analysis (Matlin, 2016). Personality-Driven Modulation.We employed BFI-based clustering for analysis rather than MBTI categories, as the latter led to sparse populations in specific groups, which cannot provide reliable statistics. We found that personality significantly acts as a filter for emotional transfer (further data analysis is available in Appendix E.2.2). First, Anx- ious Empathy: Individuals with high Agreeable- ness (A) and Neuroticism (N), such as C1, show the highest neutral-to-negative transfer rates. Here, high Aâs social sensitivity is likely hijacked by the interpretation bias of high N (McCrae and Costa Jr, 1997), as evidenced in Fig. 6a. Secondly, Neg- ativity Passivation: High Conscientiousness (C) and Extraversion (E) exhibit âNegative Passiva- tion.â C0 shows the lowest negative resonance in Fig 6b. This suggests these individuals effectively regulate distress (Gross, 1998). In contrast, low C/E (e.g., C2 and C3) led to âNegative Lockingâ due to a lack of perceived control. Finally, Neutral- ization: Illustrated in Tab. 13, we found a statistical correlation between Openness (O) and the Neutral- ization Rate (r = +0.86,p = 0.027). To align with cognitive theory, high-O individuals use cog- nitive complexity to moderate emotional activation, showing a low need for cognitive closure (Web- ster and Kruglanski, 1994). In summary, emotional elicitation is reshaped by domain effects and per- sonality traits, highlighting the move from general sentiment analysis to persona-aware modeling. C1C3C4 0 20 40 Transfer Rate (%) 45.4 27.8 24.2 (a) Anxious Empathy C0C3C2 0 10 20 30 Transfer Rate (%) (b) Negative Passivation Figure 6: (a) Transfer rate from non-negativity to neg- ativity in the news domain, and (b) Transfer rate from negativity to non-negativity, error bar means deviation among 3 domains. 5.3 RQ2. LLM Emotion Simulation Can LLMs effectively mimic personality-shaped emotional responses? We investigated whether models can predict emotional shifts with personal- ity profiles. LLMs are conditioned on three strate- gies: General Prompt, Persona Prompt (BFI person- ality vectors (Johnson, 2014)), and Persona-CoT (Chain of Thought) (Wei et al., 2022). While general performance on 100 randomly sampled events is reported in Tab. 14, we specifi- cally focus on trait-driven shifts by constructing the Subjective Divergence Subset (SDS). SDS targets scenarios where emotional responses are clear but heavily conditioned on the annotatorâs personality. Label validity is ensured via the Group Consensus Metric SourceNewsSocial MediaLife ExperienceOverall PromptGeneral PersonaCoTGeneral PersonaCoTGeneral PersonaCoTGeneral PersonaCoT Top-1 Acc. GPT5.131.825.029.518.218.227.332.435.335.329.027.031.0 LLAMA3-8B 29.529.320.59.14.813.632.432.432.426.025.023.0 QWEN3-8B18.226.925.423.121.422.720.333.726.020.028.025.0 GEMMA3-12B18.225.027.322.727.318.235.332.429.425.028.026.0 MINISTRAL3-8B18.213.620.55.09.14.536.435.335.322.020.022.0 Top-2 Acc. GPT5.154.559.159.140.950.045.547.155.955.949.056.055.0 LLAMA3-8B47.743.945.518.219.031.850.047.144.142.039.642.0 QWEN3-8B31.333.538.631.436.439.439.239.045.134.036.041.0 GEMMA3-12B 36.447.756.827.340.945.550.052.941.239.048.049.0 MINISTRAL3-8B40.929.538.635.036.422.751.550.050.043.338.039.0 Table 3: Comparison of different models performance in subset. Note: General Prompt: Prompt without personality traits, Persona: Prompt with BFI traits, CoT: CoT-Style Prompt with BFI traits. Score (S consensus ). For a probability distribution P G of group G, consensus is defined as: S consensus (G) = 1â â P p i logp i logK (2) Thus, the SDS events satisfyS consensus (G) > Îą for all personality groupsG, totaling 413 events (Life: 87, News: 257, Social: 69) forÎą = 0.3. Details of SDS are reported in Appendix E.3. Latent Emotion Understanding.As detailed in Tab. 3, experiments on the subset reveal a signifi- cant gap between Top-1 (âź25.0%) and robust Top- 2 performance (âź45.0%). This suggests that LLMs successfully map emotional contexts to relevant semantic neighborhoods (Brown et al., 2020) but fail to pinpoint the precise label. While models capture the general affective sphere, they cannot distinguish subtle sentiment shifts, indicating more human feedback is likely required in affective train- ing mechanisms (Ouyang et al., 2022). Prompt Strategies. We observed that persona prompts and CoT prompts are not universally bene- ficial; their efficacy is modulated by model capacity. Complex prompting improved the reasoning pro- cess for larger models, but degraded it in smaller architectures. For instance, GPT-5.1 increased from 29.0% to 31.0% but LLAMA3-8B declined. We attribute this to attention dilution (Kaplan et al., 2020), where the personality profiles may over- whelm the limited context of smaller models. Domain Discrepancy.Models consistently strug- gled in the social media domain compared to others (GPT-5.1: 18.2-27.3% in social for Top-1). This performance gap suggests that current LLM train- ing corpora are biased towards structured materials, failing to process the informal dynamics of vir- tual interactions. To bridge this gap in cyber-social emotional analysis, future studies must focus on the unstructured online materials. See Appendix E.3.4 for a full data breakdown. 5.4 RQ3. Cognitive Soundness Do LLMs generate human-like rationales for the emotional activation? Beyond prediction, we evaluated the cognitive plausibility of the psycho- logical reasoning process (Zhang et al., 2024; Li et al., 2024) on all 413 SDS events. Five trained re- viewers performed a best-of-three forced-choice evaluation (Kiritchenko and Mohammad, 2016) based on: a) Persona Consistency, b) Reasoning Plausibility, and c) Emotion Specificity (Defini- tions in Appendix E.4). In Tab. 4, selections were aggregated to compute win rates for each model. Personality Comparison.Tab. 4 reveals that the choice of personality directly impacts rationale quality. The BFI strategy demonstrates better align- ment with human cognitive patterns (55.4-70.4%), ahead of both MBTI and Baseline groups, par- ticularly in consistency. For instance, GPT-5.1 achieved 68.9-78.8% win rate with the BFI prompt across metrics, while the MBTI only 13.5-16.5%. In contrast, MBTI prompts show stronger consis- tency but weaker plausibility and specificity com- pared to the baseline. This suggests that the trait- based detail of BFI provides a more robust choice for simulating nuanced appraisals compared to the binary nature of MBTI (Furnham, 1996). Model Scale. Cognitive soundness exhibits a clear dependency on model capacity. GPT-5.1 consistently achieved the highest win rates on BFI prompt, indicating that complex rationale gener- ation requires large-scale models. Among open- Model ConsistencyPlausibilitySpecificity BL MBTI BFI BL MBTI BFI BL MBTI BFI GPT5.14.716.578.813.815.870.417.613.568.9 LLAMA39.419.870.821.523.155.420.518.261.3 QWEN3 10.419.470.219.720.459.923.719.456.9 GEMMA314.817.467.821.815.163.119.515.565.0 MINISTRAL312.515.572.021.619.558.922.012.565.5 Table 4: Win-Rate comparison of LLMs in the best-of- 3 selection task. Note: BL: Baseline Prompt without personality, MBTI: MBTI Prompt, BFI: BFI Prompt. source models, GEMMA-3-12B shows greater sta- bility than smaller LLMs, maintaining a >60% pref- erence in plausibility under BFI settings. This sug- gests that robust rationale generation is a capacity- intensive task for current LLMs. A detailed analy- sis of the data is provided in Appendix E.4.4. 6 Conclusion We introduced Persona-E 2 , a human-grounded dataset for personality-conditioned emotion analy- sis. Reliability experiments demonstrate that PAG reflect personality effects. Analysis of affective appraisals reveals domain-specific patterns, such as social media acts as an âEmotional Black Hole.â We also identified distinct, personality-shaped pro- cesses in emotion elicitation. While LLMs can map semantic neighborhoods in predicting emotional shifts, they currently lack precision. Our compar- isons underscore the importance of cyber-social emotional tasks, where BFI outperformed MBTI. Finally, cognitive prompt improves personalized reasoning, especially in large-scale models. Limitations We introduce Persona-E 2 , a large-scale dataset ex- plicitly grounded in personality traits to model reader-based emotional variation. However, the dataset has several limitations. We acknowledge that the diversity of event sources is constrained, and the current version is limited to Chinese and English texts, which may not fully capture emo- tional expressions in other domains, cultures, or languages. Likewise, the limited number of annota- tors should be taken into consideration when assess- ing population effectiveness. Furthermore, cross- lingual translation introduces additional challenges: differences in phrasal connotations can shift the perceived emotional meaning. Moreover, emotion perception is inherently subjective; thus, individual differences among annotators hinder the consistent reproduction of emotion labels. Moreover, there is an inherent gap between task-oriented annotators and spontaneous real-world users. The emotion label space adopts Ekmanâs six basic emotions plus neutral as a strategic trade-off, though this categor- ical scheme is coarser than dimensional or open- vocabulary alternatives. Additionally, the âGeneral Writerâ labels rely on an external classifier (Hart- mann, 2022) due to the unavailability of original author ratings, introducing potential algorithmic bias that should be considered when interpreting writer-reader emotion gaps. Ethics Considerations Subjectivity and Diversity in Annotation We explicitly acknowledge that emotional appraisal is inherently subjective and culturally situated. Un- like traditional paradigms that seek a single ground truth, our data collection protocol respects diverse interpretations. We required annotators to report their genuine emotional reactions based on their personality profiles, ensuring the dataset captures the variance of human experience rather than en- forcing a potentially biased consensus. Annotator Welfare and Consent We recruited 36 annotators through university channels. All an- notations were conducted via a self-developed and user-friendly online crowdsourcing platform. Par- ticipants were fully informed of the research pur- pose and potential risks, the public nature of the resulting dataset, and their right to withdraw at any time. We strictly adhered to fair labor practices; compensation was calculated based on task dura- tion and complexity, ensuring that it significantly exceeded the local minimum wage. Privacy and Data Anonymization The textual data used in Persona-E 2 originates from publicly available news, social media, and personal narra- tives. To protect the privacy of the original content creators, we implemented a rigorous anonymiza- tion pipeline. All Personally Identifiable Informa- tion (PII), including real names, specific locations, and user handles, was excluded from the research. Institutional ReviewThe experimental protocol, including data collection and annotator interaction, was reviewed and approved by our Institutional Review Board prior to the studyâs commencement. RepresentationWe acknowledge distinct limita- tions in coverage. Although our dataset contains ap- proximately 3,000 diverse events, it does not fully represent all events. Our pool of 36 annotators, while providing high annotation density, represents a specific demographic (university-educated, aged 18â25) that may not perfectly reflect the global population, potentially introducing demographic biases. We also utilized LLMs for initial data fil- tering, acknowledging that this automation may introduce minor biases despite human oversight. Responsible Use and Sensitivity This dataset, to be released under a license compatible with its research-only creation purpose, must be used solely for non-commercial research. Derivatives must not be deployed in real-world applications beyond re- search prototypes, especially in commercial con- texts. This dataset contains narratives that may be sensitive to specific cultural, religious, or social contexts. We urge researchers to exercise caution when deploying models trained on this data, par- ticularly in high-stakes applications such as men- tal health support or behavioral analysis. Users must be aware that the models may reproduce the specific appraisal patterns of our annotator pool, which should not be interpreted as universal cul- tural truths. Open Access and Reproducibility To foster transparency and encourage further research in personality-aware NLP, we will release the Persona- E 2 dataset under a license that permits research use while prohibiting malicious applications. We be- lieve open access is essential for the community to scrutinize, validate, and build upon our findings responsibly. Use of AI Assistants We utilized AI assistants (e.g., CHATGPT-4O) to refine the clarity and gram- mar of the manuscript. All scientific claims, exper- imental designs, and data analyses were conducted and verified by the human authors. Societal Impact Advancing Affective ComputingOur work pro- motes a paradigm shift from writer-centric senti- ment analysis to reader-based emotional appraisal. By introducing the Persona-E 2 dataset, we encour- age the research community to move beyond static, single-label classification and explore how person- ality traits shape diverse interpretations. This tran- sition is essential for building more inclusive AI systems that respect individual differences in per- ception. Bridging Cognition and AIBy integrating cog- nitive theories with LLMs, we highlight the value of psychological grounding in NLP. Our findings demonstrate that structured personality constraints (specifically BFI) enhance model reasoning, open- ing new avenues for personalized human-computer interaction. This theoretical alignment offers a foundation for developing more empathetic agents in mental health support and personalized educa- tion. Decoupling Cyber-Social DynamicsOur analy- sis reveals that emotional elicitation varies signifi- cantly across news, social media, and life narratives. This distinction suggests that future research must decouple virtual interactions from physical-world events. Understanding distinct mechanisms like the negativity bias in social media provides action- able insights for monitoring digital sentiment and mitigating online polarization. Potential RisksWhile personalized appraisal en- hances user experience, it carries risks. The abil- ity to tailor content to specific personality profiles could be misused for targeted manipulation or to reinforce echo chambers. Furthermore, without careful constraints, models might over-generalize personality traits, leading to unintended stereotyp- ing. Researchers must prioritize safety and fairness when deploying these persona-aware systems. References Mistral AI. 2025. Ministral 3 8b instruct 2512 model card.https://huggingface.co/mistralai/ Ministral-3-8B-Instruct-2512.Accessed: 2025-12-23. Elias Albouzidi. 2023. distilbert-nsfw-text-classifier. https://huggingface.co/eliasalbouzidi/ distilbert-nsfw-text-classifier.Hugging Face model. Accessed: 2025-12-23. Ebba Cecilia Ovesdotter Alm. 2008. Affect in* text and speech. University of Illinois at Urbana-Champaign. Yuqi Bai, Tianyu Huang, Kun Sun, and Yuting Chen. 2025. Scaling law in llm simulated personality: More detailed and realistic persona profile is all you need. arXiv preprint arXiv:2510.11734. Roy F Baumeister, Ellen Bratslavsky, Catrin Finkenauer, and Kathleen D Vohs. 2001. Bad is stronger than good. Review of general psychology, 5(4):323â370. Laura Ana Maria Bostan, Evgeny Kim, and Roman Klinger. 2020. GoodNewsEveryone: A corpus of news headlines annotated with emotions, semantic roles, and reader perception. In Proceedings of the Twelfth Language Resources and Evaluation Confer- ence, pages 1554â1566, Marseille, France. European Language Resources Association. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, Sandhini Agarwal, Ariel Herbert-Voss, Gretchen Krueger, Tom Henighan, Rewon Child, Aditya Ramesh, Daniel Ziegler, Jeffrey Wu, Clemens Winter, and 12 others. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877â1901. Curran Associates, Inc. BrownSweater. 2020. Smp2020-ewect: Weibo-based emotion classification dataset. Accessed: 2025-12- 09. Sven Buechel and Udo Hahn. 2017a. EmoBank: Study- ing the impact of annotation perspective and repre- sentation format on dimensional emotion analysis. In Proceedings of the 15th Conference of the Euro- pean Chapter of the Association for Computational Linguistics: Volume 2, Short Papers, pages 578â585, Valencia, Spain. Association for Computational Lin- guistics. Sven Buechel and Udo Hahn. 2017b. Readers vs. writ- ers vs. texts: Coping with different perspectives of text understanding in emotion annotation. In Pro- ceedings of the 11th Linguistic Annotation Workshop, pages 1â12, Valencia, Spain. Association for Compu- tational Linguistics. Nuo Chen, Yan Wang, Yang Deng, and Jia Li. 2024. The oscars of ai theater: A survey on role-playing with language models. arXiv preprint arXiv:2407.11484. Bao Minh Doan Dang, Laura Oberländer, and Roman Klinger. 2021. Emotion stimulus detection in Ger- man news headlines. In Proceedings of the 17th Con- ference on Natural Language Processing (KONVENS 2021), pages 73â85, DĂźsseldorf, Germany. KON- VENS 2021 Organizers. Dorottya Demszky, Dana Movshovitz-Attias, Jeongwoo Ko, Alan Cowen, Gaurav Nemade, and Sujith Ravi. 2020. GoEmotions: A dataset of fine-grained emo- tions. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 4040â4054, Online. Association for Computational Linguistics. Keyang Ding, Chuang Fan, Yiwen Ding, Qianlong Wang, Zhiyuan Wen, Jing Li, and Ruifeng Xu. 2024. Lcsep: A large-scale chinese dataset for so- cial emotion prediction to online trending topics. IEEE Transactions on Computational Social Systems, 11(3):3362â3375. Paul Ekman. 1992. An argument for basic emotions. Cognition & emotion, 6(3-4):169â200. Maxwell Forbes, Jena D. Hwang, Vered Shwartz, Maarten Sap, and Yejin Choi. 2020. Social chem- istry 101: Learning to reason about social and moral norms. In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pages 653â670, Online. Association for Computational Linguistics. Adrian Furnham. 1996. The big five versus the big four: the relationship between the myers-briggs type indicator (mbti) and neo-pi five factor model of per- sonality. Personality and Individual Differences, 21(2):303â307. Rui Gao, Bibo Hao, Shuotian Bai, Lin Li, Ang Li, and Tingshao Zhu. 2013. Improving user profile with personality traits predicted from social media con- tent. In Proceedings of the 7th ACM Conference on Recommender Systems, RecSys â13, page 355â358, New York, NY, USA. Association for Computing Machinery. Daniel T Gilbert, Elizabeth C Pinel, Timothy D Wilson, Stephen J Blumberg, and Thalia P Wheatley. 1998. Immune neglect: a source of durability bias in affec- tive forecasting. Journal of personality and social psychology, 75(3):617. Matej Gjurkovi Ě c, Vanja Mladen Karan, Iva Vukojevi Ě c, Mihaela BoĹĄnjak, and Jan Snajder. 2021. PANDORA talks: Personality and demographics on Reddit. In Proceedings of the Ninth International Workshop on Natural Language Processing for Social Media, pages 138â152, Online. Association for Computa- tional Linguistics. Aaron Grattafiori, Abhimanyu Dubey, Abhinav Jauhri, Abhinav Pandey, Abhishek Kadian, Ahmad Al- Dahle, Aiesha Letman, Akhil Mathur, Alan Schelten, Alex Vaughan, and 1 others. 2024. The llama 3 herd of models. arXiv e-prints, pages arXivâ2407. James J. Gross. 1998. The emerging field of emotion regulation: An integrative review. Review of general psychology, 2(3):271â299. FelixHamborgandKarstenDonnay.2021. NewsMTSC: A dataset for (multi-)target-dependent sentiment classification in political news articles. In Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main Volume, pages 1663â1675, Online. Association for Computational Linguistics. Pengrui Han, Rafal Kocielnik, Peiyang Song, Ramit Debnath, Dean Mobbs, Anima Anandkumar, and R Michael Alvarez. 2025. The personality illusion: Revealing dissociation between self-reports & behav- ior in llms. arXiv preprint arXiv:2509.03730. Jochen Hartmann. 2022. Emotion english distilroberta- base.https://huggingface.co/j-hartmann/ emotion-english-distilroberta-base/. Hug- ging Face model. Accessed: 2025-12-23. Jan Hofmann, Enrica Troiano, Kai Sassenberg, and Ro- man Klinger. 2020. Appraisal theories for emotion classification in text. In Proceedings of the 28th Inter- national Conference on Computational Linguistics, pages 125â138, Barcelona, Spain (Online). Interna- tional Committee on Computational Linguistics. Chao-Chun Hsu, Sheng-Yeh Chen, Chuan-Chun Kuo, Ting-Hao Huang, and Lun-Wei Ku. 2018. Emotion- Lines: An emotion corpus of multi-party conversa- tions. In Proceedings of the Eleventh International Conference on Language Resources and Evaluation (LREC 2018), Miyazaki, Japan. European Language Resources Association (ELRA). Bo Hu, Meng Zhang, Chenfei Xie, Yuanhe Tian, Yan Song, and Zhendong Mao. 2024a. RESEMO: A benchmark Chinese dataset for studying responsive emotion from social media content. In Findings of the Association for Computational Linguistics: ACL 2024, pages 16375â16387, Bangkok, Thailand. As- sociation for Computational Linguistics. Linmei Hu, Hongyu He, Duokang Wang, Ziwang Zhao, Yingxia Shao, and Liqiang Nie. 2024b. Llm vs small model? large language model based text augmenta- tion enhanced personality detection model. Proceed- ings of the AAAI Conference on Artificial Intelligence, 38(16):18234â18242. Tiancheng Hu and Nigel Collier. 2024. Quantifying the persona effect in LLM simulations. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 10289â10307, Bangkok, Thailand. Association for Computational Linguistics. Tiancheng Hu and Nigel Collier. 2025. iNews: A mul- timodal dataset for modeling personalized affective responses to news. In Proceedings of the 63rd An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 25000â 25040, Vienna, Austria. Association for Computa- tional Linguistics. S. Inoue, S. Wang, and H. Li. 2025. PersonaTAB: Pre- dicting personality traits using textual, acoustic, and behavioral cues in fully-duplex speech dialogs. In Proceedings of Interspeech 2025, pages 181â185. Oliver P John, Eileen M Donahue, and Robert L Kentle. 1991. Big five inventory. Journal of personality and social psychology. Oliver P John, Richard W Robins, and Lawrence A Pervin. 2010. Handbook of personality: Theory and research. Guilford Press. John A. Johnson. 2014. Measuring thirty facets of the five factor model with a 120-item public domain in- ventory: Development of the ipip-neo-120. Journal of Research in Personality, 51:78â89. Jared Kaplan, Sam McCandlish, T. J. Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeff Wu, and Dario Amodei. 2020. Scaling laws for neural language models. ArXiv, abs/2001.08361. Svetlana Kiritchenko and Saif M. Mohammad. 2016. Capturing reliable fine-grained sentiment associa- tions by crowdsourcing and bestâworst scaling. In Proceedings of the 2016 Conference of the North American Chapter of the Association for Computa- tional Linguistics: Human Language Technologies, pages 811â817, San Diego, California. Association for Computational Linguistics. Richard S Lazarus. 1991. Emotion And Adaptation. Oxford University Press. Wenkai Li, Jiarui Liu, Andy Liu, Xuhui Zhou, Mona T. Diab, and Maarten Sap. 2025a. BIG5-CHAT: Shap- ing LLM personalities through training on human- grounded data. In Proceedings of the 63rd Annual Meeting of the Association for Computational Lin- guistics (Volume 1: Long Papers), pages 20434â 20471, Vienna, Austria. Association for Computa- tional Linguistics. Xingxuan Li, Yutong Li, Lin Qiu, Shafiq Joty, and Li- dong Bing. 2024. Evaluating psychological safety of large language models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Lan- guage Processing, pages 1826â1843, Miami, Florida, USA. Association for Computational Linguistics. Zheng Li, Sujian Li, Dawei Zhu, Qilong Ma, and Weimin Xiong. 2025b. EERPD: Leveraging emo- tion and emotion regulation for improving person- ality detection. In Proceedings of the 31st Inter- national Conference on Computational Linguistics, pages 7721â7734, Abu Dhabi, UAE. Association for Computational Linguistics. Shengyu Mao, Xiaohan Wang, Mengru Wang, Yong Jiang, Pengjun Xie, Fei Huang, and Ningyu Zhang. 2024. Editing personality for large language mod- els. In Natural Language Processing and Chinese Computing: 13th National CCF Conference, NLPCC 2024, Hangzhou, China, November 1â3, 2024, Pro- ceedings, Part I, page 241â254, Berlin, Heidelberg. Springer-Verlag. Margaret W Matlin. 2016. Pollyanna principle. In Cog- nitive illusions, pages 315â335. Psychology Press. Robert R McCrae and Paul T Costa Jr. 1997. Concep- tions and correlates of openness to experience. In Handbook of personality psychology, pages 825â847. Elsevier. Saif Mohammad, Felipe Bravo-Marquez, Mohammad Salameh, and Svetlana Kiritchenko. 2018. SemEval- 2018 task 1: Affect in tweets. In Proceedings of the 12th International Workshop on Semantic Evaluation, pages 1â17, New Orleans, Louisiana. Association for Computational Linguistics. Isabel Briggs Myers and 1 others. 1962. The myers- briggs type indicator, volume 34. Consulting Psy- chologists Press Palo Alto, CA. OpenAI. 2025. Gpt-5.1 instant and gpt-5.1 thinking system card addendum. Technical report, OpenAI. Accessed: 2025-12-30. Long Ouyang, Jeffrey Wu, Xu Jiang, Diogo Almeida, Carroll Wainwright, Pamela Mishkin, Chong Zhang, Sandhini Agarwal, Katarina Slama, Alex Ray, John Schulman, Jacob Hilton, Fraser Kelton, Luke Miller, Maddie Simens, Amanda Askell, Peter Welinder, Paul F Christiano, Jan Leike, and Ryan Lowe. 2022. Training language models to follow instructions with human feedback. In Advances in Neural Information Processing Systems, volume 35, pages 27730â27744. Curran Associates, Inc. Cynthia L. Pickett, Wendi L. Gardner, and Megan Knowles. 2004. Getting a cue: The need to belong and enhanced sensitivity to social cues. Personality and Social Psychology Bulletin, 30(9):1095â1107. PMID: 15359014. Barbara Plank. 2022. The âproblemâ of human label variation: On ground truth in data, modeling and evaluation. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Process- ing, pages 10671â10682, Abu Dhabi, United Arab Emirates. Association for Computational Linguistics. Flor Miriam Plaza-del Arco, Alba A. Cercas Curry, Amanda Cercas Curry, and Dirk Hovy. 2024. Emo- tion analysis in NLP: Trends, gaps and roadmap for future directions. In Proceedings of the 2024 Joint International Conference on Computational Linguis- tics, Language Resources and Evaluation (LREC- COLING 2024), pages 5696â5710, Torino, Italia. ELRA and ICCL. Hannah Rashkin, Maarten Sap, Emily Allaway, Noah A. Smith, and Yejin Choi. 2018. Event2Mind: Com- monsense inference on events, intents, and reactions. In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 463â473, Melbourne, Australia. Association for Computational Linguistics. Hannah Rashkin, Eric Michael Smith, Margaret Li, and Y-Lan Boureau. 2019. Towards empathetic open- domain conversation models: A new benchmark and dataset. In Proceedings of the 57th Annual Meet- ing of the Association for Computational Linguistics, pages 5370â5381, Florence, Italy. Association for Computational Linguistics. Carley Reardon, Sejin Paik, Ge Gao, Meet Parekh, Yan- ling Zhao, Lei Guo, Margrit Betke, and Derry Tanti Wijaya. 2022. BU-NEmo: an affective dataset of gun violence news. In Proceedings of the Thirteenth Lan- guage Resources and Evaluation Conference, pages 2507â2516, Marseille, France. European Language Resources Association. Vinay Samuel, Henry Peng Zou, Yue Zhou, Shreyas Chaudhari, Ashwin Kalyan, Tanmay Rajpurohit, Ameet Deshpande, Karthik R Narasimhan, and Vish- vak Murahari. 2025. PersonaGym: Evaluating per- sona agents and LLMs. In Findings of the Associa- tion for Computational Linguistics: EMNLP 2025, pages 6999â7022, Suzhou, China. Association for Computational Linguistics. Maarten Sap, Ronan Le Bras, Emily Allaway, Chan- dra Bhagavatula, Nicholas Lourie, Hannah Rashkin, Brendan Roof, Noah A. Smith, and Yejin Choi. 2019a. Atomic: An atlas of machine commonsense for if-then reasoning. Proceedings of the AAAI Con- ference on Artificial Intelligence, 33(01):3027â3035. Maarten Sap, Hannah Rashkin, Derek Chen, Ronan Le Bras, and Yejin Choi. 2019b. Social IQa: Com- monsense reasoning about social interactions. In Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natural Lan- guage Processing (EMNLP-IJCNLP), pages 4463â 4473, Hong Kong, China. Association for Computa- tional Linguistics. Klaus R Scherer and Harald G Wallbott. 1994. Evidence for universality and cultural variation of differential emotion response patterning. Journal of personality and social psychology, 66(2):310. Lingzhi Shen, Xiaohao Cai, Yunfei Long, Imran Raz- zak, Guanming Chen, and Shoaib Jameel. 2025. Emoperso: Enhancing personality detection with self-supervised emotion-aware modelling. In Pro- ceedings of the 34th ACM International Conference on Information and Knowledge Management, pages 2577â2587. Nikita Soni, Niranjan Balasubramanian, H. Andrew Schwartz, and Dirk Hovy. 2024. Comparing pre- trained human language models: Is it better with hu- man context as groups, individual traits, or both? In Proceedings of the 14th Workshop on Computational Approaches to Subjectivity, Sentiment, & Social Me- dia Analysis, pages 316â328, Bangkok, Thailand. Association for Computational Linguistics. Henri Tajfel. 1974. Social identity and intergroup be- haviour. Social Science Information, 13(2):65â93. Gemma Team, Aishwarya Kamath, Johan Ferret, Shreya Pathak, Nino Vieillard, Ramona Merhej, Sarah Perrin, Tatiana Matejovicova, Alexandre RamĂŠ, Morgane Rivière, and 1 others. 2025. Gemma 3 technical report. arXiv preprint arXiv:2503.19786. Qwen Team. 2025. Qwen3-max: Just scale it. TostAI.2023.nsfw-text-detection-large. https://huggingface.co/TostAI/ nsfw-text-detection-large.Hugging Face model. Accessed: 2025-12-23. Enrica Troiano, Laura Ana Maria Oberlaender, Maximil- ian Wegge, and Roman Klinger. 2022. x-enVENT: A corpus of event descriptions with experiencer-specific emotion and appraisal annotations. In Proceedings of the Thirteenth Language Resources and Evaluation Conference, pages 1365â1375, Marseille, France. Eu- ropean Language Resources Association. Enrica Troiano, Laura Oberländer, and Roman Klinger. 2023. Dimensional modeling of emotions in text with appraisal theories: Corpus creation, annotation reliability, and prediction. Computational Linguis- tics, 49(1):1â72. Enrica Troiano, Sebastian PadĂł, and Roman Klinger. 2019. Crowdsourcing and validating event-focused emotion corpora for German and English. In Pro- ceedings of the 57th Annual Meeting of the Asso- ciation for Computational Linguistics, pages 4005â 4011, Florence, Italy. Association for Computational Linguistics. Yu-Min Tseng, Yu-Chao Huang, Teng-Yun Hsiao, Wei- Lin Chen, Chao-Wei Huang, Yu Meng, and Yun- Nung Chen. 2024. Two tales of persona in LLMs: A survey of role-playing and personalization. In Find- ings of the Association for Computational Linguistics: EMNLP 2024, pages 16612â16631, Miami, Florida, USA. Association for Computational Linguistics. Quan Tu, Shilong Fan, Zihang Tian, Tianhao Shen, Shuo Shang, Xin Gao, and Rui Yan. 2024. Charac- terEval: A Chinese benchmark for role-playing con- versational agent evaluation. In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 11836â11850, Bangkok, Thailand. Association for Computational Linguistics. Yan Wang, Bo Wang, Yachao Zhao, Dongming Zhao, Xiaojia Jin, Jijun Zhang, Ruifang He, and Yuex- ian Hou. 2024. Emotion recognition in conversa- tion via dynamic personality. In Proceedings of the 2024 Joint International Conference on Computa- tional Linguistics, Language Resources and Eval- uation (LREC-COLING 2024), pages 5711â5722, Torino, Italia. ELRA and ICCL. Donna M Webster and Arie W Kruglanski. 1994. Indi- vidual differences in need for cognitive closure. Jour- nal of personality and social psychology, 67(6):1049. Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Fei Xia, Ed Chi, Quoc V Le, Denny Zhou, and 1 others. 2022. Chain-of-thought prompting elic- its reasoning in large language models. Advances in neural information processing systems, 35:24824â 24837. Xuecheng Wu, Heli Sun, Junxiao Xue, Jiayu Nie, Xi- angyan Kong, Ruofan Zhai, Danlei Huang, and Liang He. 2025. Towards emotion analysis in short-form videos: A large-scale dataset and baseline. In Pro- ceedings of the 2025 International Conference on Multimedia Retrieval, ICMR â25, page 1497â1506, New York, NY, USA. Association for Computing Machinery. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, and 1 others. 2025a. Qwen3 technical report. arXiv preprint arXiv:2505.09388. Qiang Yang, Xiuying Chen, Changsheng Ma, Rui Yin, Xin Gao, and Xiangliang Zhang. 2025b. Sen- wave: A fine-grained multi-language sentiment anal- ysis dataset sourced from covid-19 tweets. ArXiv, abs/2510.08214. Jintian Zhang, Xin Xu, Ningyu Zhang, Ruibo Liu, Bryan Hooi, and Shumin Deng. 2024. Exploring collaboration mechanisms for LLM agents: A social psychology view. In Proceedings of the 62nd An- nual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14544â 14607, Bangkok, Thailand. Association for Compu- tational Linguistics. A Broad Related Work Research on event-elicited emotion analysis and personality-conditioned modeling lays the founda- tion for our work. We review these two areas below to contextualize our contributions. A.1 Event-Elicited Emotion Analysis Early work established the baseline for event- driven affective analysis. Classic psychometric efforts like ISEAR (Scherer and Wallbott, 1994) and its follow-ups (Troiano et al., 2019, 2022, 2023) collected first-person narratives of life events, treating the event as the primitive stimulus for ac- tual affective responses, rather than inferring emo- tion from lexical cues alone. Subsequent datasets expanded this scope to interpersonal and social commonsense scenarios (Rashkin et al., 2018; Sap et al., 2019b; Forbes et al., 2020). Subse- quent research enriched this view by inserting a mediating cognitive layer: appraisal (Hofmann et al., 2020). Work inspired by appraisal theory (e.g., x-enVENT (Troiano et al., 2022), crowd- enVENT (Troiano et al., 2023)) refined this view by integrating appraisal variables into the analy- sis of the relationship between events and emo- tions. In short, the community introduced multi- dimensional appraisal annotations and demon- strated that modeling appraisal variables substan- tially improves the model interpretability and pre- dictive power. A parallel line of research shifts the emphasis from the author to the audience. Datasets such as EmoBank (Buechel and Hahn, 2017a), GoodNew- sEveryone (Bostan et al., 2020), GERSTI (Dang et al., 2021) and subsequent news corpora began to highlight that readers construct responses based on context rather than inheriting the writerâs stance. Similarly, social media benchmarks (Hu et al., 2024a; Ding et al., 2024) approximate reader affect via comments and reactions. Multimodal efforts (BU-NEmo (Reardon et al., 2022), eMotions (Wu et al., 2025), iNews (Hu and Collier, 2025)) fur- ther showed that visual context and reader profiles modulate affective responses. These works mark a crucial shift from âwhat does the text express?â to âhow does the text make people feel?â Despite these advances, most reader-based re- sources rely on aggregated labels (e.g., majority vote) or weak signals (e.g., clicks, emojis), which obscure inter-individual variability. Furthermore, while they cover specific domains, few datasets span multiple event sources to study generalized emotional elicitation. As a result, it is necessary to build a wide-rangeing emotion elicitation dataset grounded in real human variability. A.2 Personality-shaped Affective Computing Research on personalityâemotion interaction has evolved along three main axes: explicitly mea- sured personality, implicitly inferred personality, and LLM-simulated persona profiles. Most per- sonality traits are measured using MBTI (Myers et al., 1962) or BFI (John et al., 2010). While the former is more prevalent in public discourse, the latter is more widely acknowledged in academic psychology. Explicit and Implicit Personality. Datasets uti- lizing explicit self-report or psychometric question- naires (e.g., PANDORA (Gjurkovi Ě c et al., 2021)) provide reliable grounding for linking traits to emo- tional tendencies. By aligning social media activity with personality traits, multimodal and conversa- tional resources such as PersonaTAB (Inoue et al., 2025) and EmotionLines (Hsu et al., 2018) have ex- tended this paradigm to dialogues and audiovisual interactions. While valuable, these datasets pri- marily capture writer-side expressions rather than reader-side elicitation. Conversely, implicit meth- ods bypass questionnaires and derive personality from textual expressions (Gao et al., 2013; Hu et al., 2024b; Shen et al., 2025; Li et al., 2025b). While scalable, these approaches lack ground truth and often struggle to distinguish between stable traits and temporary states. LLM-Simulated Personality.Recent studies ex- plore simulating diverse reactions via persona- prompted LLMs (Mao et al., 2024; Tu et al., 2024; Chen et al., 2024; Samuel et al., 2025; Li et al., 2025a; Bai et al., 2025). While richer prompts with profiles are introduced, the behavioral fidelity of simulated agents improves accordingly (Bai et al., 2025; Hu and Collier, 2024). However, emerg- ing evidence suggests a âpersonality illusionâ (Han et al., 2025): models often mimic linguistic styles (e.g., sounding âangryâ) rather than adopting the underlying appraisal mechanisms. Crucially, the field lacks a dataset that grounds these simulations in real human data, preventing rigorous verification of whether models truly capture trait-driven emo- tional diversity and reason about the underlying psychological causes. Consequently, no existing re- source offers a large-scale, human-grounded repos- itory while systematically capturing cross-person variation in emotional responses with multi-stage quality control. B Collection Details B.1 Collection Source To construct a diverse and comprehensive dataset, we aggregated data from twelve distinct sources spanning formal news, social media discussions, and specific life experience narratives. Below we provide brief descriptions and access links for each source: â˘ABC News: A collection of English-language breaking news and headlines, providing con- cise event summaries and titles to represent Western media perspectives. [Link] â˘BBC News: A global news feed offering com- prehensive coverage of international events, featuring headlines and brief abstracts useful for analyzing formal journalistic sentiment. [Link] â˘The Independent: A source of independent British journalism, supplying diverse news articles that contribute to the variation in edi- torial stance and topic coverage. [Link] â˘Today (Jinri Toutiao): An aggregate of trend- ing news titles from China. This source re- flects current domestic hot topics and utilizes popularity metrics to gauge public interest. We utilize a collection platform for conve- nience. [Link] â˘The Paper: A reputable Chinese digital me- dia outlet providing in-depth coverage of cur- rent affairs and historical events with a broad temporal span, enriching the dataset with long- tail topics. We utilize a collection platform for convenience. [Link] ⢠Weibo: Sourced from one of Chinaâs largest social media platforms, this subset includes both official news releases and real-time trend- ing topics. We selected an official account to collect trending topics. [Link] â˘WeChat (Weixin): A vast collection of arti- cles from WeChat Public Accounts. It covers a wide array of social topics and cultural nu- ances not always present in mainstream news. We utilize a collection platform for conve- nience. [Link] â˘Reddit: A large-scale aggregation of user- posted discussions covering diverse life scenarios-spanning from interpersonal con- flicts to moral dilemmas-from subreddits such as r/SocialAnxiety and r/PetPeeves. These posts naturally contain rich emotional under- currents. [Link] ⢠B.E. (Benign Existence): Sourced from r/BenignExistence, this subset contains non- dramatic, mundane life records. It serves as a crucial baseline for identifying objective and neutral emotional states. [Link] â˘FMylife: A collection of short, first-person narratives describing unfortunate or awkward daily moments. It provides specific scenarios for modeling negative, embarrassed, or self- deprecating affective responses. [Link] â˘IUTB (I Used To Believe): User submissions of childhood misconceptions and naive beliefs. This unique source captures scenarios evoking innocence, confusion, or nostalgia. [Link] â˘KindLife: Stories of authentic altruism col- lected from RandomActsOfKindness. This subset supplements the dataset with posi- tive emotional dimensions, such as gratitude, warmth, and admiration. [Link] B.2 Data processing B.2.1 NSFW Filter To ensure the safety and cleanliness of the dataset, we employed a strict dual-model filtration pipeline. Specifically, we utilizedDistilbert-NSFW(Al- bouzidi, 2023) andRoberta-large-NSFW(TostAI, 2023) to detect potential offensive content. A data sample was discarded if either model predicted it as "NSFW" (Not Safe For Work) with a confidence score exceeding the default threshold. This rigor- ous process minimizes the inclusion of explicit or harmful text. ⢠Distilbert-NSFW(Albouzidi,2023): eliasalbouzidi/distilbert-nsfw-text-classifier ⢠Roberta-large-NSFW (TostAI,2023): TostAI/nsfw-text-detection-large B.2.2 Multi-Dimensional LLM Scoring To curate a dataset capable of eliciting di- verse emotional responses, we employed QWEN3- MAX (Team, 2025) to score candidate texts based on their psychological âdifferential potential.â The scoring criteria vary slightly across domains to re- flect their unique characteristics. News Domain.We prioritized news content with significant societal impact and personal relevance. The scoring prompt focuses on the textâs ability to trigger divergent reactions based on reader person- ality. News Domain # SYSTEM ROLE Expert in Media Psychology and Personality. Ana- lyzes news headlines as psychological stimuli to elicit trait-dependent emotional responses and cognitive ap- praisals. # OBJECTIVE Evaluate headlines as emotional triggers. Core Crite- rion: High Differential Potentialâthe capacity to evoke divergent reactions across distinct personality profiles. # EVALUATION DIMENSIONS (Scale: 1-5) â˘1. Emotional Arousal: Intensity of the emo- tional provocation triggered by the headline. ⢠2. Personality Variability: Divergence in re- sponses based on traits (e.g., Neuroticism, Opti- mism/Pessimism, or Risk Preference). â˘3. Emotional Implicitness: Degree of subtlety in triggering emotions (implicit framing vs. ex- plicit emotional labels). â˘4. Personal Relevance: Perceived connection to the average readerâs daily life and immediate concerns. # SCORING GUIDELINES ⢠Prioritize Trait-Variance: Focus on divergence caused by sensitivity or orientation (e.g., Sensitive vs. Rational). ⢠Exclude External Biases: Disregard political, ideo- logical, or regional affiliations. # INPUT TEXT: "text" # OUTPUT FORMAT (JSON ONLY) "dim": "score": int, "reason": "str" Social Media Domain.We focused on content re- flecting digital social dynamics. The scoring mech- anism rewards texts that allow for multi-vocal inter- pretations (e.g., vague-booking or complex social signaling). Social Media Domain # SYSTEM ROLE Expert in Cyberpsychology and Personality. Analyzes social media content (e.g., status updates, comments) as projective stimuli to capture trait-bound emotional and behavioral variance. # OBJECTIVE Evaluate content as a psychological probe. Core Cri- terion: High Discriminant Validityâthe capacity to reveal personality differences through divergent inter- pretations and engagement styles. # EVALUATION DIMENSIONS (Scale: 1-5) â˘1. Emotional Arousal: Intensity of triggers (e.g., social exclusion, validation-seeking) and their susceptibility to personality modulation. â˘2. Personality Variability: Contrast in reac- tions based on Big Five (e.g., Introversion vs. Extroversion) and Attachment Styles (Anxious vs. Avoidant). â˘3. Emotional Implicitness: Reliance on subtext, irony, or "vague-booking" to provide a projective canvas for the reader. ⢠4. Social Ecological Validity: Alignment with common digital social dynamics (e.g., "seen but unreplied," social comparison, FOMO). # SCORING GUIDELINES ⢠Reward Multivocality: Prioritize texts that allow for multiple, trait-dependent interpretations (e.g., per- ceived as "bragging" vs. "inspiring"). ⢠Penalize Moral Universalism: Avoid high scores for content that triggers a uniform moral response (e.g., consensus on extreme injustice). # INPUT TEXT: "text" # OUTPUT FORMAT (JSON ONLY) "dim": "score": int, "reason": "str" Life Experience Domain. We emphasized sce- narios highly relevant to ordinary daily life, en- abling readers to project their own memories. The evaluation centers on ecological validity and emo- tional implicitness. Life Experience Domain # SYSTEM ROLE Expert in Personality and Social Psychology. Analyzes life narratives as projective stimuli to elicit differenti- ated responses based on traits and attachment styles. # OBJECTIVE Evaluate text as a psychological stimulus. Core Crite- rion: High Differential Potential (divergent responses across profiles). # EVALUATION DIMENSIONS (Scale: 1-5) ⢠1. Emotional Arousal: Intensity and trait-based variance. â˘2. Personality Variability: Response diver- gence based on Big Five and Attachment Styles. â˘3. Emotional Implicitness: Use of subtext and narrative gaps requiring projection. â˘4. Ecological Validity: Relatability to common life stressors. # SCORING GUIDELINES ⢠Prioritize Variance: Reward "Rorschach-like" texts. ⢠Penalize Uniformity: Avoid socially scripted re- sponses. # INPUT TEXT: "text" # OUTPUT FORMAT (JSON ONLY) "dim": "score": int, "reason": "str" B.2.3Source-dependent Filtration Thresholds Giventhediversenatureofourdata sourcesâranging from formal news articles to informal social media discussionsâthe noise levels vary significantly.To address this, we applied source-specific quality filtration thresholds. As detailed in Tab. 5, we set distinct cutoff values for different domains to balance data quality and retention rates. These Thresholds are empirical for the better efficiency of filtering. For instance, social media sources like Reddit generally required a more lenient threshold (4.7) compared to formal news sources (3.5) to accommodate their colloquial nature while still filtering out low-quality inputs. DomainSourceCutoffCount News ABC_news3.5130 BBC_news3.5248 Independent3.536 Today4.0314 The Paper4.0273 Weibo4.0478 WeChat4.037 Social Media SocialChem4.0416 Reddit4.7440 Life Experience B.E.4.0140 FMylife4.0507 IUTB4.026 KindLife4.066 Table 5: Source Distribution & Source-dependent Fil- tration Thresholds. Note: Independent (Independent News), Today (Todayâs Headlines), SocialChem (So- cialChemistry), B.E. (BenignExistence), KindLfe (Ran- domActsofKindness). B.3 Sub-category labels To establish a unified taxonomy for downstream analysis, we defined a closed set of labels for each domain. While news categories were derived from existing mainstream media sections, subcategories for life experiences and social media were syn- thesized via LLM summarization. The specific prompts used for classification are detailed below. In the news domain, we selected the most fre- quent categories to formulate the taxonomy, cover- ing Economics, Technology, Sports, Entertainment, Society, Health, International, Environment, and Education. News Classification Prompt # SYSTEM ROLE Expert News Editor specialized in precise content cat- egorization. Tasked with mapping news articles to a single, most relevant domain from a predefined taxon- omy. # TAXONOMY DEFINITIONS â˘Economy: Financial markets, corporate reports, macroeconomics, trade, and industry trends. â˘Technology: Internet, Artificial Intelligence, electronics, and scientific breakthroughs. ⢠Sports: Tournaments, athlete updates, team news, Olympics, and World Cup. ⢠Entertainment: Movies, music, celebrities, va- riety shows, and cultural activities. ⢠Social: Livelihood, crime, accidents, human in- terest stories, and community updates. â˘Health: Diseases, medical care, public health, wellness, and mental health. ⢠International: International relations, diplo- matic events, and global affairs. ⢠Environment: Climate change, conservation, natural disasters, and energy issues. ⢠Education: Schooling, education reform, aca- demic research, and admissions. # CLASSIFICATION RULES - Assign the single most relevant category. - If multiple domains overlap, prioritize the primary focus of the reporting. # INPUT NEWS CONTENT: "news_content" # OUTPUT CONSTRAINT Output ONLY the category name as a plain text string (e.g., Economic). DO NOT provide explanations, preambles, or addi- tional formatting. As for the Social Media domain, events are cat- egorized into Self-Recording, Emotional Expres- sion, Informational Sharing, Social Discussion, and Humor Expression, based on the communicative intent. Social Media Classification Prompt # SYSTEM ROLE Expert Social Media Content Analyst. Categorize posts based on communicative intent and narrative structure into five predefined taxonomies. # CLASSIFICATION HIERARCHY (Descending Priority) 1. Humor Expressionâ2. Social Discussionâ 3. Informational Sharingâ4. Self-Recording vs. Emotional Expression. # TAXONOMY DEFINITIONS â˘1. Self-Recording: Descriptions of personal life events, trajectories, or interpersonal interactions. Focuses on narrative (what happened). â˘2. Emotional Expression: Pure subjective vent- ing, state updates, or fragmented opinions. Fo- cuses on internal states/stances rather than event sequences. ⢠3. Informational Sharing: Fact-based content providing objective value (e.g., tutorials, guides, alerts). Characterized by high altruism. â˘4. Social Discussion: Topics transcending the personal sphere (e.g., societal phenomena, public policy, collective ethics). From "I" to "Society." ⢠5. Humor Expression: Jokes, parodies, or ironic content intended primarily to entertain. Takes precedence if the core intent is comedic. # INPUT:Title: "title"; Content: "content" # OUTPUT FORMAT (JSON ONLY) "category_id": int, "category_name": "string", "reasoning": "short_explanation" In the life experience domain, we defined five categories to classify events based on their behav- ioral nature: Interpersonal Interaction, Norm Trans- gression, Pursuit Consequences, Reputation Ap- praisal, and Routine Daily. Life Experience Classification Prompt # SYSTEM ROLE Expert Annotator for Behavioral Life Narratives. Tasked with mapping specific life experiences to a predefined five-class taxonomy based on their social- psychological core. # TAXONOMY DEFINITIONS ⢠1. Interpersonal Interaction: Focuses on the dynamics between⼠2parties (e.g., conflict, sup- port, intimacy). Priority is given to the process of interaction. â˘2. Norm Transgression: Focuses on ethical or procedural violations (e.g., dishonesty, rule- breaking, moral dilemmas). Priority is given to transgression over context. â˘3. Pursuit Consequences: Focuses on the out- comes of goal-oriented tasks (e.g., success/fail- ure in exams, career, or social attempts). Priority is given to achievement valence. â˘4. Reputation Appraisal: Focuses on social labeling and moral standing (e.g., gossip, being judged as "selfish" or "helpful"). Priority is given to public image. â˘5. Routine Daily: Focuses on mundane, low- tension activities (e.g., commuting, dining) with- out significant conflict or moral stakes. # CLASSIFICATION GUIDELINES - Categorize based on the primary narrative axis. - For interpersonal transgressions, prioritize Norm Transgression. - For goal-related embarrassments, prioritize Pursuit Consequences. # INPUT NARRATIVE: "text" # OUTPUT CONSTRAINT Output ONLY the exact category name string from the list above (e.g., Interpersonal Interaction). DO NOT include numbers, explanations, JSON, or punctuation. B.4 Dataset Examples English Examples from Persona-E 2 1. Dr Blaine McGraw is alleged to have secretly filmed intimate videos of patients in his care. 2. The teacher, Abby Zwerner, was shot in January 2023 in her classroom at Richneck Elementary School in Newport News, Virginia. 3. Aircraft âdisappeared from radar without transmit- ting distress signalâ minutes after entering Georgian airspace. 4. A woman sworn in as a city council member in Bangor, Maine, served time in prison for manslaughter. 5. A fire broke out at the venue hosting U.N. climate talks in Brazil, prompting evacuations as firefighters rushed to control the flames. 6. Today, we got back from our second honeymoon and went to pick the kids up from my momâs. Surprisingly, they were both sat quietly watching TV. Half jokingly, I asked my mom what her secret was. Without even a guilty pause she told me, Benadryl for chesty coughs in their juice. Youâre welcome. 7. Today, I went to the store for some pads with my dad. We got them and then went to the cashier. Thatâs when he realized that they were scented. He took one out of the box, sniffed it, made me sniff it, then insisted the cashier smell it. 8. Today, I told my boyfriend I wouldnât be able to get any time off work to go to Mexico with him, and that weâd have to get our tickets refunded, and reschedule. He said not to bother, and that he already had someone else in mind to take with him. 9. Today, my best friend on Snapchat is my mum. 10. My mom recently passed away and I miss her more than words could ever express. During the height of the pandemic she underwent awful, brutal rounds of chemo and never ever complained.I have a photo of her on my desk showing her true RandomActsofKindness, a day in the life of my mom. Sheâs walking into treatment wearing a mask, her cute bald head in a cute cap, with a big bag of sweet treats for the chemo nurses - to let them know how much she appreciated them! 11. Why do I always feel like Iâm going to die and run through disaster scenarios every time I speak publicly at work? 12. Has anyone else felt or seen âghostsâ? When other people donât notice? Iâve walked with friends and seen someone keeping pace, in my peripherally, on the sidewalk across the street. When I look over, no one is there. Iâve dreamt of people who passed away hours before or after they do. They never know they are dead, so I end up having to tell them. Looking into their eyes, taking in their scent for a moment more and letting them know why they feel so confused. 13. My coworker complained about being broke then showed up with a designer bag the next day 14. Bought the apartment across the street for my parentsâyet a bowl of soupâs distance has turned into yesterdayâs leftovers. 15. Is it normal to not give your roommate a heads up about a SO sleeping over for 5 day per week ? C Annotation Details C.1 Annotation Platform The annotation platform is designed for conve- nient online annotation and will be released in two months. The demonstration is shown in Fig. 7. C.2 Quality Control To ensure high-fidelity emotional annotations, we implemented a multi-layer quality control pipeline spanning annotator preparation, in-task monitor- ing, and post-hoc validation. All annotators un- derwent mandatory training before entering the task. The training clarified the central principle of this annotation scheme: annotators are required to report their stimulated emotion after reading the text, rather than infer the authorâs sentiment or the eventâs semantic polarity. When they realize their emotional reaction may diverge from what they perceive as the average response, annotators are explicitly instructed to record their genuine feel- ing, as inter-individual variability is essential to the study design. More specifically, each item is labeled with one primary emotion selected from 6 basic emotions plus a neutral class. Furthermore, the original English text is displayed alongside the translated Chinese version, allowing annotators to cross-reference and mitigate potential translation artifacts. All annotators possess advanced English reading proficiency. During annotation, a real-time monitoring sys- tem captures behavioral tracesâincluding la- tency, hesitation patterns, and abnormal repeti- tionâwhich supports continuous quality auditing and early correction of potential low-engagement behaviors. We also tracked per-annotator through- put statistics, flagged abnormal labeling patterns (such as repeated use of the same emotion label or unrealistically rapid completion), and monitored longitudinal fatigue trends. Annotators exhibiting notable deviations received explicit reminders, and labels associated with confirmed anomalous behav- ior were re-annotated. This combination of struc- tured training, reasoning-aligned instructions, and live supervision forms a robust quality assurance protocol and ensures that the final dataset reflects reliable, fine-grained reader-elicited emotional re- sponses. C.3 Annotator Profiles As shown in Tab. 6, all annotators recruited are from China, and possess advanced proficiency in English reading comprehension and written ex- pression. All 36 annotators (aged 18â25) are anonymized and indexed. Their personalities were profiled via MBTI and BFI questionnaires. MBTI denotes MyersâBriggs Type Indicator, where the four-letter codes represent Extraversion (E) / In- troversion (I), Sensing (S) / Intuition (N), Think- ing (T) / Feeling (F), and Judging (J) / Perceiv- ing (P) (Myers et al., 1962), obtained from the MBTI-93 Questionnaire. BFI scores correspond to the Big Five personality dimensions (Openness, Conscientiousness, Extraversion, Agreeableness, Neuroticism) (John et al., 1991), obtained from the IPIP-NEO-120 questionnaire (Johnson, 2014) and normalized to[0, 1]based on the theoretical minimum and maximum scores of each dimension. D Implementation Details On the cloud computing platform, the experiments for RQ2 and RQ3 required approximately 6 and 40 GPU hours on NVIDIA A100 GPUs, respectively. Emotion Annotation Platform Item 3, Total 3111 items Annotated: 2/3111 Today, I received my soccer team jacket that I ordered a month ago. Trying to save money, I'd selected the "No name" option to avoid an extra $20 embroidering fee. My jacket now has "NO NAME" spelled out on the side of it, and I was charged the extra $20 dollars after all. joyfearsurprisesadnessdisgustanger neutral Previous-100 Go to: 0Go+100 Next Save Data Locally System Logâź âź Context View Annotation Overview (0-99) 0123456 78910111213 14151617181920 21222324252627 28293031323334 35363738394041 42434445464748 49505152535455 56575859606162 63646566676869 70717273747576 77787980818283 84858687888990 91929394959697 9899 Figure 7: The annotation platform enables emotion labeling with seven categories, supporting live visualization, real-time data transmission, and comprehensive annotation monitoring. D.1 Hyper-parameters For both open-source and closed-source LLMs, we adjusted the generation hyper-parameters. For RQ2, we set temperature=0.2 to ensure stable emo- tion prediction. For RQ3, we set temperature=0.7 to facilitate reasoning about nuanced emotional reactions. Other hyper-parameters were kept at de- fault values (max_new_tokens=1024, top_p=0.9). D.2 Open-Source LLMs Versions The specific versions and Hugging Face identifiers for the open-source LLMs used in this study are listed below. ⢠Meta-Llama-3-8B(Grattafiori et al., 2024): meta-llama/Meta-Llama-3-8B-Instruct ⢠Qwen3-8B(Yangetal.,2025a): Qwen/Qwen3-8B ⢠Gemma-3-12B(Teametal.,2025): google/gemma-3-12b-it ⢠Ministral-3-8B(AI,2025): mistralai/Ministral-3-8B-Instruct-2512 E Experiment Details E.1 Stability Analysis of Personality Clustering To verify that the observed Personality Agreement Gap (PAG) is not sensitive to the specific choice of the clustering algorithm or the number of clus- ters (k), we evaluated three different methods: K- means, Gaussian Mixture Models (GMM), and Hi- erarchical Clustering. We varied k from 3 to 9. As shown in Table 7, the average PAG remains positive and significant across all configurations. E.2 RQ1. Dataset Affective Divergence (1) General Writer (GW):Semantic sentiment of General Writer is predicted by a pre-trained emo- tion classifier (Hartmann, 2022) that provides an identical label space (Ekmanâs six basic emotions plus neutral) to our human annotations, ensuring direct comparability. (2) General Reader (GR): majority-vote elicitation from all the annotations; and (3) Persona Reader (PR): trait-conditioned elicitation from each personality cluster. E.2.1 General Writer vs. General Reader For General Writer, the available writer-side classi- fiers or domain-adapted models are not sufficient for performance comparison, since the label space of most existing models is incompatible with the one utilized in this study. Results:We analyzed the affective transition matrices between the GW and GR (Buechel and Hahn, 2017b; Troiano et al., 2019). As demon- strated in Fig 5, the three domains exhibit distinct ID GNDRMBTIOpen.Cons.Extra.Agree.Neuro. E1MINTP0.6980.6460.6040.7400.417 E2MINFP0.6150.3540.4790.5520.677 E3FINFP0.7600.4580.4270.6880.635 E4MISTJ0.5000.8020.5000.7290.406 E5MINTJ0.6880.7920.5100.7810.312 E6FISTJ0.5730.6250.5620.7500.427 E7MENTP0.7710.8440.6670.5730.469 E8FISTJ0.6560.7500.4900.6350.448 E9MESFJ0.6350.6460.6670.6560.438 E10FINFJ0.7810.5730.3960.8540.844 E11FESFJ0.7600.7810.6460.8020.188 E12FESTJ0.3440.5310.6460.6040.594 E13 MESTP0.5620.6250.6250.5730.448 E14FINTJ0.8650.9060.5830.6040.396 E15FENFP0.6980.5620.5830.8750.552 E16FENFP0.6770.8440.6560.8330.198 E17 MISTJ0.6250.5520.4270.5100.844 E18FENTJ0.7600.9060.6670.7920.177 E19 MESTJ0.5000.6250.6460.7710.406 E20 MISTJ0.5000.5620.4380.6770.448 E21 MISTJ0.7290.8960.6250.8020.188 E22FESTJ0.5730.8750.5420.7920.229 E23 MESTP0.4790.6560.5520.5940.542 E24FISTJ0.5520.7710.5620.6670.240 E25 MINTJ0.7810.7810.5620.7190.146 E26 MENTJ0.5620.8330.6560.6150.312 E27 MENTJ0.8330.8540.8230.6880.208 E28 MESTP0.6460.6770.6040.7810.177 E29 MESTJ0.7710.8120.6150.8850.062 E30 MISFP0.6150.6460.4170.5420.646 E31 MISTJ0.5000.6460.4170.6040.396 E32FINTP0.7920.7400.4380.6770.271 E33 MISTJ0.6770.8540.5620.8120.385 E34FENTJ0.5830.8850.6560.8120.208 E35FINTP0.6350.5520.5210.7710.521 E36 MESTJ0.5620.7500.4790.6560.312 Table 6: Demographic and Personality Profiles of Partic- ipants. Annotators are indexed. BFI scores correspond to Openness, Conscientiousness, Extraversion, Agree- ableness, Neuroticism. Methodk=3k=4k=5k=6k=7k=8k=9 K-Means10.4513.4513.9216.4418.1020.0823.10 GMM12.0012.9514.8518.7020.2919.8622.24 Hierarchical10.0713.8715.2617.3517.4218.5719.72 Table 7: Sensitivity analysis of the Personality Agree- ment Gap (PAG) across different clustering methods and cluster counts (k). All values represent the average Top-1 agreement gain (%) over the global baseline. polarity patterns in affective contagion (resonance) and shift (transfer). Secondly, social media exhibits a sharp nega- tivity bias: 81.6% of neutral and 59.35% of posi- tive GW sentiments transfer into negative GR reac- tions, far exceeding news (27.6%, 25.6%) and life (30.8%, 32.1%) (Tab. 9). In contrast, life experi- ence narratives demonstrate a strong positive shift, where 43.3% of negative and 56.2% of neutral GW sentiments transfer to positive GR emotions, sig- nificantly higher than news (28.7%, 21.1%) and social media (4.7%, 11.8%) (Shown in Tab. 10). Insight AIn the News domain, a significant pro- portion of writer-expressed emotions transfers to neutral, while polar emotions maintain a relatively balanced resonance rate and high neutrality transfer (Tab. 8). Compared with the other domains, res- onance rate-48.10%, 43.66% and 56.60%-shows much more stability. Moreover, neutrality transfer- 22.33-43.66%- is significantly highest among 3 domains. News Writer\ ReaderPositiveNeutralNegative Positive48.1026.3025.60 Neutral28.7243.6627.62 Negative21.0722.3356.60 Table 8: Affective polarity transition matrices between General Writer and General Reader in news domain. Insight B Conversely, Social Media exhibits a sharp negativity bias: 81.60% of neutral and 59.35% of positive GW sentiments transfer into negative GR reactions, while the positive and neu- tral sentiment rate are largely less than 20%. Social Media Writer\ ReaderPositiveNeutralNegative Positive22.2518.4059.35 Neutral 4.7013.7081.60 Negative11.7815.3572.87 Table 9: Affective polarity transition matrices between General Writer and General Reader in social media domain. Insight C In contrast, life experience narratives demonstrate a strong positive shift, where 43.28% of negative and 56.20% of neutral GW tones trans- fer to positive GR emotions. Life Experience Writer\ ReaderPositiveNeutralNegative Positive62.405.5532.05 Neutral56.2013.0030.80 Negative 43.286.6550.07 Table 10: Affective polarity transition matrices between General Writer and General Reader in life experience domain. E.2.2 General Reader vs. Persona Reader Personality traits significantly modulate transfer patterns. We use the acronym O, C, E, A, N to rep- Startâ TargetPositiveNeutralNegative Cluster 1: High A, High N Positive45.45%17.17%37.37% Neutral11.72%36.72%51.56% Negative3.45%5.71%90.84% Cluster 3: High A, Low N Positive66.67%14.14%19.19% Neutral6.25%59.38%34.38% Negative1.50%6.31%92.19% Cluster 4: Low A, High N Positive77.78%3.03%19.19% Neutral19.53%52.34%28.12% Negative4.20%6.46%89.34% Table 11: Polarity Transfer Matrix for different clusters in the Social domain. Note: Grouped by cluster types. A: Agreement, N: Neuroticism. resent Openness, Conscientiousness, Extraversion, Agreeableness, and Neuroticism. Insight A In the social media domain, we ob- serve "Anxious Empathy" effect in Cluster 1 (High-A, High-N): neutral-to-negative transfer rate reaches 51.56% and positive-to-negative 37.4%, outstripping Cluster 3 (High-A, Low-N, 34.38%, 19.19%) and Cluster 4 (Low-A, High-N, 28.1%, 19.2%). Concrete data is shown in Tab.11. Insight B A Negative Passivation effect appears in Cluster 0 (High-C, High-E), which exhibits the lowest negative resonance (73.3-80.6% across do- mains) compared to the negative locking seen in Cluster 2 (Low-C, 87.3-92.5%) and Cluster 3 (Low- E, 87.3-92.5%). Concrete data is shown in Tab.12. Insight COpenness shows a positive correlation with the Neutralization Rate (r = +0.86,p = 0.027), where high-O individuals Cluster 0 transfer polar emotions to neutral at a rate of 88.2%, while low-O individuals Cluster 2 do so at only 23.92%. Concrete data is shown in Tab.13. E.3 RQ2. LLM Emotion Simulation Before the construction of the SDS, we also con- duct an experiment of randomly selected 100 events on GPT5.1. In doing so, the results are shown in Tab. 14. The experiment reveals a significantly higher performance in social media compared with Tab. 3. This suggests that GPT5.1 may fail to pre- dict the emotional shifts of the social media content, Startâ TargetPositiveNeutralNegative Cluster 0: High C, High E Positive66.57%15.88%17.55% Neutral27.98%53.28%18.73% Negative12.55%14.10%73.34% Cluster 2: Lower C Positive82.73%3.62%13.65% Neutral25.06%52.55%22.38% Negative4.65%2.82%92.52% Cluster 3: Lower E Positive71.59%17.83%10.58% Neutral11.19%79.08%9.73% Negative9.45%8.46%82.09% Table 12: Polarity Transfer Matrix for different clusters in the News domain. Note: Grouped by cluster types. C: Conscientiousness, E: Extraversion. MetricC0C1C2C3C4C5 Open.0.8230.7180.5530.6950.5560.612 Neu. Rate88.21%64.28%23.92%58.72%33.19%70.75% Table 13: Comparison of Openness scores and Neutral- ization Rates across different clusters. highlighting the critical need of more emphasis on the cyber-social gap. E.3.1 SDS construction E.3.2 Definition and Construction The SDS is constructed to capture scenarios where emotional responses are unambiguous within spe- cific personality groups but contradictory between them. We filter events based on a dual-criteria mechanism: 1.Intra-group Consistency: We retain events where specific personality groups demonstrate high internal agreement (S consensus (G) > Îą,Îą = 0..3), ensuring the emotional signal is not random noise. Sensitivity analysis forÎą has not been conducted. 2.Inter-group Divergence:Among high- consensus events, we select those where the dominant emotional labels differ significantly across distinct personality profiles. Following this process, the final SDS comprises 413 events, distributed across domains as 257 from News, 69 from Social Media, and 87 from Life Experience. E.3.3 Prompt Settings We conducted experiments using three distinct prompt strategies to evaluate the capabilities of LLM emotion prediction: General Prompt, Persona Prompt, and Persona-CoT. The specific templates and instructions for each strategy are detailed be- low. General Reader Emotion Prediction # SYSTEM ROLE You are a general observer representing an average human reaction. Your task is to read an event and report the most natural immediate emotional reaction. Do NOT assume any specific personality traits (like Neuroticism or Extraversion). # INPUT EVENT "[EVENT_DESCRIPTION]" # VALID LABELS â˘Emotions: Joy, Fear, Surprise, Sadness, Disgust, Anger, Neutral. ⢠Intensity: 1 (Very Weak) to 5 (Very Strong). # INSTRUCTION Return a JSON object containing the TOP 2 most likely emotions, ranked by confidence. # OUTPUT FORMAT Return exactly 2 emotions in descending order of con- fidence: "predictions": [ "Emotion": "Joy", "Intensity": 4, "Emotion": "Surprise", "Intensity": 3 ] Persona-based Emotion Prediction # SYSTEM ROLE You are a participant in a psychology experiment simu- lating a specific human persona. Your task is to read an event description and report your immediate emotional reaction based on your persona. # PERSONA PROFILE [PERSONA_DESCRIPTION] # INPUT EVENT "[EVENT_DESCRIPTION]" # VALID LABELS ⢠Positive Candidates: [Joy, Surprise] â˘Negative Candidates: [Sadness, Anger, Fear, Disgust, Surprise] ⢠Neutral Candidate: [Neutral] ⢠Intensity: 1 (Very Weak) to 5 (Very Strong). # INSTRUCTIONS Step 1 â Polarity assessment: Based on your personal- ity traits and the event description, decide whether the overall emotion is POSITIVE, NEGATIVE, or NEU- TRAL. Step 2 â Final selection: From the chosen category, pick the top two emotions that best match both your persona and the event. # OUTPUT FORMAT Return exactly 2 emotions in descending order of con- fidence: "predictions": [ "Emotion": "Joy", "Intensity": 4, "Emotion": "Surprise", "Intensity": 3 ] CoT Persona Emotion Prediction # SYSTEM ROLE You are a participant in a psychology experiment simu- lating a specific human persona. Your task is to read an event description and report your immediate emotional reaction based on your persona. # PERSONA PROFILE [PERSONA_DESCRIPTION] # INPUT EVENT "[EVENT_DESCRIPTION]" # VALID LABELS ⢠Positive Candidates: [Joy, Surprise] ⢠Negative Candidates: [Sadness, Anger, Fear, Disgust, Surprise] ⢠Neutral Candidate: [Neutral] ⢠Intensity: 1 (Very Weak) to 5 (Very Strong). # INSTRUCTIONS Step 1 â Polarity assessment: Based on your personal- ity traits and the event description, decide whether the overall emotion is POSITIVE, NEGATIVE, or NEU- TRAL. Step 2 â Final selection: From the category chosen, pick the top two emotions that best match both your per- sonality and the event. Report: Polarity, Top1 Emotion, and Top2 Emotion. # OUTPUT FORMAT Return exactly 2 emotions in descending order of con- fidence: "predictions": [ "Emotion": "Joy", "Intensity": 4, "Emotion": "Surprise", "Intensity": 3 ] Letâs think step by step. E.3.4 Data Analysis We also provide the data when events are randomly selected from all 3113 dataset As detailed in Tab. 3, the experiments on the Hard Set reveal distinct trends across model ca- pacities, prompt strategies, and domains discrep- ancy. First, while the average Top-1 accuracy hov- Metric NewsSocial MediaLife ExperienceOverall General Persona CoT General Persona CoT General Persona CoT General Persona CoT Top-1 Accuracy34.032.034.037.137.134.333.346.740.035.036.035.0 Top-2 Accuracy62.060.058.048.645.745.753.353.346.756.054.052.0 Table 14: Performance of GPT5.1 on the randomly selected set from the dataset. ers around 25.0%, Top-2 accuracy surges to ap- proximately 45.0%. Secondly, introducing per- sona constraints yields marginal gains for GPT- 5.1 (29.0-31.0%) but significantly boosts QWEN3- 8B (20.0-28.0%). Conversely, smaller models like LLAMA3-8B suffer slight performance degrada- tion under complex prompts. Finally, models con- sistently underperform in the Social Media domain (GPT-5.1: 18.2-27.3% for Top-1), lagging far be- hind News and Life Experience. E.4 RQ3. Cognitive Soundness We assess the cognitive plausibility of the reason- ing process with 3 different types of prompt in ran- dom order (Baseline/MBTI/BFI Prompt). For each test instance, reviewers performed a best-of-three forced-choice selection for three metrics: a) Per- sona Consistency, b) Reasoning Plausibility, and c) Emotion Specificity E.4.1 Baseline Prompt Baseline Reasoning Prompt # SYSTEM ROLE You are a cognitive psychology expert. Please analyze the connection between the text and the reported emo- tion solely based on the text content. Strictly adhere to these rules: 1) Output the analysis directly with- out fillers like "Okay"; 2) Do not output your thinking process; 3) Strictly follow the required format. # USER ROLE Please analyze the underlying reasons for the reported emotion and intensity when reading the text below based on the content of the text itself: [Event Text]: event_text [Reported Emotion]: emotion [Emotion Intensity]: intensity (1-5) Please output the analysis in the following format: ⢠Event Summary: (A one-sentence summary) ⢠Reasoning Process: (Brief analysis based on the text content) â˘Identified Cause: (Likely motivation derived from the text) E.4.2 MBTI Prompt MBTI-driven Reasoning Prompt # SYSTEM ROLE You are a psychology expert. Analyze the userâs emo- tions based on MBTI personality theory. [Current Sub- ject] real_name (MBTI Type: mbti). [Rules] 1) Output the analysis directly without fillers like "Okay"; 2) Do not output your thinking process; 3) Strictly fol- low the required output format. # USER ROLE Pleaseanalyzetheunderlyingreasonsfor real_nameâsreportedemotionandintensity when reading the text below: [Event Text]: event_text [Reported Emotion]: emotion [Emotion Intensity]: intensity (1-5) Please output the analysis in the following format: ⢠Event Summary: (A one-sentence summary) ⢠Reasoning Process: (Brief analysis combining mbti traits) ⢠Identified Cause: (Psychological motivation) E.4.3 BFI Prompt Personality-driven Reasoning Prompt # SYSTEM ROLE You are a cognitive psychology expert. Analyze the userâs emotional response based on the Big Five Per- sonality Traits (OCEAN) theory. [Score Interpretation] Normalized values from 0.0 to 1.0 (0.0: Lowest; 1.0: Highest; 0.5: Medium). [Di- mensions] 1. Openness (O): creativity, intelligence; 2. Conscientiousness (C): self-discipline, dutifulness; 3. Extraversion (E): sociability, optimism; 4. Agree- ableness (A): altruism, empathy; 5. Neuroticism (N): emotional instability, anxiety. [Current Subject] Name: real_name; Personality Scores: bfi_desc. [Rules] 1) Output analysis directly without fillers; 2) Explain how specific O/C/E/A/N dimensions influ- enced the emotion; 3) Do not output thinking process; 4) Follow the required format. # USER ROLE Pleaseanalyzetheunderlyingreasonsfor real_nameâsreportedemotionandintensity when reading the text below: [Event]: event_text [Emotion]: emotion [Intensity]: intensity (1-5) Please output the analysis in the following format: ⢠Event Summary: (One sentence summary) â˘Personality Analysis:(Explicitly mention which dimension (O/C/E/A/N) scores played a key role.) â˘Final Attribution: (Summary of psychological motivation) After generating reasoning chains across differ- ent styles, we recruited five annotators to conduct a comparative evaluation using a "best-of-3" se- lection protocol. During the review process, each reviewer was required to perform a forced-choice selection based on three criteria: a) Persona Consis- tency, b) Reasoning Plausibility, and c) Emotional Specificity. The formal definitions of these met- rics, used to evaluate the cognitive soundness of the LLM-generated reasoning, are provided below: â˘Cognitive Consistency: Assesses whether the rationale aligns with the specific values and behavioral patterns of the assigned persona. A rationale that exhibits thinking patterns di- vergent from those of the specified persona is considered of poor quality. â˘Reasoning Plausibility: Evaluates the logical coherence of the causal chain from the event to the elicited emotion. A post-hoc justifica- tion that merely reverse-engineers the given emotion is considered poor quality. â˘Emotional Specificity: Measures how pre- cisely the rationale is tailored to explain the specific target emotion. A vague rationale that could explain any emotional response is considered poor quality. E.4.4 Data Analysis Personality Comparison.We conducted experi- ments on a) Baseline prompts without personality, b) MBTI prompt, c) BFI prompt. Also, we in- troduced the cognitive appraisal process into the personality prompts. As shown in Tab. 4, the BFI prompting strategy demonstrates a leading perfor- mance across all metrics, particularly in Persona Consistency (55.4-70.4%), eclipsing the MBTI and Baseline groups. While MBTI marginally outper- forms the baseline group in consistency, it unex- pectedly underperforms in Reasoning Plausibility and Emotional Specificity. Model Comparison.GPT-5.1 consistently achieves the highest win rates across all dimen- sions. Among open-source models, performance correlates with parameter scale: GEMMA-3-12B significantly outperforms smaller LLMs , maintain- ing >60% preference in BFI settings. QWEN-8B lags behind, particularly in specificity.