Paper deep dive
Enhancing Debunking Effectiveness through LLM-based Personality Adaptation
Pietro Dell'Oglio, Alessandro Bondielli, Francesco Marcelloni, Lucia C. Passaro
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 3/13/2026, 1:03:21 AM
Summary
This paper proposes a methodology for enhancing the effectiveness of fake news debunking by using Large Language Models (LLMs) to generate personalized messages aligned with the Big Five personality traits. The approach uses persona-based prompting to tailor debunking content and employs an LLM-as-a-judge framework to evaluate persuasiveness, demonstrating that personalized messages are generally more effective than generic ones.
Entities (5)
Relation Signals (3)
LLM-as-a-Judge â evaluates â Personalized debunking messages
confidence 95% · we employ a separate LLM as an automated evaluator simulating corresponding personality traits
Large Language Models â generate â Personalized debunking messages
confidence 95% · This study proposes a novel methodology for generating personalized fake news debunking messages by prompting Large Language Models
Big Five personality traits â moderates â Persuasiveness
confidence 90% · If debunking messages could be tailored to resonate with an individualâs psychological profile, their persuasive impact and reception could be enhanced.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:This study proposes a novel methodology for generating personalized fake news debunking messages by prompting Large Language Models (LLMs) with persona-based inputs aligned to the Big Five personality traits: Extraversion, Agreeableness, Conscientiousness, Neuroticism, and Openness. Our approach guides LLMs to transform generic debunking content into personalized versions tailored to specific personality profiles. To assess the effectiveness of these transformations, we employ a separate LLM as an automated evaluator simulating corresponding personality traits, thereby eliminating the need for costly human evaluation panels. Our results show that personalized messages are generally seen as more persuasive than generic ones. We also find that traits like Openness tend to increase persuadability, while Neuroticism can lower it. Differences between LLM evaluators suggest that using multiple models provides a clearer picture. Overall, this work demonstrates a practical way to create more targeted debunking messages exploiting LLMs, while also raising important ethical questions about how such technology might be used.
Tags
Links
- Source: https://arxiv.org/abs/2603.09533v1
- Canonical: https://arxiv.org/abs/2603.09533v1
Trouble viewing inline? Open PDF directly â
Full Text
53,319 characters extracted from source content.
Expand or collapse full text
Enhancing Debunking Effectiveness through LLM-based Personality Adaptation Pietro DellâOglio 1â , Alessandro Bondielli 2 , Francesco Marcelloni 1 , and Lucia C. Passaro 2 1 Dipartimento di Ingegneria dellâInformazione, UniversitĂ di Pisa, Largo Lucio Lazzarino 1, Pisa, Italy pietro.delloglio@ing.unipi.it, francesco.marcelloni@unipi.it 2 Dipartimento di Informatica, UniversitĂ di Pisa, Largo B. Pontecorvo 3, Pisa, Italy alessandro.bondielli@unipi.it, lucia.passaro@unipi.it This is a preprint version of the paper accepted for publication in: Marcelloni, F., Madani, K., van Stein, N., Filipe, J. (eds) Compu- tational Intelligence. IJCCI 2025. Communications in Computer and Information Science, vol 2827. Springer, Cham. https://doi.org/10.1007/978-3-032-15632-7_23 Abstract. This study proposes a novel methodology for generating per- sonalized fake news debunking messages by prompting Large Language Models (LLMs) with persona-based inputs aligned to the Big Five per- sonality traits: Extraversion, Agreeableness, Conscientiousness, Neuroti- cism, and Openness. Our approach guides LLMs to transform generic de- bunking content into personalized versions tailored to specific personality profiles. To assess the effectiveness of these transformations, we employ a separate LLM as an automated evaluator simulating corresponding personality traits, thereby eliminating the need for costly human evalu- ation panels. Our results show that personalized messages are generally seen as more persuasive than generic ones. We also find that traits like Openness tend to increase persuadability, while Neuroticism can lower it. Differences between LLM evaluators suggest that using multiple models provides a clearer picture. Overall, this work demonstrates a practical way to create more targeted debunking messages exploiting LLMs, while also raising important ethical questions about how such technology might be used. Keywords: Personalized Debunking· Fake News· Large Language Mod- els· Prompt Engineering· Role-Playing 1 Introduction The proliferation of fake news and disinformation across digital platforms poses a significant threat to informed public discourse, social cohesion, and democratic processes [23]. This phenomenon is exacerbated by the increasing volume of â Corresponding author arXiv:2603.09533v1 [cs.AI] 10 Mar 2026 2DellâOglio et al. synthetic content, including fake news, generated by Large Language Models (LLMs) [12]. While manual fact-checking efforts are crucial, their scalability is inherently limited given the sheer volume and velocity of false information, which is often amplified by automated entities. LLMs have emerged as powerful tools with the potential to assist in countering misinformation at scale, for instance, by generating counter-narratives or debunking content [35]. However, considering that each individual has their own characteristics, cognitive styles, and pre- existing beliefs, generalized debunking may not be effective for everyone [25]. Fig. 1: Overview of the proposed methodology for generating personality-aligned fake news debunking messages. The process involves prompting an LLM with persona-based inputs corresponding to Big Five personality traits to generate tailored debunking con- tent. A separate LLM is employed as an evaluator to assess the psychological alignment and quality of the outputs. The Big Five personality framework [15] has emerged as a valuable lens for examining individual differences in susceptibility to misinformation. Prior research has identified consistent, though nuanced, associations between person- ality traits and the likelihood of believing false or misleading information [21, 1]. These traits have been shown to influence information processing, susceptibil- ity to persuasion, and communication preferences. If debunking messages could be tailored to resonate with an individualâs psychological profile, their persua- sive impact and reception could be enhanced. For example, marketing research demonstrates that targeting messages based on different levels of Extraversion can significantly improve engagement [6, 19, 22]. This study proposes a methodology for prompting LLMs to adapt generic fake news debunking messages to align with specific Big Five personality profiles. It investigates the use of persona-based prompts to guide LLMs in generating Enhancing Debunking Effectiveness through Personality Adaptation3 psychologically nuanced and personalized debunking content. To evaluate the effectiveness of the proposed approach, we employ a separate LLM as an auto- mated judge, thereby avoiding the need for a costly and time-consuming human evaluation panel. Figure 1 illustrates the proposed methodology. The remainder of this paper is organized as follows. Section 2 reviews related work. Section 3 describes our methodology. Section 4 presents the results of our assessment and analysis. Finally, Section 6 provides conclusions and outlines future directions. This investigation is guided by the following research questions: RQ1 Can an LLM personalize debunking messages to enhance persuasiveness for users with specific personality traits? RQ2 Which specific personality traits are most strongly associated with varia- tions in perceived persuasiveness scores? 2 Related Work Recent literature has increasingly focused on the application of LLMs for auto- matically identifying fake news and generating preliminary veracity assessments [5]. Various strategies have been explored, ranging from zero-shot and few-shot learning [13, 33] to fine-tuning models on domain-specific datasets [2, 29]. How- ever, the reliability of an AI-generated debunking message remains a challenge. It has been observed that even the most advanced LLM can struggle to handle nuanced factual claims [7, 8, 10, 18, 24, 30, 32], often performing better at iden- tifying opinions than at verifying hard facts [28]. Moreover, LLMs are prone to producing content that deviates from factual accuracy or includes fabricated details, a phenomenon commonly referred to as hallucination [26]. A substantial body of research indicates that the acceptance of misinfor- mation is not solely a function of analytical reasoning, but is closely linked to underlying psychological and cognitive factors [36]. Consequently, generic fact- checking messages often fall short in effectiveness, as they overlook influential cognitive biases, such as confirmation bias, the tendency to seek out and priori- tize information that aligns with oneâs pre-existing beliefs [17]. A meta-analysis on the effectiveness of fact-checking highlights that corrections are less impactful under specific conditions. In particular, (i) when they use nuanced truth scales instead of simple true/false verdicts, (i) when they only refute parts of a claim, or (i) when the claims are related to political campaigns [31]. The Big Five personality model [15] offers a robust framework for inves- tigating how individual differences influence susceptibility to misinformation. Research has demonstrated consistent, though complex, correlations between various personality traits and the tendency to believe false or misleading infor- mation [1, 21]. These findings suggest that a one-size-fits-all approach to debunking is sub- optimal, as the recipientâs personality moderates the effectiveness of the mes- sage. The concept of tailoring messages to personality traits is well-established in other domains, most notably marketing, and political communication [6, 11, 4DellâOglio et al. 19, 20, 22]. Marketers have long used the Big Five model to craft advertisements that resonate with specific consumer segments. For example, messages target- ing extraverts might emphasize social rewards and excitement, while those for introverts might focus on solitude and reflection [20]. Similarly, political micro- targeting employs psychometric profiling to craft persuasive messages aimed at influencing specific voter segments. While this practice has demonstrated effec- tiveness, it also raises important ethical concerns regarding manipulation and privacy [11]. 3 Methodology This section outlines the experimental methodology employed in our work. First, we describe the Big Five framework for psychological personalities used. Second, we detail the process of generating debunking verdicts tailored to specific per- sonality profiles. Third, we outline the persona-based evaluation framework used to assess the persuasiveness of these tailored messages via an LLM-as-a-judge approach. 3.1 The Big Five Framework The psychological framework for personalization employed in this study is the Big Five model [15], which describes personality through five fundamental traits: Extraversion, Agreeableness, Conscientiousness, Neuroticism, and Openness to Experience. To enhance the manageability of the personalization task for LLMs and simplify the experimental design, we opted for a binarization of these traits, as proposed by [14]. This approach involves considering two opposing poles for each trait, which we refer to as descriptors, as summarized in Table 1. TraitPositive Descriptor Negative Descriptor Extraversion (E)Extroverted (1)Introverted (0) Agreeableness (A)Agreeable (1)Antagonistic (0) Conscientiousness (C)Conscientious (1)Unconscientious (0) Neuroticism (N)Neurotic (1)Emotionally Stable (0) Openness to Experience (O)Open (1)Closed (0) Table 1: Binary descriptors for each Big Five personality trait, adapted from Jiang et al. (2024) [14]. The positive descriptor indicates a high presence of the trait, while the negative descriptor indicates a low presence. In total, we work with 32 distinct psychological profiles. In this context, each persona can be represented by a five-digit binary code, with each digit indicating the presence or absence of a specific Big Five personality trait. More specifically, Enhancing Debunking Effectiveness through Personality Adaptation5 a value of 1 in a specific position indicates a high level of the trait (positive descriptor), while 0 indicates a low level (negative descriptor). 3.2 Tailored Debunking The primary objective of this work is to utilize an LLM to systematically gener- ate tailored versions of a generic debunking message, each adapted to resonate with one of 32 distinct personality profiles. Our approach centers on role-play prompting [16], where we configure the LLMâs behavior and expertise before presenting the specific task to it. This method allows for greater control and consistency in the generated output [16]. Our prompting strategy involved a system prompt to establish the LLM persona, and a user prompt to execute the specific task. The system prompt, which as- signs to the LLM the role of an expert in persuasive communication, specifically primed to focus on a particular psychological target, is as follows: System Prompt You are a communication strategist with expertise in crafting persuasive messages tailored to individual personality profiles using the Big Five personality model. Your role is to reframe short factual verdicts in a way that maximally resonates with a person who exhibits the following personality traits: traits. Focus on adjusting the messageâs tone, emotional appeal, and emphasis (without altering the factual content) so that it aligns with the readerâs psychological tendencies and communication style. The traits placeholder is dynamically populated with a combination of trait descriptors corresponding to one of the 32 profiles. For instance, 10101, which represents a profile defined as Extroverted, Antagonistic, Conscientious, Emo- tionally Stable, and Open to Experience. The user prompt employed for executing the rewriting task for each specific debunking instance is as follows: User Prompt Rewrite the short âVerdictâ below to make it more persuasive for some- one described with the traits in which you are specialized. Use the âContextâ for factual accuracy and inspiration, but only rewrite the âVerdictâ. Do not explicitly mention the personality traits in the output. Claim: claim Context: context Verdict: verdict 6DellâOglio et al. The placeholders claim, context, and verdict refer, respectively, to the fake news claim being debunked, the full debunking article, and a concise summary of the debunking. In this study, we focus on personalizing the verdict, while using the context to ensure factual accuracy. 3.3 Persona-Based Evaluation For the Persona-Based Evaluation of the generated debunking content, specific LLM personas were developed to act as judges. Recent studies [4, 14] have demon- strated that LLMs can serve as effective judges [4]. In particular, they have shown superior accuracy compared to non-expert humans in specific prediction tasks that rely on pattern recognition from large amounts of text, such as predicting personality traits or behavioral outcomes from written content [14]. Although LLM judges are known to exhibit notable biasesâsuch as a preference for AI- generated contentâthis bias does not confound our evaluation, since all person- alized debunking messages under assessment are AI-generated. The instantiation of the LLM-as-a-judge personas was achieved through sys- tem prompts and user prompts, consistent with the Tailored Debunking phase (see Section 3.2). The system prompt used for this purpose, which was validated by Jiang et al. (2024) [14], is reproduced below: System Prompt You are a character who is traits. For every binarized Big Five profile used in the Tailored Debunking step, a corresponding LLM-as-a-judge persona was developed. The central task of the Persona-Based Evaluation step involves each LLM persona assessing the persuasiveness (i.e., with a 1-7 Likert scale) of multiple debunking messages related to the same fake news item. The user prompt is reproduced below: User Prompt Considering your personality, evaluate the persuasiveness of the verdict below which addresses a specific claim. Rate how persuasive you find this verdict for someone with your spe- cific personality traits using a scale from 1 to 7. Consider that: 1 = Not at all persuasive 4 = Moderately persuasive 7 = Extremely persuasive Claim: claim Verdict to Evaluate: verdict Your score: Enhancing Debunking Effectiveness through Personality Adaptation7 This approach enables a direct and controlled comparison of the effectiveness of different debunking styles. Specifically, each judge evaluates the following categories of verdicts: Matched Profile. Evaluation of a debunking content specifically personalized for its own profile. Mismatched Profile. Evaluation of two pieces of debunking content person- alized for different profiles: one similar (i.e., with a single trait changed) and one very different (i.e., with 2 to 5 traits changed) from the judgeâs own. Generic. Evaluation of non-personalized, generic debunking content. This methodology allows us not only to assess whether a personalized de- bunking is perceived as effective by the target judge but also to determine if the personalization is specific. In other words, it helps us ascertain if a matched message is preferred over a mismatched one. 4 Experiments To validate the proposed methodology, we tested it on a dataset of debunked claims using state-of-the-art open-source LLMs. As for the dataset, we used a subset of the FullFact dataset [27]. The original dataset provides urls to fact-checking articles published by FullFact 3 , a UK-based fact-checking repository, including both debunked and confirmed claims. For each url in the dataset, we extracted the following elements: the full text of the debunking article, the final verdict, the claim under review, and the correspond- ing topic. To focus exclusively on debunked (i.e., false) claims, we implemented a semi-automated filtering procedure. This involved an initial keyword-based ex- clusion of confirmed claims (i.e., those not considered fake news), followed by a manual validation. After this revision process, the final dataset comprised 933 in- stances, each containing a claim, a generic debunking article, and the associated verdict. As for the models, we used Qwen3 [34] in its 8B and 32B parameter variants (in the following, Qwen3-8B and Qwen3-32B) and Llama3-8B-Instruct (in the following Llama3) [9]. We exploited Qwen3-32B for Tailored Debunking (see Sec. 3.2) and all the models (Llama3, Qwen3-8B, and Qwen3-32B) for the Persona- Based Evaluation (Sec. 3.3). This selection was informed by an initial empirical evaluation on a subset of the data, which showed that Qwen3-32Bâs perfor- mance was comparable to other commercial models. However, we prioritized open-source models to facilitate reproducibility and ensure wider accessibility, while benefiting from greater cost-efficiency. Using multiple open-source models for evaluation allowed us to capture diverse judgment perspectives and enhance robustness. For Tailored Debunking, we set the modelâs temperature to 0.7, following [14]. This is done to introduce variability in the modelâs behavior and maintain 3 https://fullfact.org/ 8DellâOglio et al. a balance between factual consistency and creative fluency. For Persona-Based Evaluation, we set the temperature of all models to 0, as commonly done in the literature [4, 14]. All other generation parameters were left to their default settings. Given each claim-generic verdict pair in the dataset, we used Qwen3-32B with the prompt template described in Section 3.2 to generate custom verdicts aimed to persuade each psychological profile based on the Big Five framework. In practice, we fill the template with the set of traits of a profile, the claim, the generic verdict, and the complete debunking article as context, and ask the model to generate a verdict tailored to the profile. Thus, each model call is independent. Table 2 presents a sample of verdicts for a specific claim to illustrate the generated output: a generic version, and three tailored alternatives. A complete list of all 32 generated verdicts for this claim is available for review in Appendix A. . Claim-Verdict010000111111000 Claim: Pfizer CEO Al- bert Bourla said he doesnât need the vaccine because heâs healthy. Verdict: This is a mis- quote from an interview from December. Albert Bourla said he didnât want to take the vaccine ahead of more vulner- able recipients. He has since been double vacci- nated. Albert Bourla shared his thoughts in a De- cember interview, ex- plaining that he felt it was more important for others in greater need to receive the vaccine first. His words were taken from that time, and itâs worth noting that he has since chosen to be fully vaccinated, receiv- ing both doses. This is a partial repre- sentation of a Decem- ber interview where Al- bert Bourla explained his thoughtful decision not to take the vac- cine before those more in need. His choice was made with care and consideration for others, and he has since re- ceived both doses of the vaccine. Albert Bourla made it clear during a Decem- ber interview that he wanted to make sure others who were more in need got the vaccine first and thatâs exactly what he did. Since then, he has received both doses and shown full confidence in the vac- cine his company helped develop. Itâs all about doing the right thing, at the right time, for the right reason. Table 2: A sample of verdicts generated during the Tailored Debunking phase, showing a generic verdict alongside tailored alternatives for profiles 01000 (a persona who is In- troverted, Agreeable, Unconscientious, Emotionally Stable and Closed to Experience), 01111 (a persona who is Introverted, Agreeable, Conscientious, Neurotic and Open to Experience), and 11000 (a persona who is Extroverted, Agreeable, Unconscientious, Emotionally Stable and Closed to Experience) Then, for the Persona-Based Evaluation we instructed each model to im- personate a persona with a specific combination of psychological traits and to evaluate the perceived level of persuasiveness of a verdict using a 1-7 Likert scale. Details on the instantiation of personas and the evaluation method are described in Section 3.3. The results obtained from the experiments helped us answer our research questions. Both the RQ1 (Can an LLM personalize debunking messages to enhance persuasiveness for users with specific personality traits?), and RQ2 (Which specific personality traits are most strongly associated with variations in perceived persuasiveness scores?) are investigated in Section 5. Enhancing Debunking Effectiveness through Personality Adaptation9 5 Results and Discussion In this section, we present the results of our experiments through both quan- titative and qualitative analyses. The quantitative analysis evaluates the effec- tiveness of personalized verdicts compared to generic ones, while the qualitative analysis examines the behavior of individual judge profiles to explore whether specific personality traits influence judgesâ perceptions of persuasiveness. 5.1 Quantitative Analysis Figure 2 presents the mean persuasion scores assigned to each of the 32 unique judge personality profiles. For each judge, the figure reports scores for ver- dicts tailored specifically to their profile (Matched), non-tailored generic ver- dicts (Generic), and verdicts tailored to different profiles (Mismatched). The Mismatched condition is further subdivided into verdicts tailored to closed neigh- boursâprofiles differing by a single tolerance bit (i.e., one trait)âand distant neighbours, which differ by two or more traits. Throughout this work, we use the terms closed neighbours and distant neigh- bours to denote profiles differing by one or more bits in the binary represen- tation of personality traits, respectively. Each judgeâs profile is encoded as a five-digit binary string, with each bit representing one of the Big Five traits in the order: Extraversion, Agreeableness, Conscientiousness, Neuroticism, and Openness (see Table 1). We observe that Matched verdicts typically receive the highest persuasion scores. Nevertheless, there are several instancesâparticularly with Llama3âwhere a judge assigns a higher average score to a Mismatched ver- dict. In contrast, judges based on the Qwen architecture show a more consistent preference for Matched verdicts, with mismatched ones rarely outperforming them. Notably, the Generic verdict is never preferred over the other conditions. We further analyze the fact that, in some cases, the Mismatched verdicts are considered more persuasive that Matched ones. This often occurs when the Mis- matched verdict is tailored for a closed neighbour profile. This phenomenon is justifiable by the nature of the Big Five Framework, which models personality on continuous spectra rather than discrete categories. Our discretization creates artificial boundaries. We performed paired-sample t-tests for all pairwise comparisons to evaluate the statistical significance of the differences. All results were statistically signifi- cant (p-value < 0.05). Table 3 shows the t-statistics for each comparison, which are high enough to indicate that the overall effects are robust, and confirms a clear hierarchy: Matched verdicts are, on average, superior to Mismatched, and both are significantly more persuasive than Generic ones across all models. 10DellâOglio et al. Fig. 2: Mean persuasive scores assigned by LLM-based judges across three conditions: Matched (verdict tailored to the judgeâs own psychological profile), Mismatched (ver- dicts tailored to different profiles), and Generic (non-personalized). Each persona is defined by the high presence (1) or low presence (0) of each one of the Big Five traits, in order Extraversion (E), Agreebleness (A), Conscientiousness (C), Neuroticism (N), and Openness to Experience (O). Profile PairLlama3Qwen3-8BQwen3-32B ABt-statistic P-valuet-statistic P-valuet-statistic P-value MatchedMismatched14.54 4.75Ă 10 â48 29.96 1.83Ă 10 â194 37.18 9.39Ă 10 â296 MatchedGeneric92.930.088.110.078.680.0 MismatchedGeneric91.380.079.310.056.080.0 Matched Mismatched (clos.)7.57 1.87Ă 10 â14 14.27 2.47Ă 10 â46 17.37 1.53Ă 10 â67 MatchedMismatched (dist.)17.38 1.21Ă 10 â67 37.29 4.48Ă 10 â298 46.510.0 Mismatched (clos.)Mismatched (dist.)10.43 9.15Ă 10 â26 24.57 2.78Ă 10 â132 30.71 4.50Ă 10 â204 Mismatched (clos.)Generic87.540.078.990.063.740.0 Mismatched (dist.) Generic80.480.060.250.035.49 2.34Ă 10 â270 Table 3: t-test statistic for different comparisons. P-values are all < 0.05. To provide a clearer picture of model performance, we computed two aggre- gate metrics. The first metric, denoted as Accuracy p , indicates the effectiveness of the exact Profiled Debunking. Let N be the total number of verdict obser- vations. For each observation i = 1, . . . , N, let v i denote the matched verdict. Enhancing Debunking Effectiveness through Personality Adaptation11 Using dense ranking (where all verdicts tied for the highest score are assigned rank 1), the accuracy is defined as: Accuracy p = 1 N N X i=1 1 rank(v i ) = 1 where 1(·) is the indicator function, equal to 1 if the condition is true and 0 otherwise. The second metric calculates accuracy over close neighbors. More formally, we define C(v i ) as the set of verdicts created for profiles considered close neighbors. The accuracy over close neighbor profiles, denoted Accuracy cn , is defined as: Accuracy cn = 1 N N X i=1 1 rank(v i ) = 1 âš âv âC(v i ) : rank(v) = 1 The results for Accuracy p (i.e., accuracy on exact-profile debunking) and Accuracy cn (i.e., accuracy on verdicts tailored to closed neighbours) are reported in Table 4. These scores quantify how well each model distinguishes persuasive content when the verdict is optimized either for the exact profile or for a closely related one (differing by a single trait). ModelAccuracy p Accuracy cn Llama370.5986.45 Qwen3-8B 88.6496.39 Qwen3-32B 68.7886.85 Table 4: Accuracy of models on (exact) Profiled Debunking (Accuracy p ) and Close Neighbours (Accuracy cn ), expressed as percentages. Examining Accuracy p , Qwen3-8B emerges as the most accurate judge, with the Matched verdict ranked highest in 88.64% of cases. In contrast, Llama3 (70.59%) and Qwen3-32B (68.78%) are considerably less accurate. The most telling metric, however, is the Accuracy cn , which accounts for the continuous nature of personality. Here, all models perform exceptionally well, with scores of 86.45%, 96.39%, and 86.85%. This demonstrates that even when the perfectly Matched verdict does not receive the top score, the most persuasive alternative is almost always one designed for a very similar psychological profile. This confirms that the personalization is precise within a âclose neighborhoodâ of the target profile. 12DellâOglio et al. Fig. 3: Aggregate mean persuasiveness scores for profiles grouped by their number of positive descriptors activated 5.2 Qualitative Analysis Summary To explore how personality traits influence perceived persuasiveness in Persona- Based Evaluation, we conducted a detailed qualitative analysis of individual judge profiles across three LLMs. We show the results in Figure 3. The results reveal significant variability in persuasiveness scores across the 32 simulated personality profiles, indicating that both the personality traits and the LLM used for evaluation drive the effectiveness of message adaptation. A key finding is a positive correlation between the number of activated positive descriptors (e.g., Openness, Conscientiousness) in a profile and its average persuasiveness score. Profiles with more positive traits consistently rated adapted messages higher, suggesting that these traits act as stronger drivers of persuasion. Llama3 generally simulates more generous personas, assigning higher scores and showing weaker discrimination between Matched and Mismatched verdicts. Agreeableness-strong profiles, for instance, rated Mismatched verdicts higher than Matched, suggesting high susceptibility to any persuasive effort. Conversely, Neuroticism tended to reduce scores unless paired with strong positive traits, indicating some traits can override others. Qwen models, and Qwen3-8B in particular, displayed more cautious behav- ior. They assigned lower scores overall and more reliably favored Matched ver- dicts, penalizing profiles with Neuroticism. Compared to Qwen3-8B, Qwen3-32B sometimes produced higher peaks in persuasiveness and greater differentiation between verdict types, suggesting finer-grained judgment from larger models. Enhancing Debunking Effectiveness through Personality Adaptation13 Across all models, certain universal trends emerged. Highly neurotic profiles (e.g., 00010) were consistently harder to persuade, while profiles with multiple positive traits (e.g., 10101, 11101) were consistently rated as highly persuadable. The observed differences between Llama3 and Qwen evaluations likely stem from fundamental variations in their underlying architectures, and training data. Recent research has shown that different LLMs exhibit unique and distinct per- sonality profiles, even within the same model family. For instance, Bhandari et al. (2025) [3] found that OpenAI models tend to be dominant in the Agreeableness trait, making them more cooperative and friendly in their interactions, while dif- ferent versions of Llama models showed dominance in either Conscientiousness or Openness. This strongly suggests that the behavioral tendencies we observed are a direct consequence of the modelsâ inherent personalities. This highlights that the choice of an evaluator model is not neutral and that using a panel of diverse models, as done in this study, is crucial for robust findings. 6 Conclusions This paper investigated the feasibility of using LLMs to adapt debunking mes- sages based on specific personality traits. Our work introduced a methodology which exploits LLMs for both adapting profiles and evaluating the quality of the adapted profiles. Our findings demonstrate that an LLM can indeed be prompted to effectively personalize debunking messages for distinct psychological profiles. These tailored messages are perceived as more persuasive by LLM judges com- pared to generic alternatives and alternatives tailored for mismatched profiles. This was confirmed by statistically significant results showing a consistent hier- archy in which Matched verdicts outperformed both Mismatched and Generic versions on average. Furthermore, our qualitative analysis revealed that specific personality traits are strongly associated with variations in persuasiveness. A significant model-dependent effect also emerged: different LLMs simulate personality in various ways. Llama3 acted as a âgenerousâ judge, assigning high scores broadly, while the Qwen models behaved as more âcautiousâ evaluators. Furthermore, our results suggest a potential link between model scale and dis- criminative ability, with Qwen3-32B showing a greater ability to isolate the mes- sage tailored strictly for its profile. In conclusion, our work points to a meaningful step forward in spreading debunks, and persuading people about their veracity. By moving away from a one-size-fits-all schema, our approach allows for the adaptation of debunking messages to better match individual personalities, rather than addressing broad, generic audiences. This could be a valuable tool for expert fact-checkers aiming to increase the persuasive impact of their messages. Building on the promising results of this study, several avenues for future research emerge. Most importantly, it is crucial to validate the effectiveness of personalized debunking messages with human participants. Experimental stud- ies involving diverse populations will help determine whether the LLM-simulated adaptations translate into increased persuasion in real-world settings. Addi- 14DellâOglio et al. tionally, future work should explore the generalizability of this methodology. The persona-based generation and evaluation framework is domain-agnostic and could be readily applied to other areas where persuasive communication is key, such as public health messaging or educational content. Also, more nuanced and continuous models of personality should be investigated to better capture the complexity of human traits and improve the precision and effectiveness of tailored messaging strategies. 7 Limitations This study has some limitations. First, all evaluations were conducted using LLMs to simulate human judgments of persuasiveness. While this approach en- ables scalable analysis and aligns with emerging practices in AI research, it does not fully capture how real individuals respond to persuasive messaging. Future studies should include humans to validate whether the LLM-identified adapta- tions are genuinely effective for people with matching personality traits. Second, our representation of personality traits relied on a binarized ver- sion of the Big Five model. While this allowed for experimental control and interpretability, it oversimplifies the continuous nature of human personality. Incorporating more granular trait modeling would strengthen the study. Third, we experimented with only a limited set of LLMs and data. Testing a broader range of models and more diverse datasets would help to verify the generalizability and robustness of our findings. Lastly, the ethical implications of this technology warrant a deeper discus- sion. While our focus is on pro-social applications like debunking, the same techniques for crafting personalized persuasive messages could be repurposed for malicious ends. This includes not only sophisticated political micro-targeting and the spread of propaganda but also the potential to exacerbate societal po- larization by reinforcing cognitive biases within specific psychological profiles. The development of such technologies must be paired with robust research into detection methods for personalized influence campaigns and the establishment of clear ethical guidelines for their use. Acknowledgements This work has been partly funded by the PNRR - M4C2 - Investimento 1.3, Partenariato Esteso PE00000013 - âFAIR - Future Artificial Intelligence Re- search" - Spoke 1 âHuman-centered AI" under the NextGeneration EU pro- gramme, and the Italian Ministry of University and Research (MUR) in the framework of the PRIN 2022JLB83Z âPsychologically-tailored approaches to De- bunk Fake News detected automatically by an innovative artificial intelligence approach", the FoReLab and CrossLab projects (Departments of Excellence). Enhancing Debunking Effectiveness through Personality Adaptation15 A Examples of tailored verdicts for all 32 Profiles of a single claim. This appendix presents the full set of 32 personalized verdicts generated by Qwen3-32B for a specific claim. Each verdict is linked to a specific profile represented by a five-digit binary code as described in section 3.1. This provides an example of how the model adapted its tone, framing, and emphasis for each distinct personality profile. Claim: Pfizer CEO Albert Bourla said he doesnât need the vaccine because heâs healthy. 00000: Albert Bourla made his stance clear in a December interview; he felt it was important not to bypass those more in need of the vaccine at the time. His com- ments were taken out of context, and itâs worth noting that he has since followed through with receiving both doses. His actions speak louder than the isolated clip thatâs being shared. 00001: Albert Bourla made his comments in a December interview, clearly stating he didnât want to bypass those more in need of the vaccine. He later received both doses, aligning with the broader sentiment he expressed at the timeâthat confidence in the vaccine matters, and actions speak louder than words. 00010: The claim is taken out of context. In a December interview, Albert Bourla explained he wanted to wait his turn and let those more at risk go first. He has since received both doses of the vaccine. 00100: The statement is taken out of context from a December interview. Albert Bourla explained he chose to defer vaccination to ensure those at higher risk were prioritized. He later received both doses, as confirmed publicly. 01000: Albert Bourla shared his thoughts in a December interview, explaining that he felt it was more important for others in greater need to receive the vaccine first. His words were taken from that time, and itâs worth noting that he has since chosen to be fully vaccinated, receiving both doses. 10000: Hereâs the full story: In December, Albert Bourla made it clear he didnât want to skip ahead of those more in need for the vaccine. And look, he came through â heâs since gotten both shots. This wasnât about avoiding the vaccine altogether, it was about doing the right thing at the right time. Fair play for setting an example later on. 00011: Albert Bourla made his position clear in Decemberâhe believed others more at risk should come first. But donât be misled by a snippet out of context. Heâs had both doses since, and even acknowledged the power of leadership in setting an example. What matters is the outcome, and his actions speak louder than the isolated remarks critics choose to highlight. 00101: The statement in question is taken from an interview in December where Albert Bourla expressed his view that others more at risk should be prioritized for the vaccine. He made it clear that he intended to be vaccinated when appropriate and, in fact, has since received both doses. His decision was guided by a sense of responsibility and the belief that leadership through action could help build public confidence. 00110: This claim is based on an out-of-context statement from December. At the time, Albert Bourla clearly stated he did not want to receive the vaccine before those more in need, such as frontline workers and vulnerable individuals. He has since followed through and received both doses of the vaccine. 16DellâOglio et al. 01001: Albert Bourla made his comments in December, explaining at the time that he wanted to prioritize those more at risk. His words were taken out of context, and itâs reassuring to know that he has since received both doses of the vaccine, aligning with the broader message of safety and care that his company promotes. 01010: Albert Bourla shared his thoughts during an interview in December, explaining that he wanted to let those more in need go first. Itâs also important to note that he has since chosen to receive both doses of the vaccine, showing his support in a way that reflects his role. 01100: This is a partial representation of a December interview in which Albert Bourla explained his thoughtful decision to wait his turn for the vaccine, prioritizing those more in need. Since then, he has received both doses and is fully vaccinated. 10001: Albert Bourla made it clear in December he didnât want to jump ahead of people who truly needed the vaccine firstâlike frontline workers and the elderly. And letâs be honest, who better than him to set a strong example? Well, turns out he didnât wait longâby March, heâd already gotten both doses. So while the clip might paint one picture, the full story shows someone who understood the bigger picture and stepped up when the time was right. 10010: Letâs get this straightâthis clip is taken out of context. Bourla made it clear back in December that he didnât want to jump the queue and take his own vaccine before the people who really needed it. And guess what? Heâs since had both doses. Heâs doing whatâs right, even if it didnât make for the flashiest soundbite. 10100: Albert Bourla made a principled decision in December to step aside and let others with greater need come first. Since then, he has followed through and re- ceived both doses of the vaccine. The clip taken out of context misses the full pictureâhis actions speak louder than the isolated quote. 11000: Albert Bourla made it clear during a December interview that he wanted to make sure others who were more in need got the vaccine first and thatâs exactly what he did. Since then, he has received both doses and shown full confidence in the vaccine his company helped develop. Itâs all about doing the right thing, at the right time, for the right reason. 00111: Albert Bourlaâs comments were made in December, during the early stages of the vaccine rollout. He emphasized his belief that others, particularly those at higher risk, should be prioritized. While his remarks were taken out of context, they were made with a clear understanding of public perception and responsibility. Since then, Bourla has followed through with both doses of the vaccine, aligning his actions with the broader public health effort. 01011: This interview took place in December, and at the time, Albert Bourla ex- plained he wanted to prioritize others who were more in need of the vaccine. He was clear about wanting to set a positive example, and he has since received both doses of the vaccine. 01101: Albert Bourla made his comments in a December interview with CNBC, ex- plaining that he wanted to prioritize those more vulnerable during the early stages of the vaccine rollout. His remarks were made before the vaccine became widely available, and he has since received both doses of the vaccine. His approach re- flected thoughtful consideration for others, and his actions ultimately aligned with the broader public health effort. 01110: This is a partial representation of an interview from December. At that time, Albert Bourla shared that he felt it was more appropriate for others who were more vulnerable to receive the vaccine first. Since then, he has received both doses of the vaccine and publicly shared that experience. Enhancing Debunking Effectiveness through Personality Adaptation17 10011: Albert Bourla made it clear back in December that he didnât want to jump ahead of others who were more in need of the vaccineâlike frontline workers and older adults. Heâs since gotten both doses and stood by his own product. The clip taken out of context misses the bigger picture: Bourla was being thoughtful about the rollout and the message his decision might send. 10101: Hereâs the full picture: in a December interview, Albert Bourla made a princi- pled choice not to take the vaccine ahead of those more vulnerable. Heâs since taken both doses, showing leadership by example. His decision wasnât about avoiding the vaccine â it was about fairness and timing. 10110: This claim is taken out of context from a December interview where Albert Bourla made it clear he wanted to prioritize others who were more at risk. Since then, heâs followed through by getting both doses of the vaccine â proving his commitment and setting a strong example for others to follow. 11001: Albert Bourla made his comments in December, explaining he wanted to set a fair example by not getting the vaccine ahead of those more in need. Since then, he has taken both doses and shown his full support for the vaccine. 11010: Albert Bourla made his comments back in December, explaining that he wanted to prioritize others who were more in need of the vaccine at the time. He emphasized fairness and respect for everyoneâs needs. Since then, he has received both doses of the vaccine and continues to support the importance of vaccination for all. 11100: Albert Bourla made his comments in a December interview, explaining that he wanted to prioritize the vaccine for those most in need. Since then, he has received both doses and continues to support the importance of vaccination. 01111: This is a partial representation of a December interview where Albert Bourla explained his thoughtful decision not to take the vaccine before those more in need. His choice was made with care and consideration for others, and he has since received both doses of the vaccine. 10111: Look, hereâs the deal. Bourla made his stance clear back in December. He didnât want to jump the queue when others were more in need of the shot. But letâs be real â heâs since rolled up his sleeves and gotten fully vaccinated, just like the rest of us. Itâs about timing, responsibility, and leading by example â and heâs checked all the boxes. 11011: Albert Bourla shared his thoughts back in December about not wanting to take the vaccine before those who needed it most like frontline workers or older adults. He made it clear he wanted to lead by example, and in fact, heâs since gotten both doses of the vaccine. His message was always about fairness and trust, and his actions show that. 11101: This is taken out of context from a December interview where Albert Bourla made a thoughtful decision to prioritize others who were more in need of the vaccine at the time. He has since taken both doses and continues to support vaccination as a key tool in protecting public health. 11110: Albert Bourla made his comments back in December with the thoughtful idea that others who are more in need should come first. He was clear about wanting to set a good example and has since taken both doses of the vaccine, just like many others. Itâs always reassuring to see leaders showing care and responsibility. 11111: Albert Bourla made his comments during a December interview, explaining he wanted to ensure the vaccine went first to those most in need. Since then, he has received both doses and continues to support the importance of vaccination. 18DellâOglio et al. References 1. Ahmed, S., Tan, H.W.: Personality and perspicacity: Role of personality traits and cognitive ability in political misinformation discernment and sharing behavior. Personality and Individual Differences 196, 111747 (2022) 2. Alghamdi, J., Lin, Y., Luo, S.: The power of context: A novel hybrid context-aware fake news detection approach. Information 15(3), 122 (2024) 3. Bhandari, P., Naseem, U., Datta, A., Fay, N., Nasim, M.: Evaluating personality traits in large language models: Insights from psychological questionnaires. In: Companion Proceedings of the ACM on Web Conference 2025. p. 868â872 (2025) 4. Chiang, C.H., Lee, H.y.: Can large language models be an alternative to human evaluations? arXiv preprint arXiv:2305.01937 (2023) 5. Dierickx, L., Van Dalen, A., Opdahl, A.L., LindĂ©n, C.G.: Striking the balance in us- ing llms for fact-checking: A narrative literature review. In: Multidisciplinary Inter- national Symposium on Disinformation in Open Online Media. p. 1â15. Springer (2024) 6. Duong, C.D.: Big five personality traits and green consumption: bridging the attitude-intention-behavior gap. Asia Pacific Journal of Marketing and Logistics 34(6), 1123â1144 (2022) 7. Gili, J., Passaro, L., Caselli, T.: Check-IT!: A corpus of expert fact-checked claims for Italian. In: Boschetti, F., Lebani, G.E., Magnini, B., Novielli, N. (eds.) Proceedings of the 9th Italian Conference on Computational Linguistics (CLiC- it 2023). p. 227â235. CEUR Workshop Proceedings, Venice, Italy (Nov 2023), https://aclanthology.org/2023.clicit-1.29/ 8. Gili, J., Patti, V., Passaro, L., Caselli, T.: VeryfIT - benchmark of fact-checked claims for Italian: A CALAMITA challenge. In: DellâOrletta, F., Lenci, A., Mon- temagni, S., Sprugnoli, R. (eds.) Proceedings of the 10th Italian Conference on Computational Linguistics (CLiC-it 2024). p. 1116â1124. CEUR Workshop Pro- ceedings, Pisa, Italy (Dec 2024), https://aclanthology.org/2024.clicit-1.123/ 9. Grattafiori, A., Dubey, A., Jauhri, A., Pandey, A., Kadian, A., Al-Dahle, A., Let- man, A., Mathur, A., Schelten, A., Vaughan, A., et al.: The llama 3 herd of models. arXiv preprint arXiv:2407.21783 (2024) 10. Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., Steinhardt, J.: Measuring massive multitask language understanding (2021), https://arxiv.org/abs/2009.03300 11. Hersh, E.D.: Hacking the electorate: How campaigns perceive voters. Cambridge University Press (2015) 12. Hu, B., Sheng, Q., Cao, J., Li, Y., Wang, D.: Llm-generated fake news induces truth decay in news ecosystem: A case study on neural news recommendation. arXiv preprint arXiv:2504.20013 (2025) 13. Hu, B., Sheng, Q., Cao, J., Shi, Y., Li, Y., Wang, D., Qi, P.: Bad actor, good advisor: Exploring the role of large language models in fake news detection. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 38, p. 22105â 22113 (2024) 14. Jiang, H., Zhang, X., Cao, X., Breazeal, C., Roy, D., Kabbara, J.: Personallm: Investigating the ability of large language models to express personality traits. arXiv preprint arXiv:2305.02547 (2023) 15. John, O.P., Srivastava, S., et al.: The big-five trait taxonomy: History, measure- ment, and theoretical perspectives (1999) Enhancing Debunking Effectiveness through Personality Adaptation19 16. Kong, A., Zhao, S., Chen, H., Li, Q., Qin, Y., Sun, R., Zhou, X., Wang, E., Dong, X.: Better zero-shot reasoning with role-play prompting. arXiv preprint arXiv:2308.07702 (2023) 17. Kunda, Z.: The case for motivated reasoning. Psychological bulletin 108(3), 480 (1990) 18. Lin, S., Hilton, J., Evans, O.: TruthfulQA: Measuring how models mimic hu- man falsehoods. In: Muresan, S., Nakov, P., Villavicencio, A. (eds.) Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). p. 3214â3252. Association for Computational Linguis- tics, Dublin, Ireland (May 2022). https://doi.org/10.18653/v1/2022.acl-long.229, https://aclanthology.org/2022.acl-long.229/ 19. Liu, Z., Wang, Y., Mahmud, J., Akkiraju, R., Schoudt, J., Xu, A., Donovan, B.: To buy or not to buy? understanding the role of personality traits in predicting consumer behaviors. In: Social Informatics: 8th International Conference, SocInfo 2016, Bellevue, WA, USA, November 11-14, 2016, Proceedings, Part I 8. p. 337â 346. Springer (2016) 20. Matz, S.C., Kosinski, M., Nave, G., Stillwell, D.J.: Psychological targeting as an effective approach to digital mass persuasion. Proceedings of the national academy of sciences 114(48), 12714â12719 (2017) 21. Mirzabeigi, M., Torabi, M., Jowkar, T.: The role of personality traits and the ability to detect fake news in predicting information avoidance during the covid-19 pandemic. Library Hi Tech (2023) 22. Mulyanegara, R.C., Tsarenko, Y., Anderson, A.: The big five and brand personality: Investigating the impact of consumer personality on preferences towards particular brand personality. Journal of brand management 16(4), 234â247 (2009) 23. Newman, N., Dutton, W., Blank, G.: Social media in the changing ecology of news: The fourth and fifth estates in britain. International journal of internet science 7(1) (2013) 24. Passaro, L.C., Bondielli, A., DellâOglio, P., Lenci, A., Marcelloni, F.: In-context annotation of topic-oriented datasets of fake news: A case study on the notre-dame fire event. Information Sciences 615, 657â 677(2022).https://doi.org/https://doi.org/10.1016/j.ins.2022.07.128, https://w.sciencedirect.com/science/article/pii/S0020025522008167 25. Pennycook, G., Rand, D.G.: Who falls for fake news? the roles of bullshit receptiv- ity, overclaiming, familiarity, and analytic thinking. Journal of personality 88(2), 185â200 (2020) 26. Rawte, V., Sheth, A., Das, A.: A survey of hallucination in large foundation models. arXiv preprint arXiv:2309.05922 (2023) 27. Russo, D., TekiroÄlu, S.S., Guerini, M.: Benchmarking the generation of fact check- ing explanations. Transactions of the Association for Computational Linguistics 11, 1250â1264 (2023) 28. Saju, L., Bleier, A., Lasser, J., Wagner, C.: Facts are harder than opinions â a multilingual, comparative analysis of llm-based fact-checking reliability (06 2025). https://doi.org/10.48550/arXiv.2506.03655 29. Shifath, S., Khan, M.F., Islam, M.S.: A transformer based approach for fighting covid-19 fake news. arXiv preprint arXiv:2101.12027 (2021) 30. Talmor, A.e.a.: CommonsenseQA: A question answering challenge targeting com- monsense knowledge. In: Burstein, J., Doran, C., Solorio, T. (eds.) Proceed- ings of the 2019 Conference of the North American Chapter of the Associa- tion for Computational Linguistics: Human Language Technologies, Volume 1 20DellâOglio et al. (Long and Short Papers). p. 4149â4158. Association for Computational Linguis- tics, Minneapolis, Minnesota (Jun 2019). https://doi.org/10.18653/v1/N19-1421, https://aclanthology.org/N19-1421/ 31. Walter, N., Cohen, J., Holbert, R.L., Morag, Y.: Fact-checking: A meta-analysis of what works and for whom. Political communication 37(3), 350â375 (2020) 32. Wang, P., Chan, A., Ilievski, F., Chen, M., Ren, X.: PINTO: faithful language reasoning using prompt-generated rationales. In: ICLR. OpenReview.net (2023) 33. Whitehouse, C., Weyde, T., Madhyastha, P., Komninos, N.: Evaluation of fake news detection with knowledge-enhanced language models. In: Proceedings of the international AAAI conference on web and social media. vol. 16, p. 1425â1429 (2022) 34. Yang, A., Li, A., et al., B.Y.: Qwen3 technical report. arXiv preprint arXiv:2505.09388 (2025) 35. Zanartu, F., Otmakhova, Y., Cook, J., Frermann, L.: Generative debunking of climate misinformation. arXiv preprint arXiv:2407.05599 (2024) 36. Zhou, Y., Shen, L.: Processing of misinformation as motivational and cognitive biases. Frontiers in Psychology 15, 1430953 (2024)