Paper deep dive
Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech
Lorenzo Cima, Alessio Miaschi, Amaury Trujillo, Marco Avenuti, Felice Dell'Orletta, Stefano Cresci
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/4/2026, 10:23:22 AM
Summary
This paper investigates the effectiveness of AI-generated contextualized counterspeech compared to generic approaches for mitigating online toxicity. The authors propose and evaluate multiple strategies that integrate conversational context, community norms, and user history to personalize responses. Using LLaMA2-13B fine-tuned on datasets like MultiCONAN and RHSI, they conduct algorithmic evaluations (ROUGE, BLEU, BERTScore) and a crowdsourced experiment. Results indicate that lightweight strategies combining conversational context and user history improve perceived persuasiveness and adequacy, while other contextualization methods can degrade quality, suggesting that personalization is effective but not uniformly so.
Entities (12)
Relation Signals (10)
Contextualized Counterspeech â comparedto â Generic Counterspeech
confidence 98% ¡ Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech
LLaMA2-13B â usedfor â Counterspeech
confidence 95% ¡ We employ an instruction-tuned version of LLaMA2-13B for generating counterspeech responses.
MultiCONAN â usedforfinetuning â LLaMA2-13B
confidence 92% ¡ fine-tune the base model... using... MultiCONAN
RHSI â usedforfinetuning â LLaMA2-13B
confidence 92% ¡ fine-tune the base model... using... the Reddit hate-speech intervention (RHSI) dataset
Conversational Context â usedin â Contextualized Counterspeech
confidence 90% ¡ incorporating conversational context and user history improve perceived adequacy
User History â usedin â Contextualized Counterspeech
confidence 90% ¡ incorporating conversational context and user history improve perceived adequacy
BLEU â usedtoevaluate â Counterspeech
confidence 90% ¡ algorithmic measures of counterspeech quality based on... BLEU
BERTScore â usedtoevaluate â Counterspeech
confidence 90% ¡ algorithmic measures of counterspeech quality based on... BERTScore
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing approaches adopt a generic, one-size-fits-all paradigm, overlooking the conversational context and characteristics of the targeted users. Here, we propose and evaluate multiple strategies for generating contextualized counterspeech that is adapted to the moderation setting and personalized to the moderated user. In detail, we explore a range of configurations that integrate different forms of contextual information and fine-tuning techniques. We conduct a comprehensive evaluation combining quantitative indicators with a pre-registered, mixed-design crowdsourcing experiment. To ensure robustness, we implement algorithmic measures of counterspeech quality based on ROUGE, BLEU, and BERTScore, observing overall consistent results across metrics. Furthermore, we analyze which characteristics of both the generated counterspeech and the moderated toxic message most strongly influence perceived persuasiveness, yielding insights into how contextualized interventions can be made more effective. Our findings show that personalization can be effective, but not uniformly so. Lightweight strategies combining conversational context and user history improve perceived adequacy and persuasiveness, whereas several other contextualization strategies degrade human-perceived counterspeech quality. Taken together, these results provide actionable directions for developing more personalized, effective, and responsible counterspeech systems, ultimately advancing human-AI collaboration in online content moderation.
Tags
Links
- Source: https://arxiv.org/abs/2607.26236v1
- Canonical: https://arxiv.org/abs/2607.26236v1
Trouble viewing inline? Open PDF directly â
Full Text
118,142 characters extracted from source content.
Expand or collapse full text
Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech LORENZO CIMA, University of Pisa and IIT-CNR, Pisa, Italy ALESSIO MIASCHI, ILC-CNR, Pisa, Italy AMAURY TRUJILLO, IIT-CNR, Pisa, Italy MARCO AVVENUTI, University of Pisa, Pisa, Italy FELICE DELLâORLETTA, ILC-CNR, Pisa, Italy STEFANO CRESCI, IIT-CNR, Pisa, Italy AI-generated counterspeech offers a scalable and effective strategy to mitigate online toxicity by promoting more constructive dialogue. Yet, existing approaches adopt a generic, one-size-fits-all paradigm, overlooking the conversational context and characteristics of the targeted users. Here, we propose and evaluate multiple strategies for generating contextualized counterspeech that is adapted to the moderation setting and person- alized to the moderated user. In detail, we explore a range of configurations that integrate different forms of contextual information and fine-tuning techniques. We conduct a comprehensive evaluation combining quantitative indicators with a pre-registered, mixed-design crowdsourcing experiment. To ensure robustness, we implement algorithmic measures of counterspeech quality based on ROUGE, BLEU, and BERTScore, observing overall consistent results across metrics. Furthermore, we analyze which characteristics of both the generated counterspeech and the moderated toxic message most strongly influence perceived persuasiveness, yielding insights into how contextualized interventions can be made more effective. Our findings show that personalization can be effective, but not uniformly so. Lightweight strategies combining conversational context and user history improve perceived adequacy and persuasiveness, whereas several other contextualization strategies degrade human-perceived counterspeech quality. Taken together, these results provide actionable directions for developing more personalized, effective, and responsible counterspeech systems, ultimately advancing human-AI collaboration in online content moderation. .Warning: This paper contains examples that may be perceived as offensive or upsetting. Reader discretion is advised. CCS Concepts:⢠Computing methodologiesâNatural language generation;⢠Human-centered computingâ Empirical studies in collaborative and social computing. Additional Key Words and Phrases: Counterspeech; personalization; content moderation; online toxicity ACM Reference Format: Lorenzo Cima, Alessio Miaschi, Amaury Trujillo, Marco Avvenuti, Felice DellâOrletta, and Stefano Cresci. 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech. 1, 1 (July 2026), 31 pages. https://doi.org/10.1145/n.n Authorsâ Contact Information: Lorenzo Cima, lorenzo.cima@phd.unipi.it, University of Pisa and IIT-CNR, Pisa, Italy; Alessio Miaschi, alessio.miaschi@ilc.cnr.it, ILC-CNR, Pisa, Italy; Amaury Trujillo, amaury.trujillo@iit.cnr.it, IIT-CNR, Pisa, Italy; Marco Avvenuti, marco.avvenuti@unipi.it, University of Pisa, Pisa, Italy; Felice DellâOrletta, felice.dellorletta@ilc.cnr.it, ILC-CNR, Pisa, Italy; Stefano Cresci, stefano.cresci@iit.cnr.it, IIT-CNR, Pisa, Italy. Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from permissions@acm.org. Š 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM X-X/2026/7-ART https://doi.org/10.1145/n.n , Vol. 1, No. 1, Article . Publication date: July 2026. arXiv:2607.26236v1 [cs.HC] 28 Jul 2026 2Cima et al. Fig. 1. Current AI-generated counterspeech primarily relies on the content of the toxic message alone. In contrast, we generate contextualized counterspeech that integrates information about the community, the surrounding conversation, and the moderated user, aiming to produce more persuasive and targeted responses. 1 Introduction Online toxicity encompasses hateful, offensive, or otherwise harmful language that can cause emotional distress or push users to disengage from digital conversations. Its consequences are both social and economic: online, toxicity hampers participation, disrupts healthy discourse, and exacerbates divisions [1]. Offline, it has been linked to physical violence, diminished respect for norms, and significant psychological harm [28]. These risks are also difficult to anticipate from user activity alone, as harmful trajectories may not be easily predictable from observable online behavior [12]. As a result, tackling online toxicity has become a pressing issue for both regulators and platform operators [63]. To mitigate toxicity and other online harms [21,25,72], platforms employ a range of moderation tools, including user bans, content removals, demotion, and friction-based interventions [13,69,70]. One promising alternative to top-down moderation is counterspeechâuser-driven replies aimed at de-escalating or challenging toxic content [5,19,29]. Counterspeech offers the advantage of avoiding censorship and deterring toxic users from simply migrating to other spaces [45]. However, several factors limit its scalability and effectiveness. Manual counterspeech can place a significant emotional burden on users who regularly confront hostility [65], and it also raises safety concerns due to the risk of retaliation [66]. Moreover, the sheer scale of toxic content makes manual responses unfeasible. In response, researchers have turned to automated counterspeech generated by large language models (LLMs)[5,68]. However, evaluating the true effectiveness of such systems remains difficult [31,75]. Prior evaluations have often focused on surface-level properties, such as fluency or grammaticality, while deeper dimensions such as contextual fit, persuasive impact, and user-specific effectiveness remain harder to assess [51,77]. Recent work has advanced automatic and LLM-as-a-judge evaluation frameworks for counterspeech [42,56], but human-centered validation of contextualized and personalized counterspeech remains limited. A key shortcoming is that most AI-generated counterspeech is one-size-fits-all: responses are based only on the toxic message, ignoring critical contextual elements like the conversation history, the community norms, or the characteristics of the toxic user [22]. Yet, effective moderation requires context-sensitive responses [31, 46], calling for a shift from generic to tailored interventions [22]. Scientific focus. This study extends our previous work [17] to develop and evaluate a method to generate contextualized counterspeech that is adapted to the community and conversation context, and personalized to the individual being moderated. Unlike traditional approaches, our system goes beyond the toxic message and integrates broader information about the surrounding environment (community, moderated user, and specific conversation), as illustrated in Figure 1. We test whether contextualized responses are more persuasive than generic ones, conducting , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech3 experiments in politically active Reddit communities. We identify what makes counterspeech effective and design both algorithmic and human evaluations to measure these qualities. For aspects not adequately captured by existing metrics [77], we conduct a pre-registered, mixed-design crowdsourced user study. Our findings highlight both the promise and the complexity of generating contextualized counterspeech. In summary, our main contributions are: â˘We propose and test novel strategies for generating adapted and personalized counterspeech that accounts for the community, user, and conversation context. â˘We show that incorporating contextual information can enhance perceived adequacy and per- suasiveness for specific lightweight prompting configurations, while also demonstrating that many adaptation and personalization strategies do not improve over generic baselines and can even degrade human-perceived quality. â˘We identify which characteristics of the toxic message and the counterspeech (i.e., subtypes of toxicity and emotional profile) influence human perceptions of persuasiveness, providing guidelines for future developments. â˘We show that known limitations of automated evaluation metrics also apply to contextualized counterspeech, underscoring the need for evaluation frameworks that better capture human judgments of counterspeech quality. 2 Related Work 2.1 LLM-generated counterspeech Counterspeech is a moderation strategy that mitigates harmful online behavior without relying on restrictive actions, thereby preserving freedom of expression [15]. Prior work has shown its potential across observational, quasi-experimental, and controlled experimental settings, including interventions by scholars, NGOs, and individual users [16,29,34,40]. However, its effectiveness depends on message design, target audience, platform context, and intervention strategy, and there is still limited understanding of which types of counterspeech work best under which conditions [15, 43]. LLMs offer a scalable way to generate counterspeech, and recent work has explored several strategies to improve response quality. Tekiroglu et al. [67]evaluate pre-trained LLMs and decoding strategies for generating persuasive counterspeech. Wang et al. [73]instead use reinforcement learning to optimize counterspeech for factuality and faithfulness. Other approaches incorporate socio-psychological strategies, such as empathy, disapproval, appeals to moral norms, warnings about consequences, and humor [10,30]. Further work uses Retrieval-Augmented Generation to improve factual grounding and specificity [2,48,53], or emotional targeting methods, such as attribute prefix learning, to adapt the affective tone of the response [51]. Despite these advances, most LLM-based counterspeech systems generate responses mainly from the toxic message itself, with limited attention to the broader conversational setting or to the characteristics of the target user. Among the few exceptions are DoÄanç and Markov[24]who incorporate demographic attributes, Bär et al. [10]who generate counterspeech conditioned on the hateful post, and Ngueajio et al. [56]who use persona-based LLMs. Our work extends this nascent line of research [22] by systematically comparing multiple contextual signals, including user profiles, conversational dynamics, and community norms, and by evaluating how these signals affect perceived counterspeech quality. 2.2 Persuasiveness of LLM-generated messages The persuasive potential of LLMs has been studied across several domains, from LLM-to-LLM persuasion to human attitude and behavior change [8,49,62]. Much of this work focuses on whether generated messages can influence usersâ decisions or beliefs. For example, Furumai et al. , Vol. 1, No. 1, Article . Publication date: July 2026. 4Cima et al. [27]compare emotional appeals, rational arguments, and stylistic variations in donation campaigns, while Karande et al. [50]show that LLMs can act as virtual sales agents by adapting language to customer profiles. In a more adversarial setting, Goldstein et al. [35]show that AI-generated propaganda can be as persuasive as human-written one. Political and social persuasion have also received increasing attention. Hackenburg et al. [38]study how LLM-generated messages affect political attitudes, voter preferences, and civic engagement, with evidence that larger models do not necessarily produce stronger persuasive effects. Other studies examine belief change in more delib- erative settings. Costello et al. [20]use LLM-based dialogues to reduce conspiracy beliefs, finding effects that persist over time, whereas Potter et al. [58]show that biased LLMs can unintentionally shape usersâ opinions even when they are not explicitly designed to persuade. Personalization is a recurring factor in this literature. Pauli et al. [57]examine how persona-based variations affect message persuasiveness across audiences, while Gennaro et al. [30]apply social-psychological framing strategies to generate persuasive counterspeech aimed at reducing intergroup hostility. Relatedly, LLMs have been explored as moderators or facilitators of online discourse, with studies evaluating their ability to mitigate toxicity, depolarize discussions, or improve conversational outcomes through constructive responses [14,36,43,44]. However, the persuasiveness of contex- tualized and personalized counterspeech remains comparatively underexplored, especially when considering how different contextual signals affect perceived adequacy, persuasive potential, and user-specific fit. 2.3 LLM adaptation and personalization Most moderation strategies follow a one-size-fits-all approach, applying the same intervention across users and contexts. While scalable, this ignores the psychological and social dynamics underlying harmful behavior. Recent work therefore highlights the potential of contextualized and personalized interventions [18,20,22,71]. LLMs offer a scalable way to adapt messages to users and situations, but their use for personalized moderation and counterspeech remains limited. Some work aligns counterspeech with the toxic message or its affective tone. Most closely related to our work, Bär et al. [10] tested LLM-generated contextualized counterspeech in a field experiment, finding that it could backfire by increasing toxicity, while generic warning-based counterspeech reduced subsequent hateful posting. This result shows that contextualization is not necessarily beneficial by itself and that more systematic work is needed to understand which design choices make contextualized counterspeech effective. Similarly, Kumar et al. [51]adapt counterspeech to emotional tone using emotion-guided prefix learning. Other personalization strategies adapt LLM outputs to user-level traits. Some works condition messages on socio-demographic characteristics [4, 24,32,62], while others use personality traits or persona-based profiles to shape the tone, framing, and argumentative style [47,56,57]. In the political and disinformation domains, LLMs have also been used to tailor persuasive messages to political attitudes, prior beliefs, or target-group characteristics [6,37,78]. These studies suggest that personalization can affect persuasive or corrective interventions by improving audience fit, but they also raise concerns about behavioral targeting, manipulation, and unequal treatment across user groups. 3 Problem definition T=â¨í 0 ,í 1 , . . . ,í í âŠrepresents an online conversation thread, where messagesí 0 ,í 1 , . . . ,í í appear in chronological order. Given a toxic messageí í , our goal is to generate a counterspeech response Ë í í+1 =G(í í , C í ), whereGis the counterspeech generator and C í denotes the contextual information. Unlike prior work, whereGonly takesí í as input [10,14,41,44,53], here we generate adapted and personalized counterspeech by also incorporating C í as input toG. , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech5 Next, we identify two sets of properties that define high-quality counterspeech. The first includes fundamental goals of a counterspeech, while the second includes properties that are specific to AI-generated contextualized content. We intentionally exclude basic linguistic features like fluency and grammaticality, which are already reliably handled by modern LLMs [55]. 3.1 Desired properties of effective counterspeech â˘Politeness. Polite messages are more likely to engage toxic users and bystanders constructively [74]. Politeness also aligns with ethical standards by discouraging hostile interactions. â˘Adequacy. Effective counterspeech should be suitable as an intervention to the toxic message, appropriately addressing or challenging the harmful content [5]. â˘Relevance. Counterspeech must align with the specific content and context of the thread. Generic replies often fall short in terms of resonance and impact [22, 31]. â˘Diversity. A diverse set of responses can appeal to a broader range of users and situations [54], while also improving perceived authenticity and genuineness, avoiding predictable responses. â˘Truthfulness. Accurate and factual responses are essential to preserving credibility and trust within the community [5, 73]. â˘Persuasiveness. Persuasive counterspeech encourages attitude or behavior change, either by directly influencing the toxic user or by promoting prosocial norms among bystanders to have positive interactions [44]. 3.2 Relevant properties of contextualized AI-generated counterspeech ⢠Adaptation. Counterspeech should reflect the norms, tone, and content of the specific community or conversation. Adapted messages are more likely to be accepted and taken seriously, increasing their credibility and persuasive power [31]. â˘Personalization focuses on user characteristic and behavior, differently from adaptation that is based only on the conversation context. Tailoring counterspeech to the traits or history of the toxic user can enhance rapport, reduce resistance, and increase impact [22]. Personalization strengthens persuasion by showing empathy or strategic alignment. ⢠Artificiality refers to the degree to which counterspeech is perceived as being machine-generated rather than human-authored. A high level of perceived artificiality can undermine the messageâs credibility and emotional resonance. Moreover, perceived artificiality can trigger negative reac- tions, especially if users feel deceived or dismissed by automated moderation tools [33]. Recent work further shows that AI-generated and human-written counterspeech can be distinguishable by both humans and classifiers, partly due to differences in linguistic characteristics, politeness, and specificity [64]. Reducing artificiality by improving the naturalness, tone, and specificity of the message, helps counterspeech feel more genuine, increasing the likelihood that it will be taken seriously and achieve its persuasive intent [31]. 4 Methods 4.1 Generation We employ an instruction-tuned version of LLaMA2-13B 1 for generating counterspeech responses. The instruction prompts used during generation are listed in Appendix B. Then, we explore a range of generation configurations by varying the input information provided to the model and the data 1 https://huggingface.co/dfurman/Llama-2-13B-Instruct-v0.2 , Vol. 1, No. 1, Article . Publication date: July 2026. 6Cima et al. used for fine-tuning. 2 Below, we describe the experimental factors under evaluation, each identified by a unique [label]. Some factors do not involve adaptation nor personalization: â˘Base [Ba]: The unmodified instruction-tuned LLaMA2-13B model. When used alone, this con- stitutes our baseline configuration. ⢠Counterspeech fine-tuning: We fine-tune the base model for the task of counterspeech genera- tion using two publicly available datasets: [Mu] MultiCONAN [26] which comprises 500 curated pairs of hate speech and counterspeech spanning multiple target categories (e.g., race, religion, nationality, sexual orientation, disability, and gender); [Hs] the Reddit hate-speech intervention (RHSI) dataset [59] consisting of 5,020 Reddit conversations containing human-authored inter- ventions. For consistency with our generation task, we retain only instances featuring a toxic comment paired with a human-written reply, and with a maximum message length of 250 words, yielding a filtered subset of 2,974 examples. The above configurations provide reference settings inspired by prior work on automatic counter- speech generation, where responses are generated from the toxic message alone and models are optionally fine-tuned on counterspeech or intervention datasets [26,67]. In contrast, the following experimental factors introduce novel contextual information to the generator. 4.1.1 Adaptation. To improve contextual alignment with the moderation setting, we introduce adaptation strategies that tailor the modelâs output to the platform and conversation in which the toxic message occurs: â˘Community [Re]: As our study is situated within Redditâs political communities, we adapt the generator to the platformâs stylistic and linguistic norms by fine-tuning it on comment-reply pairs sampled from five high-activity political subreddits (see Section 5). This adaptation familiarizes the model with Reddit-specific language, tone, and discourse conventions, making its responses more natural and platform-appropriate. â˘Conversation [Pr]: Because toxic messagesí í appear within broader threads, we enrich the modelâs input with up to two preceding messages from the same conversation (í íâ1 ,í íâ2 ). This conversational context helps the model better understand the flow of discourse and produce counterspeech that directly responds to the ongoing exchange, rather than treating messages in isolation. 4.1.2 Personalization. To further increase the effectiveness of the counterspeech, we personalize the modelâs responses based on user-specific characteristics. This allows the generator to produce interventions that are not only context-aware but also user-tailored: â˘Comment history [Hi]: We provide the generator with a window into the userâs prior behavior by prepending each toxic messageí í with ten of the userâs previous Reddit posts. This history offers implicit cues about the userâs tone, typical topics, and style, which the model can leverage to craft more targeted and relatable responses. â˘Summary [Su]: As an alternative to providing raw message history, we distill a broader sample of the userâs activity (twenty past messages) into a compact summary. This summary is generated by an instruction-tuned LLaMA2-13B model prompted to describe the userâs writing style, preferred lexicon, and main thematic interests (prompt detailed in Appendix B). These user profiles are then used as input to the counterspeech generator, serving as an explicit, structured source of personalization. 2 All models are publicly available at https://huggingface.co/collections/alemiaschi/contextualized-counterspeech-llama-2- models-679e322e663f033e1a654f2 , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech7 We implement and evaluate 36 different generation configurations by varying the combinations of the aforementioned factors. 3 The resulting design is factorial but not fully crossed. We excluded combinations that would be conceptually meaningless or structurally redundant. In particular, factors corresponding to different model variants are mutually exclusive: [Ba] cannot be combined with [Mu] or [Hs], since these represent alternative model states rather than independent prompt- level factors. We also excluded combinations of factors that are mechanically dependent. For example, [Su] is derived from the same pool of prior user comments used for [Hi]; combining the two would duplicate the same user information and make their individual effects difficult to interpret. When combining multiple datasets for fine-tuning (e.g., [Mu] and [Hs]), we adopt a multi-task learning setup by merging the datasets during training. This comprehensive design enables us to systematically analyze the contribution of each contextual and personalization factor to the quality of the generated counterspeech. 4.2 Evaluation The goal of our evaluation is twofold: (i) to assess the degree to which the generated counterspeech messages fulfill the desired properties outlined in Section 3; and (i) to identify which genera- tion configurations are most effective. To this end, we adopt a hybrid semi-automatic evaluation framework. First, we conduct a broad algorithmic assessment across all model configurations using a set of quantitative indicators. Then, based on this automatic screening, we select a subset of configurations for in-depth human evaluation. 4.2.1 Algorithmic evaluation. Following recent work [5,41,61], we define a set of evaluation indicators, each corresponding to one or more of the targeted properties. These indicators allow us to systematically compare the outputs of different configurations: â˘Relevance: We assess the relevance as the topical alignment between each toxic messageí í and its corresponding counterspeech Ë í í+1 by computing the ROUGE score between the two texts, which captures lexical overlap and content similarity. ⢠Diversity: To evaluate the variability of counterspeech responses within each configuration, we compute an intra-configuration diversity score: Diversity= 1â 1 í(íâ 1) í âď¸ í=1 í âď¸ í=1 íâ í ROUGE( Ë í í , Ë í í ), whereROUGE( Ë í í , Ë í í )is the similarity between two counterspeech messages Ë í í and Ë í í andíis the number of generated messages in the configuration. Higher diversity indicates less repetition and more lexical variety. ⢠Readability: We measure how readable the generated messages are using the Flesch Reading Ease Score (FRES), which is based on sentence length and word complexity. ⢠Toxicity: To ensure that counterspeech does not inadvertently perpetuate harmful language, we evaluate the toxicity of each generated message using Googleâs Perspective API. 4 We use toxicity as a safety-oriented inverse proxy for impoliteness: while low toxicity does not necessarily imply that a response is polite, highly toxic counterspeech is incompatible with the goal of producing respectful and constructive interventions. â˘Adaptation: We assess how much each configuration deviates from the baseline style and content by computing the diversity (i.e., 1â ROUGE) between the counterspeech messages generated 3 For example, [BaPrHi] refers to the base LLaMA2-13B model receiving both prior conversation messages and the userâs recent comment history as input. 4 https://perspectiveapi.com/ , Vol. 1, No. 1, Article . Publication date: July 2026. 8Cima et al. by the baseline model [Ba] and those produced by the configuration under evaluation. This captures the degree of contextual adaptation introduced. ⢠Personalization lex : We measure how lexically aligned the counterspeech is with the moderated userâs typical language. Specifically, we compute the ROUGE score between each generated message Ë í í+1 and a sample of prior messages authored by the same user. â˘Personalization wri : We also assess similarity in writing style by extracting linguistic profiles for both the counterspeech and the user messages using ProfilingUD, which captures over 130 syntactic and morphological features [9]. We then compute the Spearman correlation between the stylistic profiles of each counterspeech message and the toxic userâs past writing. 4.2.2Configuration selection. While the algorithmic evaluation provides a scalable and systematic way to compare all configurations, it falls short in capturing subjective and nuanced properties such as persuasiveness and artificiality. Moreover, the accuracy and interpretability of some existing automatic indicators remain open to discussion [39,77]. To mitigate these limitations, we comple- ment our quantitative analysis with a targeted human evaluation conducted via crowdsourcing. Given the large number of configurations, a complete manual assessment is infeasible. Therefore, we adopt a principled approach to select a representative subset. First, we generate a performance ranking for each individual indicator by ordering all configurations according to their respective scores. These rankings are then aggregated into a single global ordering (i.e., a âsuper-rankingâ) by solving an optimization task that minimizes Spearmanâs footrule distance across the rankings. This super-ranking serves as the basis for configuration selection. From the aggregated ranking, we select six representative configurations: the highest- and lowest-performing configurations among those that implement only adaptation, only personalization, and a combination of both. This selection strategy allows us to compare the effectiveness of each contextualization strategy across performance extremes, while also enabling us to evaluate the alignment between automatic indicators and human judgment. In addition, we include the baseline configuration [Ba] in the evaluation set to serve as a reference point for all comparisons. To ensure that the messages assessed by human annotators are representative of each configurationâs overall behavior, we further select a subset of counterspeech messages per configuration. For each of the seven selected configurations, we compute the centroid of the configuration across all evaluation indicators. We then identify and select the 20 counterspeech messages that are closest to this centroid in terms of their indicator vectors. This ensures that the messages chosen for human evaluation are not outliers, but rather typical examples of the outputs generated by each configuration. This hybrid selection process ensures that the subsequent human evaluation is both meaningful and efficient, enabling a thorough assessment of the configurationsâ real-world effectiveness in generating counterspeech. 4.2.3 Human evaluation. We conduct an extensive human evaluation through a pre-registered, 5 mixed-design crowdsourcing experiment on Amazon Mechanical Turk. 6 The study is structured around a two-tiered design comprising both between-subjects and within-subjects components. At the start of the task, participants are randomly assigned to one of two between-subjects conditions: (i) a non-contextual condition, in which participants are shown only the toxic messageí í and the corresponding counterspeech Ë í í+1 ; or (i) a contextual condition, where participants are additionally shown the contextual information used to generate the counterspeech, including conversation his- tory and user-specific features when applicable. After this initial assignment, participants proceed to evaluate multiple(í í , Ë í í+1 )pairs in a within-subjects design. Each participant rates responses from all seven model configurations selected during the configuration selection phase. The presentation 5 https://aspredicted.org/b55z-5qy4.pdf 6 This research received ethical approval from CNRâs IRB (protocol #0306210). , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech9 order of these configurations is randomized to control for order effects and potential fatigue biases. Participants are asked to evaluate each counterspeech response on a five-point Likert scale across several dimensions: its relevance to the toxic message, its adequacy as a form of counterspeech, its truthfulness, its perceived artificiality, and its persuasiveness. We operationalize persuasiveness as perceived persuasiveness, namely participantsâ third-party assessment of the responseâs persuasive potential. Specifically, following recent work [44], persuasiveness is assessed through two distinct items: the perceived likelihood that the counterspeech (i) persuades the author of the toxic message to re-engage in a more civil manner, and that it (i) steers the broader conversation back toward civil discourse. Participants in the contextual within-subjects condition are also asked to assess how contextualized the counterspeech feels in relation to the surrounding conversation and user behavior. Finally, all participants complete a brief socio-demographic questionnaire. The full list of evaluation questions and instructions is provided in Appendix D. This mixed design enables a comprehensive analysis of both model performance and the role of contextual information. In particular, the between-subjects comparison isolates the added value of context in shaping usersâ perceptions of counterspeech quality, while the within-subjects design allows for direct comparison across generation strategies. 4.2.4Statistical analysis. To analyze the results from the within-subjects evaluation, we first assess whether differences across the evaluated configurations are statistically significant using Friedman tests. To identify specific configurations that differ significantly from the baseline, we perform pairwise comparisons using Wilcoxonâs signed-rank tests, applying a Bonferroni correction to control for multiple hypothesis testing. For each comparison, we report effect sizes and confidence intervals using the matched-pairs rank biserial correlation coefficient. In addition, we evaluate whether contextual information significantly influences participant judgments by comparing results between the two between-subjects groups. For each configuration, we perform two-sample MannâWhitney U tests with Bonferroni correction, followed by effect size estimation via Glass rank biserial correlations. 4.2.5 Power analysis. We target a sample size ofâ2,500 participants for each within-subjects experiment. This number allows us to detect small effect sizes (Cohenâsí=0.2) with strong statistical power (85%) at a 95% confidence level. 5 Data We collected Reddit comments posted over multiple years from five widely followed subreddits that regularly engage in discussions about U.S. politics. These include two ideologically polarized communities, one right-leaning (r/conservatives) and one left-leaning (r/progressive), as well as two subreddits centered around prominent political figures:r/the_donald, focused on Donald Trump (right-leaning), andr/aoc, focused on Alexandria Ocasio-Cortez (left-leaning). To ensure a more ideologically mixed setting, we also includedr/politicsâa large, general-interest political subreddit with varied viewpoints. To balance the volume of data across subreddits of different sizes, we collected comments spanning either 36 or 12 months per subreddit. Forr/the_donald, collection ended in June 2020 when the subreddit was permanently banned [11,18]. All comments were retrieved using the Pushshift dataset [3], which provides historical Reddit data. Counterspeech dataset. To create a benchmark set of toxic comments for counterspeech generation, we first computed toxicity scores for all collected comments using Googleâs Perspective API [54]. We relied on Perspective API because it is widely used for measuring online toxicity and operationalizes toxicity as language that is likely to make someone leave a discussion, which is closely aligned with our goal of identifying comments that may disrupt civil conversation and motivate counterspeech , Vol. 1, No. 1, Article . Publication date: July 2026. 10Cima et al. interventions. We retained only those comments with a toxicity scoreâĽ0.5, a threshold commonly used in literature [52,70], and that were part of conversation threads with at least two parent messages. This filtering process yielded 292 toxic comments across 49 distinct threads. From this set, we discarded comments written by users with fewer than 20 comments in the considered period in all Reddit. In such cases we could not infer sufficient user-level information to generate personalized counterspeech, as required by the personalization strategies described in the next paragraph. This additional filtering step yielded the final benchmark of 128 toxic comments across 35 distinct threads, as shown in the Appendix Table 3. Although the number of threads is relatively balanced across the considered subreddits, most toxic comments originate fromr/politics. This pattern is consistent with the larger size and broader activity ofr/politicscompared to the other communities in our dataset. At the same time, it highlights a realistic challenge for moderation systems: high-activity communities may generate a larger absolute volume of toxic content even when toxicity is distributed across multiple discussion threads. This further motivates the need for scalable counterspeech interventions that can help de-escalate toxic exchanges and support the restoration of constructive discussion. Each of the 36 counterspeech generation configurations described in Section 4.1 was tasked with generating one counterspeech response for each of the 128 toxic messages, resulting in a total of 4,608 generated counterspeech responses to evaluate. Adaptation and personalization datasets. To implement the contextual strategies outlined in Section 4.1, we constructed additional datasets tailored to each type of adaptation or personalization. For the community adaptation factor [Re], we selected a random stratified sample of approximately 7,500 commentâreply pairs from the five political subreddits. This dataset was used to fine-tune the base model to the conversational norms, tone, and linguistic style typical of political discourse on Reddit. For the conversation-level adaptation factor [Pr], we extracted the most recent parent message preceding each of the 128 selected toxic comments. Since our filtering procedure retained only toxic comments with available parent context, the dataset does not include top-level toxic comments. These parent messages were fed as contextual input via prompting at generation time. To implement personalization, we gathered twenty prior comments authored by each user who wrote one of the toxic messages, sampled from the userâs activity across Reddit rather than restricted to the target subreddit. These samples were used to generate user summaries for the [Su] condition. Additionally, ten of these user-authored comments were directly prepended to the toxic comment during counterspeech generation to implement the comment history strategy [Hi]. The selected comments were used as a practical proxy for the userâs writing style, lexical choices, and recurring topics. 6 Results 6.1 Algorithmic evaluation We evaluate all factors (N =7) and configurations (N =36) with the indicators defined in Section 4.2.1. 6.1.1Factors. We begin our analysis with the modeling of indicator values based on configuration factors. Given that there is an overlap of factors among many of the 36 configurations and that all indicator values are between 0 and 1, we employed beta regression, a robust approach to model beta distributions of values in the standard unit interval [23]. For every indicator, we used a traditional beta regression modelâhaving the logit link function and treating precision (í) as constantâbased on all the 7 single factors and their 15 pairwise interactions used in the configurations. We did not consider higher-order interactions to avoid issues with model overspecification and the additional complexity. An overview of the results of the beta regressions is depicted in Table 1, in which , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech11 adaptationâdiversityâperso. lex âperso. wri âreadabilityârelevanceâtoxicityâ termest.pest.pest.pest.pest.pest.pest.p (Int.)+1.294<.001+0.390.002-1.745<.001-0.061<.001+1.734<.001-2.122<.001-1.996<.001 Ba-2.035<.001-0.230.114-0.108.080+0.145<.001-1.433<.001+0.097.137-0.825<.001 Re+0.105<.001+0.570<.001-0.044.240+0.056<.001+0.758<.001+0.502<.001+0.569<.001 Mu-0.012.474+0.533<.001-0.024.603+0.008.303-0.277.001+0.156.003+0.015.834 Hs+0.181<.001-0.160.102-0.237<.001-0.093<.001-0.598<.001-0.329<.001+1.030<.001 Hi+0.448<.001+0.899<.001-0.188.002-0.023.029+0.014.920+0.072.270+0.104.327 Pr+0.554<.001+0.579<.001+0.012.820+0.068<.001-0.035.750+0.150.005-0.386<.001 Su+0.322<.001+0.505.001+0.011.859+0.014.195+0.030.813+0.216<.001-0.037.729 Ba :Hi+0.583<.001-0.551.002+0.256<.001+0.068<.001-0.132.362+0.118.116+0.027.865 Ba :Pr+0.145<.001-0.380.010+0.078.179-0.066<.001+0.038.746-0.114.057-0.134.271 Ba :Su+1.035<.001-0.003.988+0.013.860+0.062<.001+0.400.004+0.021.780+0.735<.001 Re :Hs-0.007.594+0.161.087+0.012.713-0.038<.001-0.356<.001-0.139<.001-0.472<.001 Re :Hi+0.060<.001-0.495<.001+0.017.687-0.001.888-0.308.001-0.158<.001-0.190.014 Re :Pr+0.081<.001-0.089.343-0.001.988-0.018.003-0.253<.001+0.020.541+0.105.095 Re :Su+0.030.071-0.422<.001+0.032.443-0.010.168-0.187.034-0.188<.001-0.048.521 Mu :Hi+0.049.022-0.357.011+0.069.223+0.057<.001+0.498<.001-0.053.384-0.466<.001 Mu :Pr-0.014.428-0.084.463+0.009.846-0.042<.001-0.017.861-0.099.0430.000.998 Mu :Su+0.049.020-0.351.010+0.013.820+0.033<.001+0.183.108-0.126.035-0.163.080 Hs :Hi-0.093<.001+0.018.878+0.227<.001+0.094<.001+0.797<.001+0.052.198-0.673<.001 Hs :Pr-0.184<.001+0.006.949+0.021.537-0.015.012+0.284<.001-0.018.578+0.039.540 Hs :Su-0.054<.001+0.169.132+0.170<.001+0.054<.001+0.699<.001+0.129.001-0.720<.001 Hi :Pr-0.562<.001-0.309<.001-0.021.528+0.015.010-0.041.561+0.001.974+0.509<.001 Pr :Su-0.552<.001-0.228.009-0.057.088-0.009.141-0.094.164-0.085.012+0.427<.001 Table 1. Overview of the beta regressions by algorithmic indicator based on all single configuration factors and allowed pairwise interactions. Only factor terms with a statistically significant p-value (í< .05) are highlighted. For estimated (est.) means, cell color represents the factorâs column-wise relative effect (i.e., at the indicator level) and its sign. All beta regressions use the standard logit link function and constant precision (í). Arrows (â/â) indicate whether higher or lower values respectively are better for each indicator. statistically significant (í< .05) estimates of means for factors are highlighted according to their relative effect on indicator values. The analysis reveals that the effect of a factor is highly dependent on the specific dimension being evaluated. No factor produces across-the-board improvements. Instead, gains in some indicators often come at the cost of degradations in others. For example, fine-tuning on the MultiCONAN dataset [Mu] and our Reddit-specific political conversations [Re] leads to noticeable improvements in relevance, diversity, and adaptation [17]. However, these same factors negatively impact toxicity and stylistic personalization, either by themselves or when interacting with other factors, suggesting a trade-off between general linguistic alignment and user-specific adaptation. This trade-off under- scores the complexity of the task: no single factor enhances all desired properties simultaneously. As such, practitioners deploying counterspeech systems may need to tailor configurations based on the specific requirements of a given use caseâprioritizing readability or relevance in some scenarios, and personalization or toxicity in others. Some factors, however, perform poorly across most metrics. Fine-tuning on the RHSI dataset [Hs] consistently degrades performance in terms of relevance, diversity, toxicity, and writing style personalization, which is sometimes exacerbated or alleviated by interacting with other factors. On the personalization front, [Su] yields improvements in both lexical and stylistic indicators when interacting with [Hi], whereas [Hi] degrades writing style and lexical personalization by itself. Interestingly, the base factor [Ba], which appears only in configurations without any fine-tuning, shows relatively favorable results in the stylistic person- alization indicator. This suggests that fine-tuning on large-scale datasets such as MultiCONAN, RHSI, or even Reddit-specific conversations may dilute the modelâs ability to align with individual usersâ writing patterns. , Vol. 1, No. 1, Article . Publication date: July 2026. 12Cima et al. 6.1.2Configurations. We now turn to a fine-grained analysis of the algorithmic evaluation results for each individual configuration. These results are reported in the Appendix Table 4, where con- figurations are grouped into four categories, from top to bottom: those with neither adaptation nor personalization, those with adaptation only, those with personalization only, and those with both. For each evaluation indicator, the best-performing configuration is shown in bold, while the remaining top five are underlined. Although many configurations yield similar scores, each group contains at least some configurations that perform well across selected indicators. Nonethe- less, key differences emerge across the groups. Notably, several configurations that incorporate both adaptation and personalization achieve top results across multiple indicators. Representative examples include [MuRePrHi], [MuHsReHi], and [MuReHi]. Configurations relying exclu- sively on adaptation also perform strongly in several cases, for instance, [MuRe], [MuRePr], and [MuHsRePr] consistently attain competitive results. In contrast, configurations that rely solely on personalization exhibit weaker and less consistent performance. Interestingly, a few configurations that involve neither adaptation nor personalization still yield solid results. The base- line configuration [Ba], as well as the version fine-tuned solely on MultiCONAN [Mu], perform comparably to more complex strategies in several indicators. This underscores that model quality is not determined purely by the number of added factors, but also by how well these are integrated. In sum, while overall performance differences across configurations tend to be modest, the results suggest that applying adaptationâor adaptation in combination with personalizationâtends to improve the quality of generated counterspeech more consistently than using personalization alone or applying no contextualization at all. In the Appendix Section D.4, we present a comparative analysis with the counterspeech generated by another model, Qwen3-8B. 6.1.3 Extended evaluation of algorithmic indicators. Several algorithmic indicators defined in Section 4.2.1 rely on ROUGE to compute text similarities. However, multiple such metrics exist in the literature. To assess the robustness of our findings, we re-implemented all ROUGE-based indicators using BLEU and BERTScore, and compared the results across the three implementations. Figure 2 reports linear and rank correlations for the four indicators originally based on ROUGE. Overall, we observe strong agreement between implementations. We measured perfect consistency for lexical personalization and high correlations for both adaptation and diversity. In contrast, measuring relevance with BERTScore yields results largely uncorrelated with those obtained via ROUGE or BLEU, which remain closely aligned. These results indicate a relatively strong invariance of the considered indicators with respect to the underlying text similarity metric used. 6.1.4 Configuration selection. Next, we select a subset of configurations from Table 4 to further evaluate manually. The selection is carried out via the methodology described in Section 4.2.2. Specifically, we identify the best- and worst-performing configurations within each of the three main strategies: adaptation-only, personalization-only, and the combined approach. These six con- figurations are highlighted in Table 4. In addition to these, we also include the baseline configuration [Ba ] to provide a meaningful reference point for comparison. 6.2 Human evaluation We recruitedí=2,444 andí=2,353 participants on Amazon Mechanical Turk for the non- contextual and contextual between-subjects conditions, respectively. These numbers exclude par- ticipants whose responses were rejected due to excessively fast completion times, fewer than <100% correct answers to control questions, or equal responses across all items. We report no deviations from our pre-registered protocol. We further assessed the reliability and similarity of crowdworkersâ judgments by computing Krippendorff âsíź, pairwise Spearman correlation, pairwise agreement, and normalized match distance, reported in the Appendix Table 5. Results show very , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech13 0.84â0.02 â0.12 RG BL BS 0.830.80 0.89 RG BL BS 0.940.96 0.93 RG BL BS 1.001.00 1.00 RG BL BS 0.660.02 0.06 RG BL BS 0.650.63 0.71 RG BL BS 0.460.70 0.53 RG BL BS 1.001.00 1.00 RG BL BS 0.820.01 0.09 RG BL BS 0.820.75 0.86 RG BL BS 0.600.86 0.71 RG BL BS 1.001.00 1.00 RG BL BS Pearson Kendall Spearman relevance personalization lex diversityadaptation Fig. 2. Linear and rank correlations (y axis) between four algorithmic indicators (x axis) when computed with different text similarity metrics: ROUGE (RG), BLEU (BL), and BERTScore (BS). lowíźand correlation values, indicating limited consistency in annotatorsâ relative judgments, but moderate pairwise agreement and low normalized match distance, reflecting that ratings are often concentrated in a narrow range around the middle-high values of the Likert scale. 6.2.1Non-contextual experiment. Participants in the non-contextual condition evaluated pairs con- sisting only of a toxic message and a counterspeech response. A Friedman test reveals statistically significant differences across the evaluated configurations. Figure 3a reports the effect sizes, confi- dence intervals, and statistical significance for each configuration compared to the baseline [Ba ]. The results show two distinct groups of performance. Configurations such as [MuRe], [HsHi], [MuHsHi], and [MuRePrHi] consistently perform worse than the baseline across all evaluated aspects, except for artificiality, where they are perceived as more human-like. This pattern is consistent with recent evidence that AI-generated counterspeech can remain distinguishable from human-written counterspeech in terms of linguistic characteristics, politeness, and specificity [64]. These results suggest that while these models generate responses that seem less machine-generated, they nonetheless produce counterspeech perceived as less relevant, adequate, truthful, or persuasive than that of the baseline. In contrast, configurations [BaPr] and [BaPrHi] match or exceed the baseline in some dimensions, with statistical significance. Notably, [BaPrHi] performs sig- nificantly better in terms of adequacy and its perceived capacity to persuade the author of the toxic message. It also scores higher (though not significantly) in relevance, truthfulness, and the perceived ability to persuade bystanders. Configuration [BaPr] performs similarly to the baseline, but yields slightly higher scores in both persuasiveness questions. We emphasize that these results concern perceived persuasiveness: participants judged the likely persuasive effect of each response on another user or on the broader conversation, rather than reporting their own attitude change. Therefore, the findings should be interpreted as evidence of perceived persuasive potential, not as direct evidence of actual persuasion. Overall, results from the non-contextual within-subjects condition indicate that [BaPrHi] improves upon the baseline in terms of perceived adequacy and persuasiveness, supporting the value of contextualized counterspeech. Interestingly, the best-performing configurations in this experimentâ[BaPr] and [BaPrHi]âwere among the worst ranked by the algorithmic indicators in Table 4, suggesting a potential misalignment between automated and human evaluations. , Vol. 1, No. 1, Article . Publication date: July 2026. 14Cima et al. BaPr MuRe HsHi MuHsHi BaPrHi MuRePrHi (a) Non-contextual condition. BaPr MuRe HsHi MuHsHi BaPrHi MuRePrHi (b) Contextual condition. Fig. 3. Human evaluation results in terms of effect sizes (blue dots) and confidence intervals (black bars) for the scores assigned to various configurations compared to the baseline. Statistical significance levels are indicated as follows: ***: í< 0.01, **: í< 0.05, *: í< 0.1. Ba BaPr MuRe HsHi MuHsHi BaPrHi MuRePrHi Fig. 4. Differences in human evaluation results between the contextual and non-contextual conditions. Statis- tical significance: ***: í< 0.01, **: í< 0.05, *: í< 0.1. 6.2.2Contextual experiment. Participants in the contextual condition received additional informa- tion alongside the toxic message and counterspeech, including the subreddit name, the previous message in the thread, and a user summary obtained as described in Section 4.1.2. Once again, a Friedman test reveals statistically significant differences across configurations. Detailed results, shown in Figure 3b, largely mirror those from the non-contextual experiment. Configurations [BaPr] and [BaPrHi] consistently achieve the highest scores. In particular, [BaPr] yields a statistically significant improvement over the baseline in its ability to persuade the author of the toxic message. While most other improvements are not statistically significant, these configurations clearly outperform alternatives such as [MuRe], [HsHi], [MuHsHi], and [MuRePrHi], which again show significantly worse performance than the baseline, except for marginal improvements in artificiality. This experiment also included an additional item with respect to the non-contextual one, evaluating the perceived contextualization of each response. Figure 3bG shows that [BaPrHi] scores markedly better than the baseline, whereas [MuRePrHi] performs worse, reinforcing previous findings about the benefits of the [BaPrHi ] configuration. Our study design allows for direct comparisons between the contextual and non-contextual experiments. Figure 4 shows that all configurationsâincluding the baselineâachieve higher ratings in the contextual condition across all dimensions except adequacy and truthfulness. However, the magnitude of improvement varies. The baseline and configurations that already performed well, such as [BaPr] and [BaPrHi], show relatively modest gains. In contrast, weaker configurations benefit more substantially from contextual information. These differences may be partly attributable , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech15 Fig. 5. Aggregated rankings of the selected configura- tions, based on algorithmic and human evaluations. Rank- ings based on algorithmic indicators are negatively cor- related with human evaluations, highlighting a marked mismatch between the two. 0.90 0.71 0.62 â0.05 0.05 â0.33 â0.43 â0.33 â0.71 0.62 aRG aBL aBS hNC hWC â1.0 â0.5 0.0 0.5 1.0 Fig. 6. Rank correlation coefficients for the se- lected model configuration rankings between algorithmic indicators based on ROUGE (aRG), BLEU (aBL) and BERTScore (aBS), and human ones without (hNC) and with context (hWC). to the evaluation setting itself: providing raters with conversation- and user-level context may help them interpret both the toxic message and the counterspeech response more coherently, leading to generally higher judgments. In addition, the contextual condition included an explicit question about whether the response was personalized rather than generic, which may have made the study hypothesis more salient to participants and potentially influenced their other ratings. 6.2.3Comparing algorithmic and human evaluations. We compare the performance of the selected configurations across algorithmic and human evaluations. For each evaluation method (i.e. ROUGE- based quantitative indicators, human assessments with and without context) Figure 5 presents the aggregated rankings of configurations across all considered aspects, with the best-performing configurations at the top. Rankings are aggregated using the method described in Section 4.2.2. As shown, the rankings from the two human evaluations are broadly consistent with each other, yielding a Kendall rank correlationí=0.62. In contrast, the ranking derived from the quan- titative indicators diverges markedly from both human evaluations, with negative correlations í=â0.05 andí=â0.43 relative to the non-contextual and contextual experiments, respectively. Figure 6 extends this analysis by incorporating algorithmic indicators computed with BLEU and BERTScore, in addition to those based on ROUGE. The figure reports Kendall rank correlations between every evaluation method. Consistent with Figure 2, the different implementations of the algorithmic indicators remain strongly correlated with one another, and the two human-based evaluations also show strong internal agreement. However, correlations between algorithmic and human evaluations are either negligible or, in most cases, moderately to strongly negative. Notably, BERTScore-based evaluations exhibit a strong negative correlation with human judgments in the contextual experiment, withí=â0.71. Our findings provide domain-specific evidence that the well-known evaluation problem of automatic metrics in NLP generation also arises in contextualized counterspeech generation [77], where key properties such as adequacy, perceived persuasiveness, contextualization, and artificiality are highly subjective and difficult to capture with surface-level similarity measures. At the same time, the existence of strongâalbeit negativeâcorrelations opens up intriguing opportunities for developing predictive models, for instance via classification or regression, to better approximate human judgments of counterspeech effectiveness, an area that has received little attention so far [7]. 6.2.4 Impact of toxicity types and emotional profiles on counterspeech persuasiveness. Here we refine the evaluation of the generated counterspeech messages by investigating persuasiveness in relation to the type of toxic speech that the counterspeech corrects, and their emotional profile. , Vol. 1, No. 1, Article . Publication date: July 2026. 16Cima et al. identity_attack insult obscene sexually_explicit threat toxicity type 0.0 0.2 0.4 0.6 0.8 1.0 score toxic message counterspeech identity_attack insult obscene sexually_explicit 3.4 3.5 3.6 3.7 3.8 3.9 4.0 4.1 persuasiv. (toxic user) identity_attack insult obscene sexually_explicit 3.4 3.5 3.6 3.7 3.8 3.9 4.0 4.1 persuasiv. (conversation) Fig. 7. Distribution of toxicity subtypes in toxic messages and the generated counterspeech (left-side panel) and persuasiveness of the counterspeech with respect to the subtypes of toxic speech in the toxic message (right-side panels). anger anticipation disgust fear joy sadness surprise trust emotion type 2 1 0 1 2 3 4 score toxic message counterspeech anticipation trust disgust joy anger fear 3.4 3.5 3.6 3.7 3.8 3.9 4.0 4.1 persuasiv. (toxic user) ** trust anticipation fear disgust anger joy 3.4 3.5 3.6 3.7 3.8 3.9 4.0 4.1 persuasiv. (conversation) ****** Fig. 8. Distribution of emotions subtypes in toxic messages and the generated counterspeech (left-side panel) and persuasiveness of the counterspeech with respect to the dominant emotion of the counterspeech message (right-side panels). Types of toxicity. In addition to the overall toxicity score used so far, Perspective API also provides scores for the following subtypes of toxic speech: identity attack, insult, obscene, sexually explicit, and threat. The left-side panel of Figure 7 provides a quantitative characterization of the toxicity profile of the 128 toxic messages in our dataset by reporting their subtype scores. The same panel also reports the corresponding scores for the generated counterspeech. As expected, counterspeech messages exhibit overall low toxicity scores, while the original toxic messages display substantially higher values across several toxicity dimensions. Among the toxic messages, obscene content is the most prevalent subtype, followed by insult and sexually explicit content. This distribution clarifies the scope of our benchmark and indicates that the evaluation is mostly shaped by obscene and insulting toxic language, whereas other toxicity subtypes, such as threat, are less represented. Based on these scores, we then identified which toxicity subtypes are more effectively mitigated by counterspeech. For this purpose, we define the dominant toxicity subtype of a message as the one with the highest score among those assigned to that message by Perspective API. The right- side panels of Figure 7 show the persuasiveness scores obtained by the generated counterspeech messages with respect to the subtypes of toxic speech of the corresponding toxic message. 7 As shown, the counterspeech that we generated appears most effective when responding to identity attacks. In contrast, responses to sexually explicit content achieved the lowest median ratings. While this pattern suggests that certain types of toxicity may be more amenable to persuasive intervention, the differences are not statistically significant, likely due to the relatively small size of our sample, which amounts to 140 counterspeech instances. Emotions. We obtain the emotional profile of both toxic and counterspeech messages with EmoAtlas, 8 a framework to detect the presence of eight core emotions: anticipation, anger, dis- gust, fear, joy, sadness, surprise, and trust. EmoAtlas returns Z-scores representing the intensity 7 Threat is excluded from this analysis due to its limited presence in our data. 8 https://github.com/massimostel/emoatlas , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech17 203040506070 age 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 persuasiv. (user) r: -0.192 203040506070 age 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 persuasiv. (conversation) r: -0.202 203040506070 age 1.0 1.5 2.0 2.5 3.0 3.5 4.0 4.5 5.0 artificiality r: -0.167 Fig. 9. Relationship between respondentsâ age and their judgments: persuasiveness with respect to the toxic user (left panel), persuasiveness with respect to the conversation (middle panel), and artificiality of the counterspeech (right panel). Panels also report Pearson correlations í (statistically significant at í< 0.01). of each emotion relative to its training corpus. The left-side panel of Figure 8 shows the distri- butions of these emotions. We note that both toxic messages and counterspeech overall present more positive scores than negative ones, indicating that both types of messages generally exhibit stronger emotions than the EmoAtlas baseline. Furthermore, the emotions distributions of toxic and counterspeech messages are very similar, implying that, overall, the counterspeech mirrored the emotional profile of the comments it addressed. Next, we identified the dominant emotion in each toxic and counterspeech message as the one with the largest Z-score. This allowed to assess possible variations in persuasiveness of the counterspeech with respect to the dominant emotion of the counterspeech itself or the toxic message it responds to. Results shown in Appendix Figure 13 reveal that the dominant emotion in a toxic message does not significantly affect the perceived persuasiveness of the counterspeech response. Instead, the right-side panels of Figure 8 show marked differences based on the dominant emotion of the counterspeech message. 9 Counterspeech conveying anticipation and trust exhibit the highest median persuasiveness scores for both the user- and conversation-oriented evaluations. In contrast, counterspeech expressing anger is consistently and significantly less persuasive than others (í<0.05, Wilcoxon test). This result suggests that crafting counterspeech that conveys positive emotions may be a promising strategy to enhance persuasiveness. Interestingly however, counterspeech expressing joy also scores relatively low, especially for persuasiveness toward the conversation (í<0.01), possibly indicating that excessive positivity may feel out of context or less credible in confrontational settings. 6.2.5 Age-related differences in perceptions of counterspeech. Our questionnaire comprised a set of socio-demographic questions, 10 including the age of the participant, as that might influence user evaluations. To investigate its possible impact, we analyzed whether increases in respondentsâ age correlate with differences in their assessments. We grouped together the non-contextual and contextual responses and we computed the average score that each participant assigned across the 14 randomly selected counterspeech messages they evaluated. As shown in Figure 9, there is a noticeable tendency for older respondents to assign lower scores. This trend is reflected in the moderately negative yet significant (í<0.01) correlations observed between age and three key dimensions: persuasiveness with respect to the toxic user (í=â0.192), persuasiveness with respect to the conversation (í= â0.202), and artificiality (í= â0.167). Several factors can explain this phenomenon. Older individuals may be more wary of AI-generated content, perceiving messages as less trustworthy or authentic. They might also be less accustomed to the informal or emotionally charged communication style often used in social media and counterspeech, which could make such messages feel less persuasive or even intrusive. Additionally, generational differences in digital 9 Sadness and surprise are excluded from this analysis due to their limited presence in our data. 10 The socio-demographic profile of the respondents is shown Appendix Figure 11. , Vol. 1, No. 1, Article . Publication date: July 2026. 18Cima et al. configuration toxicâ degenerateâ unframedâ length Ba.000.008.00056 BaPr.000.000.00059 BaPrHi.000.008.00074 MuRe.133.266.80519 HsHi.055.273.35912 MuHsHi.031.195.65614 MuRePrHi.086.406.90617 Table 2. Automatic failure indicators for the seven configurations submitted to human evaluation, computed over all 128 generations of each configuration. Values are proportions. Length is the mean generation length in tokens. Lower values are better for all indicators. literacy and familiarity with automated systems could influence expectations and evaluations, leading to a more critical stance toward the perceived quality or intent of AI interventions [76]. 6.3 Failure modes and fine-tuning effects The human evaluation identifies which configurations degrade perceived counterspeech quality, but does not explain the source of the degradation. Here we characterize the failure modes of the gen- erated messages through automatic indicators computed over all 6,912 generations. Operationally, we label a generated counterspeech message as: â˘Toxic, if its Perspective API score>0.5. This is the same threshold used in Section 5 to select the toxic comments. â˘Degenerate, if it is shorter than ten tokens, truncated, or copies a verbatimí-gram (í âĽ8) from the toxic message. ⢠Unframed, if it contains no marker of a moderation act. That is, no second-person address combined with a directive form, no explicit normative appeal, and no imperative opening. The unframed indicator captures a property of register rather than of pragmatic function, and we therefore validated it before use, as per Appendix Section D.5. Failure rates and relation to human evaluation. Table 2 reports the resulting failure rates for the seven configurations submitted to human evaluation, while the Appendix Table 6 illustrates each failure mode with representative outputs. The three configurations based on the non-fine-tuned model never exceed the toxicity threshold and produce no unframed output over the 128 toxic messages, whereas all fine-tuned configurations produce both. Note that this is a threshold-based count: base configurations do produce some mildly toxic outputs, with maximum toxicity=0.43, but none of them reaches the threshold used to select the toxic comments themselves, which is consistent with the mean toxicity scores reported in Table 4. In [MuRePrHi], by contrast, 90.6% of the responses contain no moderative framing and 8.6% are themselves toxic. For the fine-tuned configurations, the absence of a moderation act is thus the prevalent outcome rather than an occasional error. Notably, [MuRePrHi] ranks among the best configurations under the algorithmic indicators (Section 6.1), which further illustrates the misalignment between automatic metrics and both human judgment and functional adequacy. The Appendix Table 6 additionally documents a failure type that our indicators do not capture, namely factually incorrect statements, which remains a target for future automatic evaluation. , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech19 The messages submitted to annotators were the twenty closest to each configurationâs centroid in indicator space (Section 4.2.2), which raises the question of whether this procedure excluded the failures reported above. Upon inspection, we verified that it did not exclude the prevalent one: out of the centroid-proximal generations of [MuRePrHi] and [MuRe], 85% and 70% respectively are unframed, so participants did rate failing outputs. The selected sample is nonetheless cleaner than the non-selected one (pooled unframed rate 0.329 against 0.401; stratified permutation test, í=0.015), and none of the 140 evaluated messages exceeds the toxicity threshold (maximum toxicity =0.41), whereas between 3.1% and 13.3% of the full output pools of the fine-tuned configurations do. This shows that toxic generations are extreme in indicator space and were removed by centroid selection. The degradation reported in Section 6.2 should therefore be read as a conservative estimate, and the most severe failure mode of the pipeline is not observable within the human study, which motivates conducting the present analysis on the full generation pool. Fine-tuning failure. We consider three explanations for why fine-tuning degrades performance. The first is instruction overload: the model may struggle to integrate multiple concurrent signals [4, 32]. Our factorial design tests this directly, since [BaPrHi] and [MuRePrHi] receive the same toxic message, conversational context, and user history, but differ only in their weights. The former is never unframed, whereas the latter is unframed in 90.6% of cases (McNemarâs test,í â3Ă10 â31 ; 116 messages fail only under fine-tuning, and none in the opposite direction). Moreover, within the non- fine-tuned family, adding conversational context and user history never produces unframed outputs ([Ba], [BaPr], [BaHi], and [BaPrHi] are all at 0.000). 11 Instruction load is therefore not the operative factor. However, a related source of failure may lie in the modelâs difficulty in identifying and using the relevant contextual information needed to produce suitable counterspeech [60]. The second explanation is toxic priming: fine-tuning on Reddit data, especially when combined with user history, may lead the model to reproduce the targetâs linguistic behaviour [17]. A logistic model of output toxicity over the full design (í=6,912) supports this account only in part. Reddit fine-tuning increases toxicity as a main effect (OR=4.7,í<0.001), but its interaction with the comment-history factor is negative and significant (í˝=â0.80,í=0.001): toxicity is higher for [Re] without user history (9.9%) than with user history (7.6%). 12 Thus, toxicity is attributable to the fine-tuning corpus rather than to conditioning on the moderated user. This is consistent with Appendix Figure 10: 19.36% of the Reddit commentâreply pairs used for community adaptation exceed the Perspective API toxicity threshold, compared with 4.95% in MultiCONAN and 1.46% in RHSI. Reddit fine-tuning therefore exposes the model to toxic conversational material, plausibly contributing to toxic outputs and to the shift toward the toxic register. The third explanation, which the data support, concerns the register of the generated text. A factorial logistic model over the full design shows that the loss of moderative framing is driven primarily by MultiCONAN fine-tuning (OR=7.3,í<0.001) and, to a lesser extent, by Reddit fine-tuning (OR=2.1,í<0.001), whereas RHSI has no significant effect (OR=0.9,í=0.13). Toxicity, by contrast, is driven by Reddit fine-tuning alone. MultiCONAN contains argumentative counter-narratives that rebut the toxic message without addressing its author, thereby removing the moderative frame ([Mu] is unframed in 82.8% of cases, against 0.0% for [Ba]). Reddit commentâreply pairs instead model ordinary conversational participation and introduce both further loss of framing 11 The summary factor [Su] is an exception, as it produces truncated outputs in over half of the cases when combined with the base model. Since no [Su ] configuration entered the human evaluation, we do not analyse it further. 12 Since [Re] never occurs without [Mu] in our design, its coefficient is to be read as the additional effect of Reddit fine-tuning given MultiCONAN fine-tuning. , Vol. 1, No. 1, Article . Publication date: July 2026. 20Cima et al. and toxicity (13.3% toxic for [MuRe] against 0.8% for [Mu]). The configuration [MuRePrHi] combines both sources. This register shift is also visible linguistically. Non-fine-tuned configurations produce longer, multi-sentence, and explicitly normative interventions (mean length 72.8 tokens; 96% with a norma- tive appeal), whereas fine-tuned ones produce shorter responses (16.8 tokens; 20%). ProfilingUD [9] confirms this shift: imperative mood decreases from 32.7 to 6.1, subordinate clauses from 8.9 to 2.2, and sentences per message from 5.6 to 1.4, with all differences significant after FDR correction. Moreover, only the Reddit-fine-tuned configurations move toward the toxic register when compared with the 128 toxic messages (mean distance 25.4, against 29.0 for non-fine-tuned configurations and 28.8 for fine-tuned ones without Reddit data). Thus, stylistic mirroring is specific to Reddit fine-tuning. This also explains why fine-tuned configurations are perceived as less artificial but judged less adequate and persuasive: their outputs resemble ordinary conversational turns, which lowers artificiality but conflicts with the function of a moderation intervention. Best configurations and implications. A complementary question is whether the higher ratings obtained by [BaPrHi] reflect contextual fit or, more simply, the fact that without domain fine- tuning the model produces neutral-sounding responses that evaluators find credible. Our data support a qualified version of the latter interpretation. Non-fine-tuned configurations are markedly formulaic: their generations are more similar to one another than those of fine-tuned configurations (mean pairwise similarity 0.383 against 0.207), and 90% contain at least one canonical moderation phrase. They are not, however, generic in content, since their lexical overlap with the toxic message is higher than for fine-tuned configurations (0.092 against 0.064), while the [MuRe] family reaches the highest content specificity (0.147) without producing any moderative frame. Participants therefore appear to reward not vagueness, but the recognizable register of a moderation act, despite these configurations receiving the highest artificiality scores. This effect holds at the configuration level: register features correlate with mean adequacy across the seven evaluated configurations (í=0.93 for imperative mood,í=7), but not within configurations (all pooled within-configuration correlations|í|<0.20). Thus, the register hypothesis accounts for configuration-level rankings, but not for message-level rating variation. Overall, these results indicate that fine-tuning is not neutral with respect to the pragmatic function of the generated text, and that the choice of the fine-tuning corpus determines which aspect of that function is lost. Notably, the best-performing configuration is itself contextualized. What fails is fine-tuning as a vehicle for contextualization, not contextualization as such, suggesting that contextual information is more safely supplied at inference time than encoded in the model weights. Moreover, safety and quality degrade jointly: the toxic outputs of the [Re] configurations are the extreme of a distributional shift whose most common outcome is a fluent, non-toxic response that performs no moderation act at all, and a toxicity filter applied downstream would remove the 8â13% of harmful outputs while leaving the 80â90% that do not function as counterspeech. 7 Discussion and Conclusions Generation. Our extensive analysis of adaptation and personalization strategies, evaluated through both algorithmic metrics and large-scale human studies, highlights both the promise and the fragility of generating effective contextualized counterspeech. The human evaluation shows that contextualization is not uniformly beneficial: only a small subset of configurations, especially [BaPr] and [BaPrHi], matched or improved over the baseline on adequacy and perceived persuasiveness, whereas several other adaptation and personalization strategies signifi- cantly underperformed. This finding suggests that lightweight contextual prompting can improve , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech21 counterspeech quality when conversational context and user-history information are incorporated in a controlled way. At the same time, the weaker performance of several other configurations indicates that adding more contextual information or applying additional fine-tuning does not necessarily lead to better counterspeech. Current difficulties in consistently generating high-quality contextualized responses may stem from the load placed on LLMs when they must integrate multi- ple concurrent instructions or layers of information [4,32]. These results should be interpreted in light of the ecological-validity gap between crowdsourced perception ratings and actual behavioral change. Our evaluation captures third-party judgments of counterspeech quality, including per- ceived adequacy and perceived persuasiveness, but it does not measure whether the author of the toxic message would actually change their behavior after receiving the response. This distinction is crucial because LLM-generated contextualized counterspeech can backfire, as highlighted in other works [10]. Finally, our error analysis shows that some configurations can produce inadequate, incorrect, or even toxic responses. Thus, our results identify contextualized counterspeech as a promising but challenging direction: personalization can be effective in specific settings, but it requires careful design, controlled conditioning, and human-centered evaluation. Future advances in larger and more capable LLMs may reduce some of these limitations, but they will still need to be accompanied by safeguards and rigorous evaluation [22]. Evaluation. Our findings also carry important implications for how counterspeech systems should be evaluated. While many studies rely on algorithmic metrics to assess the quality of generated counterspeech, our results show that these indicators correlate poorly with human judgments, an observation that aligns with other recent findings [42,77]. This mismatch suggests that automatic indicators and human evaluators may attend to different, and sometimes orthogonal, qualities of a response. The low inter-rater reliability further reinforces this point: even when ratings are numerically close, annotators often differ in how they rank or interpret counterspeech quality. This heterogeneity suggests that counterspeech effectiveness may depend not only on the toxic message itself, but also on the expectations, background, and preferences of the people evaluating or receiving the intervention, further motivating personalized and audience-aware approaches. This discrepancy underscores the importance of developing more comprehensive and nuanced evaluation protocols that combine both algorithmic and human-centered assessments [7]. Relying exclusively on either type may lead to incomplete or misleading conclusions about a systemâs true effectiveness. Future work should prioritize this endeavor to ensure more accurate assessments of real-world impact. Meanwhile, our findings also show that contextualized, AI-generated counterspeech can be persuasive and impactful according to human evaluation when appropriately adapted and personalized. This suggests a promising direction for scalable, AI-driven interventions aimed at curbing online toxicity. At the same time, this study highlights the limitations of current evaluation methods and points to the need for human-AI collaboration in both the design and assessment of such tools. By combining the scale and adaptability of AI with the nuance of human judgment, future systems could be more effective in promoting healthy online discourse. Limitations and future work. While our study provides extensive empirical insights, it is con- strained by several limitations. Firstly, we experimented with two large language models and a limited set of adaptation and personalization strategies; alternative architectures or contextualiza- tion methods could lead to different results. In addition, the toxic-message benchmark is numerically limited and restricted to five U.S.-politics-oriented Reddit communities. Its toxicity profile is also skewed toward obscene and insulting content, as shown in Figure 7, which may limit generalization to other platforms, cultural contexts, languages, and toxicity types. Similarly, our algorithmic evaluation is restricted to a limited set of quantitative indicators, and to one generation for each toxic message. Several of these indicators rely on ROUGE-based lexical similarity to approximate , Vol. 1, No. 1, Article . Publication date: July 2026. 22Cima et al. constructs such as relevance, diversity, adaptation, and personalization, raising construct-validity concerns. Moreover, since automatic indicators were also used in the selection pipeline, some configurations or messages preferred by human evaluators may have been excluded. Future work should therefore consider random, stratified, or human-in-the-loop selection strategies and stronger validation of automatic metrics, and should consider multiple generations to capture stochastic variance. Human evaluations are also subject to crowdsourcing limitations, including sample rep- resentativeness, cultural bias, and response variability. In particular, our persuasiveness scores capture third-party perceptions of persuasive potential rather than actual attitude or behavior change, and may be influenced by the demographic and political composition of the crowdworker sample. Our non-parametric tests also do not jointly model participant- and stimulus-level vari- ability in ordinal ratings; future analyses could complement them with cumulative link mixed models. These limitations call for further research on contextualized counterspeech using broader models, datasets, and evaluation designs. Future work should also investigate fairness and bias in generated counterspeech, as well as longitudinal or field studies assessing its effects on toxic users, bystanders, and subsequent conversational behavior. 8 Ethics, Risks, and Unintended Harms We investigate AI-generated counterspeech as a tool to support healthier online conversations, but we acknowledge that the same technical components may introduce certain ethical risks and unintended harms. In particular, our models rely on user histories to generate personalized responses, which may raise concerns about behavioral profiling and privacy. However, our goal is limited to capturing topics of discussion and writing style from the target usersâ prior comments, without inferring sensitive or protected attributes. Moreover, our data collection relies exclusively on publicly available Reddit comments. The proposed approach also has dual-use potential. Although our intended use is moderation support and harm reduction, techniques for contextualized and personalized counterspeech could be repurposed to generate manipulative, deceptive, or politically targeted persuasion at scale [35,78]. A further risk concerns bias. Since LLMs are known to be bias-prone [32], generated counterspeech may reproduce stereotypes, treat communities unevenly, or respond differently depending on the political, cultural, or linguistic characteristics of users and conversations. For these reasons, any real-world deployment should include safeguards such as transparency about AI involvement, human review, limits on user profiling, audits for biased or manipulative outputs, and mechanisms for users and communities to contest or opt out of automated interventions. Furthermore, deployment should be assessed against the current data-access policies and terms of service of the target platform, as well as applicable privacy and data-protection regulations. In light of these risks, we remark that our experiments were pre-registreted and received ethical approval from CNRâs IRB (protocol #0306210). Participants were provided detailed information about the studyâs purposes, including being informed that participation entailed exposure to toxic comments. All participants gave their informed consent. 9 Acknowledgments This work is partially supported by the European Union â NextGenerationEU within the ERC project DEDUCE (Data-driven and User-centered Content Moderation) under grant #101113826; the PRIN 2022 project PIANO (Personalized Interventions Against Online Toxicity) under CUP B53D23013290006; and the the PNRR MUR project FAIR: Future AI Research (PE00000013). Partial support was also received by the MUR in the framework of the FoReLab project (Departments of Excellence) and by the project âAdvancing Italian Language Processing with Small-Scale Training and Preference Modelingâ (IsCb8_AILP), funded by CINECA under the ISCRA initiative, for the availability of HPC resources and support. , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech23 References [1]Ana Aleksandric, Sayak Saha Roy, Hanani Pankaj, Gabriela Mustata Wilson, and Shirin Nilizadeh. 2024. Usersâ behavioral and emotional response to toxicity in Twitter conversations. In AAAI ICWSM. [2] Anirban Saha Anik, Xiaoying Song, Elliott Wang, Bryan Wang, Bengisu Yarimbas, and Lingzi Hong. 2025. Multi-Agent Retrieval-Augmented Framework for Evidence-Based Counterspeech Against Health Misinformation. In COLM. [3]Jason Baumgartner, Savvas Zannettou, Brian Keegan, Megan Squire, and Jeremy Blackburn. 2020. The pushshift Reddit dataset. In AAAI ICWSM. [4]Tilman Beck, Hendrik Schuff, Anne Lauscher, and Iryna Gurevych. 2024. Sensitivity, performance, robustness: Deconstructing the effect of sociodemographic prompting. In EACL. [5]Helena Bonaldi, Yi-Ling Chung, Gavin Abercrombie, and Marco Guerini. 2024. NLP for counterspeech against hate: A survey and how-to guide. In NAACL. [6]Angana Borah, Rada Mihalcea, and VerĂłnica PĂŠrez-Rosas. 2026. Persuasion at play: Understanding misinformation dynamics in demographic-aware human-LLM interactions. In EACL. [7] Nimet Beyza Bozdag, Shuhaib Mehri, Gokhan Tur, and Dilek Hakkani-Tur. 2026. Persuade me if you can: A framework for evaluating persuasion effectiveness and susceptibility among large language models. In ACM CAIS. [8]Simon Martin Breum, Daniel VĂŚdele Egdal, Victor Gram Mortensen, Anders Giovanni Møller, and Luca Maria Aiello. 2024. The persuasive power of large language models. In AAAI ICWSM. [9]Dominique Brunato, Andrea Cimino, Felice DellâOrletta, Giulia Venturi, and Simonetta Montemagni. 2020. Profiling-UD: A tool for linguistic profiling of texts. In LREC. [10] Dominik Bär, Abdurahman Maarouf, and Stefan Feuerriegel. 2024. Generative AI may backfire for counterspeech. arXiv:2411.14986 (2024). [11]Aldo Cerulli, Lorenzo Cima, Benedetta Tessa, Serena Tardelli, and Stefano Cresci. 2026. The Big Ban Theory: A pre-and post-intervention dataset of online content moderation actions. In AAAI ICWSM. [12] Aldo Cerulli, Benedetta Tessa, Giuseppe La Selva, Oronzo Mazzeo, Lorenzo Cima, Lucia Monacis, and Stefano Cresci. 2026. Dark personality traits and online toxicity: Linking self-reports to reddit activity. Computers in Human Behavior (2026), 109085. [13]Eshwar Chandrasekharan, Shagun Jhaver, Amy Bruckman, and Eric Gilbert. 2022. Quarantined! Examining the effects of a community-wide moderation intervention on Reddit. ACM TOCHI 29, 4 (2022). [14]Hyundong Cho, Shuai Liu, Taiwei Shi, Darpan Jain, Basem Rizk, Yuyang Huang, Zixun Lu, Nuan Wen, Jonathan Gratch, Emilio Ferrara, and Jonathan May. 2024. Can language model moderators improve the health of online discourse? NAACL (2024). [15]Yi-Ling Chung, Gavin Abercrombie, Florence Enock, Jonathan Bright, and Verena Rieser. 2024. Understanding Counterspeech for Online Harm Mitigation. Northern European Journal of Language Technology 10, 1 (2024). [16]Yi-Ling Chung, Serra Sinem TekiroÄlu, and Marco Guerini. 2021. Towards knowledge-grounded counter narrative generation for hate speech. In ACL-IJCNLP. [17] Lorenzo Cima, Alessio Miaschi, Amaury Trujillo, Marco Avvenuti, Felice DellâOrletta, and Stefano Cresci. 2025. Contextualized counterspeech: Strategies for adaptation, personalization, and evaluation. In ACM W. [18] Lorenzo Cima, Benedetta Tessa, Amaury Trujillo, Stefano Cresci, and Marco Avvenuti. 2025. Investigating the heterogeneous effects of a massive content moderation intervention via Difference-in-Differences. Online Social Networks and Media 48 (2025), 100320. [19] Jordi Guillem Condom Tibau, Angelina Voggenreiter, JĂźrgen Pfeffer, et al.2025. Prevalence, Substance and Responses to Hate Speech Against LGBTQ Communities on TikTok. In AAAI ICWSM. [20]Thomas H Costello, Gordon Pennycook, and David G Rand. 2024. Durably reducing conspiracy beliefs through dialogues with AI. Science 385 (2024). [21] Stefano Cresci, Roberto Di Pietro, Marinella Petrocchi, Angelo Spognardi, and Maurizio Tesconi. 2014. A criticism to society (as seen by Twitter analytics). In IEEE ICDCS Workshops. [22]Stefano Cresci, Amaury Trujillo, and Tiziano Fagni. 2022. Personalized interventions for online moderation. In ACM Hypertext. [23] Francisco Cribari-Neto and Achim Zeileis. 2010. Beta regression in R. Journal of statistical software 34 (2010), 1â24. [24] Mekselina DoÄanç and Ilia Markov. 2023. From generic to personalized: Investigating strategies for generating targeted counter narratives against hate speech. In ACL CS4OA. [25]Cynthia Dwork, Chris Hays, Jon Kleinberg, and Manish Raghavan. 2024. Content moderation and the formation of online communities: A theoretical framework. In ACM W. [26]Margherita Fanton, Helena Bonaldi, Serra Sinem TekiroÄlu, and Marco Guerini. 2021. Human-in-the-Loop for data collection: A multi-target counter narrative dataset to fight online hate speech. In ACL-IJCNLP. [27]Kazuaki Furumai, Roberto Legaspi, Julio Vizcarra, Yudai Yamazaki, Yasutaka Nishimura, Sina J Semnani, Kazushi Ikeda, Weiyan Shi, and Monica S Lam. 2024. Zero-shot persuasive chatbots with LLM-generated strategies and information , Vol. 1, No. 1, Article . Publication date: July 2026. 24Cima et al. retrieval. EMNLP (2024). [28]John D Gallacher, Marc W Heerdink, and Miles Hewstone. 2021. Online engagement between opposing political protest groups via social media is linked to physical violence of offline encounters. Social Media + Society 7, 1 (2021). [29]Joshua Garland, Keyan Ghazi-Zahedi, Jean-Gabriel Young, Laurent HĂŠbert-Dufresne, and Mirta Galesic. 2022. Impact and dynamics of hate and counter speech online. EPJ Data Science 11, 1 (2022). [30]Gloria Gennaro, Laurenz Derksen, Aya Abdelrahman, Emma Broggini, Mariya Alexandra Green, Victoria Andrea Haerter, Elia Heer, Isabel Heidler, Fiona Kauer, Han-Nuri Kim, et al.2025. Counterspeech encouraging users to adopt the perspective of minority groups reduces hate speech and its amplification on social media. Scientific Reports 15, 1 (2025). [31] Tarleton Gillespie. 2020. Content moderation, AI, and the question of scale. Big Data & Society 7, 2 (2020). [32]Tommaso Giorgi, Lorenzo Cima, Tiziano Fagni, Marco Avvenuti, and Stefano Cresci. 2025. Human and LLM biases in hate speech annotations: A socio-demographic analysis of annotators and targets. AAAI ICWSM (2025). [33] Natasha Goel, Thomas Bergeron, Blake Lee-Whiting, Thomas Galipeau, Danielle Bohonos, Sarah Lachance, Sonja Savolainen, Clareta Treger, and Eric Merkley. 2024. Artificial influence? Comparing AI and human persuasion in reducing belief certainty. (2024). https://doi.org/10.31219/osf.io/2vh4k. [34] Pierpaolo Goffredo, Valerio Basile, Bianca Cepollaro, Viviana Patti, et al.2022. Counter-TWIT: An Italian corpus for online counterspeech in ecological contexts. In ACL WOAH. [35] Josh A Goldstein, Jason Chao, Shelby Grossman, Alex Stamos, and Michael Tomz. 2024. How persuasive is AI-generated propaganda? PNAS Nexus 3, 2 (2024). [36]Jarod Govers, Eduardo Velloso, Vassilis Kostakos, and Jorge Goncalves. 2024. AI-Driven Mediation Strategies for Audience Depolarisation in Online Debates. In ACM CHI. [37] Kobi Hackenburg and Helen Margetts. 2024. Evaluating the persuasive influence of political microtargeting with large language models. PNAS 121, 24 (2024). [38]Kobi Hackenburg, Ben M Tappin, Paul RĂśttger, Scott A Hale, Jonathan Bright, and Helen Margetts. 2025. Scaling language model size yields diminishing returns for single-message political persuasion. PNAS 122, 10 (2025). [39]Sadaf MD Halim, Saquib Irtiza, Yibo Hu, Latifur Khan, and Bhavani Thuraisingham. 2023. WokeGPT: Improving counterspeech generation against online hate speech by intelligently augmenting datasets using a novel metric. In IEEE IJCNN. [40] Sabit Hassan and Malihe Alikhani. 2023. DisCGen: A framework for discourse-informed counterspeech generation. In IJCNLP-AACL. [41]Bing He, Mustaque Ahamad, and Srijan Kumar. 2023. Reinforcement learning-based counter-misinformation response generation: A case study of COVID-19 vaccine misinformation. In ACM W. [42] Amey Hengle, Aswini Kumar Padhi, Anil Bandhakavi, and Tanmoy Chakraborty. 2025. CSEval: Towards Automated, Multi-Dimensional, and Reference-Free Counterspeech Evaluation using Auto-Calibrated LLMs. In NAACL. [43]Daniel Hickey, Daniel MT Fessler, Matheus Schmitz, Paul Smaldino, Kristina Lerman, Goran MuriÄ, and Keith Burghardt. 2026. Assessing How Hate, Counterspeech, and Toxicity Affect Hate Group Newcomers. In AAAI ICWSM. [44]Lingzi Hong, Pengcheng Luo, Eduardo Blanco, and Xiaoying Song. 2024. Outcome-constrained large language models for countering hate speech. EMNLP (2024). [45] Manoel Horta Ribeiro, Shagun Jhaver, Savvas Zannettou, Jeremy Blackburn, Gianluca Stringhini, Emiliano De Cristofaro, and Robert West. 2021. Do platform migrations compromise content moderation? Evidence from r/The_Donald and r/Incels. In ACM CSCW. [46]Evey Jiaxin Huang, Abhraneel Sarma, Sohyeon Hwang, Eshwar Chandrasekharan, and Stevie Chancellor. 2024. Opportunities, tensions, and challenges in computational approaches to addressing online harassment. In ACM DIS. [47]Hang Jiang, Xiajie Zhang, Xubo Cao, Cynthia Breazeal, Deb Roy, and Jad Kabbara. 2024. PersonaLLM: Investigating the ability of large language models to express personality traits. NAACL (2024). [48]Shuyu Jiang, Wenyi Tang, Xingshu Chen, Rui Tang, Haizhou Wang, and Wenxian Wang. 2025. ReZG: Retrieval- augmented zero-shot counter narrative generation for hate speech. Neurocomputing 620 (2025). [49] Cameron Jones and Benjamin Bergen. 2026. Lies, damned lies, and language statistics: a comprehensive review of risks from manipulation, persuasion, and deception with large language models. Artificial Intelligence Review 59, 4 (2026), 116. [50] Shirish Karande, V Santhosh, and Yash Bhatia. 2024. Persuasion games with large language models. In ACL ICON. [51] Aswini Kumar, Anil Bandhakavi, and Tanmoy Chakraborty. 2025. Counterspeech the ultimate shield! Multi-Conditioned Counterspeech Generation through Attributed Prefix Learning. In ACL. [52] Nihal Kumarswamy, Mohit Singhal, and Shirin Nilizadeh. 2025. Causal Insights into Parlerâs Content Moderation Shift: Effects on Toxicity and Factuality. In ACM W. [53]Rohan Leekha, Olga Simek, and Charlie Dagli. 2024. War of Words: Harnessing the Potential of Large Language Models and Retrieval Augmented Generation to Classify, Counter and Diffuse Hate Speech. In AAAI FLAIRS. , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech25 [54]Alyssa Lees, Vinh Q Tran, Yi Tay, Jeffrey Sorensen, Jai Gupta, Donald Metzler, and Lucy Vasserman. 2022. A new generation of perspective API: Efficient multilingual character-level transformers. In ACM KDD. [55] Junyi Li, Tianyi Tang, Wayne Xin Zhao, Jian-Yun Nie, and Ji-Rong Wen. 2024. Pre-trained language models for text generation: A survey. ACM Computing Surveys 56, 9 (2024). [56]Mikel K Ngueajio, Flor Miriam Plaza-del Arco, Yi-Ling Chung, Danda B Rawat, and Amanda Cercas Curry. 2025. Think Like a Person Before Responding: A Multi-Faceted Evaluation of Persona-Guided LLMs for Countering Hate. In ACL WOAH. [57]Amalie Brogaard Pauli, Isabelle Augenstein, and Ira Assent. 2024. Measuring and Benchmarking Large Language Modelsâ Capabilities to Generate Persuasive Language. In NAACL. [58]Yujin Potter, Shiyang Lai, Junsol Kim, James Evans, and Dawn Song. 2024. Hidden Persuaders: LLMsâ Political Leaning and Their Influence on Voters. In ACL EMNLP. [59] Jing Qian, Anna Bethke, Yinyin Liu, Elizabeth Belding, and William Yang Wang. 2019. A benchmark dataset for learning to intervene in online hate speech. In EMNLP-IJCNLP. [60]Emanuele Ricco, Elia Onofri, Lorenzo Cima, Stefano Cresci, and Roberto Di Pietro. 2026. A Geometric Analysis of Small-sized Language Model Hallucinations. In Forty-third International Conference on Machine Learning. [61] Punyajoy Saha, Kanishk Singh, Adarsh Kumar, Binny Mathew, and Animesh Mukherjee. 2022. CounterGeDi: A controllable approach to generate polite, detoxified and emotional counterspeech. In IJCAI. [62] Francesco Salvi, Manoel Horta Ribeiro, Riccardo Gallotti, and Robert West. 2025. On the conversational persuasiveness of GPT-4. Nature Human Behaviour 9, 8 (2025), 1645â1653. [63]Gautam Kishore Shahi, Benedetta Tessa, Amaury Trujillo, and Stefano Cresci. 2025. A Year of the DSA Transparency Database: What it (Does Not) Reveal About Platform Moderation During the 2024 European Parliament Election. In ICWSM Workshops. [64]Xiaoying Song, Sujana Mamidisetty, Eduardo Blanco, and Lingzi Hong. 2025. Assessing the human likeness of AI-generated counterspeech. In ACL COLING. [65]Miriah Steiger, Timir J Bharucha, Sukrit Venkatagiri, Martin J Riedl, and Matthew Lease. 2021. The psychological well-being of content moderators: The emotional labor of commercial moderation and avenues for improving support. In ACM CHI. [66] Madiha Tabassum, Alana Mackey, Ashley Schuett, and Ada Lerner. 2024. Investigating moderation challenges to combating hate and harassment: The case of Mod-Admin power dynamics and feature misuse on Reddit. In USENIX. [67]Serra Sinem Tekiroglu, Helena Bonaldi, Margherita Fanton, and Marco Guerini. 2022. Using pre-trained language models for producing counter narratives against hate speech: A comparative study. In ACL. [68]Serra Sinem TekiroÄlu, Yi-Ling Chung, and Marco Guerini. 2020. Generating counter narratives against online hate speech: Data and strategies. In ACL. [69]Benedetta Tessa, Lorenzo Cima, Amaury Trujillo, Marco Avvenuti, and Stefano Cresci. 2025. Beyond trial-and-error: Predicting user abandonment after a moderation intervention. Engineering Applications of Artificial Intelligence 162 (2025), 112375. [70]Amaury Trujillo and Stefano Cresci. 2022. Make Reddit Great Again: Assessing community effects of moderation interventions on r/The_Donald. In ACM CSCW. [71] Amaury Trujillo and Stefano Cresci. 2023. One of many: Assessing user-level effects of moderation interventions on r/The_Donald. In ACM WebSci. [72] Amaury Trujillo, Tiziano Fagni, and Stefano Cresci. 2025. The DSA Transparency Database: Auditing self-reported moderation actions by social media. In ACM CSCW. [73]Haiyang Wang, Yuchen Pan, Xin Song, Xuechen Zhao, Minghao Hu, and Bin Zhou. 2024. F2rl: Factuality and faithfulness reinforcement learning framework for claim-guided evidence-supported counterspeech generation. In ACL EMNLP. [74] Xinchen Yu, Eduardo Blanco, and Lingzi Hong. 2024. Hate cannot drive out hate: Forecasting conversation incivility following replies to hate speech. In AAAI ICWSM. [75]Yi Zheng, BjĂśrn Ross, and Walid Magdy. 2026. Validating Automatic Evaluation of Controllable Counterspeech Generation: Rankings Matter More Than Scores. In EACL. [76]Nawaal Zubair, Shazia Hashmat, and Ume Aimen. 2025. Artificial Intelligence and the Generational Divide: A Study on Trust and Acceptance. Annual Methodological Archive Research Review 3, 6 (2025), 19â44. [77] Irune Zubiaga, Aitor Soroa, and Rodrigo Agerri. 2024. A LLM-based ranking method for the evaluation of automatic counter-narrative generation. In ACL EMNLP. [78]Aneta Zugecova, Dominik Macko, Ivan Srba, Robert Moro, Jakub Kopal, Katarina Marcincinova, and Matus Mesarcik. 2025. Evaluation of LLM vulnerabilities to being misused for personalized disinformation generation. In ACL. , Vol. 1, No. 1, Article . Publication date: July 2026. 26Cima et al. Appendix A Dataset For counterspeech generation, we retrieved 128 toxic comments from 49 Reddit threads. Table 3 reports the distribution of these toxic comments across the corresponding subreddits. subredditthreads toxic comments r/AOC78 r/conservatives915 r/politics885 r/progressive34 r/the_donald816 TOTAL35128 Table 3. Distribution of the toxic comments and relative threads through the five political subreddits analyzed. For fine-tuning, we used three datasets. For counterspeech fine-tuning, we relied on Multi- CONAN [26], which comprises 500 curated hate-speechâcounterspeech pairs, and the Reddit Hate-Speech Intervention (RHSI) dataset [59], which consists of 5,020 Reddit conversations con- taining human-authored interventions. For Reddit conversation tuning, we used approximately 5K commentâreply pairs sampled from five high-activity political subreddits. Since fine-tuning on conversational or intervention data may expose the model to toxic or confrontational language, it could also introduce undesirable stylistic patterns into the generated counterspeech. To assess this risk, Figure 10 reports the Perspective API toxicity-score distributions for the three fine-tuning datasets. redditrhsiconan 0.0 0.2 0.4 0.6 0.8 1.0 toxicity Fig. 10. Toxicity distributions for the fine-tuning datasets. B Prompts This section reports the LLM prompts that we used to generate counterspeech responses and user summaries for the different factors described in Section 4.1: â˘Default counterspeech generation:"You are a moderator of a subreddit and you come across a comment that exhibits hate speech. Generate a brief response that aims to counteract this comment in a respectful manner." â˘Counterspeech generation (Conversation [Pr] strategy):"You are a moderator of a subreddit and you come across a comment that exhibits hate speech. Generate a brief response that aims to counteract the comment at the end of this conversation in a respectful manner." , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech27 â˘Counterspeech generation (Comment history [Hi] strategy):"You are a moderator of a subreddit and you come across a comment that exhibits hate speech. Given the following 10 comments from the same reddit user: <comments>, generate a brief response that aims to counteract this comment in a respectful manner, using these comments to understand the userâs style and personalize your response." â˘Counterspeech generation (Summary [Su] strategy):"You are a moderator of a subreddit and you come across a comment that exhibits hate speech. Given the following summary describing the reddit user that made the comment: <summary>, generate a brief response that aims to counteract this comment in a respectful manner, using the userâs summary to understand his/her style and personalize your response." ⢠User summary generation (Summary [Su] strategy):"Given the following comments written by the same Reddit user:<comments>, generate a concise and schematic summary describing the user, following this schema: 1) Writing style and lexicon: Identify and describe the predominant writing style of the user; 2) Interests: Describe the interests and topics generally covered by the user. Do not add any other information or infer details about the userâs age, gender, or any other personal information." C Algorithmic evaluation We report on Table 4 a fine-grained analysis of the algorithmic evaluation results for each individual configuration. Configurations are split in four groups depending on their use of adaptation factors, personalization factors, neither, or both. D Crowdsourcing questionnaire D.1 Task description Each participant in our crowdsourcing experiment was allowed to complete the questionnaire only once and received $0.70 as compensation, which, given the average completion time, is above the US minimum wage. Upon providing their informed consent to take part in the experiment, participants received the following description of the task:"Your task is to evaluate a set of counterspeech responses to toxic messages posted on social media based on several criteria. With counterspeech we mean a response that addresses or challenges harmful, offensive, or toxic content with the aim to encourage a more respectful and constructive communication. Consider the toxic post and corresponding response below, then rate the following statements from strongly disagree (1) to strongly agree (5). D.2 Counterspeech questions The following questions were asked for each pair of toxic message and corresponding counterspeech response: ⢠Relevance: The response is relevant to the toxic post. ⢠Adequacy: The response is suitable as counterspeech. ⢠Truthfulness: The response is truthful (i.e., honest, sincere). ⢠Persuasiveness (toxic user): The response would persuade the author of the toxic post to re-engage in the conversation in a civil manner. ⢠Persuasiveness (conversation): The response would steer the overall conversation back to civil discourse. ⢠Artificiality: The response was generated by AI. Participants assigned to the contextual between-subjects condition (see Section 4.2.3) also received the following question: , Vol. 1, No. 1, Article . Publication date: July 2026. 28Cima et al. evaluation indicators personalization configurationrelâ divâ readâ tox â adaâ lexâwriâ Ba.117.562.576.050â.138.519 Mu.129.754.805.119.855.153.488 Hs.083.545.753.142.858.113.463 MuHs.086.615.688.285.857.114.461 BaPr.120.566.574 .040.490.144.524 MuPr.122.808.808.086.862.141.490 HsPr.086.706.805.207.867.133.473 MuRe.181.794.906.176.882.146.503 MuHsPr.101.802.765.208.866.128.471 MuHsRe.130.851.799.315 .908.113.469 MuRePr.207.849.869.177.886.137.503 MuHsRePr.128 .876.753.233.901.122.464 BaHi.139.604.544.063.583.143.534 MuHi.121.801.889.088.874.133.495 HsHi.085.768.872.172.884.125.478 BaSu.142.657.676.113.671.137.540 MuSu.128.732.837.112.860.137.499 HsSu.109.716.874.151.869.150.479 MuHsHi.102.815.892.127.884.134.498 MuHsSu.116.772.874.112.878.144.488 BaPrHi.141.618.538.063.607.153.535 BaPrSu.139.656.655.092.683.144.538 MuPrHi.135.836.851.090.878.131.506 MuPrSu.134.781.833.101.856 .155.504 MuReHi.172.826.919.135.893.125.506 MuReSu.180.779.900.158.877.148.510 HsPrHi.105.797.900.207.882.127.498 HsPrSu.113.776.875.162.874.131.489 MuHsPrHi.096.809 .924.137.887.127.498 MuHsPrSu.102.798.842.159.873.136.489 MuHsReHi.116.821.906.098.892.128.500 MuHsReSu.125.784.858.146.872.140.491 MuRePrHi.173.850.893.144.901.130.517 MuRePrSu.165.823.861.174.885.142.509 MuHsRePrHi.125.831.896.147.890.133.503 MuHsRePrSu.132.820.891.162.880.137.487 Ba: LLaMa2 baseline;Mu: Multi-CONAN fine-tuning;Hs: RHSI fine-tuning; Re: political subreddits fine-tuning;Pr: previous comments;Hi: user comment history;Su : user summary. Table 4. Algorithmic evaluation results of each configuration. For each indicator, the best value is in bold font and the remaining top-5 are underlined. Configurations are split in four groups depending on their use of adaptation factors, personalization factors, neither, or both. Icons highlight the overall bestand worst configurations of each group. ⢠Contextualization: The counterspeech response is personalized (as opposed to being generic) with respect to the postâs context. Table 5 reports agreement and similarity statistics for the human-evaluation dimensions in the non-contextual condition, the contextual condition, and both conditions pooled together. We report Krippendorffâsíźand average pairwise Spearman correlation to assess consistency in annotatorsâ judgments, together with pairwise agreement and normalized match distance to capture how close the assigned Likert scores are. , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech29 questionKrippendorffâs íź Spearman í pairwise accuracy norm. match distance relevance0.0040.0020.3040.263 adequacy0.0050.0050.3010.267 truthfulness0.0040.0040.3050.259 persuasiveness (toxic user)0.0020.0020.2860.278 persuasiveness (conversation)0.0030.0020.2900.275 artificiality0.0020.0020.2850.281 contextualization0.0080.0020.2850.282 AVERAGE0.0040.0030.2940.272 Table 5. Inter-rater reliability and response-similarity statistics for the seven human-evaluation dimensions, reported separately for the non-contextual condition, the contextual condition, and both conditions pooled together. For each condition, we report Krippendorffâsíź, pairwise Spearman correlation, pairwise agreement, and normalized match distance (NMD). The last row reports the average across the seven evaluated dimen- sions. 20-2525-3030-3535-4040-4545-5050-5555-6060-6565-7070-75 0 200 400 600 800 1000 AGE non-contextual contextual High school or less Some college College or more 0 250 500 750 1000 1250 1500 1750 2000 EDUCATION Asian/As. American Black/Afr. American Hispanic/Latino White/Caucasian Other 0 500 1000 1500 2000 ETHNICITY Female Male Non-binary Not disclose 0 200 400 600 800 1000 1200 1400 1600 GENDER Democratic Lean Democratic Lean Republican Republican 0 200 400 600 800 1000 1200 POLITICS Never Rarely Sometimes Often Very Often 0 100 200 300 400 500 600 700 800 SOCIAL FREQUENCY 01 2-34-5 >5 0 200 400 600 800 1000 1200 1400 SOCIAL QUANTITY Fig. 11. Socio-demographic characteristics of the participants in our crowdsourcing experiment, separately for the non-contextual and contextual evaluation tasks. D.3 Socio-demographic questions The following questions were asked once for each participant, at the end of the questionnaire: ⢠Age: [free text, numeric] ⢠Gender: [Female, Male, Non-binary or gender diverse, I prefer not to disclose] ⢠Education: [High school or less, Some college, College graduate or more] ⢠Which of the following describes your race/ethnicity? [Asian/Asian American, Black/African American, Hispanic/Latino, White/Caucasian, Other] ⢠Which of the following describes best your political affiliation? [Democratic, Lean Democratic, Lean Republican, Republican] ⢠How frequently do you use social media (e.g., Facebook, Twitter/X, Instagram, Reddit, etc.)? [Never. Rarely (less than once a week). Sometimes (once a week to several times a week). Often (daily). Very often (multiple times a day)] ⢠How many different social media do you actively use (at least once a week)? [None, 1, 2-3, 4-5, 5+] Figure 11 shows the distribution of socio-demographic characteristics of the participants. , Vol. 1, No. 1, Article . Publication date: July 2026. 30Cima et al. â â m=0.874 m=0.888 â â m=0.789 m=0.781 â â m=0.125 m=0.107 â â m=0.854 m=0.782 â â m=0.136 m=0.125 â â m=0.498 m=0.508 â â m=0.145 m=0.129 adaptation â diversity â relevance â readability â personalization lex â personalization wri â toxicity â llamaqwenllamaqwenllamaqwenllamaqwenllamaqwenllamaqwenllamaqwen 0.00 0.25 0.50 0.75 1.00 Fig. 12. Algorithmic indicator values averaged by configuration and large language model (LLM), either LLaMA2-13B or Qwen3-8B. Each dot represents one of the 36 configurations. Lines connecting configuration dots show average change across LLMs. The median (m) value per indicator and LLM is marked by a red line. Arrows (â /â) indicate whether higher or lower values respectively are better for each indicator. sadness disgust trust anger anticipation surprise fear joy 3.4 3.5 3.6 3.7 3.8 3.9 4.0 4.1 persuasiv. (toxic user) anger sadness trust disgust joy fear surprise anticipation 3.4 3.5 3.6 3.7 3.8 3.9 4.0 4.1 persuasiv. (conversation) Fig. 13. Persuasiveness of the counterspeech with respect to the dominant emotion of the toxic message replied to. D.4 Algorithmic Indicators Based on an Alternative Large Language Model To assess whether our results depend on the specific generation model, we conducted an additional robustness analysis comparing LLaMA2-13B with Qwen3-8B. LLaMA2-13B was used as the main model in the original study, because it offers a practical balance between generation quality, openness, reproducibility, and feasibility for systematic fine-tuning and inference across a large configuration space. In our extension, we added Qwen3-8B as a more recent open-weight model with stronger instruction-following capabilities. Figure 12 shows the average indicator values with LLaMA2-13B and Qwen3-8B. Differences between LLMs are statistically significant for all 7 indicators (based on a Wilcoxon signed-rank test for paired configurations per indicator). Qwen3-8B tends produce values with a lower spread, but is in average better than LlaMA2-13B only in some of the indicators, not all. For instance, for the indicator readability LLaMA2-13B is noticeably better. D.5 Validation of the Unframed indicator When investigating the failure modes of the generated counterspeech message in Section 6.3, we labeled them as unframed if they contain no marker of a moderation act. That is, no second- person address combined with a directive form, no explicit normative appeal, and no imperative opening. Since the unframed indicator captures a property of register rather than of pragmatic function, we validated it before use. Specifically, we verified that it is not a by-product of message length, since within fine-tuned configurations the correlation between length and the indicator is negligible (í=â0.08) and fine-tuned outputs of at least thirty tokens still trigger it in 82% of cases. Furthermore, we also found that it is consistent with human judgment, since flagged messages receive lower adequacy ratings from crowdworkers (Cliffâs íż=â0.44, í< 0.001). , Vol. 1, No. 1, Article . Publication date: July 2026. Contextualized Counterspeech Can Be More Persuasive Than Generic Counterspeech31 configuration toxic messagegenerated counterspeechissue BaPrHiProbably true considering the Christian god is a genocidal lunatic that condones slavery and expects misogyny. Hi there! I understand that you may have strong feelings about religion, but please remember that all beliefs and opinions de- serve respect and kindness. Name-calling and spreading hate only serves to divide us and perpetuate harmful stereotypes. Instead, I invite you to share your thoughts and engage in con- structive dialogue. Letâs work together to create a welcoming and inclusive community for everyone. BaPrAhh ya sorry! Forgot about that idiot haha.Hi there! I understand that you may have forgotten about Saskatchewan, but please refrain from using derogatory lan- guage to refer to its residents. Everyone deserves respect and kindness, regardless of where they are from. Letâs focus on having a productive and inclusive conversation. Is there any- thing specific you would like to discuss? HsHiYou donât have to shit in everyone elseâs cereal because they arenât digesting the news exactly the way you want. I think itâs important to remember that people can have differ- ent opinions and still get along. unframed MuReresult of this administration. Now is not the time to engage in forgive and forget healing bullshit. Iâm sorry but I donât see how your statement is hateful.unframed MuRePrHiFor a while, red tide was the slogan to counter Blue wave. Completely ignoring the fact that the comparison is to an abundance of dead fish. Canât make this shit up. I think itâs more about the color than anything else.unframed MuRePrHiObama and Osama are different by two letters, B and S. BS = bullshit! Theyâve been mocking us all along, in plain sight. I think youâre confusing âBSâ with âBSâ.degenerate MuReElon reinstated Trump but he hasnât used it. Or as I call it, throwing more dog shit into the camp fire. I donât think Elon has ever said he would reinstate Trump.factually incorrect MuRePrHiWonât someone please think of the children that we acciden- tally forgot to murder? I thought about them, but then I realized that if we didnât kill them, they would grow up to be adults who will vote for democrats. toxic Table 6. Examples of toxic messages and the corresponding counterspeech generated by some of the evaluated configurations. Examples are selected via purposeful sampling, to illustrate the failure modes quantified in Table 2. The issue column reports the automatic indicators triggered by each message. The example of a factually incorrect statement illustrates a failure type that our indicators do not capture. Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009 , Vol. 1, No. 1, Article . Publication date: July 2026.