Paper deep dive
How Annotation Trains Annotators: Competence Development in Social Influence Recognition
Maciej Markiewicz, Beata Bajcar, Wiktoria Mieleszczenko-Kowszewicz, Aleksander SzczÄsny, Tomasz Adamczyk, Grzegorz Chodak, Karolina Ostrowska, Aleksandra Sawczuk, Jolanta Babiak, Jagoda Szklarczyk, PrzemysĆaw Kazienko
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/10/2026, 2:03:09 AM
Summary
This study investigates how the process of human data annotation influences the competence of annotators, specifically in the context of social influence recognition. Using a dataset of 1,021 dialogues annotated by 25 participants (experts and non-experts) over three rounds, the researchers found that annotators experienced significant increases in self-perceived competence, confidence, and annotation efficiency. Qualitative and quantitative analyses revealed that annotators' work became more detailed and nuanced over time, with expert groups showing more pronounced improvements. These shifts in annotator competence were found to impact the performance of Large Language Models (LLMs) trained on the resulting data.
Entities (6)
Relation Signals (3)
Maciej Markiewicz â authored â How Annotation Trains Annotators: Competence Development in Social Influence Recognition
confidence 100% · Paper title and author list
Annotation Process â enhances â Annotator Competence
confidence 90% · observed changes in data quality suggest that the annotation process may enhance annotator competence
Annotator Competence â impacts â LLM performance
confidence 90% · The observed shifts in annotator competence have a visible impact on the performance of LLMs
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Human data annotation, especially when involving experts, is often treated as an objective reference. However, many annotation tasks are inherently subjective, and annotators' judgments may evolve over time. This study investigates changes in the quality of annotators' work from a competence perspective during a process of social influence recognition. The study involved 25 annotators from five different groups, including both experts and non-experts, who annotated a dataset of 1,021 dialogues with 20 social influence techniques, along with intentions, reactions, and consequences. An initial subset of 150 texts was annotated twice - before and after the main annotation process - to enable comparison. To measure competence shifts, we combined qualitative and quantitative analyses of the annotated data, semi-structured interviews with annotators, self-assessment surveys, and Large Language Model training and evaluation on the comparison dataset. The results indicate a significant increase in annotators' self-perceived competence and confidence. Moreover, observed changes in data quality suggest that the annotation process may enhance annotator competence and that this effect is more pronounced in expert groups. The observed shifts in annotator competence have a visible impact on the performance of LLMs trained on their annotated data.
Tags
Links
- Source: https://arxiv.org/abs/2604.02951v1
- Canonical: https://arxiv.org/abs/2604.02951v1
Trouble viewing inline? Open PDF directly â
Full Text
39,550 characters extracted from source content.
Expand or collapse full text
How Annotation Trains Annotators: Competence Development in Social Influence Recognition Maciej Markiewicz 1[0009â0004â2882â6741] , Beata Bajcar 1[0000â0001â5044â4070] , Wiktoria Mieleszczenko-Kowszewicz 1[0000â0002â3948â268X] , Aleksander SzczÄsny 1[0009â0003â6808â2321] , Tomasz Adamczyk 1[0009â0005â9703â4630] , Grzegorz Chodak 1[0000â0002â9604â482X] , Karolina Ostrowska 2[0009â0004â9959â2487] , Aleksandra Sawczuk 2[0009â0002â0677â7905] , Jolanta Babiak 1[0000â0002â6604â7763] , Jagoda Szklarczyk 3[0009â0004â6784â8273] , and PrzemysĆaw Kazienko 1[0000â0001â5868â356X] 1 WrocĆaw University of Science and Technology, WrocĆaw, Poland name.surname@pwr.edu.pl 2 University of Silesia in Katowice, Katowice, Poland 3 SWPS University, WrocĆaw, Poland Abstract. Human data annotation, especially when involving experts, is often treated as an objective reference. However, many annotation tasks are inherently subjective, and annotatorsâ judgments may evolve over time. This study investigates changes in the quality of annotatorsâ work from a competence perspective during a process of social influence recognition. The study involved 25 annotators from five different groups, including both experts and non-experts, who annotated a dataset of 1,021 dialogues with 20 social influence techniques, along with intentions, re- actions, and consequences. An initial subset of 150 texts was annotated twice â before and after the main annotation process â to enable compar- ison. To measure competence shifts, we combined qualitative and quan- titative analyses of the annotated data, semi-structured interviews with annotators, self-assessment surveys, and Large Language Model training and evaluation on the comparison dataset. The results indicate a signif- icant increase in annotatorsâ self-perceived competence and confidence. Moreover, observed changes in data quality suggest that the annotation process may enhance annotator competence and that this effect is more pronounced in expert groups. The observed shifts in annotator compe- tence have a visible impact on the performance of LLMs trained on their annotated data. Keywords: annotation· learning process· LLMs· social influence· competence development Preprint: Accepted to AIED 2026: The 27th International Conference on Arti- ficial Intelligence in Education arXiv:2604.02951v1 [cs.CL] 3 Apr 2026 2M. Markiewicz et al. 1 Introduction Human annotation is foundational for creating high quality datasets and has been extensively studied. Expert annotation is often treated as objective ground truth, and aggregating multiple annotators typically yields satisfactory results despite individual errors. However, in subjective or partially-subjective tasks, such as the recognition of social influence, or when annotators are not experts, even human annotation may not be accurate or may change over the course of the process [7]. The aim of this study is to identify the changes in annotatorsâ work and analyze their origins; whether it is the effect of increased competence, random noise, or a varying worldview. We then want to assess the impact of these changes on data quality and AI model training. The study is conducted in conjunction with our unpublished social influence recognition dataset but focuses solely on the annotation process. We formulate the following research questions: RQ1 How does the competence of annotators regarding social influence recog- nition change over the course of the annotation process? RQ2 How does the quality of the annotatorsâ work change throughout the an- notation process? RQ3 Does using data from the beginning or the end of the annotation process for AI model training lead to changes in model performance? 2 Related Work Recent work on subjective tasks, such as the recognition of social influence, challenges the ground truth assumption, proposing perspectivist approaches that preserve individual annotator viewpoints rather than aggregating them [7,14]. To evaluate the reliability of these viewpoints, [2] emphasize using intra-annotator agreement to distinguish valid subjective interpretations from noise, rather than relying solely on inter-annotator agreement, which has been found to be a weak quality measure in subjective tasks [3]. However, annotator performance is rarely stable. [1] observe that intra-annotator agreement diminishes over time, while [3] show that annotatorsâ behavior has evolved significantly under supervision. The observed evolution in annotator behavior can be framed through learning science. [15] argue that example-based learning is a reliable method of knowledge acquisition, similar to learning from oneâs own experience [11]. Similarly, crowd- sourced annotation tasks can be treated as learning environments where workers actively gain knowledge [6,9]. This aligns with the definition of competence as an integration of knowledge and skills that develop over time [17]. Empirical studies demonstrate that non-experts, such as language learners, can contribute high-quality data as their skills improve [18], while [12] show that annotation curricula allow annotators to implicitly learn task schemes. Our study extends this by investigating how this natural competence acquisition affects the utility of the resulting data for model training. Competence Development in Social Influence Recognition3 3 The annotation process The studied annotation process aimed to create a dataset of 1,021 AI-generated dialogues depicting social influence targeted at adolescents, Figure 1. Each text has been verified and adjusted to sound natural, both by experts and by a single super-annotator. Each dialogue was annotated by 5 annotators, one from each group. The groups included experts: psychologists and communication experts, and non-experts: adolescents, parents, and teachers. The annotatorsâ task was to label texts with the following: (a) the degree to which social influence occurs in a text; (b) one or more of 20 social influence techniques; (c) the possible intentions of the person exerting influence and the clarity of those intentions; (d) possible consequences of giving in to the influence and their severity; (e) the reaction of the influenced person, the level of resistance against the influence, and whether they eventually submitted; (f) the annotatorâs certainty level; and (g) optional comments. Submission was a binary decision (submitted/not submitted). Reactions, in- tentions, and consequences (RICs) were entered as free text, with each item recorded separately. The remaining questions had a 5-level scale, and all ques- tions had a "hard to determine" option. Every annotator annotated 204 or 205 texts in three rounds: first, they re- ceived a set of 30 texts (the Pre set) for one week. Then, they annotated the main part of 174 or 175 texts for 2 weeks (the Main set). Finally, they received the same 30 texts as at the beginning and labeled them from scratch (Post set), again for one week. The Pre/Post set of 30 texts is the main source for comparison. Before starting, the annotators participated in a technical training session on the online annotation tool (Argilla 4 ), as well as a training session on social influence and the definitions used. After the first round of 30 texts, a calibration session was organized to answer questions and clarify technical mis- understandings. The session did not use any text as an example to ensure that it did not interfere with specific judgments. Our findings address both the Main and Pre/Post phases. Detailed annotator instructions, guidelines, answers from the calibration session, code, and definitions are available online 5 . 3.1 Annotator demographics The annotation process was conducted by a group of 25 annotators, including 14 women and 11 men, with a mean age of 34 years. The annotator group was intentionally structured to represent diverse social and professional perspectives. It consisted of five educators (pedagogues), five psychologists, five adolescents, five parents of adolescents, and five communication experts. Most had completed higher education (n = 19), five were students, and one had primary education. This diversity was intended to reduce perspective-related bias. 4 https://argilla.io 5 https://github.com/MaciejMarkiewicz/annotator-competence-growth 4M. Markiewicz et al. Pre annotation (30 texts) Main annotation (175 texts) Post annotation (same 30 texts) Initial training Calibration meeting SIR-SC questionnaire (I) NASA-TLX questionnaire (I) SIR-SC questionnaire (I) NASA-TLX questionnaire (I) 1 week1 week2 weeks1 week Observed competence development Analyses: quantitative Pre/Post comparison qualitative Pre/Post comparison LLM training on Pre/Post data extended interviews SIR-SC comparison NASA-TLX comparison annotation speed inter- and intra- annotator agreement group differences Fig. 1. Overview of the studied annotation process with competence shift analyses. 4 Quantitative analysis methodology We assess annotator competence through data quality, defining increased compe- tence as higher-quality work or equivalent quality in less time. We analyze quan- titative differences in the identification of social influence between annotation rounds. Changes include technique labels, the number of "hard to determine" answers, reported answer certainty, and the number of free text RIC responses. Since there are no gold standard labels and deriving them from the dataset itself would be biased (this is captured by agreement), we cannot treat an increase or decrease as desirable without qualitative confirmation. Specifically, we employed the following criteria: 4.1 Inter- and intra-annotator agreement Agreement for technique identification was assessed using Krippendorffâs alpha with Jaccard distance between sets, where each annotation was treated as a set of identified techniques. Inter-annotator agreement was calculated across all annotators and texts, while intra-annotator agreement measured the consistency between each annotatorâs Pre and Post labels for the same texts. Statistical significance was evaluated via bootstrap resampling of texts (10,000 iterations), preserving the dependency structure within the texts. 4.2 Per-text annotation time change The per-text time change was computed as the difference between consecutive submission timestamps for each annotator, with annotations ordered chrono- logically. To exclude outliers (very short times indicating skipped texts or very long times indicating breaks), a 1â20 minute window was applied. This analysis was conducted on both the Pre/Post and Main datasets. Statistical compar- isons used Mann-Whitney U tests. Due to the nature of the observations, this method may not be fully accurate for precise time measurements but is relevant for capturing the trend. Competence Development in Social Influence Recognition5 4.3 LLM training A direct measurement of data quality is the performance of AI models trained with it. To assess this, we evaluated the performance of LLMs in detecting social influence techniques by comparing the use of Pre and Post datasets for training and using the Main dataset for testing. Social influence technique detection was formulated as a multi-label classifi- cation problem. We tested two aggregation methods to establish ground truth labels: majority voting (AC MV ) and 2 annotator consensus (AC 2 ). Ground truth labels consisted of techniques identified by 3 or 2 out of 5 annotators, respectively. Some texts were labeled as lacking social influence, indicating that annotators did not detect any social influence in the text. For the sake of this experiment, instances where consensus could not be established were removed from the data. This approach reduced the Main set from 871 to 869 (AC 2 ) or 800 (AC MV ) ex- amples, and the Pre/Post sets from 150 to 149 (AC 2 ) or 139 (AC MV ) examples. We evaluated both in-context learning (ICL) and supervised fine-tuning (SFT). ICL was conducted on DeepSeek-V3.2 (with the temperature set to 0) for n â 0, 3, 10, 30 shots, averaged over 10 runs with distinct seeds. SFT was per- formed on Llama-3.1-8B-Instruct (batch size=1, LR=1e â5 , 2 epochs). Model performance was measured using Jaccard similarity. 4.4 Annotatorsâ self-perception of competence Two questionnaires captured changes in self-perceived competence. Along with the last one, after completing Post annotations, the annotators were asked an ad- ditional open-ended question: "What did you learn by taking part in the study?". SIR-SC â Social Influence Recognition Self-Competence questionnaire was developed, consisting of two sub-scales: Perceived Competence (8 items on self-evaluated recognition skills) and Self-Confidence (4 items on judgment cer- tainty). Responses used a five-point Likert scale. Responses were rated on a five-point Likert scale (from 1 â strongly disagree to 5 â strongly agree). In both sub-scales, total scores were averaged. Higher scores indicate a higher level of perceived competence or self-confidence, respectively. NASA-TLX [8] was used to assess participantsâ subjective workload associated with the annotation task. The instrument evaluates perceived workload across six dimensions: mental demand, physical demand, temporal demand, performance, effort, and frustration. On each dimension, participants provided a separate rat- ing for a single, dedicated question on a 0â20 scale, reflecting their subjective assessment of the respective aspect of workload. In the present study, the raw NASA-TLX score (Raw-TLX) was applied; no weighting procedure was used. The overall workload score was calculated by aggregating the ratings across the six dimensions. NASA-TLX was selected due to its established reliability and validity in assessing subjective workload in human-system interaction and task- based performance studies. 6M. Markiewicz et al. 5 Qualitative study methodology Two qualitative analyzes were performed to gain a deeper insight into the find- ings on the development of competences among annotators in identifying social influence. To this aim, two independent analyzes were conducted: a content anal- ysis of annotatorsâ free-text responses, followed by in-depth interviews. 5.1 Content analysis of free-text responses For content analysis, we adopted a conventional content analysis methodology based on [10]. The analysis of annotatorsâ responses primarily focused on iden- tifying differences in RICs at the semantic level and recurring change patterns between two time points (Pre and Post annotations). Based on the data, we manually defined qualitative criteria for changes in RICs. For all three, it included the number of semantically different observations, a breadth of perspective (narrowing, broadening, no change, reformulation), the level of detail in descriptions (more, less, no change), the style of language (for- mal, casual), and thematic categories of items specific to each. For the consequences of social influence, the following additional criteria were defined: reversibility (more focus on reversible, more focus on irreversible, no change), inevitability (more focus on potential, more focus on certain, no change), reach (individual, group, social), and time horizon (shorter, longer, no change). Thematic categories were psychological, social, economic, health, moral, and legal. For the intentions of a person exerting influence, the additional criteria included a change in focus concerning thematic categories: intentions related to behaviors, actions, and facts, or related to motivations, goals, and mental pro- cesses. For the reactions of the influenced person, we also assessed the change in focus. The additional thematic categories are emotional, cognitive, and behav- ioral. Detailed definitions and codes are available in the supplementary material. Next, following [16], we performed LLM-assisted coding (using gpt-4o-mini) of the above criteria, followed by expert verification. An LLM was only involved in the coding task and was not provided with information about the source (Pre or Post dataset) of each response. The final analysis and interpretation of the results were performed by the authors. 5.2 A thematic analysis of in-depth interviews Next, a semi-structured interview was conducted after the annotation process with five annotators, one from each group. Semi-structured interviews allow for balancing comparability between participants with the flexibility to pursue novel themes that emerge during the conversation [5]. The interview consisted of 13 open-ended questions to provide detailed information and reflections on the changes occurring in the process of double annotation of the same texts in terms of identifying and justifying social influence. To analyze the interview responses, we adopted thematic analysis as de- scribed by Braun and Clarke [4], which provides a systematic framework for Competence Development in Social Influence Recognition7 identifying, analyzing, and reporting patterns within qualitative data. The re- sponses of the respondents were recorded, transcribed, and analyzed in four described categories of the annotation process: (i) changes in the interpretation of texts, (i) changes in speed, efficiency, and confidence in recognized techniques, (i) changes in critical attitudes toward AI-generated texts, and (iv) self-perceived ability to explain the social influence mechanisms to others. In addition, we re- ceived information about annotatorsâ strategies for memorizing techniques and difficulties in the annotation process. This framework supports the transparency and credibility of the findings de- rived from semi-structured interview transcripts. We applied a hybrid inductive- deductive thematic analysis. The initial codes were grounded in the interview guide, which led to the emergence of thematic categories. At least two researchers- experts independently conducted the coding, and disagreements were resolved through discussion. Finally, the themes were organized into higher-level cate- gories that reflected the growth in competence. 6 Results and analyses A high number of annotation changes was encountered in the Pre/Post data. Table 1 presents the changes. Expert groups showed significant increases (Mann- Whitney U test, p < 0.05 for reactions, p < 0.001 for intentions and conse- quences), while non-experts did not. Table 1. Pre, Post, and change (â = PostâPre) in counts of techniques, reactions, intentions, and consequences by group. The biggest absoluteâ is marked in bold. Group TechniquesReactionsIntentionsConsequences Pre Post â Pre Post â Pre Post â Pre Post â Expert2.06 2.29 0.23 1.40 1.83 0.43 1.49 1.91 0.42 2.28 3.27 0.99 Non-expert 2.68 2.51 -0.17 1.42 1.42 0.00 1.69 1.62 -0.07 2.45 2.51 0.06 Overall2.43 2.42 -0.01 1.41 1.58 0.17 1.61 1.73 0.12 2.38 2.81 0.43 Qualitative analysis pointed to differences in how the annotators articulated social influence qualities. Post responses were generally more detailed and offered a broader perspective, as presented in Table 2. This corresponds well with an observed increase in the number of RICs, suggesting that new elements contain novel observations. This is further evaluated by assessing semantically different concepts and their corresponding thematic categories. Unlike quantitative an- alyzes, this effect is observed for all groups, not only experts, but it is more prevalent among them (particularly among the communication experts group). 6.1 Changes in the formulation of free-text responses Figure 2 shows thematic category distributions across RICs. Post annotations ex- hibited more balanced distributions across categories, consistent with increased 8M. Markiewicz et al. Table 2. Distribution of categorical changes between Pre and Post annotations across consequences, reactions, and intentions. Change TypeConsequences Reactions Intentions Breadth: Broadening 275 (44.2%) 230 (46.7%) 261 (39.7%) Breadth: Narrowing142 (22.8%) 103 (20.9%) 183 (27.9%) Breadth: No change196 (31.5%) 159 (32.3%) 199 (30.3%) Detail: More details341 (54.8%) 256 (52.0%) 395 (60.1%) Detail: Less details201 (32.3%) 153 (31.1%) 231 (35.2%) Detail: No change80 (12.9%)83 (16.9%)31 (04.7%) thematic coverage. Initial Pre annotations described the intentions behind social influence mainly as motivations or goals rather than behaviors or actions (861 vs 354 distinct intentions). Post responses added slightly more of the latter (923 vs 449), resulting in a more balanced distribution. Analyzes of language change indicate the adoption of a more formal (194 shifts from casual to formal, 74 vice versa) language, which may be a sign of greater fluency in describing the intentions. Motivations/Goals/MentalBehaviors/Actions/Facts 0 200 400 600 800 Count Intentions - Thematic Categories Cognitive BehavioralEmotional 0 50 100 150 200 250 300 350 Count Reactions - Thematic Categories PsychologicalSocialEconomicHealthMoral 0 100 200 300 400 Count Consequences - Thematic Categories Pre AnnotationsPost Annotations Fig. 2. Shifts in the number of individual RICs by thematic categories. For reactions to social influence, a similar change in the balance of the dis- tribution of thematic categories may be observed, with an increase in the least common categories. Language analyzes support this, with a more common use of casual language (154 changes from formal to casual, 74 vice versa), which may be associated with a higher number of concrete behavioral reactions. In describing the consequences of social influence, all participant groups tended to preserve a similar temporal horizon, with no marked shifts in the perceived time span of effects. A slight shift towards more formal language (118 vs 35 changes) was observed. In terms of thematic categories, a slight increase in the least common categories was observed, but there was no significant change in the overall distribution. A more noticeable change emerged in the dimensions of inevitability and range. Post annotations more often framed the consequences of social influence as potential, hypothetical, or conditional rather than as definite Competence Development in Social Influence Recognition9 outcomes (335 changes towards potential, 74 towards definite, 223 no change), indicating a move toward greater interpretative caution and openness in conse- quence formulation. The perception of range has broadened, with a shift from noticing mostly consequences concerning one person (163 individual, 43 group, 1 systematic in Pre) to more people (604 individual, 198 group, 7 systematic in Post). This indicates a higher awareness of the outcomes. In general, qualitative analyzes suggest that changes usually involve the en- richment, refinement, or reconsideration of the emphasis between categorical components, rather than a significant conceptual change. These observations may indicate the increased competence of the annotators. 6.2 Annotation Speed and Efficiency Analysis of annotation time revealed a significant decrease between the first and the repeated rounds. In the first round, annotators spent an average of 6.56 minutes per text (median = 5.49, SD = 3.99), while in the repeated round, it decreased to an average of 5.86 minutes (median = 4.69, SD = 3.74). This difference was statistically significant according to the Mann-Whitney U test (U = 171438.0, p = 0.0009). Trend analysis across the entirety of the Main annotation revealed a consis- tent pattern of decreasing annotation time. The overall combined trend showed a significant negative slope (ÎČ = â0.0084, R 2 = 0.224, p < 0.001, Figure 3), indicating that annotators became progressively faster throughout the process. 0255075100125150175 Sequence number of the annotation of a given annotator 4 5 6 7 8 Time (minutes) Raw mean MA (window=20) Trend (slope=-0.0084, RÂČ=0.224) Fig. 3. Annotation time change over the course of the process. 6.3 Annotator agreement For technique classification, inter-annotator agreement (Krippendorffâs α with Jaccard distance) improved slightly from α = 0.319 to α = 0.327, though this difference was not statistically significant. Intra-annotator agreement, measured as the pairwise α between each annotatorâs Pre and Post labels, was higher than inter-annotator agreement (mean α = 0.442, SD = 0.121), which is consis- tent with the characteristics of subjective tasks but is still quite low, suggesting 10M. Markiewicz et al. difficulty [2]. However, perfect intra-annotator agreement would exclude the pos- sibility of competence increase. Group-Level Differences Expert annotators demonstrated both higher inter- (α = 0.383 â 0.405) and intra-annotator agreement (mean α = 0.514, SD = 0.069) compared to non-experts (inter- α = 0.290 â 0.286; intra- mean α = 0.394, SD = 0.124). Notably, inter-annotator agreement improved for the expert group (+0.023) while remaining essentially unchanged for non-experts (â0.004), suggesting that the annotation process might have reinforced convergence pri- marily among those with prior domain knowledge. 6.4 Technique Stability The analysis revealed systematic shifts in annotation patterns. While the total number of technique labels remained stable (1821 vs 1816), specific assignments changed substantially, with an average retention rate of 65% across techniques. The most consistent techniques were those with clear behavioral markers: Door- in-the-face and Show disappointment (75% retention), Flattery (74.5%), and Give to take (71.2%). In contrast, abstract affective techniques showed lower stability, with Liking retained only 40.8% of the time. The most frequent substitution patterns â Liking â Labeling (15 instances), We are exceptional â The âWeâ rule (11), and Gratitude â Give to take (10) â suggest that annotators refined broad, intuitive categorizations into more spe- cific, behaviorally-defined techniques. This is reflected at the aggregate level: Labeling showed the largest net increase (+46.4%), while Liking showed the largest decrease (â25.2%). These patterns indicate that the annotation process might have enhanced annotatorsâ ability to distinguish between conceptually overlapping techniques. 6.5 LLM training Post data consistently improved model performance across all settings (Table 3. For ICL, performance scaled positively with the number of few-shot examples (n), and the Pre/Post gap widened as n increased. All ICL results were above a zero-shot baseline of 0.3833. While the difference was negligible at n = 3 (â = +0.0011), the Post set demonstrated a distinct advantage at n = 30 (â = +0.0149). For SFT, Llama-3.1-8B-Instruct improved from an untrained baseline of 0.1235. Post data yielded a small but positive delta of (â = +0.0069), consistent with the trajectory observed in ICL experiments. The size of the improvement is limited due to a small number of training samples (149). All presented data used the AC MV consensus. We found that using the looser AC 2 consensus yielded a much worse overall performance, with insignificant differences observed between conditions (â < SD). Competence Development in Social Influence Recognition11 Table 3. Model performance difference when using Pre and Post data for training. Model and settingPre (Jaccardâ)Post (Jaccardâ) â Jaccard DeepSeek-V3.2 (3-shot)0.4523 (± 0.0058) 0.4534 (± 0.0058)+0.0011 DeepSeek-V3.2 (10-shot)0.4874 (± 0.0044) 0.4962 (± 0.0014)+0.0088 DeepSeek-V3.2 (30-shot)0.5214 (± 0.0048) 0.5363 (± 0.0033)+0.0149 Llama-3.1-8B-Instruct (SFT)0.26840.2753+0.0069 6.6 Workload perception The analysis of NASA-TLX scores revealed statistically significant differences between the measurements. The mean TLX score increased from 50.80 (SD = 15.95) after Pre annotations to 60.56 (SD = 19.21) after Post annotations. At the dimensional level, the second measurement showed higher scores for mental demand, physical demand, temporal demand, effort, and frustration. Conversely, the performance score significantly decreased, indicating a lower subjective assessment of task performance despite a higher workload (see 4). According to the literature, this is a characteristic of the learning phase or "con- scious incompetence", where gaining knowledge makes learners more aware of their own mistakes [13]. Mental DemandPhysical DemandTemporal DemandPerformanceEffortFrustration 0 5 10 15 20 Measurement After pre annotations After post annotations Fig. 4. Comparison of mean scores of NASA-TLX dimensions between measurements. 6.7 Annotatorsâ perception of competence development and confidence Figure 5 presents the Perceived Competence and Self-Confidence SIR-SC scales at both measurement points. Perceived Competence showed a notable increase with a large effect size (Cohenâs d = 0.567), indicating a substantial impact on participantsâ self-assessed mastery. Self-Confidence also improved significantly, with a medium effect size (Cohenâs d = 0.388). Survey results aligned with answer certainty from annotation data: certainty scores increased significantly across all groups (p < 0.001, Mann-Whitney U test) from M = 3.67,SD = 0.73 to M = 3.92,SD = 0.70. Qualitative analysis of open responses revealed that the most frequently men- tioned gains related to distinguishing social influence (31.7%), followed by tech- 12M. Markiewicz et al. 012345 Perceived competence Self confidence After Pre annotationsAfter Post annotations Fig. 5. Box Plots of Perceived Competence and Self-confidence across measurements. nique detection, increased awareness, and the ability to name specific mecha- nisms (each 19.5%). Less common were awareness of consequences (7.3%) and knowledge of assertive responses (2.4%). These answers are reflected and elabo- rated upon in the analysis of extended interviews: Changes in the interpretation of texts were described as the reconstruc- tion of additional context, re-interpretation of irony as manipulation, and re- assessment of previously neutral statements for hidden intent and consequences. One annotator framed this as a deepened analysis rather than a revision. The main challenge was differentiating conceptually similar techniques, especially those tied to group identity, and deciding when multiple techniques were ap- plied to one text. With continued annotation, the boundaries became clearer. These answers suggest an increase in sensitivity to subtle cues. However, some changes between rounds were not fully understood, despite being noticed. Development of speed, confidence, and detection ability over time was reported by most annotators, consistent with previous findings. They no- ticed their judgments shifting from moderate to more polarized high-confidence ratings, reporting that competence developed ânaturallyâ through continual ref- erence to definitions and examples. One interviewee diverged, stating that later texts felt more ambiguous and cognitively demanding, leading to longer deliber- ations and more frequent uncertainty. Increasing critical awareness of AI-generated content was mentioned by all participants, especially regarding AI-generated texts. Texts perceived as unrealistic elicited negative emotions. One interviewee noted a pre-existing crit- ical stance due to professional experience but still reported increased awareness of the scale and pervasiveness of influential content. The ability to explain manipulation to others has improved, as re- ported by all participants. Getting familiar with formal terminology was re- peatedly described as enabling clearer explanations of previously intuitive im- pressions, thereby increasing confidence and ease of communication. Some had already discussed techniques with their children or peers. Several also noted that annotation is likely one component of greater education to counteract social in- fluence, which directly references the scope of this study. Additional remarks included developing deliberate learning strategies de- scribed by four participants and the experienced spontaneous detection of ma- Competence Development in Social Influence Recognition13 nipulation in everyday communication. Annotators relied mainly on repeated ex- posure to guidelines and extensive practice. Prior domain familiarity (in expert groups) supported faster learning, while unfamiliar techniques required more in- tentional practice and consultation of training materials. One participant noted recurring lexical/situational patterns (e.g., repeated brand-name cues) that fa- cilitated classification, while others did not report this. 7 Discussion Annotation of social influence is difficult and highly subjective, as reflected by moderate inter- and intra-annotator agreement [2]. This is further supported by the fact that both measures are higher in the expert group compared to non- experts. Although changes between annotation rounds are necessary to observe improvement, groups with higher intra-annotator agreement (fewer changes) tended to exhibit more characteristics of an increase in data quality. All annotators learned to perform the task more quickly, and their self- perceived competence and confidence increased significantly. Annotators also began to notice the effects of learning about social influence extending beyond the scope of their work. An improved perception of the intentions of a person exerting social influence, its possible consequences, and a greater awareness of the reactions of the target are supported by an increase in the mean number of these observations. The better quality of these was confirmed in a qualitative analysis for all groups, with a broadened perspective, a higher level of detail, and wider thematic coverage. Finally, our intention was for annotators to be active participants in designing the system aimed at supporting adolescents in resisting social influence. There- fore, high annotator competence and their involvement in defining educational goals and processes, further implemented with the use of LLMs trained on the annotated data, are essential. 8 Conclusions Regarding RQ2 and RQ3, improved model performance and increased annotator agreement suggest higher data quality over time. In response to RQ1, the ob- served effects may indicate that annotator competence increases over the course of the annotation process, particularly among annotators who are already ex- perts. It is therefore necessary to monitor annotatorsâ competence development, as it has a direct impact on the performance of models trained on their data. The process of annotating social influence can be seen as an effective ed- ucational intervention. It strengthens awareness of persuasive intent, improves recognition of subtle manipulation, and increases participantsâ confidence in com- municating about manipulation with others. Our findings support the integration of annotation-based activities into educational programs addressing AI literacy and protecting youth from harmful, persuasive content. 14M. Markiewicz et al. Directions for future work in this area include investigating larger participant groups across multiple subjective tasks. The additional incorporation of non- subjective tasks could enable external, quantitative competence assessment and provide a more direct interpretation of agreement measures. Acknowledgments. This work was financed by (1) the National Science Centre, Poland, project no. 2021/41/B/ST6/04471; (2) the statutory funds of the Department of Artificial Intelligence, WrocĆaw University of Science and Technology; (3) the Pol- ish Ministry of Education and Science within the programme âInternational Projects Co-Fundedâ; (4) the European Union under the Horizon Europe, grant no. 101086321 (OMINO). However, the views and opinions expressed are those of the author(s) only and do not necessarily reflect those of the European Union or the European Research Executive Agency. Neither the European Union nor European Research Executive Agency can be held responsible for them. References 1. Abercrombie, G., et al.: Temporal and second language influence on intra-annotator agreement and stability in hate speech labelling. In: Proc. 17th Linguistic An- notation Workshop (LAW-XVII). p. 96â103. ACL, Toronto, Canada (2023). https://doi.org/10.18653/v1/2023.law-1.10 2. Abercrombie, G., et al.: Consistency is key: Disentangling label variation in nlp with intra-annotator agreement. In: Proc. 4th Workshop on Perspectivist Ap- proaches to NLP. p. 63â74. ACL, Suzhou, China (2025). https://doi.org/10. 18653/v1/2025.nlperspectives-1.6 3. Bassi, D., et al.: Annotating the annotators: Analysis, insights and modelling from an annotation campaign on persuasion techniques detection. In: Findings of ACL. p. 17918â17929. ACL, Vienna, Austria (2025). https://doi.org/10.18653/v1/ 2025.findings-acl.922 4. Braun, V., Clarke, V.: Using thematic analysis in psychology. Qualitative Research in Psychology 3(2), 77â101 (2006) 5. DeJonckheere, M., Vaughn, L.M.: Semistructured interviewing in primary care research. Family Medicine and Community Health 7(2) (2019), e000057 6. Doroudi, S., et al.: Toward a learning science for complex crowdsourcing tasks. In: Proc. CHI. p. 2623â2634. ACM, New York, NY, USA (2016). https://doi.org/ 10.1145/2858036.2858268 7. Fleisig, E., et al.: The perspectivist paradigm shift: Assumptions and challenges of capturing human labels. In: Proc. NAACL-HLT. p. 2279â2292. ACL, Mexico City, Mexico (2024). https://doi.org/10.18653/v1/2024.naacl-long.126 8. Hart, S.G., Staveland, L.E.: Development of the NASA-TLX (Task Load Index). In: Human Mental Workload. Advances in Psychology, vol. 52, p. 139â183. Elsevier, Amsterdam (1988). https://doi.org/10.1016/S0166-4115(08)62386-9 9. Hata, K., et al.: A glimpse far into the future: Understanding long-term crowd worker quality. In: Proc. CSCW. p. 889â901. ACM (2017). https://doi.org/ 10.1145/2998181.2998248 10. Hsieh, H.F., Shannon, S.: Three approaches to qualitative content analysis. Qual- itative Health Research 15, 1277â1288 (11 2005). https://doi.org/10.1177/ 1049732305276687 Competence Development in Social Influence Recognition15 11. Kolb, D.A.: Experiential Learning: Experience as the Source of Learning and De- velopment. Prentice-Hall, Englewood Cliffs, NJ (1984) 12. Lee, J.U., et al.: Annotation curricula to implicitly train non-expert annotators. Comput. Linguist. 48, 343â373 (2022). https://doi.org/10.1162/coli_a_00436 13. Mohamed, R., et al.: Validation of the NASA-TLX to evaluate the learning curve for endoscopy training. Can. J. Gastroenterol. Hepatol. 28, 892476 (2014). https: //doi.org/10.1155/2014/892476 14. Mokhberian, N., et al.: Capturing perspectives of crowdsourced annotators in sub- jective learning tasks. In: Proc. NAACL-HLT. p. 7337â7349. ACL, Mexico City, Mexico (2024). https://doi.org/10.18653/v1/2024.naacl-long.407 15. Renkl, A.: Toward an instructionally oriented theory of example-based learning. Cognit. Sci. 38, 1â37 (2014). https://doi.org/10.1111/cogs.12086 16. Tai, R.H., et al.: An examination of the use of large language models to aid analysis of textual data. Int. J. Qual. Methods 23, 16094069241231168 (2024). https:// doi.org/10.1177/16094069241231168 17. Vitello, S., et al.: What is competence? A shared interpretation of competence to support teaching, learning and assessment. Cambridge University Press (2021). https://doi.org/10.17863/CAM.110829 18. Yoo, H., et al.: Rethinking annotation: Can language learners contribute? In: Proc. ACL. p. 14714â14733. ACL, Toronto, Canada (2023). https://doi.org/ 10.18653/v1/2023.acl-long.822