Paper deep dive
LLMs for Qualitative Data Analysis Fail on Security-specificComments in Human Experiments
Maria Camporese, Fabio Massacci, Yuanjun Gong
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 95%
Last extracted: 4/14/2026, 2:25:42 AM
Summary
This paper investigates the efficacy of Large Language Models (LLMs) as automated annotators for thematic analysis of security-specific technical comments in human experiments. By comparing LLM performance against human annotators using Cohen's Kappa, the study finds that while detailed codebooks improve performance, LLMs are currently insufficient to reliably replace human annotators for complex security-relevant coding tasks.
Entities (7)
Relation Signals (4)
Maria Camporese → authored → LLMs for Qualitative Data Analysis Fail on Security-specific Comments in Human Experiments
confidence 100% · Authors: Maria Camporese, University of Trento (Italy) Fabio Massacci...
Fabio Massacci → authored → LLMs for Qualitative Data Analysis Fail on Security-specific Comments in Human Experiments
confidence 100% · Authors: Maria Camporese... Fabio Massacci...
Yuanjun Gong → authored → LLMs for Qualitative Data Analysis Fail on Security-specific Comments in Human Experiments
confidence 100% · Authors: ... Yuanjun Gong, University of Trento (Italy)
LLMs → performedtask → Thematic Analysis
confidence 90% · Explore whether LLMs can act as automated annotators for technical security comments
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:[Background:] Thematic analysis of free-text justifications in human experiments provides significant qualitative insights. Yet, it is costly because reliable annotations require multiple domain experts. Large language models (LLMs) seem ideal candidates to replace human annotators. [Problem:] Coding security-specific aspects (code identifiers mentioned, lines-of-code mentioned, security keywords mentioned) may require deeper contextual understanding than sentiment classification. [Objective:] Explore whether LLMs can act as automated annotators for technical security comments by human subjects. [Method:] We prompt four top-performing LLMs on LiveBench to detect nine security-relevant codes in free-text comments by human subjects analyzing vulnerable code snippets. Outputs are compared to human annotators using Cohen's Kappa (chance-corrected accuracy). We test different prompts mimicking annotation best practices, including emerging codes, detailed codebooks with examples, and conflicting examples. [Negative Results:] We observed marked improvements only when using detailed code descriptions; however, these improvements are not uniform across codes and are insufficient to reliably replace a human annotator. [Limitations:] Additional studies with more LLMs and annotation tasks are needed.
Tags
Links
- Source: https://arxiv.org/abs/2604.10834v1
- Canonical: https://arxiv.org/abs/2604.10834v1
Trouble viewing inline? Open PDF directly →
Full Text
76,715 characters extracted from source content.
Expand or collapse full text
by LLMs for Qualitative Data Analysis Fail on Security-specific Comments in Human Experiments Authors: Maria Camporese, University of Trento (Italy) Fabio Massacci, University of Trento (Italy), Vrije Universiteit Amsterdam (The Netherlands) Yuanjun Gong, University of Trento (Italy) This work has been partly supported by the European Union (EU) under Horizon Europe grant n . 101120393 (Sec4AI4Sec), by the Nederlandse Organisatie voor Wetenschappelijk Onderzoek (NWO) under grant n. KIC1.VE01.20.004 (HEWSTI), and by the Italian Ministry of University and Research (MUR), under the P.N.R.R. – NextGenerationEU grant n. PE00000014 (SERICS). This paper reflects only the author’s view and the funders are not responsible for any use that may be made of the information contained therein. Cybersecurity for AI-Augmented Systems (Sec4AI4Sec) . As artificial intelligence (AI) becomes omnipresent, even integrated within secure software development, the safety of digital infrastructures requires new technologies and new methodologies, as highlighted in the EU Strategic Plan 2021-2024. To achieve this goal, the EU-funded Sec4AI4Sec project will develop advanced security-by-design testing and assurance techniques tailored for AI-augmented systems. These systems can democratise security expertise, enabling intelligent, automated secure coding and testing while simultaneously lowering development costs and improving software quality. However, they also introduce unique security challenges, particularly concerning fairness and explainability. Sec4AI4Sec is at the forefront of the move to tackle these challenges with a comprehensive approach, embodying the vision of better security for AI and better AI for security. More information at https://sec4ai4sec.eu. Hybrid Explainable Workflows for Security and Threat Intelligence (HEWSTI) In research into threats to safety and security, people and AI collaborate to obtain actionable intelligence. Their sources and methods often have significant uncertainties and biases. Experts are aware of these limitations, but lack the formal means to handle these uncertainties in their daily work. This project will invent a ‘metadata of uncertainty’ for threat intelligence (in both machine-readable and also human-interpretable forms) and validate it empirically. Intelligence agencies will then be able to explicitly consider the trade-off between the accuracy, proportionality, privacy, and cost-effectiveness of investigations. This will contribute towards the responsible use of AI to create a safer, more secure society. In searCh Of eVidence of stEalth cybeR Threats (COVERT) AT 3 aims to analyze emerging attack methodologies and develop advanced methods for detecting attacks and identifying guidelines for designing IT systems that ensure reduced vulnerability to new attack categories. The detailed objectives can be divided into four macro categories: (i) Development of advanced tools for analyzing malware and software aimed at identifying vulnerabilities that could be exploited by malware; (i) Development of tools for analyzing network traffic to identify communications related to ongoing attacks; (i) Development of machine learning systems that are robust to attacks and through which it is possible to extract knowledge aimed at creating more advanced tools for timely analysis and early identification of attacks; (iv) Analysis of the ”human factors” involved in an attack with the development of tools for analyzing and correlating information from OSINT (open sources intelligence) and for the defense and prevention of attacks based on social engineering techniques. Maria Camporese (MSc 2022) is a PhD student at the University of Trento, Italy. Her research interests include security and machine learning. Contact her at maria.camporese@unitn.it. Fabio Massacci (Phd 1997) is a professor at the University of Trento, Italy, and Vrije Universiteit Amsterdam, Fabio Massacci is a professor at the University of Trento, Trento, Italy, and Vrije Universiteit Amsterdam, 1081 HV Amsterdam, The Netherlands. His research interests include empirical methods for the cybersecurity of sociotechnical systems. For his work on security and trust in sociotechnical systems, he received the Ten Year Most Influential Paper Award at the 2015 IEEE International Requirements Engineering Conference. He is named co-author of CVSS v4. He leads the Horizon Europe Sec4AI4Sec project and the Dutch National Project HEWSTI. He is past chair of the Security and Defense Group of the Society for Risk Analysis, and IEEE CertifAIEd Lead Assessor. Contact him at fabio.massacci@ieee.org. Yuanjun Gong (PhD 2025) is a postdoc researcher at the University of Trento, Italy. Her research interests include static analysis, software security and machine learning. Contact her at yuanjun.gong@unitn.it. How to cite this paper: • Camporese, M., Massacci, F. and Gong, Y. LLMs for Qualitative Data Analysis Fail on Security-specific Comments in Human Experiments. Proceedings of the 2026 IEEE/ACM 34st International Conference on Program Comprehension (ICPC). IEEE Computer Society. License: • This article is made available with a perpetual, non-exclusive, non-commercial license to distribute. LLMs for Qualitative Data Analysis Fail on Security-specific Comments in Human Experiments Maria Camporese University of Trento, IT maria.camporese@unitn.it , Fabio Massacci University of Trento, IT, Vrije Universiteit Amsterdam, NL fabio.massacci@ieee.org and Yuanjun Gong University of Trento, IT yuanjun.gong@unitn.it (2026) Abstract. [Background:] Thematic analysis of free-text justifications in human experiments provides significant qualitative insights. Yet, it is costly because reliable annotations require multiple domain experts. Large language models (LLMs) seem ideal candidates to replace human annotators. [Problem:] Coding security-specific aspects (code identifiers mentioned, lines-of-code-mentioned, security keywords mentioned) may require deeper contextual understanding than sentiment classification. [Objective:] explore whether LLMs can act as automated annotators for technical security comments by human subjects. [Method:] We prompt four best LLMs on LiveBench to detect nine security-relevant codes in free-text comments by human subjects analyzing vulnerable code snippets. Outputs are compared to the human annotators along with Cohen’s Kappa (chance-corrected accuracy). We test different prompts mimicking annotation best practices: emerging codes, a detailed codebook with examples, and conflicting examples. [Negative Results:] We observed marked improvements only with the code descriptions, but they are not uniform across codes and not sufficient to reliably replace a human annotator. [Limitations:] Additional studies with more LLMs and annotation tasks are needed. Qualitative Data Analysis, Thematic Analysis, Large Language Models, Security-Specific Annotation †journalyear: 2026†copyright: c†conference: 34th IEEE/ACM International Conference on Program Comprehension; April 12–13, 2026; Rio de Janeiro, Brazil†booktitle: 34th IEEE/ACM International Conference on Program Comprehension (ICPC ’26), April 12–13, 2026, Rio de Janeiro, Brazil†doi: 10.1145/3794763.3798172†isbn: 979-8-4007-2482-4/2026/04 1. Introduction Who has run a software engineering experiment with human subjects and not received a reviewer’s comment “You should manually analyze the participants’ explanations”? The process for doing this qualitative analysis is called Thematic Analysis (TA) (Dai et al., 2023). TA is a method for identifying, coding, and interpreting patterns of meaning (themes) within qualitative data (Clarke and Braun, 2017). The aim is to distill insight from large volumes of qualitative material to allow for quantitative analysis (Zhang and Wildemuth, 2009). Fragments of textual, audio, or video recordings are annotated with labels called codes. Codes represent the smallest meaningful units, capturing noteworthy aspects of the recorded material. These codes are used for constructing broader themes, i.e. recurring patterns of meaning. To ensure reliable analysis, the same piece of data is assigned to at least one coder and one reviewer. Further, human coders develop and deepen their data interpretation and coding over multiple iterations, making TA labor-intensive and time-consuming. In the realm of empirical software and security engineering, thematic analysis is used in a variety of settings. For example, opinion mining over software repositories’ comments is a popular research area (Lin et al., 2022). In experiments with humans, qualitative analysis is used to validate the rationale behind their choices (Papotti et al., 2024, 2025). When discussing coding for opinion mining, one typically focuses on the Sentiment polarity identification (e.g., positive, neutral, or negative (Lin et al., 2022)), viewpoints and perspectives identification (e.g., general reasons behind forking a repository (Jiang et al., 2017)), or other knowledge extraction (e.g. usefulness of code reviews (Rahman et al., 2017)). This can be done both manually or automatically, e.g., the cited work by Jiang (Jiang et al., 2017) had both a manual annotation of open-ended questions and automated analysis of developer responses. More recent works proposed to use LLMs (Dai et al., 2023; Mæhlum et al., 2024). Using LLMs for replacing at least one human annotator seems, therefore, a low-hanging fruit. We are interested in more specific, security-relevant annotations in which the LLM has to extract more domain-specific answers, such as whether the vulnerability type is mentioned in the human explanations, or whether a possible exploit is mentioned, or some security keywords are mentioned. They are further detailed in Table 1. Detection of such qualitative codes has also used by SAP to build their heuristic approach to identify security fixes among thousands of commits starting from vulnerability descriptions (Sabetta et al., 2024). We therefore raise two research questions to explore the potential of LLMs for automatic technical comment thematic annotation: RQ1.: Can LLMs replace a human annotator in security technical annotations? RQ2.: What is the impact of different prompts mimicking the best practice of code refinements on LLMs’ performance on security technical annotations? We propose and instantiate an experimental methodology where LLMs receive a dataset of human comments, partially annotated by other humans, and must generate the annotations of security-relevant codes for a fraction not yet annotated comments (the ones assigned to the replaced human). The design of the prompting strategy reflects the human coding procedure (emerging coding at first, followed by a detailed codebook with examples, finally addressing conflicting examples). For the actual experiments, we then identified four state-of-the-art LLMs and asked them to generate the security-relevant codes instead of human annotators. When looking for a dataset for annotations, we realized that LLMs are likely trained on public datasets. Hence, we reached out to the authors of a security study where security technical comments have been manually annotated (Papotti et al., 2024, 2025) and asked if they had additional, not yet published annotations. They provided a thematic analysis of a different, unpublished experiment in which they had multiple annotators. This provided us with a dataset of n=263n=263 human-generated comments where 13 security-relevant codes (c=9c=9 security-relevant codes are tested) have been annotated by four people (see some examples in Table1). Two key observations are important to evaluate such replacement: first, we should not hold an LLM to a higher standard than a human annotator. Second, for a code to be useful, it has to be at once recurrent but sparse to truly capture different concepts (Armborst, 2017). Hence, an LLM could obtain a good accuracy just by chance because a meaningful code is typically not applicable to a random comment. So we measure whether the replaced LLM can achieve the same Cohen’s Kappa as the (replaced) human annotator and their reviewer. The same formula used for Cohen’s Kappa can also be interpreted as chance-corrected accuracy (Barnston, 1992), a metric proposed in weather forecasting to test for rare events. Indeed, when using traditional accuracy, we obtained over 75% in our experiment, when correcting for chance, the results dropped to half that value. Negative Result: LLMs cannot reliably replace human annotators for security-relevant annotations, not even with a detailed code-book and examples from other annotators. 1.1. Artifact Availability The data and the code implementation are open-sourced in https://zenodo.org/records/18742065. 2. Human Annotations - Problem Specification Our research focuses on one of the simplest protocols for generating reliable annotations, as shown in Figure 1. In the qualitative data collection from Papotti et al. (Papotti et al., 2025), the authors describe the process they used for annotation, which is quite standard for applied Thematic Analysis (Guest et al., 2011; Gregory et al., 2015): first, a subset of authors jointly reviewed a sample of justifications for each vulnerable scenario to identify a first set of emerging codes; second, the codes are consolidated into a codebook (Table 8 in (Papotti et al., 2025)). After an additional phase of consultation, the author marked all remaining justifications, and an independent researcher checked them. Finally, additional conflicts were discussed by a subset of the authors and resolved. With human annotators (A,B,C,D,…)(A,B,C,D,...), the procedure should be split down into four steps: Figure 1. Annotation by human annotators and LLMs This figure shows the four stages of the human annotation process. Table 1. Example of Codes and Related Comments Abbr. Code Definition Code applicable Comments Code inapplicable Comments Var Variable/Method identifiers are mentioned I think it had something to do with inv and debug. header buffer capacity is not checked. Lin Lines number mentioned In this line, the memory at constant address is copied, this memory will most likely not stay consistent among systems. The null value may cause the function crash. Key Relevant Keywords mentioned If the file attribute in the code is tampered with for an instance of the BlockDriverState it might happen that the code runs an infinite recursion, essentially stopping the machine. I don’t see any vulnerabilities. Vul Text is related to the specific vulnerability Might modify environment variable with strtol. I think that negative checkin triggers the vulnerability for the ML, not sure. Sec Text related to security A check should be added which checks if the buffer is not overflowing. the datatype of debug wont support strtol. Exp Potential exploit is mentioned The memcpy() function might be injected with a size that is too big for the program to handle, therefore freezing the computer’s resources. stderr is not created. NoVul I don’t find vulnerability Since I did not find the code to be vulnerable, I have not marked any sections. The confidence however is low, as I am struggling to find any vulnerable line. I am not sure. Unsr I’m not sure The written buffer may overflow like said before, but I’m not really sure. memcpy buffer overflow. Unclr Unclear answer Address. peer might be null. (0) Emergence of codes (a) Annotators individually read some samples and make up their own codes (principle of emergence). (b) The annotators gather up the codes and agree on code names (Column 1 of Table1). (1) First annotation with shared codes (names) (a) The annotators get an initial understanding from the background information and the codebook. The annotators individually mark a large fraction of their comments111The golden standard is that all comments should be annotated to avoid influencing each other. In the experience of one author, annotators tend to stop and call a meeting when too many cases accumulate in which the annotator is unsure.. (b) At this point, a formal definition of the codebook is agreed (Column 1 and 2 of Table1). (2) Completed annotation with codebook and first review (a) The annotators get a deeper understanding of the task from the explicit definition for every code; They return to their comments and consolidate their codes on all assigned comments. (b) Each annotator reviews the annotations from another annotator, marking agreement and conflicts. Reviewers are randomly assigned.222To reduce cost, one might consider a single annotator and a single reviewer, but this choice reduces the reliability too much. Also in this scenario, the human reviewer would have to review all comments anyhow, and then they could directly mark them. (c) By the end of this step, the disagreements are discussed and Column 1-4 of Table1 are generated as examples of representative do and don’t. (3) Final consolidate coding with do-s and don’t-s (a) The annotators get a comprehensive understanding of the task from the examples of other annotators, as well as the conflicting samples; Annotators individually finalize the annotations along the codebook, bearing in mind the conflicting examples. (b) By the end of this step, the final review discuss and collectively resolve all standing disagreements. For the application scenario, our research focuses on the technical comments annotation. The technical comments are provided by human subjects in response to a request to perform a security analysis of code snippets or patches. The background information is the summary information on the type of vulnerabilities and evidence shown to the commenter and known to the annotator. The codes are thematic aspects applied to these comments. The annotators are expected to label a code as “1” if the code applies to the participant’s comment and “0” if not. Table1 describes the definition of codes and some representative examples illustrating the annotation tasks, including positive samples and negative samples. For example, in the comment “I think it had something to do with inv and debug”, the code “Variable/Method identifiers are mentioned” (abbreviated as Var) is labeled as “1”, since the comment mentions the identifier “inv”. In our dataset, the four annotators divided the comments into four sets and followed the protocol above, each of them randomly annotating a fourth of the comments and reviewing the annotations of a different annotator. The LLM should start after step 0b, when at least the code names are defined by brainstorming among humans. It should then be able to replace any of the human annotators (not as a reviewer), with an acceptable success in each of the annotation steps (1a, 2a, 3a) using the corresponding information resulting from the previous steps (0b, 1b, 2c). To apply the LLM support, we follow the process of the cognition progress of human annotators, and design the prompts to provide the corresponding information to the LLM. If the LLM has the ability to substitute human annotators, it should have a similar performance after getting the same prompt they got. For the last steps (2a) we consider also a favourable version of the task for the LLM, as we show in the first half of the bottom part of Figure 1: we give it the comments of the other annotators who later would be reviewers of its own task. Finally, for step 3a, we ask it to annotate only the comments on which the humans agreed, while all comments on which humans disagreed are provided as examples. Both ‘helping hands’ should actually boost both Cohen’s Kappa and the chance-corrected accuracy. This favorable condition can be acceptable if the LLM is intended to replace the N-th annotator after the original N-1 humans have already run the protocol on their part and want to delegate the completion of the remaining 1/Nth part of the comments to the LLM. In this scenario, it makes no sense to withhold from the LLM the consensus they have already reached. Still, the LLM is favored as in the actual replacement scenario, it would also have to mark the difficult comments on which there was a human disagreement. 3. Methodology The next section provides a broad overview of the main stages of our study to address our research questions. Some additional steps are further refined, including data collection, input preparation, annotation generation, prompt refinement, and result analysis. 3.1. Methodology Overview Data Collection and human annotation. We initially planned to use data from published experiments (e.g. (Papotti et al., 2025)) when we realized that annotations published on the internet might be known to the LLM. Other experiments on commits and code analysis either did not make available the codes or were not on security relevant (e.g. (Jiang et al., 2017; Papotti et al., 2024)). So we contacted the authors of (Papotti et al., 2024, 2025) and obtained a different dataset that is not yet published. The human annotation procedure is done by four annotators, following the process in § 2. The dataset is represented as a spreadsheet and contains the labels of code snippets or patches, the corresponding comments, and the annotator-labeled codes. We also obtain the reviewer-corrected code for each annotated comment (which we used as the final ground truth in the best prompt). The security-specific comments collection and the annotation will be further described in §3.2. Input Preparation. To be consistent with the human protocol, the input to LLMs is the same spreadsheet where the codes labeled by the selected human annotator are cleared. The remaining information in the spreadsheet, such as the codes of the remaining annotators, can provide extra information for the LLMs to understand the task. As mentioned in Section 2, this is a significant help for the LLM: the human annotators had to perform the work independently. We believe it is justified as the idea is to use the LLM to “finish off” the work for the first #A−1\#A-1 annotators Prompt Refinement. To guide the model effectively, imitating the human annotation process (§2), details and examples are progressively added to the prompts (research goal and experimental design, and format and semantics of the spreadsheet used for annotation, examples of both successes and failures). A more detailed explanation is provided in §3.4. LLM Annotation Generation. The LLM under evaluation will be instructed by the prompts to fill in the empty cells of the input spreadsheet generated from the previous step. The LLM will be prompted to identify the present or absent labels of each code in each participant’s comment, and generate the annotations based on the remaining information from the spreadsheet and the knowledge provided by the prompts. The result of LLM auto-annotation should also be reported in a spreadsheet, generated by LLMs. Analysis. The LLM annotation experiments will be conducted individually, each time employing one #L\#L LLMs and one of the #P\#P prompt types to replace one of the #A\#A annotators. The results for each run will be presented across #C\#C codes. The annotations from LLMs are analyzed against the annotators’ and the reviewers’ annotations, to compute the suitable performance metrics (most notably Cohen’s κ) and identify statistically significant differences among LLMs or prompts. The analysis methods are refined in §3.5. 3.2. Comments Collection and Annotation A realistic dataset requires underlying software engineering experiments with raw comments of the human participants and the code labels from human annotators. The natural production of the data is that, in an underlying software or security engineering experiment, the participants were asked to give security-related comments on the provided code snippet or patches to explain the rationale for the choices of the experiment. Then, the human-annotators analyzed the comments following the steps in §2. The Underlying Experiment. The annotation dataset used in this research came from an unpublished experiment from the researchers behind Papotti et al. (Papotti et al., 2025). In their experiment, a group of researchers conducted a controlled experiment to assess whether AI-generated suggestions could support Computer Science Master’s students in identifying vulnerabilities within source code snippets. In that experiment, participants were shown six code snippets, each in one of three possible versions: a plain (control) version, or a version enhanced with highlights generated via explainability methods. The highlights correspond to codes deemed most relevant by VulDeePecker (Li et al., 2018), a DL-based vulnerability detection model. The highlighted versions were generated using either a white-box method (Gradient × Input (Zeiler and Fergus, 2014)) or a black-box method (LEMNA (Guo et al., 2018)). The code snippets included common vulnerability types: buffer overflows, format strings, and NULL pointer dereferences. The experimenters collected both quantitative metrics and qualitative comments from participants justifying their vulnerability detection decisions, resulting in a dataset of 263 free-text comments, each corresponding to a participant’s rationale for detecting (or failing to detect) a vulnerability in a given code snippet. They were further annotated by other annotators. 3.3. LLM Selection M1: Models must be able to process spreadsheet files with annotations. Annotations are supplied with several different codes at once, and the model must be able to mimic the use. M2: Different best models on LiveBench (White et al., 2024). At the time of designing, the best-performing proprietary models for Reasoning Average score are ChatGPT-5 (OpenAI, 2025a), o3 (OpenAI, 2025b) (Both from OpenAI), and Claude Sonnet 4 (Anthropic, 2025) (from Anthropic). M3: Open Source vs Proprietary Models. Some of the models should be open source to provide a replicable experiment. The best-performing open-source models for Reasoning Average score are DeepSeek-V3.2 (DeepSeek-AI et al., 2025), Kimi K2 Thinking (Team et al., 2025) from Moonshot AI, and Qwen3-Max (Yang et al., 2025) (from Alibaba). M4: Models must produce results of the expected size. In preliminary experiments, some models failed to return the correct number of entries, either omitting part of the input or generating too many rows. Because this made it impossible to reliably match their outputs to the ground truth, we discarded models that did not produce the required number of rows. For the proprietary models, we use the chat interface and the professional subscription, as we assume it is the best available model. We also investigated Grok (xAI, 2024) (from xAI), but it is not available in Italy at the time of designing. Besides DeepSeek-V3.2, Kimi K2 Thinking, and Qwen3-Max, we also examined Mistral (Mistral AI, 2024) and Llama (Meta AI, 2024) through the Amazon Bedrock interface. 3.4. Prompt Engineering To answer RQ1, we have crafted a series of prompts to enable the LLMs to function as an annotator. At the same time, to answer RQ2, the prompts are applied sequentially, involving the progressive enrichment of knowledge on code definitions, annotated examples, targeted clarifications, and interactive feedback mechanisms. Prompt Selection Prompts are compared to check whether increasing prompt complexity yields significant improvements (RQ2): P1 P_1: The prompt contains only the task description and the name of the codes. P1 P_1 describes research context, experiment design, spreadsheet structure (with code names), and annotation task, and it provides the file without the applicable code of the other annotators. To generate a correct annotation without clues, the model can only rely on having seen something similar in the past. P2 P_2: The prompt includes the formal definition of the codes. P2 P_2 describes the expected behavior and provides an example for each column to be filled. This stage corresponds to the development of a formal codebook (Table 1). P3 P_3: The prompt includes the examples of the other annotators. P3 P_3 provides the file with the labels of other annotators on other comments. Few-shot examples are considered a significant help for the LLMs, but they are not considered a good practice for human annotators, as they can be a source of bias and lack of independence. P3pilot P_3pilot: The prompt asks for conflicting examples and provides clarifications from the researcher.P3pilot P_3pilot asks the pilot LLM to go through the human-annotated labels and extract two conflicting examples for each code, and ask what is unclear. A researcher then compiles the response to LLM’s questions. This is done only once in the interaction. P3+ P_3+: The prompt shows conflicting examples and the clarifications from researchers. P3+ P_3+ shows LLM the conflicting examples obtained in the pilot interaction, as well as the researchers’ answers to those ambiguities. The selection of prompts is based on a pilot experiment elaborated below, e.g., the prompt P3+ P_3+ is designed based on P3pilot P_3pilot. Pilot Experiment User (P3pilot P_3pilot) …Try to find two examples for each column that you find similar but have opposite labeling, so I can clarify subtle differences for you. GPT-4o I’ve found contrastive examples from your annotations for each label, where student responses are similar in phrasing or content but were marked differently. Below are two examples for each column, showing the kind of subtle distinctions your reviewers made. You can now help me improve by clarifying these distinctions. Variable/Method identifiers are mentioned: - ‘I think it had something to do with inv and debug’ → marked 1 - ‘I did not select anything because […] I am not that experienced with C’ → marked 0 Clarify: Are only clear technical identifiers enough, or should any code-ish word be counted? [Omit the contrastive examples for other codes …] User (P3pilot P_3pilot) - Variable/Method identifiers are mentioned: The first sentence uses “inv” and “debug” in a way that you can understand they are not part of a natural language sentence, so they must be code identifiers. So the mark is 1. The second sentence is a natural language sentence, the only part for which you could have doubts is “C”, but in this case it’s part of the natural sentence expressing the student did not notice any vulnerability, maybe because he does not have enough experience in C programming [Omit the clarifications for other codes …] Figure 2. The pilot interaction where P3pilot P_3pilot is constructed For the exploration of the initial design of the prompts, we employed OpenAI’s GPT-4o model (Hurst et al., 2024) (interactive version) to conduct the pilot study for the prompts refinement. We chose GPT-4o since this model is not used in the main experiments and can therefore reduce the bias. The output of each turn from the pilot LLM is manually reviewed for insights. Based on these outputs, we iteratively refined the input context and prompt content, incorporating targeted feedback and clarifications to address systematic errors or ambiguities observed in the model’s predictions. Prompts P1 P_1 ~P3+ P_3+ were progressively specified and refined throughout the pilot experiment. In particular, during the specification of P3+ P_3+ the user first asked the pilot LLM to find contrastive examples from the provided samples, then provided clarification addressing those ambiguities. Figure 2 shows an illustrative example of P3pilot P_3pilot, and Figure 3 shows the crafted P3+ P_3+ based on the pilot interaction. The pilot LLM was asked to find contrastive examples. LLMs in the main experiment are not asked to label these contrastive examples to ensure that the corresponding answers are not inadvertently disclosed to the LLMs when providing prompt P3+ P_3+. User (P3+ P_3+) For each column, there are two examples that you might find similar but have opposite labeling. I can clarify the subtle differences for you. Variable/Method identifiers are mentioned - ‘I think it had something to do with inv and debug’ → marked 1 - ‘I did not select anything because […] I am not that experienced with C’ → marked 0 Clarify: Are only clear technical identifiers enough, or should any code-ish word be counted? Clarification: The first sentence uses inv” and “debug” in a way that you can understand they are not part of a natural language sentence, so they must be code identifiers. So the mark is 1. […] [Omit the contrastive examples and the clarifications for other codes …] Figure 3. The defined prompt P3+ P_3+ based on P3pilot P_3pilot 3.5. Formal Analysis Success Metrics For each run, across #L\#L LLMs, #P\#P prompts, #C\#C codes, and #A\#A annotators, we obtain a tuple which summarizes the result of the annotation of the N comments. ⟨ℓ,P,c,a,TPℓPca,FPℓPca,,FNℓPca,FNℓPca⟩ , P,c,a,TP_ Pca,FP_ Pca,,FN_ Pca,FN_ Pca , where: • TPℓPcaTP_ Pca is the number of comments where c=1c=1 by the annotator a and c=1c=1 by the LLM ℓ with prompt P; • FPℓPcaFP_ Pca is the number of comments where c=0c=0 by the annotator a and c=1c=1 by the LLM ℓ with prompt P; • FNℓPcaFN_ Pca is the number of comments where c=1c=1 by the annotator a and c=0c=0 by the LLM ℓ with prompt P; • TNℓPcaTN_ Pca is the number of comments where c=0c=0 by the annotator a and c=0c=0 by the LLM ℓ with prompt P. We define by correctℓ,P,c,a=TPℓ,P,c,a+TNℓ,P,c,acorrect_ , P,c,a=TP_ , P,c,a+TN_ , P,c,a the number of correctly coded comment by the LLM ℓ for code c, we denote by posℓ,P,c,a=TPℓ,P,ca+FNℓ,P,c,apos_ , P,c,a=TP_ , P,ca+FN_ , P,c,a the number of positives according to the human annotator, and by alertℓ,P,c,a=TPℓ,P,c,a+FPℓ,P,c,aalert_ , P,c,a=TP_ , P,c,a+FP_ , P,c,a we denote the number of comments marked as 1 for code c. Cohen’s Kappa κ and Chance Corrected Accuracy As mentioned in the introduction, the first step is to make sure that we do not hold LLMs to higher standards than human annotators. Cohen’s Kappa is a robust indicator of agreement in which the percentage of codes in agreement (TPTP and TNTN) between two coders (in our case, the LLM and one of the annotators) is normalized by the percentage of the agreement that could have been obtained by chance. For the sake of simplicity, we remove in the subsequent equations the indexes describing the LLM, the column, and the annotator. (1) E[Agree] E[Agree] = = alertN⋅posN+(1−alertN)⋅(1−posN) alertN· posN+ (1- alertN )· (1- posN ) (2) Agree(=Acc) Agree(=Acc) = = correctN correctN (3) κ(=Acc∗) κ(=Acc^*) = = Agree−E[Agree]1−E[Agree] Agree-E[Agree]1-E[Agree] The definition of Agreement (2) is actually identical to accuracy if we consider each annotator as the ground truth, and Cohen’s Kappa (3) is actually identical to the definition of chance corrected accuracy. This indicator is quite common in the weather prediction literature (Barnston, 1992): In the presence of an unbalanced dataset, some of the performance metrics might be obtained by pure chance, just because we have too many (or too few) events of interest in the system. To avoid such mistakes, we need to compute the expected values of true positives or true negatives that we would have obtained by chance and subtract them from the values we obtain with the experiment. The chance corrected accuracy is also known in the weather forecasting literature as the Heidke skill index (HSS). Common guidelines (Warrens, 2015) state that an interval of [0.0,0.2][0.0,0.2] indicates slight agreement, (0.2,0.4](0.2,0.4] fair agreements, (0.4,0.6](0.4,0.6] moderate agreement, (0.6,0.8](0.6,0.8] substantial agreement, and (0.8,1.0](0.8,1.0] indicates almost perfect agreement. We can therefore use the same guidelines to compare chance-corrected accuracy against the bands of Cohen’s κ. We also have to account for the possibility that chance-corrected accuracy is negative, namely that the LLM is actually worse than randomly guessing human annotations. We summarize these levels in Table 2 and adapt it to the LLMs. Table 2. Proposed Thresholds for Human Replaceability κ(Acc∗)κ(Acc^*) Replaceability and Rationale <0<0 Hallucinate (✗) - The LLM has worse performance than random guessing in controlled experiments. It is not recommended to use it. [0.0,0.2][0.0,0.2] Unusable (★) - The LLM will rarely generate the expected annotations and will almost always disagree with the human annotator. (0.2,0.4](0.2,0.4] Limited (★), The LLM only partly generates the annotations of a human coder. It might not converge when resolving conflicts with other annotators. (0.4,0.6](0.4,0.6] Moderate (★), The LLM is not able to generate human annotation for around half the time. It might or might not converge when discussing conflicts. (0.6,0.8](0.6,0.8] Substantial (★) - The LLM will generate codes close to a human annotator in practice. Multiple interactions over conflicts are likely to converge. (0.8,1.0](0.8,1.0] Almost Perfect (★) - The LLM would generate almost the same annotations as human coders, or at least will disagree with them only in a few cases. Since we have multiple annotators, when reporting a value, we take the average across annotators. Analysis and Statistical Tests To answer RQ1, we first compute each metric and we report the average of the results across A annotators. Since an LLM might be better suited to some types of codes than others (e.g., better at understanding the presence of program identifiers than security keywords), we also consider the individual coders as a second detailed comparison. Then compare the relative effectiveness of LLM ℓi _i globally, we perform a Wilcoxon paired test as a statistical test to compare the set of kappa’s value κℓ,P,c,a _ , P,c,a for each prompt, considering each code and annotator as pairs. (4) W(⟨κℓ1,P,c,a,κℓ2,P,c,a⟩|P,c,a) for ℓ1 and ℓ2 W ( \ . _ _1, P,c,a~,~ _ _2, P,c,a | P,c,a \ ) for _1 and _2 To answer RQ2, first, we evaluate the effectiveness of the prompts irrespective of the LLMs. We aggregate the results over the codes and annotators for each LLM ℓ and prompt P. We perform the pairwise Wilcoxon test to compare the prompts with increasing level of effort, i.e., comparing prompt Pi−1 P_i-1 with prompt Pi P_i. (5) W(⟨κℓ,Pi−1,c,a,κℓ,Pi,c,a⟩|ℓ,c,a) W ( \ . _ , P_i-1,c,a, _ , P_i,c,a | ,c,a \ ) Since prompts are compared in increasing order (Pi−1 P_i-1 is compared with Pi P_i, which is then compared with Pi+1 P_i+1, etc. they are independent tests, and we do not need to perform a Bonferroni correction. We can refine the analysis by comparing each LLM while keeping the prompt constant, essentially repeating the test in Equation (4) for all prompts. In this setting, we need to perform a Bonferroni correction by considering as statistically significant only those tests with a p-value p≤5%#Lp≤ 5\%\#L where #L\#L is the number of LLMs compared because we are repeating the previous test again now with a different metric each time. 4. Results 4.1. RQ1: LLMs Ability on Automatic Annotation Table 3 reports the summary with the average precision, recall, accuracy, and Cohen’s Kappa (chance-corrected accuracy) for our experiments where we tried to replace a human annotator with an LLM for the best prompt P3+ P_3+ that we have presented in §3.4. Values are averaged across annotators and columns. e.g. κ^ℓ=1#A1#C∑a,cκℓ,P3+,c,a κ_ = 1\#A 1\#C _a,c _ , P_3+,c,a to provide a first global assessment. As we can see from the Table the result is negative. We cannot reliably assume that an LLM can generally replace a human annotator. Table 3. Average Results with the Final Prompt Based on the qualitative levels from Table 2 for Cohen’s κ guidelines (Warrens, 2015), we see that LLMs cannot reliably replace human annotators. We also see the importance of using robust metrics not affected by unbalanced datasets. LLM Precision Recall Acc κ/Acc∗κ/Acc^* Replace GPT-5 0.62 0.47 0.74 0.26 ★ Claude-4 0.57 0.51 0.74 0.30 ★ DeepSeek-V3.2 0.79 0.64 0.83 0.53 ★ Qwen3-Max 0.74 0.73 0.86 0.61 ★ Table 2: Substantial - ★, Moderate -★, Limited - ★. Human annotators are also error-prone and might disagree with each other. Therefore, a too high bar might not correspond to practical use. The next step investigate whether the poor (overall) performance is only due to some specific codes that are also hard for human annotators. Table 4 shows the summary of the results for the chance-corrected accuracy of each code for the final prompt. In this set-up we have given LLMs an easier task as we only consider the comments on which the human annotators agree with the reviewers. So LLMs were given codes where there is 100% agreement with humans. Once again, we only consider our best prompt P3+ P_3+. Table 4. The Cohen’s κ of the human annotator/LLM’s agreement with the reviewers across codes In spite of giving the LLM an easy task (by supplying it only the comments on which there was almost 100% agreement among the humans), it could not successfully annotate most codes. Only on comments that describe the absence of a vulnerability (NoVul) DeepSeek-V3.2 and Qwen3-Max are able to get close to the human agreement. GPT-5 and Claude-4 exhibit very low performance on certain columns, but other models also have column-specific limitations. Human* GPT-5 Claude-4 DeepSeek-V3.2 Qwen3-Max κ κ Agreement κ κ Replace κ κ Replace κ κ Replace κ κ Replace Var 0.98 ★ 0.00 ★ 0.52 ★ 0.79 ★ 0.79 ★ Lin 0.94 ★ 0.57 ★ 0.46 ★ 0.31 ★ 0.42 ★ Key 0.98 ★ 0.30 ★ 0.14 ★ 0.34 ★ 0.56 ★ Vul 0.98 ★ 0.50 ★ 0.44 ★ 0.39 ★ 0.50 ★ Sec 0.98 ★ 0.11 ★ 0.15 ★ 0.50 ★ 0.66 ★ Exp 0.98 ★ 0.12 ★ 0.20 ★ 0.56 ★ 0.64 ★ NoVul 0.98 ★ 0.30 ★ 0.42 ★ 0.82 ★ 0.85 ★ Unsr 0.93 ★ 0.38 ★ 0.16 ★ 0.66 ★ 0.59 ★ Unclr 0.95 ★ 0.00 ★ 0.13 ★ 0.44 ★ 0.41 ★ Total 0.97 ★ 0.25 ★ 0.29 ★ 0.54 ★ 0.60 ★ Table 2: Almost Perfect - ★, Substantial - ★, Moderate -★, Limited - ★, Unusable - ★, Hallucinate - ✗). As we can see, on individual codes, the results are partly consistent across models with some notable exceptions. All models perform very poorly on some codes, such as identifying the unclear answer (Unclr). They have a better success on spotting the negative answer ‘I don’t find vulnerability’ (Vul) and to identify whether there are Variable/Method identifiers mentioned (Var). Negative Answer to RQ1: The tested LLMs are not generally able to match the human annotators except for some specific code and only for some specific models. Surprisingly, open source models perform generally better than proprietary LLMs. LLMs are not yet ready to replace human annotators in security-relevant comments. 4.2. RQ2: Impact of Prompt Efforts Table 5 shows the results per prompt across the four LLMs as we are interested in determining whether there is a significant difference between the prompts, irrespective of the LLMs. Table 5. Cohen’s κ (Chance Corrected Accuracy) per Prompt The only significant difference is obtained by providing the LLMs with the definition of the codes (from P1 P_1 to P2 P_2). The large effort to further improve the prompts does not bring significant gain across different LLMs and also seems to confuse the model. Prompt Precision Recall Acc κ(Acc∗)κ(Acc^*) Replaceability P1 P_1 0.63 0.51 0.78 0.37 ★ P2 P_2 0.67 0.56 0.79 0.42 ★ P3 P_3 0.59 0.54 0.76 0.35 ★ P3+ P_3+ 0.68 0.58 0.79 0.42 ★ Table 2: Substantial - ★, Moderate -★. Once again, we do not see a major difference. After the first prompt, where the chance-corrected accuracy moves from 37% to 42%, the values linger in the area. Refining the prompts does not yield major changes. Table 6 reports the results of Wilcoxon paired tests between increasingly sophisticated prompts. While the comparison between P1 P_1 and P2 P_2 yields a p-value lower than 5%, such a value is not statistically significant after the Bonferroni correction, where we divide by the number of multiple comparisons p>0.05/(P−1)=0.017p>0.05/(P-1)=0.017 where P is the number of compared prompts in order of sophistication. Table 6. Wilcoxon paired tests between increasingly sophisticated prompts Contrast Wilcoxon’s W p-value P1 P_1 P2 P_2 181 0.037 P2 P_2 P3 P_3 292 0.890 P3 P_3 P3+ P_3+ 30 0.004 So we cannot really conclude that spending time refining the prompts yields any better outcome by the LLMs. This finding persists even if separating the results by model, see Table 7. Table 7. Wilcoxon paired tests across LLMs and prompts LLM Contrast Wilcoxon’s W p-value GPT-5 P1 P_1 P2 P_2 3 0.04 P2 P_2 P3 P_3 45 1.0 P3 P_3 P3+ P_3+ 0 0.004 Claude-4 P1 P_1 P2 P_2 9 0.06 P2 P_2 P3 P_3 31 0.85 P3 P_3 P3+ P_3+ 16 0.42 DeepSeek-V3.2 P1 P_1 P2 P_2 12 0.23 P2 P_2 P3 P_3 3 0.63 P3 P_3 P3+ P_3+ 0 0.5 Qwen3-Max P1 P_1 P2 P_2 28 0.75 P2 P_2 P3 P_3 0 0.001 P3 P_3 P3+ P_3+ 3 1.0 These findings are particularly interesting as the different prompts have different budgets in terms of tokens. Figure 4 shows the results in terms of chance-corrected accuracy against the tokens used to obtain those results. Since the models have different ways to count tokens, and it is unclear how the Excel files are precisely counted, we opted for a third-party software (AWS Bedrock), which can be used to query several models to obtain a token count. Figure 5 shows the results for each code. (a) GPT-5 (b) Claude-4 (c) DeepSeek-V3.2 (d) Qwen3-Max Figure 4. Average Cohen’s κ (Chance Corrected Accuracy if human were ground truth) vs Effort (Size of Prompts) (a) GPT-5 (b) Claude-4 (c) DeepSeek-V3.2 (d) Qwen3-Max Figure 5. Chance corrected accuracy against the required effort in tokens, with each curve corresponding to a different code. Negative Answer to RQ2: Prompt refinements with increasing efforts in terms of tokens do not uniformly guarantee a good performance across LLMs for security-specific annotations. There is only a significant gain at the beginning when a formal definition of the codes is provided with the positive examples. Even proprietary models might be significantly confused by moire data. 4.3. Qualitative Analysis LLMs’ elaboration of tabular data. Several LLMs initially considered were excluded due to technical limitations encountered during experimentation. We first evaluated models via the Amazon Bedrock (Bhattacharjee, 2025) platform. Although Claude 4 Sonnet, LLaMA, and DeepSeek could accept spreadsheet inputs, they were limited to textual outputs and failed to reliably produce tabular data with the expected number of rows, even when explicitly prompted to do so. Additional model-specific constraints further reduced feasibility: Mistral did not support document inputs on the tested endpoint, while models in the LLaMA family repeatedly exceeded the document ingestion time limit. Consequently, we abandoned the Bedrock-based setup and re-ran the experiments using each model’s publicly available chatbot interface. Grok was unavailable in Italy at the time of the experiment and was therefore excluded. Among the remaining models, Kimi K2 Thinking failed to consistently generate valid tabular outputs, producing either an incorrect number of entries or malformed files. For GPT, DeepSeek, and Qwen, we successfully prompted the generation of spreadsheet files. This approach did not generalize to Claude, which instead produced a non-functional interactive web application. However, requesting tabular data as structured lists yielded outputs with the correct number of entries, allowing their inclusion in subsequent analyses. Case studies To investigate the sources of classification errors, we qualitatively analyzed a subset of LLM-generated annotations. This analysis revealed recurring patterns underlying the models’ limited performance. For example, in prompt 3, ChatGPT-5 systematically labeled the Variable/Method identifiers are mentioned feature as present for all annotators, regardless of comment content. A common challenge across models was distinguishing general technical terminology from code-specific identifiers. Some models appeared overly sensitive to keywords (e.g., GPT labeling all occurrences of “vulnerable” as positive), while others were misled by generic technical terms such as “debug” (e.g., Qwen). In general, it seems models can get confused by the literal meaning of words and miss the high-level significance, for example ”By elimination it seems that there is no other lines that can be assessed as vulnerables” for human annotators was just positive for I don’t find vulnerability, but DeepSeek also marked it as positive for Text is related to the specific vulnerability. Models also varied in their interpretation of uncertainty; for instance, GPT and Claude occasionally inferred uncertainty from modal expressions that human annotators did not consider indicative of uncertainty (”I would be inclined to think” or ”it is likely that”). Finally, certain misclassifications remained difficult to explain, for example, DeepSeek marked ”I am not very sure it is a vulnerable line” as positive for Relevant Keywords mentioned, but ”I don’t think any of the lines are vulnerable” as negative, even if they include almost all the same words. 5. Discussion Our exploration revealed both the potential and the limits of using LLMs for security-specific annotations of subjects’ justifications for security choices in experiments where humans are asked to identify security issues in code snippets. While the model demonstrated the ability to parse structured input and apply simple binary classification tasks, its performance varied widely across codes, especially those involving nuanced reasoning or implicit context. Most importantly, it is not possible to achieve even a moderate accuracy of the given human annotators. Therefore, the use of LLMs as an automatic annotator is still not recommended for security-relevant annotations. Prompt engineering showed diminishing returns. Adding detailed explanations and examples improved results marginally, but not to a degree that would significantly offset the manual effort required to craft such prompts. Despite increased prompt complexity and refinement, performance improvements were limited. These outcomes indicate that the model’s initial limitations in understanding security-specific annotation tasks were not easily overcome by progressively more elaborate prompting alone. These findings suggest the following key insights: • (some) LLMs can assist with straightforward code detection (e.g., mentions of line numbers or identifiers), but struggle with interpretive codes like confusion. • Prompt quality matters, but beyond a point, iterative refinement without architectural adaptation (e.g., fine-tuning) may yield diminishing returns. • Interactive prompting strategies may not be sufficient to bridge the semantic gap between surface-level understanding and security-informed reasoning. 6. Related Work Opinion mining for Software development Recent research has increasingly applied opinion mining and sentiment analysis to understand developers’ perspectives across various software engineering (SE) tasks (Lin et al., 2022). Various Natural Language Processing (NLP) and Machine Learning (ML) techniques have been applied to analyze qualitative text in software engineering contexts. Sentiment polarity detection—using lexicon-based tools like SentiStrength (Thelwall et al., 2010), ML models like Senti4SD (Calefato et al., 2018), and Deep Learning classifiers like SentiCR (Ahmed et al., 2017)—has been used to classify developer emotions in issue reports, Stack Overflow posts, and code reviews. Topic modeling methods such as LDA (Blei et al., 2003) and TwitterLDA (Zhao et al., 2011) have supported the discovery of latent concerns in developer discussions and user reviews. Supervised classifiers have been employed to extract structured insights from unstructured text, such as identifying bug reports, code requests, or usability issues in app reviews (e.g., MARC 3.0 (Jha and Mahmoud, 2018, 2019), Ticket-Tagger (Kallis et al., 2019)). More recently, hybrid techniques combining rule-based and statistical models have been used for emotion detection (e.g., DEVA (Islam and Zibran, 2018), EmoTxT (Calefato et al., 2017)), and trust inference in developer collaboration (da Cruz et al., 2016). Despite these advances, many studies report performance limitations when applying tools trained on general-domain data to software-specific texts, underscoring the need for domain-adapted approaches and annotated datasets curated for SE tasks (Jongeling et al., 2017). Qualitative data analysis in SE experiments. Qualitative data analysis plays a vital role in SE experiments by uncovering insights from textual sources such as interviews, open-ended surveys, and observational data (Wohlin et al., 2012). The goal is to generate credible findings while maintaining a transparent chain of evidence, linking conclusions clearly to the original data (Runeson and Höst, 2009). The process involves coding segments of text to capture recurring ideas and developing themes or hypotheses through techniques like constant comparison or cross-case analysis (Seaman, 1999). These are followed by hypothesis confirmation strategies, such as triangulation or replication. To reduce bias, multiple researchers often code independently and reconcile results collaboratively. These methods have supported research into developer rationale in code review comments (Mäntylä and Lassenius, 2008), selection of patches for vulnerability repair (Papotti et al., 2024), vulnerability identification with program slicing (Papotti et al., 2025), barriers to adopting automated tools (Johnson et al., 2016), and developer sentiments in socio-technical systems (Ford et al., 2016). Such studies highlight the importance of rigorous qualitative analysis for exploring the human and collaborative aspects of software engineering. LLMs for qualitative data analysis. With the development of large language models, their potential to assist humans in qualitative research has been discussed increasingly. Several studies have explored the application of LLMs in qualitative analysis across different domains. For example, in healthcare, GPT-4 demonstrates moderate agreement with human qualitative analysis (Li et al., 2024); in psychology, LLMs have been tested for deductive coding tasks in analyzing children’s curiosity-driven questions, showing moderate to high agreement with experts on certain metrics (Xiao et al., 2023); in processing, LLMs have also been applied for historical literature annotation (Dunivin, 2024), music shuffle preference annotation and password management annotation (Dai et al., 2023), and show the potential ability. In software engineering, the usage of LLMs for qualitative data analysis is a relatively new topic. (Rasheed et al., 2024) proposes a multi-agent strategy to study the LLMs’ ability in different qualitative research tasks, involving thematic annotation for GitHub or Stack Overflow discussions. While showing potential in qualitative research, some limitations of LLMs have also raised concerns, such as model biases, illusions, and ethical problems (Schroeder et al., 2025). However, LLMs’ ability to analyze technical comments and to annotate different types of themes remains to be explored. 7. Threats to Validity Selected task might not represent all annotation scenarios Our evaluation focuses on a single type of security experiment and a fixed set of nine security-relevant codes. The comments dataset is relatively representative of vulnerability assessment comments, coming from CS master’s students. However, the tasks and code can be different in the scope of security-specific annotation. For instance, the comments might cause bias since the students’ comments can vary a lot from those of the developers in the real-world. Moreover, in this paper, the task is for LLMs strictly as full replacements for human annotators, while the ability of LLMs to suggest candidate code, pre-filter, or detect disagreement remains to be defined. Selected LLMs might not represent future performance The selection of large language models was guided by the widely used benchmark LiveBench. All LLMs were utilized in their full versions under subscription plans. However, the models were chosen at a specific point in time and may not reflect future developments. The impact of prompt engineering Our prompt design aimed to steer LLM outputs toward reliable annotations by embedding clear task descriptions, structured guidelines, and representative examples, thus leveraging in‑context learning and few‑shot prompting to reduce ambiguity. Additionally, we iteratively refined prompts,P3+ P_3+ in particular, through pilot interactions with GPT‑4o, adjusting phrasing and examples based on model behavior to mitigate obvious misinterpretations. These practices align with common prompt engineering techniques such as template structuring and exemplar selection identified in recent literature on prompting methodologies (Marvin et al., 2023) and systematic taxonomies of prompting strategies (Schulhoff et al., 2024). Nonetheless, prompt engineering lacks well‑established, universally optimal procedures, and the effectiveness of particular formulations can vary substantially across tasks and models. Our design choices, therefore, reflect subjective decisions and limited expertise, and alternative prompt structures could lead to different model behaviors, limiting generalizability and posing a threat to validity. 8. Conclusion Recent research has shown that large language models (LLMs) can replicate human annotators for extracting sentiment analysis from text. In this paper, we have investigated whether code capturing domain-specific aspects in security and software engineering, such as code identifiers mentioned, lines-of-code-mentioned, security keywords mentioned, can be achieved by LLMs to reduce the manual effort involved in the qualitative analysis of technical comments by acting as automated annotators. We have prompted the four best proprietary and open source LLMs on LiveBench (GPT-5, Claude-4, DeepSeek-V3.2, and Qwen3-Max) to identify the presence of 9 security-relevant codes in free-text comments from humans analyzing code snippets for vulnerabilities. The LLMs’ outputs were compared against a ground truth annotated by expert coders using precision, recall, and the Heidke Skill Score (a chance-corrected accuracy measure). We refined the prompts by mimicking the process of human annotators: emerging codes, a codebook with examples, and conflicting examples. We observed small improvements after providing detailed descriptions and examples for each code, but we did not see a statistically significant gain across codes or models. Most importantly, the overall performance is too low to replace a human annotator. Given these findings, we argue that the role of LLMs in qualitative security coding should, at present, remain assistive rather than autonomous. More experiments on larger corpora of experiments are needed to go beyond generating a first draft of annotations, highlighting potential codes for human verification, or accelerating low-complexity tagging tasks. These results motivate future work on structured prompting interfaces, model calibration, and the design of hybrid pipelines where LLMs complement rather than replace human insight in technically demanding qualitative analyses. Acknowledgments This work was partly funded by the EU under the Horizon Europe Program with n. 101120393 (Sec4AI4Sec), by the Italian Ministry of University and Research (MUR) under the P.N.R.R. – NextGenerationEU grant n. PE00000014 (SERICS subproject, CUP E63C24000590001). CRediT Conceptualization: MC, YG, FM; Methodology: MC, YG, FM; Software Programming: MC; Validation: YG, FM; Formal analysis: FM; Investigation: MC, YG; Resources: FM; Data Curation: MC; Writing - Original Draft: MC, YG; Writing - Review & Editing: Visualization: MC , YG; Supervision: YG, FM; Project administration: FM; Funding acquisition: FM. References (1) Ahmed et al. (2017) Toufique Ahmed, Amiangshu Bosu, Anindya Iqbal, and Shahram Rahimi. 2017. SentiCR: A customized sentiment analysis tool for code review interactions. In 2017 32nd IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 106–111. Anthropic (2025) Anthropic. 2025. Claude 4: Claude Opus 4 & Claude Sonnet 4 Large Language Models. System card / technical documentation. https://w.anthropic.com/news/claude-4 Includes Opus 4 and Sonnet 4, released May 22, 2025. Armborst (2017) Andreas Armborst. 2017. Thematic proximity in content analysis. Sage Open 7, 2 (2017), 2158244017707797. Barnston (1992) Anthony G Barnston. 1992. Correspondence among the correlation, RMSE, and Heidke forecast verification measures; refinement of the Heidke score. Weather and Forecasting 7, 4 (1992), 699–709. Bhattacharjee (2025) Avik Bhattacharjee. 2025. Introduction to Amazon Bedrock. In A Practical Guide to Generative AI Using Amazon Bedrock: Building, Deploying, and Securing Generative AI Applications. Springer, 79–112. Blei et al. (2003) David M Blei, Andrew Y Ng, and Michael I Jordan. 2003. Latent dirichlet allocation. Journal of machine Learning research 3, Jan (2003), 993–1022. Calefato et al. (2018) Fabio Calefato, Filippo Lanubile, Federico Maiorano, and Nicole Novielli. 2018. Sentiment polarity detection for software development. In Proceedings of the 40th International Conference on Software Engineering. 128–128. Calefato et al. (2017) Fabio Calefato, Filippo Lanubile, and Nicole Novielli. 2017. Emotxt: a toolkit for emotion recognition from text. In 2017 seventh international conference on Affective Computing and Intelligent Interaction Workshops and Demos (ACIIW). IEEE, 79–80. Clarke and Braun (2017) Victoria Clarke and Virginia Braun. 2017. Thematic analysis. The journal of positive psychology 12, 3 (2017), 297–298. da Cruz et al. (2016) Guilherme A Maldonado da Cruz, Elisa Hatsue Moriya Huzita, and Valéria D Feltrim. 2016. Estimating Trust in Virtual Teams-A Framework based on Sentiment Analysis. In International Conference on Enterprise Information Systems, Vol. 2. SciTePress, 464–471. Dai et al. (2023) Shih-Chieh Dai, Aiping Xiong, and Lun-Wei Ku. 2023. LLM-in-the-loop: Leveraging large language model for thematic analysis. arXiv preprint arXiv:2310.15100 (2023). DeepSeek-AI et al. (2025) DeepSeek-AI, Aixin Liu, Aoxue Mei, Bangcai Lin, Bing Xue, Bingxuan Wang, Bingzheng Xu, Bochao Wu, Bowei Zhang, Chaofan Lin, Chen Dong, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenhao Xu, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Erhang Li, Fangqi Zhou, Fangyun Lin, Fucong Dai, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Hanwei Xu, Hao Li, Haofen Liang, Haoran Wei, Haowei Zhang, Haowen Luo, Haozhe Ji, Honghui Ding, Hongxuan Tang, Huanqi Cao, Huazuo Gao, Hui Qu, Hui Zeng, Jialiang Huang, Jiashi Li, Jiaxin Xu, Jiewen Hu, Jingchang Chen, Jingting Xiang, Jingyang Yuan, Jingyuan Cheng, Jinhua Zhu, Jun Ran, Junguang Jiang, Junjie Qiu, Junlong Li, Junxiao Song, Kai Dong, Kaige Gao, Kang Guan, Kexin Huang, Kexing Zhou, Kezhao Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Wang, Liang Zhao, Liangsheng Yin, Lihua Guo, Lingxiao Luo, Linwang Ma, Litong Wang, Liyue Zhang, M. S. Di, M. Y. Xu, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingxu Zhou, Panpan Huang, Peixin Cong, Peiyi Wang, Qiancheng Wang, Qihao Zhu, Qingyang Li, Qinyu Chen, Qiushi Du, Ruiling Xu, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runqiu Yin, Runxin Xu, Ruomeng Shen, Ruoyu Zhang, S. Wu, Z. Z. Ren, Zehua Zhao, Zehui Ren, Zhangli Sha, Zhe Fu, Zhean Xu, Zhenda Xie, Zhengyan Zhang, Zhewen Hao, Zhibin Gou, Zhicheng Ma, Zhigang Yan, Zhihong Shao, Zhixian Huang, Zhiyu Wu, Zhuoshu Li, Zhuping Zhang, Zian Xu, Zihao Wang, Zihui Gu, Zijia Zhu, Zilin Li, Zipeng Zhang, Ziwei Xie, Ziyi Gao, Zizheng Pan, Zongqing Yao, Bei Feng, Hui Li, J. L. Cai, Jiaqi Ni, Lei Xu, Meng Li, Ning Tian, R. J. Chen, R. L. Jin, S. S. Li, Shuang Zhou, Tianyu Sun, X. Q. Li, Xiangyue Jin, Xiaojin Shen, Xiaosha Chen, Xinnan Song, Xinyi Zhou, Y. X. Zhu, Yanping Huang, Yaohui Li, Yi Zheng, Yuchen Zhu, Yunxian Ma, Zhen Huang, Zhipeng Xu, Zhongyu Zhang, Dongjie Ji, Jian Liang, Jianzhong Guo, Jin Chen, Leyi Xia, Miaojun Wang, Mingming Li, Peng Zhang, Ruyi Chen, Shangmian Sun, Shaoqing Wu, Shengfeng Ye, T. Wang, W. L. Xiao, Wei An, Xianzu Wang, Xiaowen Sun, Xiaoxiang Wang, Ying Tang, Yukun Zha, Zekai Zhang, Zhe Ju, Zhen Zhang, and Zihua Qu. 2025. DeepSeek-V3.2: Pushing the Frontier of Open Large Language Models. arXiv preprint arXiv:2512.02556 (2025). https://arxiv.org/abs/2512.02556 Dunivin (2024) Zackary Okun Dunivin. 2024. Scalable qualitative coding with llms: Chain-of-thought reasoning matches human performance in some hermeneutic tasks. arXiv preprint arXiv:2401.15170 (2024). Ford et al. (2016) Denae Ford, Justin Smith, Philip J Guo, and Chris Parnin. 2016. Paradise unplugged: Identifying barriers for female participation on stack overflow. In Proceedings of the 2016 24th ACM SIGSOFT International symposium on foundations of software engineering. 846–857. Gregory et al. (2015) Robert Wayne Gregory, Mark Keil, Jan Muntermann, and Magnus Mähring. 2015. Paradoxes and the nature of ambidexterity in IT transformation programs. Information Systems Research 26, 1 (2015), 57–80. Guest et al. (2011) Greg Guest, Kathleen M MacQueen, and Emily E Namey. 2011. Applied thematic analysis. sage publications. Guo et al. (2018) Wenbo Guo, Dongliang Mu, Jun Xu, Purui Su, G. Wang, and Xinyu Xing. 2018. LEMNA: Explaining Deep Learning based Security Applications. Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security (2018). Hurst et al. (2024) Aaron Hurst, Adam Lerer, Adam P Goucher, Adam Perelman, Aditya Ramesh, Aidan Clark, AJ Ostrow, Akila Welihinda, Alan Hayes, Alec Radford, et al. 2024. Gpt-4o system card. arXiv preprint arXiv:2410.21276 (2024). Islam and Zibran (2018) Md Rakibul Islam and Minhaz F Zibran. 2018. DEVA: sensing emotions in the valence arousal space in software engineering text. In Proceedings of the 33rd annual ACM symposium on applied computing. 1536–1543. Jha and Mahmoud (2018) Nishant Jha and Anas Mahmoud. 2018. Using frame semantics for classifying and summarizing application store reviews. Empirical Software Engineering 23, 6 (2018), 3734–3767. Jha and Mahmoud (2019) Nishant Jha and Anas Mahmoud. 2019. Mining non-functional requirements from app store reviews. Empirical Software Engineering 24 (2019), 3659–3695. Jiang et al. (2017) Jing Jiang, David Lo, Jiahuan He, Xin Xia, Pavneet Singh Kochhar, and Li Zhang. 2017. Why and how developers fork what from whom in GitHub. Empirical Software Engineering 22, 1 (2017), 547–578. Johnson et al. (2016) Brittany Johnson, Rahul Pandita, Justin Smith, Denae Ford, Sarah Elder, Emerson Murphy-Hill, Sarah Heckman, and Caitlin Sadowski. 2016. A cross-tool communication study on program analysis tool notifications. In Proceedings of the 2016 24th ACM SIGSOFT International Symposium on Foundations of Software Engineering. 73–84. Jongeling et al. (2017) Robbert Jongeling, Proshanta Sarkar, Subhajit Datta, and Alexander Serebrenik. 2017. On negative results when using sentiment analysis tools for software engineering research. Empirical Software Engineering 22 (2017), 2543–2584. Kallis et al. (2019) Rafael Kallis, Andrea Di Sorbo, Gerardo Canfora, and Sebastiano Panichella. 2019. Ticket tagger: Machine learning driven issue classification. In 2019 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 406–409. Li et al. (2024) Kevin Danis Li, Adrian M Fernandez, Rachel Schwartz, Natalie Rios, Marvin Nathaniel Carlisle, Gregory M Amend, Hiren V Patel, and Benjamin N Breyer. 2024. Comparing GPT-4 and human researchers in health care data analysis: qualitative description study. Journal of Medical Internet Research 26 (2024), e56500. Li et al. (2018) Zhen Li, Deqing Zou, Shouhuai Xu, Xinyu Ou, Hai Jin, Sujuan Wang, Zhijun Deng, and Yuyi Zhong. 2018. Vuldeepecker: A deep learning-based system for vulnerability detection. arXiv preprint arXiv:1801.01681 (2018). Lin et al. (2022) Bin Lin, Nathan Cassee, Alexander Serebrenik, Gabriele Bavota, Nicole Novielli, and Michele Lanza. 2022. Opinion mining for software development: a systematic literature review. ACM Transactions on Software Engineering and Methodology (TOSEM) 31, 3 (2022), 1–41. Mæhlum et al. (2024) Petter Mæhlum, David Samuel, Rebecka Maria Norman, Elma Jelin, Øyvind Andresen Bjertnæs, Lilja Øvrelid, and Erik Velldal. 2024. It’s Difficult to be Neutral–Human and LLM-based Sentiment Annotation of Patient Comments. arXiv preprint arXiv:2404.18832 (2024). Mäntylä and Lassenius (2008) Mika V Mäntylä and Casper Lassenius. 2008. What types of defects are really discovered in code reviews? IEEE Transactions on Software Engineering 35, 3 (2008), 430–448. Marvin et al. (2023) Ggaliwango Marvin, Nakayiza Hellen, Daudi Jjingo, and Joyce Nakatumba-Nabende. 2023. Prompt engineering in large language models. In International conference on data intelligence and cognitive informatics. Springer, 387–402. Meta AI (2024) Meta AI. 2024. The LLaMA 3 Herd of Models. arXiv preprint (2024). https://arxiv.org Most recent LLaMA model family at time of writing. Mistral AI (2024) Mistral AI. 2024. Mixtral of Experts. arXiv preprint (2024). https://arxiv.org Sparse mixture-of-experts language model. OpenAI (2025a) OpenAI. 2025a. GPT-5: Technical Overview and Capabilities. Model documentation and release notes. https://openai.com Large multimodal foundation model released by OpenAI. OpenAI (2025b) OpenAI. 2025b. OpenAI o3: Reasoning-Focused Large Language Model. Technical documentation and model card. https://openai.com Reasoning-optimized model in the OpenAI o-series. Papotti et al. (2024) Aurora Papotti, Ranindya Paramitha, and Fabio Massacci. 2024. On the acceptance by code reviewers of candidate security patches suggested by Automated Program Repair tools. Empirical Software Engineering 29, 5 (2024), 1–35. Papotti et al. (2025) Aurora Papotti, Katja Tuma, and Fabio Massacci. 2025. On the effects of program slicing for vulnerability detection during code inspection. Empirical Software Engineering 30, 3 (2025), 1–37. Rahman et al. (2017) Mohammad Masudur Rahman, Chanchal K Roy, and Raula G Kula. 2017. Predicting usefulness of code review comments using textual features and developer experience. In 2017 IEEE/ACM 14th International Conference on Mining Software Repositories (MSR). IEEE, 215–226. Rasheed et al. (2024) Zeeshan Rasheed, Muhammad Waseem, Aakash Ahmad, Kai-Kristian Kemell, Wang Xiaofeng, Anh Nguyen Duc, and Pekka Abrahamsson. 2024. Can large language models serve as data analysts? a multi-agent assisted approach for qualitative data analysis. arXiv preprint arXiv:2402.01386 (2024). Runeson and Höst (2009) Per Runeson and Martin Höst. 2009. Guidelines for conducting and reporting case study research in software engineering. Empirical software engineering 14 (2009), 131–164. Sabetta et al. (2024) Antonino Sabetta, Serena Elisa Ponta, Rocio Cabrera Lozoya, Michele Bezzi, Tommaso Sacchetti, Matteo Greco, Gergő Balogh, Péter Hegedűs, Rudolf Ferenc, Ranindya Paramitha, et al. 2024. Known vulnerabilities of open source projects: Where are the fixes? IEEE Security & Privacy 22, 2 (2024), 49–59. Schroeder et al. (2025) Hope Schroeder, Marianne Aubin Le Quéré, Casey Randazzo, David Mimno, and Sarita Schoenebeck. 2025. Large Language Models in Qualitative Research: Uses, Tensions, and Intentions. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems. 1–17. Schulhoff et al. (2024) Sander Schulhoff, Michael Ilie, Nishant Balepur, Konstantine Kahadze, Amanda Liu, Chenglei Si, Yinheng Li, Aayush Gupta, HyoJung Han, Sevien Schulhoff, et al. 2024. The prompt report: a systematic survey of prompt engineering techniques. arXiv preprint arXiv:2406.06608 (2024). Seaman (1999) Carolyn B. Seaman. 1999. Qualitative methods in empirical studies of software engineering. IEEE Transactions on software engineering 25, 4 (1999), 557–572. Team et al. (2025) Kimi Team, Yifan Bai, Yiping Bao, Guanduo Chen, Jiahao Chen, Ningxin Chen, Ruijue Chen, Yanru Chen, Yuankun Chen, Yutian Chen, et al. 2025. Kimi k2: Open agentic intelligence. arXiv preprint arXiv:2507.20534 (2025). Thelwall et al. (2010) Mike Thelwall, Kevan Buckley, Georgios Paltoglou, Di Cai, and Arvid Kappas. 2010. Sentiment strength detection in short informal text. Journal of the American society for information science and technology 61, 12 (2010), 2544–2558. Warrens (2015) Matthijs J Warrens. 2015. Five ways to look at Cohen’s kappa. Journal of Psychology & Psychotherapy 5 (2015). White et al. (2024) Colin White, Samuel Dooley, Manley Roberts, Arka Pal, Ben Feuer, Siddhartha Jain, Ravid Shwartz-Ziv, Neel Jain, Khalid Saifullah, Sreemanti Dey, et al. 2024. LiveBench: A challenging, contamination-limited LLM benchmark. arXiv preprint arXiv:2406.19314 (2024). Wohlin et al. (2012) Claes Wohlin, Per Runeson, Martin Höst, Magnus C Ohlsson, Björn Regnell, Anders Wesslén, et al. 2012. Experimentation in software engineering. Vol. 236. Springer. xAI (2024) xAI. 2024. Grok: A Large Language Model by xAI. Model documentation and release announcement. https://x.ai Latest publicly released Grok model family at time of writing. Xiao et al. (2023) Ziang Xiao, Xingdi Yuan, Q Vera Liao, Rania Abdelghani, and Pierre-Yves Oudeyer. 2023. Supporting qualitative analysis with large language models: Combining codebook with GPT-3 for deductive coding. In Companion proceedings of the 28th international conference on intelligent user interfaces. 75–78. Yang et al. (2025) An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388 (2025). Zeiler and Fergus (2014) Matthew D Zeiler and Rob Fergus. 2014. Visualizing and understanding convolutional networks. In Computer Vision–ECCV 2014: 13th European Conference, Zurich, Switzerland, September 6-12, 2014, Proceedings, Part I 13. Springer, 818–833. Zhang and Wildemuth (2009) Yan Zhang and Barbara M Wildemuth. 2009. Qualitative analysis of content. Applications of social research methods to questions in information and library science 308, 319 (2009), 1–12. Zhao et al. (2011) Wayne Xin Zhao, Jing Jiang, Jianshu Weng, Jing He, Ee-Peng Lim, Hongfei Yan, and Xiaoming Li. 2011. Comparing twitter and traditional media using topic models. In Advances in Information Retrieval: 33rd European Conference on IR Research, ECIR 2011, Dublin, Ireland, April 18-21, 2011. Proceedings 33. Springer, 338–349.