Paper deep dive
Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Adoption in Cybersecurity Operations on Reddit
Souradip Nath, Chih-Yi Huang, Aditi Ganapathi, Kashyap Thimmaraju, Jaron Mink, Gail-Joon Ahn
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 94%
Last extracted: 4/14/2026, 1:59:22 AM
Summary
This paper presents a mixed-methods analysis of 892 Reddit posts from cybersecurity-focused forums to understand how security practitioners use, perceive, and adopt Large Language Models (LLMs) in Security Operations Center (SOC) workflows. The study finds that while general-purpose LLMs like ChatGPT are dominant, practitioners are cautious, favoring low-risk productivity tasks over autonomous high-stakes decision-making due to concerns regarding reliability, verification overhead, and security risks.
Entities (5)
Relation Signals (3)
Security Practitioners â discusson â Reddit
confidence 100% · we analyzed 892 posts... from three cybersecurity-focused forums on Reddit
Large Language Models â augment â Security Operations Center
confidence 95% · LLMs have recently emerged as promising tools for augmenting Security Operations Center (SOC) workflows
Security Practitioners â use â Large Language Models
confidence 90% · practitioners report meaningful gains in efficiency and effectiveness in LLM-assisted workflows
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Large language models (LLMs) have recently emerged as promising tools for augmenting Security Operations Center (SOC) workflows, with vendors increasingly marketing autonomous AI solutions for SOCs. However, there remains a limited empirical understanding of how such tools are used, perceived, and adopted by real-world security practitioners. To address this gap, we conduct a mixed-methods analysis of discussions in cybersecurity-focused forums to learn how a diverse group of practitioners use and perceive modern LLM tools for security operations. More specifically, we analyzed 892 posts between December 2022 and September 2025 from three cybersecurity-focused forums on Reddit, and, using a combination of qualitative coding and statistical analysis, examined how security practitioners discuss LLM tools across three dimensions: (1) their stated tools and use cases, (2) the perceived pros and cons of each tool across a set of critical factors, and (3) their adoption of such tools and the expected impacts on the cybersecurity industry and individual analysts. Overall, our findings reveal nuanced patterns in LLM tools adoption, highlighting independent use of LLMs for low-risk, productivity-oriented tasks, alongside active interest around enterprise-grade, security-focused LLM platforms. Although practitioners report meaningful gains in efficiency and effectiveness in LLM-assisted workflows, persistent issues with reliability, verification overheads, and security risks sharply constrain the autonomy granted to LLM tools. Based on these results, we also provide recommendations for developing and adopting LLM tools to ensure the security of organizations and the safety of cybersecurity practitioners.
Tags
Links
- Source: https://arxiv.org/abs/2604.09998v1
- Canonical: https://arxiv.org/abs/2604.09998v1
Trouble viewing inline? Open PDF directly â
Full Text
105,609 characters extracted from source content.
Expand or collapse full text
Like a Hammer, It Can Build, It Can Break: Large Language Model Uses, Perceptions, and Adoption in Cybersecurity Operations on Reddit Souradip Nath 1 , Chih-Yi Huang 1 , Aditi Ganapathi 1 , Kashyap Thimmaraju 2 , Jaron Mink 1 , Gail-Joon Ahn 1 1 Arizona State University 2 Technische UniversitĂ€t Berlin Abstract Large language models (LLMs) have recently emerged as promising tools for augmenting Security Operations Center (SOC) workflows, with vendors increasingly marketing au- tonomous AI solutions for SOCs. However, there remains a limited empirical understanding of how such tools are used, perceived, and adopted by real-world security practitioners. To address this gap, we conduct a mixed-methods analysis of discussions in cybersecurity-focused forums to learn how a diverse group of practitioners use and perceive modern LLM tools for security operations. More specifically, we analyzed 892 posts between December 2022 and September 2025 from three cybersecurity-focused forums on Reddit, and, using a combination of qualitative coding and statistical analysis, ex- amined how security practitioners discuss LLM tools across three dimensions: (1) their stated tools and use cases, (2) the perceived pros and cons of each tool across a set of critical factors, and (3) their adoption of such tools and the expected impacts on the cybersecurity industry and individual analysts. Overall, our findings reveal nuanced patterns in LLM tools adoption, highlighting independent use of LLMs for low-risk, productivity-oriented tasks, alongside active interest around enterprise-grade, security-focused LLM platforms. Although practitioners report meaningful gains in efficiency and effec- tiveness in LLM-assisted workflows, persistent issues with reliability, verification overheads, and security risks sharply constrain the autonomy granted to LLM tools. Based on these results, we also provide recommendations for developing and adopting LLM tools to ensure the security of organizations and the safety of cybersecurity practitioners. 1 Introduction SOCs play a critical role in protecting organizations against an increasingly complex and fast-moving threat landscape. They combine people, processes, and technology to support a wide range of defensive functions, including continuous monitoring, real-time detection, alert triage, and incident re- sponse [1â3]. While many SOC processes have become in- creasingly automated through tools like SIEM, SOAR [3], and traditional machine-learning (ML)âbased techniques [4â8], evidence from both industry and academia indicates that SOC teams remain under significant strain. Industry reports show that SOCs receive thousands of alerts daily, with persistent false positives contributing to severe alert fatigue, stress, and burnout among analysts [9, 10]. Academic studies echo these findings, highlighting the limits of existing automation and the continued need for more effective decision support in SOC workflows [1, 11, 12]. Against this backdrop, LLMsâ generative reasoning capabil- ities have recently emerged as promising tools for augmenting SOC workflows [13â15]. Over the past years, research efforts have increasingly explored the application of LLM-powered automation to support a range of SOC tasks, from routine analysis to more advanced investigative workflows [16â20]. Capitalizing on this momentum, cybersecurity vendors have begun introducing autonomous solutions, such as Microsoft Copilot for Security [21] and CrowdStrikeâs Charlotte AI [22], to augment alert triage and incident response processes. While these advances consistently highlight the transforma- tive potential of LLM tools within SOCs [23], understanding how these tools are used in practice and perceived by practi- tioners remains critical for assessing their real-world impact and for guiding the effective integration of LLMs into SOCs. Prior to the rise of LLMs, researchers examined the percep- tions and challenges of security practitioners regarding tra- ditional MLâbased security tools [24, 25]. More recent work has shifted toward task-focused investigations of practition- ersâ interactions with LLM tools, examining issues such as explainability, autonomy, trust, and humanâAI collaboration within specific SOC contexts [13, 14, 26, 27]. While these studies provide valuable insights into particular tools, tasks, or LLM capabilities, a broader understanding is still needed regarding how LLM tools are used across diverse SOC use cases, what motivating factors and barriers shape their adop- tion, and the expected impacts on the cybersecurity industry and individual analysts. 1 arXiv:2604.09998v1 [cs.CR] 11 Apr 2026 To address this gap, we turn to large-scale discourse in online cybersecurity forums. Reddit has been widely used as a rich source of community discussion in secu- rity and privacy research [28â31], hosting several large and active cybersecurity-focused communities, with the largest r/cybersecurityhaving over one million members as of January 2026. In light of this, to capture natural, practitioner- driven discussions at scale around the uses, perceptions, and adoption of LLM tools in security operations, we conducted a mixed-method analysis of 892 posts drawn from three active cybersecurity-focused forums on Reddit between December 2022 and September 2025 (inclusive), with the majority orig- inating fromr/cybersecurity. Based on this analysis, we seek to answer the following research questions: RQ1What LLM tools and use cases are mentioned in security practitionersâ discussions of SOC work? RQ2 What benefits and drawbacks of LLM tools do security practitioners discuss in the context of SOC workflows? RQ3How do security practitioners discuss LLM adoption in SOCs and its potential implications for future practice? Our findings reveal key patterns in how security practition- ers discuss uses, perceptions, and adoption of LLM tools in cybersecurity operations: First, although general-purpose LLMs are not explicitly designed for security operations, they dominate tool-related discussions. Tools such as ChatGPT and Microsoft Copi- lot are referenced far more frequently than security-specific commercial tools, such as Security Copilot, Dropzone, or In- tezer, that are designed to augment security operations. While practitioner discussions reveal a fragmented awareness of the rapidly expanding ecosystem of security-focused LLM plat- forms, discussions of active adoption are significantly more concentrated around a small set of widely adopted general- purpose models. Second, LLM use cases cluster around a range of opera- tional SOC activities, with incident response and triage most frequently discussed, followed by productivity-oriented tasks in scripting and reporting, while knowledge support and threat analysis are less prevalent. Notably, when focusing on active adoption, attention shifts from core incident response toward lower-risk self-scoped activities that afford greater analyst control and easier verification. Third, practitioner discussions reveal clear gradients of au- tonomy: LLMs are most commonly used as decision support tools, less frequently embedded in human-in-the-loop triage pipelines, and only rarely granted fully autonomous control over end-to-end mitigation. Fourth, practitionersâ perceptions of LLM tools are shaped by a balance of benefits and concerns across multiple factors. While many report positive experiences with the capabilities and efficiency of LLM tools in augmenting analyst workflows and reducing workload in specific contexts, these benefits are tempered by strong concerns related to insufficient reliability, security and privacy risks, limited autonomy, and unjustifiable cost. Overall, practitioners view LLM tools as promising yet insufficiently trustworthy to warrant broad delegation of high- stakes SOC tasks without sustained human oversight. Based on our findings, we derive implications for LLM adoption in SOCs as a sociotechnical process, the design of trustworthy systems for security operations, and the long-term sustainability of the SOC workforce. 2 Related Work and Background AI and Cybersecurity Operations. Prior to the emergence of LLMs, research examined the use of traditional ML tech- niques for intrusion detection, anomaly identification, predic- tive analysis, and decision support [4â8]. More recent works have shifted toward applying LLMs to augment SOC work- flows, including LLM-based query generation from natural- language descriptions [17, 18], explanation of alerts and de- tection logic [27], generation of incident summaries [13, 19], creation of cybersecurity exercises [32], and the reduction of analyst cognitive load through natural-language reason- ing [15, 20, 33]. Likewise, recent studies have investigated LLM-powered systems for more autonomous or collabora- tive security operations [23, 34]. These works propose and evaluate agentic LLMs for automated detection, triage, and response, including IRCopilot [16], CORTEX [35], Cyber- SOCEval [36], Locus [37], and FuzzGPT [38]. Human Factors in AI-augmented SOC. Complementing technical, tool-oriented research, a growing body of work has examined the human factors of AI-augmented security operations [1,11,39â42]. Recent works have investigated prac- titionersâ interactions with MLâbased security tools through interview-based studies [24, 25]. Mink et al. [24] conducted a qualitative study examining how practitioners perceive ML- driven detection systems and explanation mechanisms, identi- fying where such tools are considered effective and where they fall short compared to traditional rule-based approaches. With the emergence of LLMs and reasoning models, recent work has shifted toward understanding augmented intelligence and humanâAI collaboration in decision-making contexts within cybersecurity operations [43â45]. Roch et al. [26, 46] con- ducted in-depth interviews with senior cybersecurity leaders to explore humanâAI collaboration, emphasizing factors such as autonomy levels, transparency, and the evolving division of labor between analysts and AI systems. Rastogi et al. [27] ex- amined the role of explainability in AI-driven SOCs through a mixed-methods study, highlighting how AI-generated ex- planations influence trust and actionability in high-stakes en- vironments. Other recent studies have explored task-specific uses of LLMs within SOC, including longitudinal analyses of query-based interactions by SOC analysts [14] and evalua- tions of LLM-assisted incident response summarization using real-world incidents [13]. 2 Our work differs from prior studies in scope and method- ology. Rather than focusing on a specific tool, task, or AI capability, we conduct a large-scale discourse analysis of prac- titioner discussions across online security communities. This approach captures focused yet candid conversations about LLM tools in diverse operational and organizational contexts within security operations, spanning adoption, perceived ben- efits and risks, and future expectations. Terminology. Next, we define the key terms related to data analytics and insights that are used throughout the paper. LLM Tools. In this paper, the term âLLM Toolsâ is used as an umbrella term to refer to a group of standalone tools, compound systems, or commercial platforms in which LLMs play a central role, enabling generative to agentic capabilities, such as reasoning, planning, and learning. Consequently, tra- ditional rule-based automation or ML models fall outside the scope of the tools considered in this work. Practitioners. In this paper, we refer to Reddit posters col- lectively as âsecurity practitionersâ, as the communities stud- ied in this work are cybersecurity-focused. When available, we report more specific, self-identified roles (e.g., SOC man- ager) explicitly stated in a post or indicated by user flairs. Al- though we cannot verify these self-reported roles (§ 3.3), we remain faithful to the data by reporting roles only when they are explicitly self-identified; otherwise, we use the generic term âpractitionerâ to describe participants in the discussions. Reddit Terms. Throughout this paper, we use three Reddit- specific terms: subreddit, thread, and post. Reddit is organized into topic-specific communities called âsubredditsâ. Within each subreddit, users initiate and participate in multiple dis- cussion threads. A âthreadâ refers to a user-initiated discussion consisting of a title, a conversation prompt, and all subsequent replies associated with that submission. Following the termi- nology used by prior works [29, 30], we use the term âpostâ to refer to any individual contribution within a thread, includ- ing both the original submission that starts the thread (often called the âoriginal postâ) and any subsequent replies made in response to it or to other posts within the thread. 3 Methodology To investigate how LLMs are used and perceived in practice, we perform a large-scale qualitative analysis across online cybersecurity forums discussing LLM tools. We detail the eth- ical considerations for our IRB-approved study in Section 3.1. 3.1 Data Collection To accurately capture real-world security practitionersâ uses and perceptions of LLM tools, we gathered and analyzed 1,703 posts made across online cybersecurity forums. Discovering Subreddits. Inspired by prior work analyzing Reddit data [28â31], we selected a set of popular and active SubredditThreads (Pulled) Threads (Relevant) Posts (Pulled) Posts (Relevant) r/cybersecurity83069 (8.31%)1,665863 (51.83%) r/Information_Security965 (5.21%)2115 (71.43%) r/ciso252 (8.00%)1714 (82.35%) r/blueteamsec1870â r/CyberSecurityAdvice760â r/cybersecurity_help360â r/ComputerSecurity350â r/SecurityBlueTeam150â r/SIEM100â Total1,31076 (5.80%)1,703892 (52.38%) Table 1: Collected Reddit Dataset â number of threads retrieved per subreddit, number and share of relevant threads, and the number (and share) of posts/comments coded as relevant. subreddits that enable discussions among security practition- ers. To discover security-centric discussions on emerging LLM tools, we began with informal keyword-based searches on Reddit to identify subreddits that yielded relevant practi- tioner discussions, and then expanded this set by including closely related, active communities with similar topical fo- cus. We selected active forums covering general cybersecurity discussions (r/cybersecurity,r/ComputerSecurity), cy- bersecurity leadership (r/ciso), blue-team and SOC op- erations (r/blueteamsec,r/SIEM), and cyberspace Q&A (r/CyberSecurityAdvice). Furthermore, we read and ad- hered to each forumâs posted community guidelines and avoided any communities that prohibited the use of forum data for academic research. In total, we gathered a set of nine subreddits that spanned between4.5Kâ1.4M users (Table 1). Gathering Threads and Posts. Within each subreddit, we looked at how LLMs were perceived and used by its mem- bers. First, we performed a keyword-based search to col- lect and filter threads; with a set of13AI-focused terms including âSOC AIâ, âAI Security Operationsâ, and âLLM in cybersecurityâ (see Appendix A), we retrieved 1,310 threads on September 30, 2025, collecting the submissionâs ti- tle, content, anonymized username, posting date, and number of posts. We further refined the set of collected threads using AI-assisted classification. UsingGPT-4.1-miniAPI [47], we created a few-shot prompt to guide the LLM in performing a relevance classification task, supplying it with the thread title and the original post content (see Appendix A). To establish reliability, the primary author manually reviewed a random sample of100threads, labeled each as ârelevantâ or âirrele- vantâ. Out of100threads, only2were incorrectly coded, and no relevant threads were missed, yielding a high inter-rater reliability (Îș=0.96). Having found reliability, we then used the LLM to classify the remaining threads at scale, yielding 80 relevant and 1,230 irrelevant threads. All relevant threads were also further manually reviewed, and 4 were deemed irrelevant, giving us 76 relevant threads. We also randomly 3 sampled50threads classified as irrelevant, and manually ver- ified that all the classifications were correct; these threads focused on topics outside the scope of this study, including frameworks for securing AI systems 1 , cybersecurity career pathways, and interview preparation, or syntactically similar but unrelated topics, e.g., SoC (System-on-chip) Security. Within each thread, we retrieved all posts underneath it, producing a dataset of 1,703 posts as shown in Table 1. Simi- larly, these posts were then also manually labeled as ârelevantâ or âirrelevantâ (α = 0.88). A post was deemed irrelevant if it had been deleted by the creator, removed by moderators or the platform, or if it did not pertain to our research ques- tions. Additionally, redundant replies from the same user that repeated already shared opinions without contributing new information were also labeled irrelevant. Of the 1,703 posts coded, 892 were considered relevant (see Table 1). Ethics Statement. This study does not involve direct interac- tions with human subjects, as it relies exclusively on data that are publicly available [48, 49]. Nevertheless, we consulted the Institutional Review Board (IRB) at the leading institution and obtained an IRB exemption. This said, we actively recognize the ethical implications of analyzing online public discourse and took steps to mitigate potential harms, adhering to the eth- ical standards set in prior published works [28â31]. First, we acknowledge that Reddit users may not have anticipated their posts being used for research purposes, and that neither com- munity members nor moderators explicitly consented to such use. To address this, we followed the best practices of Reddit data collection and limited our corpus to publicly accessible subreddits that do not prohibit research use in their terms of service. Unlike some subreddits that explicitly restrict aca- demic research due to the sensitivity of their content [50], the communities we studied impose no such restrictions. Sec- ond, although the analyzed data is likely not sensitive, we implemented additional safeguards to protect user privacy. All identifying metadata (e.g., usernames and organization names) were removed prior to analysis, and contributors are referenced using anonymized identifiers (e.g.,PXXX). To fur- ther reduce the risk of re-identification through reverse search, all excerpts included in this paper were carefully paraphrased in accordance with established ethical guidelines [51, 52]. Data Availability Statement. Consistent with research prac- tices involving Reddit data [28â31], we will not publicly re- lease the dataset to protect the privacy and avoid any potential for deanonymization of the Reddit posters. Instead, interested researchers may contact the author to request an archived copy. To promote reproducibility, we provide the list of key- words and the few-shot prompt used in our data collection process in Appendix A. 1 We distinguish between the application of LLM tools for cybersecurity (our focus) and the security of such tools, which is an equally important but complementary concern not addressed in this study. 3.2 Data Analysis To analyze online discussions, we employed a mix of quali- tative and quantitative methods [53] to gain insight into the discussions and statistically compare the prevalence of differ- ent LLM tools uses and opinions. Qualitative Analysis of Posts. To analyze the gathered posts, we first conducted a hybrid coding methodology [54]. To ensure proper interpretation, all posts were analyzed in the context of any posts that they were in reply to. If any posts contained available external links, as in other works [28], we reviewed all linked content as well. To construct the codebook, a single coder first began with a set of codes informed by the use cases, factors for tool adop- tion, and opinions of AI noted in prior works [24, 26]; this initial codebook was inductively expanded upon through the coding of 350 posts by the single author. To assess inter-rater reliability (IRR) for this initial codebook, a second coder was introduced and familiarized with the codebook through co- coding of25posts. After developing a shared understanding, the two coders then independently coded a new set of50 posts, after which IRR was computed using Krippendorffâs Alpha (α) [55, 56] for every code. The authors then resolved coding disagreements and adjusted the codebook to clarify and refine definitions. If at least one code did not obtain sub- stantial agreement (α℠0.8), this process was repeated. After eight rounds of double coding, agreement for all codes was achieved, and the primary coder then independently coded the remaining 1,278 posts. The final codebook and resulting reliability are presented in Appendix B. Through the coding process and after creating these codes, the primary coder, along with the full research team, con- ducted a reflexive thematic analysis [57] on the coded dataset. Through data exploration and routine meetings with the full research team, we distilled patterns and trends in the codes into large-scale themes that we present in our results. Quantitative Analysis of Codes. Having established the re- liability of codes, we also conducted quantitative analysis to measure and statistically compare their prevalence. In particu- lar, to assess the prevalence of certain LLM tools (§ 5.1) and use cases (§ 5.2) discussed, we conducted pairwise compar- isons using two-sample tests for equality of proportions [58]. Also, to understand whether certain opinions about LLM fac- tors were more positive or negative mentioned in posts, we conducted a chi-square test of independence between factors and opinions (§ 6), and post hoc tests using adjusted standard- ized residuals (z) [59]. To account for multiple comparisons, we applied the Bonferroni correction [60] to our tests. 3.3 Limitations Similar to other works that analyze anonymous community forums [28â31], our study shares a number of limitations that we account for when interpreting our results. First, our 4 Topics of Discussion# Total Posts (Out of 892) LLM Uses (§ 5) Tools Mentioned410 (45.96%) Use Cases Mentioned325 (36.44%) Perceptions of LLM Tools (§ 6) Opinions Shared406 (45.52%) LLM Factors459 (51.46%) Implications of LLM Adoption (§ 7) LLM Adoption373 (41.82%) Vision for the Future276 (30.94%) Table 2: Distribution of Discussion Topics â Identified across 892 relevant posts. Individual posts often addressed multiple themes. selection of forums and keywords, while carefully curated, may reflect biases and capture only a subset of relevant dis- cussions about practitionersâ experiences with LLM tools in SOCs. Second, we rely on self-reported information provided by Reddit users. Although we report specific roles when ex- plicitly stated in posts, we cannot independently verify the roles, professional identities, or organizational contexts of the contributors, which may substantially influence how certain tools or factors are perceived and prioritized. Third, Reddit users possibly represent only a subset, and possibly a biased section of the security community. Prior studies have noted that Reddit users tend to be more engaged with emerging technologies than the general practitioner population [61]. As such, these discussions may reflect an upper bound on awareness of LLM tools rather than typical practice. Despite these limitations, our work surfaces nuanced prac- titioner perspectives and timely insights into emerging LLM adoption trends in SOCs, which can be further validated through empirical and design-oriented methodologies. 4 Overview of Forum Threads and Posts The threads in our dataset came from three subred- dits:r/cybersecurity(n=69),r/Information_Security (n=5), andr/ciso(n=2). These relevant threads were posted between December 2022 and August 2025 (Figure 2). We can see that conversations around LLM tools have increased over time, with 75% posts made within the last year of data collec- tion (Sepâ24-Augâ25). This aligns with the recent emergence and broader visibility of SOC-focused LLM tools around mid- 2024, including Microsoft Copilot for Security [21], Prophet Securityâs AI SOC Analyst [62], and the introduction of Gem- ini in Google SecOps [63]. Relevant threads contained an average of 22 (±32.8) posts, ranging from 1-193 posts. Across these 76 threads, we ana- lyzed 892 (52.38%) relevant posts. As shown in Table 2, these relevant posts touched on our topics of inquiry: Tools (n=410), Use Cases (n=325), Opinions (n=406), LLM Factors (n=459), LLM Adoption (n=373), and Vision for the Future (n=276). We discuss each topic in detail in the following sections. Categories of LLM Tools# (Out of 410) By Purpose General-Purpose LLM Tools248 (60.49%) Security-Specific LLM Tools180 (43.90%) By Origin Commercially Available398 (97.07%) In-House Built14 (3.42%) Table 3: Reported LLM Tools â Categorized by purpose and origin. Note that individual posts may mention multiple tools. 5 Tools and Use Cases in Practice (RQ1) We find that security practitioners discuss a wide range of LLM tools and uses within SOC workflows. In particular, these discussions reflect a dominant mention of general- purpose LLMs alongside a fragmented ecosystem of com- mercial tools for security. Moreover, LLM use cases span a diverse set of SOC activities, with discussions most heavily concentrated on incident response and triage, followed by scripting and reporting tasks, and less frequently on knowl- edge support, threat analysis, and training-related functions. 5.1 LLM Tools Across 410 posts, practitioners discussed a wide-range of LLM tools used within SOCs. To understand the tools used, we code each by whether it is (1) general-purpose, such as ChatGPT, or security-specific, such as Microsoft Security Copilot, and (2) commercially available or built in-house. More General-purpose LLMs Are Referenced Than Security-Focused LLMs. As shown in Table 3, while not explicitly designed for security operations, 60.5% (n=248) of posts referenced general-purpose LLMs, such as ChatGPT, Microsoft Copilot, Claude, along with broader terms such as âLLMs,â or âGenerative AI (GenAI)â. Conversely, only 43.9% (n=180) of tools references mentioned security-specific tools, including Microsoft Security Copilot, Dropzone AI, and In- tezer. To evaluate whether these differences are significant, we ran a two-sample test for equality of proportions, and found that general LLM tools are significantly more likely to be discussed by security practitioners (Ï 2 (1) = 21.94,p<.001, Ï = 0.23). Across these discussions, practitioners described diverse integration patterns of these tools into SOC work- flows (§ 5.2), alongside sharing hands-on reflections on their benefits and limitations (§ 6). A Long Tail of Security-Specific Tools Exists. While general-purpose LLMs were centered around a few key players, a long list of security-specific tools was referenced (see Appendix C.1). Practitioners mentioned only 9 distinct general-purpose tools, with discussion dominated by Chat- GPT (n=83), followed by MS Copilot (n=11), and others. In contrast, security-focused LLMs had a long list of 30 dis- tinct commercial platforms, yet only four of these tools were 5 Types of AI Use CasesDescription# Posts (of 325) Triage & Incident ResponseSupport for alert triage, investigation, correlation, mitigation, and response workflows.139 (42.77%) Scripting & Query SupportGeneration and refinement of security scripts and detection queries.88 (27.08%) Reporting & DocumentationDrafting and summarizing investigation reports, policies, and threat briefings, etc.84 (25.85%) Threat AnalysisSupport for risk and vulnerability analysis, threat hunting, and modeling workflows.54 (16.62%) Knowledge SupportConcept explanation, information retrieval, and document-based knowledge extraction.53 (16.31%) Training, Compliance, and OthersDrafting training exercises, reviewing compliance materials, and other non-core tasks.18 (5.54%) Table 4: Reported LLM Tools Use Cases and Descriptions â Note that individual posts mention multiple use cases. mentioned more than five times: Security Copilot (n=40), Dropzone AI (n=10), Intezer (n=8), and Cortex XSIAM (n=6). The remaining appeared three times or fewer, with half of the 30 tools mentioned only once, indicating fragmented aware- ness across a rapidly expanding AI-for-cybersecurity vendor ecosystem. Notably, half of these tools are marketed as âAI SOC analystsâ (further discussed in § 7.1) positioned as au- tonomous assistants capable of performing Tier 1/2 tasks [23]. In-house Custom Workflows Are Rare, but Used. In-house developments accounted for only 3.4% (n=14) of posts, com- pared to the commercially available solutions referenced in 97% (n=398) of posts. Across this smaller subset, practition- ers described building custom LLM workflows internally to augment security workflows. For example, in a thread about LLMs for SIEM, one practitioner (P310) described leverag- ing open-source agent frameworks to enhance SIEM search workflows: âI have developed custom Python scripts using open-source agent frameworks to feed SIEM query results into an agent to iteratively process the data and generate a consolidated review of the search output.â Similarly, P013 mentioned, âWe developed our own agent and connected it to SOAR to automatically handle and escalate incidents.â This suggests that although off-the-shelf solutions are popular, such custom workflows may be valuable for closing integration gaps and supporting organization-specific security processes. 5.2Uses of LLM Tools in Security Operations As shown in Table 4, across 325 posts, LLM use cases were one of six broad activities: triage and incident response, script- ing support, reporting, threat analysis, knowledge support, and miscellaneous activities. To understand the relative prevalence across use cases, we conducted pairwise two-sample tests of proportions with a Bonferroni correction and found that the prevalence of these uses differs significantly. In summary, âTriage and IRâ > all others (p<.001for all), âScripting & Query Supportâ and âReporting & Documentationâ > others (p< 0.05for all), âThreat Analysisâ & âKnowledge Supportâ > âTraining, Compliance, and Othersâ (p < .001 for all). For full statistical results, see Appendix C.2. Incident Response and Investigation Support. Across all LLM cybersecurity use cases, 39.7% (n=139) of posts fo- cused on improving or automating tasks relating to incident response, including triaging, investigating, responding, and mitigating alerts and security threats. Practitioners described a varied set of ways for LLM tools to aid in reducing their workload and making investigations more effective. LLM Investigation Buddies: Similar to common uses of LLMs in non-cybersecurity contexts, practitioners described using LLMs as decision support tools during investigations, operating in a âthink and assistâ mode that performs subtasks such as correlation, hypothesis generation, and sensemaking, while analysts retain full control over interpretation and final decision making. As P452 shared: âDuring a complex investi- gation, I utilize LLMs like a buddy, throwing ideas at them. I let them examine the data to help with correlation, and I han- dle the critical thinking.â Similarly, P757 described providing contextual artifacts to LLMs for exploratory analysis, noting: âI provide logs, screenshots, or event timelines to LLMs to help me piece together what might be happening, either to validate my findings, or help me zoom in on a problem area.â LLM-Driven Triaging: Beyond providing investigative sup- port, practitioners described using LLMs to partially automate alert triage through human-in-the-loop workflows. In these settings, LLMs may perform an initial analysis of incom- ing alerts, proposing likely scenarios and potential mitigation actions for human analysts to interpret and resolve. Such LLM-driven triage was most commonly discussed for high- volume âTier-1â alerts, where rapid filtering and escalation decisions are required, though some accounts also described its use in more in-depth analytical contexts. For instance, one L3 analyst (P725) described an LLM-driven pipeline that au- tomatically extracts context from EDR alerts and presents a structured breakdown for analysis to ultimately review and take action on: âWhile we use it to augment the analysis and get more clarity on things, we do not allow it to take actions. Ultimately, trust but verify!â Others, however, described al- lowing LLMs to fully handle and resolve low-impact alerts, escalating to humans only when necessary: âOur team has been using LLMs to assess high-noise/high-volume alerts that are low-payoff/low-impact. Only if there are any outliers, the alert is flagged for additional human evaluationâ (P864). Fully Autonomous LLM Mitigation: Lastly, a small fraction of posts (n=5) described delegating entire classes of routine high-confidence tasks to fully autonomous LLM pipelines. Commonly integrated with SOAR systems, these workflows independently triage alerts, correlate incidents, enrich find- 6 ings, and initiate actions with minimal human involvement (P013, P066, P247). In a few cases, this autonomy extended into active responses; for instance, a CISO (P086) detailed how their LLM-powered correlation engine âdetermines inci- dent severity and autonomously dispatches remediation tasks, like isolating endpoints, running full scans, or disabling suspi- cious accounts.â These use cases demonstrate that, although limited, some organizations might be experimenting with fully autonomous workflows to scale high-volume alert workloads. Scripting & Query Support. Consistent with recent ad- vances in LLM-based code generation [64â66], 27.08% (n=88) of use cases described practitioners using LLMs to generate scripts and queries to augment their daily workflows. Code Generation: Several posts (n=46) discussed using tools such as ChatGPT to draft scripts in Python, PowerShell, or Bash, typically to create boilerplate code or outlining logic that the analyst then adapts. As P753 explained, âI find "write code to do ABC in language X" prompts are useful for get- ting started with boilerplate code instead of writing it manu- ally.â Others mentioned leveraging LLMs to enhance or debug scripts they wrote themselves (P742). Query Support: Practitioners also mentioned LLM-assisted query generation and refinement (n=51), particularly for SIEM and detection engineering workflows. These cases in- cluded assistance with crafting or troubleshooting KQL, SQL, or regex queries. As P309 noted, âIâve explored a number of LLM chatbots to assist me fix SIEM queries, offer recommen- dations, or craft highly targeted queries for specific needs.â Reporting & Documentation. Across 25.85% (n=84) of mentioned tasks, practitioners described incorporating LLMs into their technical reporting and documentation workflows, supporting tasks ranging from general writing assistance (n=20) to SOC-specific activities such as policy and SOP de- velopment (n=19), investigation summaries (n=13), threat in- telligence briefings (n=11), and risk assessment reports (n=9). As one described using ChatGPT to rewrite technical CVE definitions into an accessible language that âa business man- ager without a lot of technical skill can understand,â (P920). Others mentioned using LLMs to offload time-intensive re- porting tasks: âI use ChatGPT to ingest articles and generate summaries of CTI briefings. This allows me to save time and concentrate on threat hunting for any IOCsâ (P823). Learning & Knowledge Support. Practitioners also de- scribed using LLMs as versatile knowledge-support tools (n=53, 16.31%), reflecting a broader shift in how analysts retrieve, learn, and navigate complex technical information. Learning and Explanation: Across 20 posts, practitioners mentioned relying on LLMs to clarify obscure command-line behavior or provide context for vague error messages, high- lighting the value of LLM-generated explanations for learning something new or complex. As P921 noted, âhaving an LLM explain something as if you were five may seem trivial, but can be incredibly helpful when approaching complex topics.â Information Retrieval: In 17 posts, LLMs were framed as a more efficient path from questions to actionable directions, often contrasted with a conventional Google search. As P442, a threat hunter, described âthe role of an analyst is to know the right question to get the answer. Google has been a part of our lives for many years. LLMs are simply far more efficient.â Knowledge Extraction: Finally, practitioners discussed uti- lizing LLMs to extract knowledge from documents (n=13) to receive targeted responses. For example, P722 described querying documents via NotebookLM, and P740 described embedding internal materials (policies, risk reports, acronym lists) into an LLM agent: âI created an agent loaded with our client policies, risk reports, and acronym lists to allow our analysts to ask questions or perform light analysisâ (P740). Threat Analysis. Around 16.62% (n=54) of posts discussed the use of LLMs to support analytical reasoning in activi- ties such as risk and vulnerability assessment, threat model- ing, and threat hunting. For instance, P879 described using LLMs to help plan and review risk assessments, noting: âWe have been augmenting LLMs extensively to our risk assess- ment planning to ensure completeness.â Others reported using LLMs to sanity-check penetration testing findings, such as asking: âWhat title and CVSS score would you assign to a discovery I have from a pentest related to X?â (P751), to save time by offloading the analytical thinking to the LLM. Training, Compliance, and Others. A smaller subset of posts (n=18, 5.54%) described LLM use for miscellaneous tasks, including drafting cyber training exercises and developing scenarios to test defensive controls (n=6), reviewing policies or guideline reports (n=6), and other non-technical activities. As P823 noted: âI train junior analysts on IR, I use ChatGPT to draft scenarios and questions to answer.â 6 Perceptions of LLM Tools (RQ2) Across 406 posts, practitioners shared opinions and experi- ences with LLM tools for cybersecurity operations. Similar to prior work [24], we find that practitioners framed their ex- periences primarily around six factors: capabilities, efficiency, reliability, security and privacy, autonomy, and cost (Table 5). By analyzing the factors and sentiments expressed, we find both statistical and qualitative evidence that, although prac- titioners are increasingly satisfied with the capabilities and efficiency of LLM tools, persistent concerns around reliability, autonomy, security, and cost temper their trust and limit their willingness to delegate sensitive tasks. Statistical Analysis of Factors and Perceptions. To exam- ine whether sentiment toward LLM tools differed across the discussed factors, we conducted a chi-square test of indepen- dence. The test revealed a significant association between factor type and sentiment (Ï 2 (5) = 172.21, p<.001), indi- cating that practitionersâ perceptions differed significantly across factors. As observed in Figure 1, post-hoc tests with 7 FactorsDescription CapabilitiesSpecific tasks and SOC use cases where LLM tools are perceived as helpful EfficiencyTime-related aspects, such as speed of analysis and de- ployment, workload reduction ReliabilityTrustworthiness, accuracy, and consistency of LLM- generated insights Security & PrivacySecurity of LLM tools, organizational data governance, and privacy considerations AutonomyThe ability of LLM tools to operate independently with minimal human supervision CostFinancial aspects, e.g., operational costs, perceived return on investment Table 5: Definition of LLM Factors â the six dimensions practitioners mentioned when sharing opinions on LLM tools. adjusted residuals find that comments around LLMâs capabili- ties(z = 8.11, p<.001), and efficiency(z = 6.38, p<.001) were significantly more likely to be positive, while reliability (z =â6.77, p<.001), security and privacy(z = â5.41,p< .001), level of independent autonomy(z =â4.94, p<.001), and cost(z =â4.02, p<.001)were significantly more likely to be negative. We now discuss each factor in detail. 6.1 Capabilities Practitionersâ discussions around LLMâs capabilities were significantly more positive than negative(p<.001). Notably, nearly three-quarters of these posts (n=130) also included spe- cific SOC use cases (§ 5.2), highlighting that LLM-powered tools can effectively augment real-world SOC tasks. Better Contextualization and Visibility of Alerts. Practi- tioners frequently described LLM tools as effective for inci- dent enrichment and surfacing low-visibility behaviors (n=13). While existing ML tools can correlate well-defined alerts, practitioners noted that LLM tools can effectively interpret a multitude of varied signals to better answer an investigation, which can also provide essential context for analyst interpre- tation. For instance, reflecting on experiences with Agentic LLM tools such as Purple AI, one practitioner noted that, âLLMs are going to transform incident enrichment. Compared to scripting or googling, LLMs can spit out the who, what, when, where, and why of an incidentâ (P044). Similarly, P038 noted that these tools not only dramatically reduce their work- load, but also provide important context and remain easily accessible with natural language queries: I use Agentic LLMs, and theyâre incredible. Not only does it turn half a million alerts/events a month into 1-3 relevant daily alerts, but also gives impor- tant context. During an investigation, I can ask, âDoes this event fire every time a user logs in or is this a new alert?â or, âDid they get challenged for MFA?â... Itâs not just reducing my workload, itâs finding things I couldnât see. 0 50 100 150 200 Positive Negative CapabilitiesEfficiencyReliabilityS&PAutonomyCost *** *** *** *** *** *** 174 97 54 10 3 55 6 46 6 41 5 30 #Posts Figure 1: Opinions of LLM Tools by Factor â All factors exhibit statistically significant differences in sentiment (p<.001); LLM Capabilities and Efficiency are discussed more positively, while other factors are discussed more negatively. Improved Interpretability of Signals. In addition to better contextualization of signals, practitioners (n=10) also noted that LLMs made previously hard-to-understand signals read- ily interpretable. In particular, practitioners emphasized how LLMs can translate verbose or opaque log data into human- readable narratives through querying, summarization, and ex- planations. As P294 explained, âWhen you feed ChatGPT a few log lines, it can assist in explaining whatâs happening in simple language. For example, Windows events can be diffi- cult to interpret due to the event IDs and combinations you need to memorize.â Reduced False Positives. Perhaps due to their effective use and analysis of many signals, practitioners also consistently framed LLM tools as effective in performing accurate clas- sification decisions and automating alert triage and manage- ment pipelines (n=18). In particular, while concerns of high false negatives were notable in traditional alert classification systems [1, 11, 41], and thought to be further exasperated in traditional ML systems [24], LLM tools were surprisingly perceived by analysts as broadly effective in detecting and re-emptively removing false positive alerts, and instead pro- viding a smaller amount of high-confidence signals. For in- stance, P017 noted: âThe âautonomousâ aspect of LLM tools is most evident in their ability to eliminate high-confidence false positives that are not worth wasting an analystâs time.â Inabilities of LLMs. Beyond the predominantly positive dis- cussions of LLM capabilities, 97 posts (38.65%) expressed reservations about their practical value in SOC contexts. A recurring concern was whether the probabilistic nature of LLM tool outputs was fundamentally susceptible to incorrect conclusions, particularly around novel and unseen examples. As P910 argued, âCybersecurity mostly depends on dealing with outliers and anomalies... LLMs are awful at even com- prehending these things, let alone acting upon.â Skepticism also stemmed from firsthand experiences with performance breakdowns in complex or structured tasks, including fail- ures to correctly interpret alert data or generate functional 8 detection queries in complex languages, such as KQL. Some posts emphasized systemic data challenges affecting LLM use at both training and inference stages. As P114 noted, âThe quality of the tools is heavily dependent on the training data,â while others emphasized that alert data already suffers from poor signal-to-noise ratios, suggesting that, â...with AI, the problem of garbage in, garbage out still persistsâ (P263). Furthermore, these issues were compounded by vendors over- promising LLM toolsâ abilities, the perception that existing tools may be good enough (§ 7.2), and concerns that even outputs for tasks that LLMs can readily perform can also still suffer from hallucinations and inaccuracies (§ 6.3). 6.2 Efficiency Practitioners were significantly more likely to speak positively about the efficiency of LLM tools(p<.001), frequently not- ing that LLM tools speed up their analysis, while also noting that they may incur overhead to verify their outputs. LLM Tools Speed Up Analysis and Deployment. Within core SOC workflows such as alert triage, investigation, and IR, practitioners frequently reported measurable efficiency gains from using LLM tools in their workflow. For instance, an L3 analyst (P725) described how using an LLM tool in their workflows helped them automate their triaging and âreduced MTTT (mean-time-to-triage) from around 45 minutes to less than 2 minutes.â Similarly, the use of Agentic LLM tools dras- tically reduced how long investigations took by pre-emptively filtering and analyzing alerts they needed to respond to: By delivering end-to-end investigations, these tools condense the investigation from âhereâs 500 things to look atâ, to âhereâs what happened and what you should probably doâ (P085). Beyond triage and during investigation, practitioners also noted that, compared with human analysts in particular, the detection and response time with LLM tools were often much faster for processing and resolving large volumes of alerts (P011, P870, P430), with P870 noting: âAI can process mas- sive volumes of data quickly and detect threats that would take a human analyst hours to identify.â Despite this, issues remain with how quickly certain tools resolve alerts. While discussing experiences with Agentic LLM tools, analysts, in- cluding P010 and P788, noted that response times can be very high. Furthermore, beyond the speed of the toolâs operation, practitioners also noted that the speed to effectively deploy LLM tools was lower than that of others. This was discussed in both how fast analysts could learn to use (P002) and set up (P001) such tools. LLM Tools Introduce Verification Overhead. In contrast to their time-saving abilities in performing operations, a small set of posts (n=10) highlighted that LLM tools introduce new overhead as analysts may often need to correct, tune, or other- wise verify their outputs. For example, P206 expressed frustra- tion after their SOAR was replaced with an LLM tool-driven workflow, and noted that, âWe spend nearly four times the amount of effort correcting and changing LLM behavior as we do actually addressing incidents.â Similarly, P289 reflected after using LLMs for generating regular expressions for fire- wall rules: âI had to verify and correct the output, so did I end up saving time? Not sure.â 6.3 Reliability When discussing the accuracy and trustworthiness of LLM predictions, practitioners were significantly more likely to discuss negative aspects(p<.001)such as their inconsistent behaviors, and the false confidence of hallucinated answers. Hallucinations and Non-determinism Cause Concern. Most posts discussing reliability (n=25) focused on LLM toolsâ tendencies to hallucinate [67]. Often, these were discov- ered through personal experiences; for instance, P171 noted that an LLM toolâs confident but incorrect conclusion still makes them worry about relying on such tools: âWhenever Iâve tried LLMs for security work, itâs produced pure garbage. It made up descriptions of an imaginary malware I named. I do not trust any of it.â Beyond arriving at incorrect con- clusions, practitioners noted that LLMs could also fabricate justifications and âhallucinate evidence to prove itâ (P120). For many, this felt antithetical to the idea of cybersecurity; as noted by P924, âLLMs can be so confident in wrong an- swers... in the security space, these âsmallâ mistakes can lead to larger problems by introducing new holes in the Swiss cheese model of defense.â Furthermore, for others (n=4), their concerns werenât just based around incorrect answers, but the lack of determinism around their answers. Consistent with previous studies highlighting this issue [68, 69], practitioners such as P058 noted: âThe problem with implementing au- tonomous AI solutions for security is the unpredictable nature of the outputs.â Practitioners also described how this reliabil- ity is not even constant, but can vary based on the specific task, and its specific context; for instance, the particular language that the code is generated: âWhen writing code, outcomes get increasingly unreliable as the code language gets more complex, like with Powershellâ (P175). 6.4 Privacy and Security When discussing privacy and security risks of LLM tools, practitioners were significantly more likely to bring up nega- tive viewpoints(p<.001). These concerns often focused on unintentional data leakage (n=31), and the expanded attack surfaces LLMs may introduce (n=17). Difficulties in Data Governance and Possible Leakages. Practitioners frequently expressed concern that users may inadvertently enter organizational information into commer- cial LLMs, raising the risk that public models could retain 9 or learn from sensitive inputs (e.g., P783, P827, P789). As P294 cautioned: âLLMs may learn from the information you provide, so be cautious about what you prompt; otherwise, sensitive organizational information may be inadvertently ex- posed.â These concerns were particularly pronounced around integrated tools that require access to internal enterprise data. As P863 shared, while their teams were excited by the success of early LLM tools, they remained âuncertain about the toolâs access to internal dataâ and felt a tension between giving the LLM tool data access to be âuseful without that access be- ing a huge risk to themselves.â Some practitioners, like P049, questioned the necessity and scope of data access, and became worried that it may even be misused or stolen: âIt sounds like what all AI companies are trying to do: get as much data to train models that they can sell to other customersâ (P885). LLMs Increase Attack Surfaces. Practitioners also raised concerns that LLM tools themselves become new attack sur- faces. In the excitement to produce and sell LLM tools for security, practitioners noted that security concerns of the tools themselves may not be prioritized: âBe careful, security with new technology often lagsâ (P705), Indeed, using simple tech- niques like jailbreaking [70] and prompt injections [71], prac- titioners like P926, became worried that LLM tools can be eas- ily attacked: âI play with around LLMs a lotâ it is alarmingly easy to get a model violate their guardrails.â After break- ing these boundaries, practitioners became worried that LLM tools themselves could be used to conduct attacks against their own company: âyou could keep rephrasing open-ended questions to bypass its âno bad guy stuff filterâ (safeguards).â 6.5 Autonomy In discussing autonomous task execution without human su- pervision, practitioners were significantly more likely to hold negative perceptions(p<.001), noting that, despite vendor claims, LLM tools still required heavy oversight. While Improving, Human Oversight Is Still Needed. Par- ticipants noted that while LLM tools are often marketed as fully autonomous and can act as replacement analysts, they were not often reliable enough to fully delegate tasks (P047, P257, P925). While practitioners discussed a few cases of fully autonomous LLM pipelines (§ 5.2), most believed that current systems were better suited to decision support than to autonomous operations. As P047 mentioned, âI work with a wide range of tools that span the Gartner quadrant 2 , they are still deeply flawed and require human intervention and over- sight.â SOC work, in particular, was viewed as too nuanced, context-dependent, and constantly evolving for LLMs to fully handle (P011). Thus, while LLMs were viewed as useful and in some cases essential, practitioners would not allow them to 2 The Gartner Magic Quadrant [72] is a market research framework that evaluates technology vendors based on their ability to execute and complete- ness of vision, positioning them into four categories: Leaders, Challengers, Visionaries, and Niche Players, to illustrate the competitive landscape. perform actions autonomously: âIn my professional circle, we all agree that we will need to leverage AI, but only for advice, not for taking autonomous actionsâ (P089). Instead, they rec- ommended to others that they âUse it like an intern. Allow it to gather information and raise its hand when it notices something, but donât let it touch anything criticalâ (P089). 6.6 Cost Practitioners were significantly more likely to hold negative opinions(p<.001)when discussing the cost of LLM tools. While practitioners do not typically control organizational budgets, their negative perceptions regarding LLM-related costs reflect operational feasibility concerns and expected return on value within SOC workflows. These practitioner- level perceptions are critical, as they influence how tools are piloted, adopted, or resisted within SOC environments. Query Costs Add Up Fast. Several practitioners noted that the LLM inference cost [73] can become prohibitively expen- sive at scale, raising concerns about the economics of scaling LLM-driven analysis to SOC-sized datasets (e.g., P227, P233, P409, P895). As P227 commented, âThe issue here is the cost. Since each inference will cost a few cents, it might not be the most economical option for processing a large number of events.â Furthermore, several practitioners, such as P056 and P071, were worried that future costs could be even worse given their increasing reliance. Moreover, because of data gov- ernance concerns (§ 6.4), some believe that privacy-concerned SOCs will need to invest heavily in computing infrastructure: âItâs going to be so expensive to locally train and operate these high-quality LLMsâ you will need data centersâ (P895). Returns Are Limited. Several participants were skeptical that given this high cost, whether LLM tools were practical. Costs were so high that practitioners, such as P790, questioned whether they could not simply hire just as many workers with the same money: âAlthough Security Copilot has several in- teresting features, it is still not worth the price. Even minimal usage can be equivalent to one full-time employeeâs salary.â Because of this, several participants, such as P362, believed that âAt this stage of its development, I donât think LLM tools can provide enough value or a real return on investment.â Other practitioners, such as P048 and P919, explicitly noted that they did not adopt LLM tools solely because of the cost. 7 Adoption of LLM Tools (RQ3) We now analyze how practitioners discuss adopting LLM tools within their operations and the concerns they hold. 7.1 Trajectories of LLM Adoption By analyzing posts (n=373) for whether practitioners adopted LLM tools within their works, we found that 200 posts 10 (53.62%) reported actively using and adopting LLM tools, 123 posts (32.98%) were interested or in the process of eval- uating LLM tools for adoption, and 50 posts (13.40%) were not yet adopting or evaluating. This suggests that while in- terest continues to grow, most practitioners may be actively adopting LLM tools in real-world security operations. General Purpose LLMs Are Often Adopted, But Commer- cial Tools Drive Interest. Across the posts describing active use of LLM tools, general-purpose LLMs were referenced most frequently and appeared in 103 posts (51.5%), compared to only 37 posts (18.5%) mentioning security-specific LLMs (and mirrors the pattern that general-purpose LLMs are more often discussed as well; § 5.1). However, when looking at posts that discuss either actively evaluating or considering LLM-use, security-focused LLM tools were referenced in 58 posts (47.16%), more than double the frequency of general LLMs, which appeared in only 28 posts (22.76%). In general, this exploratory discussion came from people asking whether the claims made by security vendors were real. For instance, several SOC managers (e.g., P001, P094, P305), were actively exploring whether security-specific tools that promise near- automation of tasks were practical: âWe have started looking into AI SOC Analysts. Our team still spends a significant amount of time on L1/L2-type work, which should have been automated by nowâ (P094). community input on several re- lated questions, including whether these LLM tools delivered value (P008, P055, P094, P860), introduced overhead (P055), integrated well with existing technology (P786), compared favorably with traditional SOAR solutions (P001, P008, P316, P411, P890), and justified their cost (P305, P890). Tools Are Often Adopted For Non-Core Security Opera- tions. While broad discussions around use cases were often focused around core security operations like triage and IR (§ 5.2), posts that described actively adopting LLM tools of- ten focused on smaller, productivity-oriented tasks. Of posts that described adopted uses for LLM tools, 28% focused on scripting & query support and 26% on reporting and documen- tation; in comparison, 18% discussed triage & IR account, and only 6% for threat analysis. This may indicate that while core operational tasks dominate discussions around LLM tools, the tasks for which these tools are actively being adopted are likely to concentrate on self-scoped tasks that afford greater analyst control and easier verification. 7.2 Cautions against LLM Adoption Despite reported adoption and curiosity, practitioner discus- sions highlighted recurring barriers that shaped how they eval- uated the appropriateness of LLM tools for SOC. Inflated Vendor Promises. Practitioners frequently inter- preted LLM adoption through the lens of a long-standing history of overpromising technologies in cybersecurity, lead- ing to skepticism toward marketing claims of âautonomous SOCâ capabilities. As P859 emphasized, âThe marketing people will claim their product can do everything. However, without hands-on experiments, such products can be anything and everything at once, or nothing.â Consistent with this view, across20posts, practitioners reported that they or their or- ganizations had previously evaluated tools, such as Security Copilot, that did not result in adoption, primarily due to inef- ficient performance and unjustifiable costs. Sufficiency of Traditional Solutions. Practitioner skepticism toward LLM adoption was also shaped by a perceived ten- dency to overuse, or âlook for places to cram in AIâ (P226). As P463 argued that âmany organizations lack the use case, scale, or resources to justify LLMs or agentic security so- lutions, yet they are forcing AI into workflows.â Across 11 posts, practitioners expressed a clear preference for traditional automation, emphasizing that many SOC tasks, particularly those that are deterministic and well defined, are already ef- fectively addressed by existing approaches. From this per- spective, LLMs were often characterized as overkill: âItâs like bringing a tank to a knife fight when using LLMs to determin- istic security problems. Stick to traditional automation, itâs simpler, cheaper, and gets the job doneâ (P460). Organizational Restrictions and Policies. Another cluster of narratives (n=10) highlighted that organizations often restrict the use of public LLMs primarily due to data-loss prevention or confidentiality concerns (§ 6.4). As P922 explained, âSince people continued to carelessly pour company data into public LLMs, we incorporated them into our company block list.â Beyond explicit access restrictions, one practitioner described a broader uncertainty within organizations about how to assess the risks associated with LLMs, driven by the absence of a shared vocabulary or established frameworks for evaluating such tools. As P784 noted, âSince Gen AI is so new, nobody even understands how to discuss it. . . the CISO is unsure what the security risks are, the Chief Risk Officer doesnât know how to characterize them on the risk matrix.â Anti-LLM Sentiment. Lastly, a smaller subset of posts (n=9) reflected a clear reluctance to LLM adoption, rooted in per- sonal or ideological opposition to the technology. These posts were rarely supported by detailed reasoning; instead, they expressed dismissiveness through statements such as âNo, I have a brainâ (P768), and âIf people could stop talking about their AI solutions, I would pay for itâ (P005). One practi- tioner (P840), however, articulated inhibition grounded in the broader implications of LLMs for SOC workforce (further discussed in § C.3). As they explained, âCompanies would kill to replace us with AI, so I donât want to contribute to my own obsolescence... Itâs similar to being excited about offshoring when their ultimate goal is to eliminate you.â 11 8 Discussion In this section, we reflect on our findings and outline direc- tions for future research to understand the nuances of LLM adoption, design trustworthy systems for security operations, and sustain SOC workforce development. LLM Adoption is Shaped by Control and Commitment. Our findings suggest that LLM adoption in SOCs is not mono- lithic but rather a layered process that differs across tools and stakeholders. As discussed in § 7.1, general-purpose LLMs are often adopted by analysts through independent experimen- tation to improve productivity on self-scoped tasks. In con- trast, much of the curiosity surrounding security-specific en- terprise tools is expressed by self-identified decision-makers, reflecting organizational priorities. This split helps explain why SOCs may exhibit high re- ported adoption of general LLMs while still showing friction around enterprise adoption, despite active interest. LLMs af- ford analysts greater control and reversibility, allowing them to selectively apply the tools to low-risk tasks without deep in- tegration or organizational commitment. In contrast, adopting enterprise security-focused platforms is a procurement and governance-level decision [74, 75] that often entails broader data access and tighter coupling with existing SOC tooling. Research Opportunity. While recent work has begun to ex- plore the adoption of LLMs for specific SOC tasks [13,14], fu- ture research should examine adoption as a multi-stakeholder phenomenon. Prior work has already identified inherent mis- matches between how managers and analysts evaluate SOC work [1]; adopting a stakeholder-specific lens is therefore crit- ical to understand the divergent priorities that shape adoption decisions. Future work should also investigate how informal analyst-level LLM use can be secured under governed de- ployments, such as [76], without sacrificing the low-friction interaction patterns that make LLMs inherently valuable. Reliability is the Hard Ceiling for Autonomy. Our findings indicate that practitionersâ reluctance to grant autonomy to LLM tools is not rooted in abstract distrust, but in hands-on experience with unreliable system behavior. Reliability issues (§ 6.3) necessitate manual verification before LLM-generated outputs can be applied, directly constraining the autonomy these systems can be granted. As a result, LLM-integrated workflows remain largely operator-driven with LLMs serv- ing primarily as decision support, while fully autonomous deployments are rare (§ 5.2). When LLM judgments cannot be confidently trusted, delegation becomes a source of risk rather than benefit, echoing prior work linking low trust in automation to reversion to manual control [77]. Moreover, the verification overhead offsets reported efficiency gains (§ 6.2), reinforcing that LLM autonomy in SOC contexts depends less on technical capability and more on whether outputs can be acted upon without extensive verification [26]. Research Opportunity. This framing highlights the need for future research that treats trust and autonomy as interde- pendent properties, and examines how LLM systems commu- nicate uncertainty, limitations, and confidence in ways that allow analysts to assess reliability in situ. Prior work shows that generic uncertainty disclaimers, such as, âthis model may make mistakes,â have minimal impact on the credibility threshold in high-stakes expert settings [78]. Our findings reinforce the importance of further research into how LLM tools can communicate uncertainty in actionable and task- specific ways, as a necessary step toward more trustworthy and autonomous integration of LLMs within SOC workflows. The SOC Workforce Development Crisis. Our findings surface an important tension with the implications for the long-term sustainability of the SOC workforce. As noted in § C.3, practitioners widely expect LLMs to reduce or re- place entry-level (L1) responsibilities, while higher-skill roles in incident response, threat hunting, and governance remain human-driven. This shift positions analysts as supervisors of LLM-mediated workflows, presupposing substantial domain expertise. However, prior research emphasized that such ex- pertise develops incrementally through hands-on operational exposure and mentorship [79â82]. Lee et al. [83] observed that generative AI redirects cognitive effort toward verifica- tion and oversight rather than direct problem-solving. Our findings point to a similar transformation within SOCs, where analysts are increasingly positioned as reviewers and gover- nors of LLM-mediated workflows (§ 5.2, § C.3). This creates a circular dependency: effective oversight of LLMs requires experienced analysts, yet the experiential learning pathways that produce such expertise through entry- level tasks are now being automated. As a result, questions emerge about how future SOC analysts will acquire the experi- ential knowledge needed to critically evaluate LLM decisions, particularly in ambiguous or high-stakes scenarios. Research Opportunity. Sustaining expertise in LLM- augmented SOCs may require rethinking training as a process of co-learning between humans and LLMs. While machines have traditionally been framed as learning from humans, re- cent work has begun to formalize co-learning paradigms in which humans and AI mutually adapt through interaction, feedback, and shared problem solving [84, 85]. Early indus- try efforts are beginning to explore this space explicitly. For example, COACH by Dropzone [86] is an LLM-powered security mentor that exposes junior analysts to investigative reasoning, and contextual decision making, through real-time mentorship during live investigations [87]. In parallel, re- cent academic research demonstrates how LLM-powered tutoring can be tightly integrated into hands-on cybersecurity training environments [88â90], where learners develop skills through guided practice, feedback, and exposure to progres- sively complex scenarios. Future research can explore how co- learning approaches can be grounded in long-standing cyber workforce development principles [81], while supporting scaf- folded learning through LLM-mediated skill development. 12 References [1] Faris Bugra Kokulu, Ananta Soneji, Tiffany Bao, Yan Shoshitaishvili, Ziming Zhao, Adam DoupĂ©, and Gail- Joon Ahn. Matched and Mismatched SOCs: A Quali- tative Study on Security Operations Center Issues. In Proc. of the 2019 ACM SIGSAC Conf. on Computer and Communications Security, pages 1955â1970, Lon- don United Kingdom, November 2019. ACM. [2]Manfred Vielberth, Fabian Böhm, Ines Fichtinger, and GĂŒnther Pernul. Security Operations Center: A Sys- tematic Study and Open Challenges. IEEE Access, 8:227756â227779, 2020. [3] Jenny Hofbauer and Kevin Mayer. Blue Team Fun- damentals: Roles and Tools in a Security Operations Center. In The 18th International Conf. on Emerging Security Information, Systems and Technologies, pages 176â184. IARIA, 2024. [4]S Sreelakshmi, A Aalan Babu, C Lakshmipriya, LA Anto Gracious, M Nalini, and R Siva Subrama- nian. Enhancing intrusion detection systems with ma- chine learning. In 2024 2nd International Conf. on Self Sustainable Artificial Intelligence Systems (ICSSAS), pages 557â564. IEEE, 2024. [5] Giuseppina Andresini, Feargus Pendlebury, Fabio Pier- azzi, Corrado Loglisci, Annalisa Appice, and Lorenzo Cavallaro. Insomnia: Towards concept-drift robustness in network intrusion detection. In Proc. of AISecâ21, pages 111â122, 2021. [6]Nicola Capuano, Giuseppe Fenza, Vincenzo Loia, and Claudio Stanzione. Explainable artificial intelligence in cybersecurity: A survey. Ieee Access, 10:93575â 93600, 2022. [7] Farid Binbeshr, Muhammad Imam, Mustafa Ghaleb, Mosab Hamdan, Mussadiq Abdul Rahim, and Moham- mad Hammoudeh. The Rise of Cognitive SOCs: A Systematic Literature Review on AI Approaches. IEEE Open Journal of the Computer Society, 2025. [8]Jalal Ghadermazi, Ankit Shah, and Sushil Jajodia. A machine learning and optimization framework for effi- cient alert management in a cybersecurity operations center. Digital Threats: Research and Practice, 5(2):1â 23, 2024. [9]2023 Voice of the SOC.https://w.tines.com/ reports/voice-of-the-soc-2023/. [10] 2023Stateof Threat Detection.https: //w.vectra.ai/resources/2023-state- of-threat-detection. [11]Bushra A Alahmadi, Louise Axon, and Ivan Marti- novic. 99% false positives: A qualitative study of SOCanalystsâ perspectives on security alarms. In 31st USENIX Security 22, pages 2783â2800, 2022. [12]Wajih Ul Hassan, Shengjian Guo, Ding Li, Zhengzhang Chen, Kangkook Jee, Zhichun Li, and Adam Bates. Nodoze: Combatting threat alert fatigue with auto- mated provenance triage. In network and distributed systems security symposium, 2019. [13]Diana Kramer, Lambert Rosique, Ajay Narotam, Elie Bursztein, Patrick Gage Kelley, Kurt Thomas, and Al- lison Woodruff. Integrating large language models into security incident response. In SOUPS 2025, pages 133â148, 2025. [14]Ronal Singh, Shahroz Tariq, Fatemeh Jalalvand, Mo- han Baruwal Chhetri, Surya Nepal, Cecile Paris, and Martin Lochner. LLMs in the SOC: An empirical study of human-ai collaboration in security operations centres. arXiv:2508.18947, 2025. [15]Alsharif Abuadbba, Chris Hicks, Kristen Moore, Vasil- ios Mavroudis, Burak Hasircioglu, Diksha Goel, and Piers Jennings. From Promise to Peril: Rethinking Cybersecurity Red and Blue Teaming in the Age of LLMs. arXiv:2506.13434, 2025. [16] Xihuan Lin, Jie Zhang, Gelei Deng, Tianzhe Liu, Xi- aolong Liu, Changcai Yang, Tianwei Zhang, Qing Guo, and Riqing Chen.IRCopilot: Automated Incident Response with Large Language Models. arXiv:2505.20945, 2025. [17]Xinye Tang, Amir H Abdi, Jeremias Eichelbaum, Ma- han Das, Alex Klein, Nihal Irmak Pakis, William Blum, Daniel L Mace, Tanvi Raja, Namrata Padmanabhan, et al. Nl2kql: From natural language to kusto query. arXiv:2404.02933, 2024. [18] Saleha Muzammil, Rahul Reddy, Vishal Kamalakrish- nan, Hadi Ahmadi, and Wajih Ul Hassan. Towards Small Language Models for Security Query Genera- tion in SOC Workflows. arXiv:2512.06660, 2025. [19]Charupriya Bisht and Anurag Jain. Improving Cyberse- curity Decision-Making Through Text Summarization Addressing Key Applications and Overcoming Chal- lenges. In 2025 International Conf. on Networks and Cryptology, pages 1544â1550. IEEE, 2025. [20]Walaa Saber Ismail. Threat detection and response using AI and NLP in cybersecurity. J. Internet Serv. Inf. Secur, 14(1):195â205, 2024. 13 [21]Microsoft Security. Microsoft Copilot for Security is generally available on April 1, 2024, with new capabilities.https://w.microsoft.com/en- us/security/blog/2024/03/13/microsoft- copilot-for-security-is-generally- available-on-april-1-2024-with-new- capabilities/, 2024. [22] Crowdstrikecharlotteai.https://w. crowdstrike.com/en-us/platform/charlotte- ai/. [23] Nir Kshetri. Transforming cybersecurity with agentic ai to combat emerging cyber threats. Telecommunica- tions Policy, page 102976, 2025. [24]Jaron Mink, Hadjer Benkraouda, Limin Yang, Arrid- hana Ciptadi, Ali Ahmadzadeh, Daniel Votipka, and Gang Wang. Everybodyâs got ML, tell me what else you have: Practitionersâ perception of ML-based secu- rity tools and explanations. In IEEE Symp. on Security and Privacy (SP), pages 2068â2085. IEEE, 2023. [25] Sean Oesch, Robert Bridges, Jared Smith, Justin Beaver, John Goodall, Kelly Huffer, Craig Miles, and Dan Scofield. An assessment of the usability of ma- chine learning based tools for the security operations center. In 2020 International Conf.s on Internet of Things (iThings), pages 634â641. IEEE, 2020. [26] Neele Roch, Hannah Sievers, Lorin Schöni, and Verena Zimmermann. Navigating autonomy: unveiling secu- rity expertsâ perspectives on augmented intelligence in cybersecurity. In SOUPS 2024, pages 41â60, 2024. [27] Nidhi Rastogi, Devang Dhanuka, Amulya Saxena, Pranjal Mairal, and Le Nguyen. Measuring the Security and Cognitive Impacts of Explainability in AI-Driven SOCs. arXiv:2503.02065, 2025. [28] Elijah Bouma-Sims, Hiba Hassan, Alexandra Nisenoff, Lorrie Faith Cranor, and Nicolas Christin. "It was honestly just gambling": Investigating the Experiences of Teenage Cryptocurrency Users on Reddit. In Symp. on Usable Privacy and Security, pages 333â352, 2024. [29]Rajvardhan Oak and Zubair Shafiq. Victims, vigilantes, and advice givers: An analysis ofScam-Relateddis- course on reddit. In SOUPS 2025, pages 57â71, 2025. [30]Jaakko VĂ€kevĂ€, Perttu HĂ€mĂ€lĂ€inen, and Janne Lindqvist. " donât you dare go hollow": How dark souls helps players cope with depression, a thematic analysis of reddit discussions. In Proc. of the 2025 CHI Conf. on Human Factors in Computing Systems, pages 1â20, 2025. [31]Elijah Bouma-Sims, Mandy Lanyon, and Lorrie Faith Cranor. "Is this a scam?": The Nature and Quality of Reddit Discussion about Scams. In Proc. of the 2025 ACM SIGSAC Conf. on Computer and Communica- tions Security, pages 2444â2458, 2025. [32]Muhammad Mudassar Yamin, Ehtesham Hashmi, Mo- hib Ullah, and Basel Katt. Applications of llms for generating cyber security exercise scenarios. IEEE Access, 2024. [33]Mohamed Amine Ferrag, Fatima Alwahedi, Ammar Battah, Bilel Cherif, Abdechakour Mechri, and Norbert Tihanyi. Generative ai and large language models for cyber security: All insights you need. Available at SSRN 4853709, 2024. [34]Massimiliano Albanese, Xinming Ou, Kevin Lybarger, Daniel Lende, and Dmitry Goldgof. Towards ai-driven human-machine co-teaming for adaptive and agile cy- ber security operation centers.arXiv:2505.06394, 2025. [35]Bowen Wei, Yuan Shen Tay, Howard Liu, Jinhao Pan, Kun Luo, Ziwei Zhu, and Chris Jordan.Cortex: Collaborative llm agents for high-stakes alert triage. arXiv:2510.00311, 2025. [36]Lauren Deason, Adam Bali, Ciprian Bejean, Diana Bolocan, James Crnkovich, Ioana Croitoru, Krishna Durai, Chase Midler, Calin Miron, David Molnar, et al. Cybersoceval: Benchmarking llms capabilities for malware analysis and threat intelligence reasoning. arXiv:2509.20166, 2025. [37]Jie Zhu, Chihao Shen, Ziyang Li, Jiahao Yu, Yizheng Chen, and Kexin Pei. Locus: Agentic predicate synthe- sis for directed fuzzing. arXiv:2508.21302, 2025. [38]Yinlin Deng, Chunqiu Steven Xia, Chenyuan Yang, Shizhuo Dylan Zhang, Shujing Yang, and Lingming Zhang.Large language models are edge-case fuzzers: Testing deep learning libraries via fuzzgpt. arXiv:2304.02014, 2023. [39]Limin Yang, Zhi Chen, Chenkai Wang, Zhenning Zhang, Sushruth Booma, Phuong Cao, Constantin Adam, Alexander Withers, Zbigniew Kalbarczyk, Rav- ishankar K Iyer, et al. True attacks, attack attempts, or benign triggers? an empirical measurement of network alerts in a security operations center. In 33rd USENIX Security, pages 1525â1542, 2024. [40] Sathya Chandran Sundaramurthy, Alexandru G Bar- das, Jacob Case, Xinming Ou, Michael Wesch, John McHugh, and S Raj Rajagopalan. A human capi- tal model for mitigating security analyst burnout. In SOUPS 2015, pages 347â359, 2015. 14 [41]Kashyap Thimmaraju, Sybe Izaak Rispens, and Gail- Joon Ahn. Human performance in security opera- tions: a survey on burnout, well-being and flow state among practitioners. In Proc. 2025 Workshop on Secu- rity Operations Center Operations and Construction (WOSOC 2025), pages 2â4, 2025. [42]Subigya Nepal, Javier Hernandez, Robert Lewis, Ahad Chaudhry, Brian Houck, Eric Knudsen, Raul Rojas, Ben Tankus, Hemma Prafullchandra, and Mary Czer- winski. Burnout in cybersecurity incident responders: exploring the factors that light the fire. Proc. of ACM Human-Computer Interaction, pages 1â35, 2024. [43]Mohan Baruwal Chhetri, Shahroz Tariq, Ronal Singh, Fatemeh Jalalvand, Cecile Paris, and Surya Nepal. To- wards human-ai teaming to mitigate alert fatigue in security operations centres. ACM Transactions on In- ternet Technology, 24(3):1â22, 2024. [44]Ahmad Mohsin, Helge Janicke, Ahmed Ibrahim, Iqbal H Sarker, and Seyit Camtepe. A unified frame- work for human ai collaboration in security operations centers with trusted autonomy. arXiv:2505.23397, 2025. [45]Masike Malatji. Augmented intelligence framework for humanâartificial intelligence teaming in cybersecu- rity. Human-Centric Intelligent Systems, 2025. [46]Neele Roch. Building trust bridges: The dynamics of autonomy and transparency in expert-ai collaboration for cybersecurity. In Companion Proc. of the 30th In- ternational Conf. on Intelligent User Interfaces, pages 202â204, 2025. [47]GPT-4.1 mini.https://platform.openai.com/ docs/models/gpt-4.1-mini. [48]Ivor A Pritchard. Searching for "research involving human subjects": What is examined? what is exempt? what is exasperating? IRB: Ethics & Human Research, 23(3):5â13, 2001. [49] Shruti Sannon, Billie Sun, and Dan Cosley. Privacy, surveillance, and power in the gig economy. In Proc. of the 2022 CHI Conf. on human factors in computing systems, pages 1â15, 2022. [50]r/Drugs Subreddit.https://w.reddit.com/r/ Drugs/. Accessed: 01-13-2026. [51] Casey Fiesler, Michael Zimmer, Nicholas Proferes, Sarah Gilbert, and Naiyan Jones. Remember the hu- man: A systematic review of ethical considerations in reddit research. Proc. of the ACM on Human-Computer Interaction, 8(GROUP):1â33, 2024. [52]Joseph Reagle. Disguising reddit sources and the effi- cacy of ethical research. Ethics and Information Tech- nology, 24(3):41, 2022. [53] James Mattei, Christopher Pellegrini, Matthew Soto, Marina Sanusi Bohuk, and Daniel Votipka. "Iâm trying to learn... and Iâm shooting myself in the foot": Be- ginnersâ Struggles When Solving Binary Exploitation Exercises. In USENIX, 2025. [54]Theophilus Azungah. Qualitative Research: Deductive and Inductive Approaches to Data Analysis. Qualita- tive Research Journal, 18(4):383â400, 2018. [55]Andrew F Hayes and Klaus Krippendorff. Answering the call for a standard reliability measure for coding data. Communication methods and measures, 1(1):77â 89, 2007. [56] Nora McDonald, Sarita Schoenebeck, and Andrea Forte. Reliability and inter-rater reliability in quali- tative research: Norms and guidelines for cscw and hci practice. Proc. of the ACM on human-computer interaction, 3(CSCW):1â23, 2019. [57]Virginia Braun and Victoria Clarke. Thematic analysis: A practical guide. 2021. [58]PennState Eberly College of Science. Comparing two proportions.https://online.stat.psu.edu/ stat415/lesson/9/9.4. [59]Cornell Statistical Consulting Unit. Using adjusted standardized residuals for interpreting contingency ta- bles, 2020. [60]Eric W Weisstein. Bonferroni correction.https: //mathworld.wolfram.com/, 2004. [61]Jan H Klemmer, Stefan Albert Horstmann, Nikhil Pat- naik, Cordelia Ludden, Cordell Burton Jr, Carson Pow- ers, Fabio Massacci, Akond Rahman, Daniel Votipka, Heather Richter Lipford, et al. Using ai assistants in software development: A qualitative study on security practices and concerns. In Proceedings of the 2024 on ACM SIGSAC Conf. on Computer and Communica- tions Security, pages 2726â2740, 2024. [62] Prophet Security launches with an Agentic AI SOC An- alyst.https://w.prophetsecurity.ai/blog/ announcing-prophet-security, 2024. [63]Gemini in Google SecOps.https://docs.cloud. google.com/chronicle/docs/secops/release- notes#March_26_2024, 2024. [64]Juyong Jiang, Fan Wang, Jiasi Shen, Sungju Kim, and Sunghun Kim. A survey on large language models for code generation. arXiv:2406.00515, 2024. 15 [65]Jianxun Wang and Yixiang Chen. A review on code generation with llms: Application and evaluation. In 2023 IEEE International Conf. on Medical Artificial Intelligence (MedAI), pages 284â289. IEEE, 2023. [66]Andrei Sobo, Awes Mubarak, Almas Baimagambe- tov, and Nikolaos Polatidis. Evaluating llms for code generation in hri: A comparative study of chatgpt, gemini, and claude. Applied Artificial Intelligence, 39(1):2439610, 2025. [67] Jia-Yu Yao, Kun-Peng Ning, Zhen-Hui Liu, Mu-Nan Ning, Yu-Yang Liu, and Li Yuan. Llm lies: Hallucina- tions are not bugs, but features as adversarial examples. arXiv:2310.01469, 2023. [68]Berk Atil, Sarp Aykent, Alexa Chittams, Lisheng Fu, Rebecca J Passonneau, Evan Radcliffe, Guru Rajan Rajagopal, Adam Sloan, Tomasz Tudrej, Ferhan Ture, et al. Non-determinism of deterministic llm settings. arXiv:2408.04667, 2024. [69] Yifan Song, Guoyin Wang, Sujian Li, and Bill Yuchen Lin. The good, the bad, and the greedy: Evaluation of llms should not ignore non-determinism. In Proc. of the 2025 Conf. of the Nations of the Americas Chapter of the Association for Computational Linguistics: Hu- man Language Technologies (Volume 1: Long Papers), pages 4195â4206, 2025. [70] Junjie Chu, Yugeng Liu, Ziqing Yang, Xinyue Shen, Michael Backes, and Yang Zhang. Jailbreakradar: Comprehensive assessment of jailbreak attacks against llms. In Proc. of the 63rd Annual Meeting of the Asso- ciation for Computational Linguistics (Volume 1: Long Papers), pages 21538â21566, 2025. [71] Yi Liu, Gelei Deng, Yuekang Li, Kailong Wang, Zihao Wang, Xiaofeng Wang, Tianwei Zhang, Yepang Liu, Haoyu Wang, Yan Zheng, et al. Prompt injection attack against llm-integrated applications. arXiv:2306.05499, 2023. [72]Gartner Magic Quadrant.https://w.gartner. com/en/research/methodologies/magic- quadrants-research. [73] LLM Inference Benchmarking: How Much Does Your LLM Inference Cost?https://developer.nvidia. com/blog/llm-inference-benchmarking-how- much-does-your-llm-inference-cost/, 2025. [74] Jens Opdenbusch, Jonas Hielscher, and M Angela Sasse. "Where Are We On Cyber?" A Qualitative Study On Boardsâ Cybersecurity Risk Decision Mak- ing. In NDSS, 2025. [75]Stef Schinagl and Abbas Shahim. What do we know about information security governance? âfrom the base- ment to the boardroomâ: towards digital security gov- ernance. Information & Computer Security, 28(2):261â 292, 2020. [76]Georgios Syros, Anshuman Suri, Jacob Ginesin, Cristina Nita-Rotaru, and Alina Oprea. Saga: A se- curity architecture for governing ai agentic systems. arXiv:2504.21034, 2025. [77]John D Lee and Katrina A See. Trust in automation: Designing for appropriate reliance. Human factors, 46(1):50â80, 2004. [78] Yavuz Selim Kıyak, Ăzlem Co ̧skun, and I ̧sıl Ì Irem Bu- dako Ì glu. âchatgpt can make mistakesâ warnings fail: A randomized controlled trial. Medical Education, 2025. [79]Jessica Dawson and Robert Thomson. The future cy- bersecurity workforce: Going beyond technical skills for successful cyber performance. Frontiers in psychol- ogy, 9:744, 2018. [80]John R Goodall, Wayne G Lutters, and Anita Kom- lodi. Developing expertise for network intrusion detec- tion. Information Technology & People, 22(2):92â108, 2009. [81]Marie Baker. Striving for effective cyber workforce development. Software Engineering Institute, pages 1â26, 2016. [82]Kashyap Thimmaraju, Duc Anh Hoang, Souradip Nath, Jaron Mink, and Gail-Joon Ahn. Before the vicious cy- cle starts: Preventing burnout across soc roles through flow-aligned design. arXiv preprint arXiv:2602.14598, 2026. [83]Hao-Ping Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks, and Nicholas Wil- son. The impact of generative ai on critical thinking: Self-reported reductions in cognitive effort and con- fidence effects from a survey of knowledge workers. In Proc. of the 2025 CHI Conf. on human factors in computing systems, pages 1â22, 2025. [84] Jinwei Lu, Yikuan Yan, Keman Huang, Ming Yin, and Fang Zhang. Do we learn from each other: Under- standing the human-ai co-learning process embedded in human-ai collaboration. Group Decision and Nego- tiation, 34(2):235â271, 2025. [85]Yi-Ching Huang, Yu-Ting Cheng, Lin-Lin Chen, and Jane Yung-jen Hsu. Human-ai co-learning for data- driven ai. arXiv:1910.12544, 2019. 16 [86]COACH: AI-Powered Security Alert Mentor for SOC Analysts.https://w.dropzone.ai/coach. Ac- cessed: 01-21-2026. [87] AI Is Automating SOC, But Can It Train the Next Gen- eration of Analysts?https://w.dropzone.ai/ blog/ai-soc-training-junior-analysts .Ac- cessed: 01-21-2026. [88]Connor Nelson, Adam DoupĂ©, and Yan Shoshitaishvili. Sensai: Large language models as applied cybersecu- rity tutors. In Proceedings of the 56th ACM Technical Symposium on Computer Science Education V. 1, pages 833â839, 2025. [89]Tianyu Wang, Nianjun Zhou, and Zhixiong Chen. Cy- bermentor: Ai powered learning tool platform to ad- dress diverse student needs in cybersecurity education. arXiv:2501.09709, 2025. [90]Chola Chhetri.Exploring large language model- powered pedagogical approaches to cybersecurity edu- cation. In Proc. of the 25th Annual Conf. on Informa- tion Technology Education, pages 163â166, 2024. [91] Who Were the Luddites?https://w.history. com/articles/who-were-the-luddites. [92] Chatgpt. https://chatgpt.com/. [93]Microsoft copilot.https://securitycopilot. microsoft.com/. [94] Claude. https://claude.ai/. [95] Gemini. https://gemini.google.com/app. [96] Llama. https://w.llama.com/. [97] Perplexity. https://w.perplexity.ai/. [98] Notebooklm. https://notebooklm.google/. [99] Grok. https://grok.com/. [100] Amazon q. https://aws.amazon.com/q/. [101] Microsoftsecuritycopilot.https:// securitycopilot.microsoft.com/. [102] Dropzone ai. https://w.dropzone.ai/. [103] Intezer. https://intezer.com/. [104] Cortex xsiam.https://w.paloaltonetworks. com/cortex/cortex-xsiam. [105] Prophet security.https://w.prophetsecurity. ai/. [106]Purple ai.https://w.sentinelone.com/ platform/purple/. [107] Cmd zero. https://w.commandzero.ai/. [108] Abnormal. https://abnormal.ai/. [109] Google secops.https://cloud.google.com/ security/products/security-operations. [110] Darktrace. https://w.darktrace.com/. [111] Torq socrates. https://torq.io/socrates/. [112] Qevlar ai. https://w.qevlar.com/. [113] Arcanna ai. https://w.arcanna.ai/. [114] Vectra ai. https://w.vectra.ai/. [115] Whiterabbitneo. https://w.deephat.ai/. [116] D3morpheus.https://d3security.com/ morpheus/. [117] Tandemtrace. https://tandemtrace.ai/. [118] Radiant security. https://radiantsecurity.ai/. [119] Exaforce. https://w.exaforce.com/. [120] 7ai. https://7ai.com/. [121] Rapid7. https://w.rapid7.com/. [122] Whistic. https://w.whistic.com/. [123]Splunk enterprisesecurity.https://w. splunk.com/en_us/products/enterprise- security.html. [124] Sirp. https://w.sirp.io/. [125] Hackerai (pentestgpt). https://hackerai.co/. [126] Xbow. https://xbow.com/. [127] Nebula ai.https://berylliumsec.github.io/ nebula/. [128]Gradient cyber.https://w.gradientcyber. com/. [129] Reliaquest. https://reliaquest.com/. 17 12-202203-202306-2023 09-2023 12-202303-202406-2024 09-2024 12-202403-202506-2025 09-2025 0 2 4 6 8 10 2 11 0 3 00 2 11 000 1 2 11 2 00 1 2 1 44 3 4 10 8 6 8 4 3 0 Month-Year #Threads Figure 2: Number of Threads in Our Dataset Over Time. A Data Collection and Curation The following is the list of keywords used to find potential relevant threads (§ 3.1): âSOC AIâ,âSOC AI Agentsâ,âSOC LLMâ,âAI-powered SOCâ, âAgentic AI SOCâ,âautonomous SOCâ,âAI SOC Analystâ, âAI Security Operationsâ,âAI cybersecurityâ,âLLM in cybersecurity â, âAI augmenting cybersecurityâ, âAI agents cybersecurityâ, âLLM for threat detectionâ As discussed in § 3.1, the following few-shot prompt was used to guide thegpt-4.1-minimodel, with temperature set to 0, in performing the post-level relevancy check task. You are an expert annotator working on a research project about the use of AI tools in Security Operations Centers (SOCs). You are given the title and content of a Reddit post. Your task is to classify each post as either: **Relevant:** if the post reflects a personâs opinion, experience, discussion, or inquiry about the use of AI tools in SOC workflows. **Not Relevant:** if the post is about anything else, e.g., AI security frameworks, general AI for cybersecurity trends, SOC compliance, hiring, or any topic not centered on the use of AI tools in SOCs. Use only the title and post content. Do not hallucinate missing details. Output should be just: **Relevant** or **Not Relevant**. Here are some examples: <Two relevant, and one irrelevant example> Now classify the following: **Title:** <title> **Content:** <content> **Output:** Figure 2 shows the temporal frequency distribution of the threads between December 2022 and September 2025. 0 50 100 150 #Posts AI Use Cases *** p<.001, ** p<.01, * p<.05, NS n.s. (Bonferroni) 139 88 84 54 53 18 *************** NS ******* ***** NS *** *** Triage & IR Scripting Reporting Threat Analysis Knowledge Support Misc. Figure 3: Relative Prevalence across Reported AI Use Cases â Triage & IR (n = 139), Scripting & Query Support (n = 88), Reporting & Documentation (n = 84), Threat Analysis (n = 54), Knowledge Support (n = 53), and Miscellaneous (n = 18). B Codebook Table 6 displays the complete codebook used in our thematic analysis, documenting all primary codes and subcodes, along with their definitions and observed frequencies. C Additional Results and Discussion C.1 Commercial LLM Tools Table 7 provides a comprehensive list of all commercial LLM tools, along with their purpose (general-purpose or security- specific), observed frequencies and brief description, men- tioned in the practitioner discussions. C.2Relative Prevalence of Reported Use Cases As briefly discussed in Section 5.2, to understand the rel- ative prevalence across the six types of use cases (triage and incident response, scripting, reporting, threat analysis, and misc.), we conducted pairwise two-sample tests of pro- portions with a Bonferroni correction. As depicted in Fig- ure 3, triage and IR was significantly more prevalent than all other reported AI use cases (Ï 2 (1)â„ 16.92, allp<.001). Following triage and IR, writing-oriented tasks, including scripting, query generation, and reporting, were the next most frequently discussed AI use cases. Statistically, scripting (n = 88) and reporting (n = 84) did not differ significantly from one another (Ï 2 (1) = 0.07, p =.79), but were both sig- nificantly higher than threat analysis and knowledge support (Ï 2 (1)â„ 7.74, p< 0.05). Lastly, the discussions around threat 18 Primary CodeSub CodeFreq.Description Tools Mentioned (α = 0.87) General-Purpose LLM Tools248General-purpose LLM tools not specifically designed for SOC (e.g., ChatGPT). Security-Focused LLM Tools182Commercial or custom-built LLM-powered security products (e.g., Sec Copilot). Use Cases Mentioned Triage & Incident Response (α= 0.84)139Use of LLM tools to support alert triage, investigation, detection & response. Scripting & Query Support (α = 0.95)88Use of LLMs to generate, or debug scripts, detection logic, or data queries. Reporting & Documentation (α= 0.87)84Use of LLMs for drafting reports, summaries, and procedural documentation. Threat Analysis (α = 1.00)54Use of LLM tools for threat modeling, and vulnerability assessment. Knowledge Support (α = 0.84)53Use of LLMs for learning, explanation, and information retrieval. Miscellaneous (α = 0.79)18Use of LLMs for other tasks such as compliance, audit, and training simulation. LLM Factors Effectiveness (α = 0.86)272Discussions on whether LLM tools can effectively augment SOC workflows. Efficiency (α = 1.00)77Discussions on time savings or inefficiencies arising from LLM-use in SOC. Security & Privacy (α = 1.00)74Discussions on privacy, and security implications of LLM tools. Reliability (α = 0.81)69Discussions on the trustworthiness of LLM-generated outputs. Autonomy (α = 0.88)65Discussions around the degree to which LLM tools act independently. Cost (α = 0.88)49Discussions about financial considerations of LLM tools. Opinions on LLM (α = 0.83) Positive214 Expressions of support, enthusiasm, or favorable experiences using LLM tools. Negative243Expressions of concerns, or unfavorable experiences with LLM tools. LLM Adoption Stage (α = 0.84) Already Using200Mentions indicating active use of LLM tools within SOC workflows. Curious or Actively Evaluating123Mentions reflecting interest in or active evaluation of LLM tools for SOC use. Not Using50Explicit statements of non-adoption of LLM tools for SOC purposes. Future Predictions (α = 1.00)â276Discussions about anticipated future impacts of LLM on SOC workflows. Irrelevant (α = 0.88)â811Comments excluded from analysis due to the lack of relevance. Table 6: Analysis Codebook â summarizing the codes used in the thematic analysis, along with their frequencies and descriptions. analysis and knowledge support did not differ significantly (Ï 2 (1) = 0.00, p = 1.00), but both were greater than misc. C.3 Impact of LLMs on Security Workforce Building on observed adoption patterns and barriers (§ 7.1, § 7.2), practitionersâ discussions also reflect on the implica- tions of LLMs for the SOC workforce. Despite vendor claims of autonomy and replacement, practitioner discussions reveal a more nuanced view of how LLMs reshape roles, responsi- bilities, and skill demands within SOCs. Where LLMs Can Meaningfully Augment Humans. Across practitioner discussions (n=28), a widely shared be- lief was that if LLM tools were to replace any SOC role, L1 responsibilities are the most vulnerable. Practitioners ar- gued that L1 positions were already being reduced prior to the LLM hype and that the widespread adoption of LLMs is likely to accelerate this trend. For example, as one prac- titioner, P250, explained how their company âlet go of all eight L1 SOC members, because the SOAR playbooks han- dled almost everything, and phishing, the only task manually reviewed, was later handed off to an AI tool.â Moreover, sev- eral posts frame L1 workflows as âbutton-clickingâ (P916), âbrain-dead workâ (P244), âbarely a security roleâ (P237), or âsimply processing routine tasks shown by the SIEMâ (P304), justifying that LLM tools need not human-level reasoning to affect L1 staffing. Building on this reasoning, while some practitioners predicted substantial reductions in entry-level positions, e.g., âthe lowest tier of analysts will not exist in 5 yearsâ (P242), others anticipated a more moderate contrac- tion, with smaller L1 teams retained for budgetary, workload, or escalation needs. As P227 estimated, âAI could impact headcount by 15â20% . . . but replacing half of L1 staff or more is doubtful.â Where Human Expertise Remains Critical. Despite con- cerns about the future of L1 roles, practitioners overwhelm- ingly rejected the idea that SOCs are close to becoming fully autonomous. Building on the challenges with LLM autonomy (§ 6.5), discussions consistently emphasized that within secu- rity operations, human oversight remains indispensable, even as LLM tools become deeply embedded in workflows. High-skill SOC Responsibilities: Across 20 posts, practi- tioners emphasized that many high-skill SOC responsibilities inherently require humans, regardless of future LLM advance- ments. Digital forensics and IR, along with threat hunting and penetration testing (P711, P367, P539, P927), responsibilities typically handled by L2âL3 teams, were especially cited as remaining human-driven for the foreseeable future. These tasks were described as complex, requiring âcreativity, and intuitionâ (P212), which practitioners viewed as enduringly human. As P279 emphasized, âAny task that requires com- plex reasoning, logical synthesis and judgment, I see people having a strong presence in handling.â This perspective was further reinforced by stressing the criticality of human exper- tise: âEven the most advanced AI tools are ineffective without a skilled security team to implement them or without strong executive backing from a CISO (or equivalent)â (P631). LLM Supervision and Governance: Another set of posts (n=18) highlighted that with LLMs in the picture, humans are needed more than ever. Practitioners consistently pointed 19 out that organizations will still need humans to âsupervise and verify LLM outputsâ (P005, P243, P401, P390). As P243 summarized, âEven with LLMs, there will still be a need for certain levels of verification, which in itself could be an L1 responsibility.â Practitioners further stressed that âcertain re- sponsibilities, such as governance and compliance, cannot be delegated to AI, as doing so would be the fox watching the hen houseâ (P241). They also argued that within SOC, AI cannot operate without human-provided context. As P354 explained, even someone skilled at prompting requires substantial prior experience to guide the LLM: If a company hires someone who can generate solutions through effective prompting, that alone does not make them a replacement for skilled ana- lysts. Meaningful use of LLMs still requires domain knowledge and expertise of an experienced analyst. Accountability: Lastly, across a small set of posts (n=6), practitioners highlighted humans will still be needed for ac- countability, arguing that organizations cannot solely rely on LLM tools for decisions that may have legal, regulatory, or financial consequences (P048, P237, P285). This was espe- cially evident for incident response workflows. As P347 noted, âUnsure what a breach response would look like if the âAI em- ployeeâ overlooked anything and it caused harm to people... some âhumanâ will ultimately need to be held responsible.â Overall, these arguments support that the role of LLM tools in the SOC will be augmentative, reshaping how analysts work rather than removing them from the equation. How Analysts Must Adapt. Beyond the debate between replacement and augmentation, a few posts (n = 11) high- lighted that avoiding LLMs or dismissing its relevance by âburying head in the sandâ (P017) should be an untenable stance. Practitioners argued that early adopters of LLM tools are already integrating them into daily workflows, and there- fore dismissing this overall shift by being one of the âdeniers and luddites [91]â (P292) would be unwise. As noted by P275, âWe would be naĂŻve to overlook the scale and speed of change underway, no one can predict with certainty how the field will evolve.â Consequently, practitioners offered explicit advice to fellow peers on how to navigate this shift. Developing AI Literacy: A commonly repeated suggestion was developing AI literacy, followed by the sentiment that âAI itself is not the reason for job-threat; rather, it is the peers who learn to use it effectivelyâ (P248). Several, including P246, P818, and P843, described âAI as a bell that cannot be unrung,â (P246) stressing that analysts who fail to build fluency will be the first to fall behind as organizations increas- ingly seek people who can work confidently with AI-enabled tooling. P292 contextualized this urgency by pointing out to the unprecedented pace of LLM adoption, noting that, âonly few technologies in recent history have achieved such rapid global familiarity and enterprise adoption.â More Depth in Security Reasoning and Response: Another category of advice emphasized the importance of strengthen- ing fundamentals and continuous upskilling. As one practi- tioner, P304, advised, âTo anyone aspiring, make upskilling a part of your DNA. Go beyond simply responding to alerts to actually understand why detections fire and how systems operate.â Others also underscored the ongoing importance of âhands-on experience with traditional SOC tools such as SIEM and EDRâ (P034), as well as familiarity with emerging areas such AI security threats, e.g., ISO 27090 3 (P271). P292 further reiterated this importance by describing how a deep grounding in security concepts and system behavior forms the foundation for understanding why recent advances in LLMs represent such a dramatic shift. 3 https://w.iso.org/standard/56581.html 20 Tool NameCategoryFrequencyDescription ChatGPT [92]General-Purpose83A general-purpose conversational LLM developed by OpenAI Microsoft Copilot [93]General-Purpose11A generative AI assistant integrated across Microsoft products Claude [94]General-Purpose8A conversationsal LLM developed by Anthropic Gemini [95]General-Purpose6A multimodal LLM developed by Google Llama [96]General-Purpose5A family of open-source large language models released by Meta Perplexity [97]General-Purpose3A generative AI-powered web search assistant NotebookLM [98]General-Purpose3A research and note-taking online tool powered by Google Gemini Grok [99]General-Purpose2A conversational LLM developed by xAI Amazon Q [100]General-Purpose1 Amazonâs enterprise-focused generative AI assistant for AWS customers Microsoft Security Copilot [101]Security-Specific40Microsoftâs LLM-powered agentic security automation platform Dropzone AI [102]Security-Specific10Autonomous LLM-powered agentic âAI SOC Analystâ platform Intezer [103]Security-Specific8Autonomous LLM-powered agentic âAI SOC Analystâ platform Palo Altoâs Cortex XSIAM [104]Security-Specific6Extended security intelligence and automation management platform Prophet Security [105]Security-Specific4Autonomous LLM-powered agentic âAI SOC Analystâ platform Purple AI [106]Security-Specific4Autonomous LLM-powered agentic âAI SOC Analystâ platform CMD Zero [107]Security-Specific3Autonomous & AI-assisted Cyber Investigation Platform Abnormal [108]Security-Specific3AI-Native Platform for Human Behavior Security Google SecOps [109]Security-Specific3Googleâs intelligence-driven security operations platform Darktrace [110]Security-Specific3AI-powered proactive cybersecurity platform for enterprise security Torq Socrates [111]Security-Specific2Autonomous LLM-powered agentic âAI SOC Analystâ platform Qevlar AI [112]Security-Specific2Autonomous LLM-powered agentic âAI SOC Analystâ platform Arcanna AI [113]Security-Specific2Trustworthy agentic âAI SOC Analystâ platform Vectra AI [114]Security-Specific2AI-powered platform for network, identity, and cloud security WhiterabbitNeo [115]Security-Specific2Cybersecurity model built for offensive reasoning D3 Morpheus [116]Security-Specific1Autonomous LLM-powered agentic âAI SOC Analystâ platform TandemTrace [117]Security-Specific1Autonomous LLM-powered agentic âAI SOC Analystâ platform Radiant Security [118]Security-Specific1Autonomous LLM-powered agentic âAI SOC Analystâ platform CrowdStrike Charlotte AI [22]Security-Specific1Autonomous LLM-powered agentic âAI SOC Analystâ platform Exaforce [119]Security-Specific1Autonomous LLM-powered agentic âAI SOC Analystâ platform 7ai [120]Security-Specific1Autonomous LLM-powered agentic âAI SOC Analystâ platform Rapid7 [121]Security-Specific1AI-powered MDR platform for business resilience Whistic [122]Security-Specific1AI-First Platform for Comprehensive Third-Party Risk Management Splunk Enterprise Security [123]Security-Specific1AI-powered threat detection, investigation, and response platform SIRP [124]Security-Specific1Autonomous LLM-powered agentic âAI SOC Analystâ platform HackerAI (PentestGPT) [125]Security-Specific1LLM-powered penetration testing platform XBOW [126]Security-Specific1LLM-powered penetration testing platform Nebula AI [127]Security-Specific1LLM-powered penetration testing platform Gradient Cyber [128]Security-Specific1AI-assisted MXDR designed for mid-market organizations ReliaQuest [129]Security-Specific1Autonomous LLM-powered agentic âAI SOC Analystâ platform Table 7: Comprehensive list of named, commercially available LLM tools mentioned in the dataset, categorized as general- purpose or security-focused, along with their observed frequencies. 21