Paper deep dive
Supporting Artifact Evaluation with LLMs: A Study with Published Security Research Papers
David Heye, Karl Kindermann, Robin Decker, Johannes Lohmöller, Anastasiia Belova, Sandra Geisler, Klaus Wehrle, Jan Pennekamp
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 93%
Last extracted: 3/13/2026, 12:28:33 AM
Summary
The paper introduces an LLM-driven toolkit designed to automate and support the Artifact Evaluation (AE) process in cybersecurity research. The toolkit consists of three stages: 'Rate' (text-based reproducibility scoring), 'Prepare' (autonomous sandboxed execution environment setup), and 'Assess' (detection of methodological pitfalls). The system achieves over 72% accuracy in reproducibility classification and significantly reduces reviewer workload by automating environment setup and identifying common research flaws.
Entities (6)
Relation Signals (4)
Large Language Models â supports â Artifact Evaluation
confidence 95% · In this work, we demonstrate that Large Language Models (LLMs) can provide powerful support for AE tasks
Rate â partof â Artifact Evaluation Toolkit
confidence 90% · We introduce an Large Language Model-driven toolkit that analyzes paper texts and accompanying artifacts
Prepare â partof â Artifact Evaluation Toolkit
confidence 90% · We introduce an Large Language Model-driven toolkit that analyzes paper texts and accompanying artifacts
Assess â partof â Artifact Evaluation Toolkit
confidence 90% · We introduce an Large Language Model-driven toolkit that analyzes paper texts and accompanying artifacts
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Artifact Evaluation (AE) is essential for ensuring the transparency and reliability of research, closing the gap between exploratory work and real-world deployment is particularly important in cybersecurity, particularly in IoT and CPSs, where large-scale, heterogeneous, and privacy-sensitive data meet safety-critical actuation. Yet, manual reproducibility checks are time-consuming and do not scale with growing submission volumes. In this work, we demonstrate that Large Language Models (LLMs) can provide powerful support for AE tasks: (i) text-based reproducibility rating, (ii) autonomous sandboxed execution environment preparation, and (iii) assessment of methodological pitfalls. Our reproducibility-assessment toolkit yields an accuracy of over 72% and autonomously sets up execution environments for 28% of runnable cybersecurity artifacts. Our automated pitfall assessment detects seven prevalent pitfalls with high accuracy ($F_1$ > 92%). Hence, the toolkit significantly reduces reviewer effort and, when integrated into established AE processes, could incentivize authors to submit higher-quality and more reproducible artifacts. IoT, CPS, and cybersecurity conferences and workshops may integrate the toolkit into their peer-review processes to support reviewers' decisions on awarding artifact badges, improving the overall sustainability of the process.
Tags
Links
- Source: https://arxiv.org/abs/2603.06862v1
- Canonical: https://arxiv.org/abs/2603.06862v1
Trouble viewing inline? Open PDF directly â
Full Text
58,002 characters extracted from source content.
Expand or collapse full text
Supporting Artifact Evaluation with LLMs: A Study with Published Security Research Papers David Heye1,2, Karl Kindermann1,2, Robin Decker1,2, Johannes Lohmöller1, Anastasiia Belova3, Sandra Geisler3, Klaus Wehrle1, Jan Pennekamp1 2 These authors contributed equally to this work. Abstract Artifact Evaluation is essential for ensuring the transparency and reliability of research, closing the gap between exploratory work and real-world deployment is particularly important in cybersecurity, particularly in IoT and CPSs, where large-scale, heterogeneous, and privacy-sensitive data meet safety-critical actuation. Yet, manual reproducibility checks are time-consuming and do not scale with growing submission volumes. In this work, we demonstrate that Large Language Models can provide powerful support for AE tasks: (i) text-based reproducibility rating, (i) autonomous sandboxed execution environment preparation, and (i) assessment of methodological pitfalls. Our reproducibility-assessment toolkit yields an accuracy of over 72 and autonomously sets up execution environments for 28 of runnable cybersecurity artifacts. Our automated pitfall assessment detects seven prevalent pitfalls with high accuracy (> 92F_1> 92). Hence, the toolkit significantly reduces reviewer effort and, when integrated into established AE processes, could incentivize authors to submit higher-quality and more reproducible artifacts. IoT, CPS, and cybersecurity conferences and workshops may integrate the toolkit into their peer-review processes to support reviewersâ decisions on awarding artifact badges, improving the overall sustainability of the process. I Introduction ©2025 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses, in any current or future media, including reprinting/republishing this material for advertising or promotional purposes, creating new collective works, for resale or redistribution to servers or lists, or reuse of any copyrighted component of this work in other works. The rapid evolution of cyber threats poses a significant challenge for maintaining resilience across many networked domains, including Internet of Things and Cyber-Physical System deployments, industrial control systems, connected vehicles, and smart cities. Recent reports by the World Economic Forum [36] and the European Union Agency for Cybersecurity [10] demonstrate that adversaries not only refine established attack vectors but also exploit emerging technologies, particularly Artificial Intelligence, to evade traditional defenses. In response, the volume and complexity of security research have grown rapidly [26]. Yet a critical gap persists between proof-of-concept prototypes for research evaluation and solutions that are mature and robust enough for real-world deployment [26]. This gap undermines confidence in published results and obstructs the translation of academic advances into practical solutions. To foster trust and accelerate the technological transfer, the academic community increasingly adopts reproducibility badges and performs Artifact Evaluation within the peer-review process [3, 26, 21]. These processes require authors to submit code, data, and instructions, which independent reviewers use to verify claimed results. However, Artifact Evaluation is labor-intensive, depends on volunteers with specialized expertise, and struggles to keep pace with the rising submission rate in cybersecurity conferences and workshops [26]. Large Language Models have demonstrated remarkable natural language understanding, code synthesis, and knowledge extraction capabilities. In cybersecurity contexts, Large Language Models have been applied to intrusion and anomaly detection [40], secure coding assistance [30], and automated penetration testing [14]. Simultaneously, some researchers are looking into improving conventional peer review with Large Language Models [6, 22, 42, 38, 17, 32, 16]. In this paper, we explore a new dimension of their utility: supporting and automating Artifact Evaluation for cybersecurity research. In light of the growing number of submissions at cybersecurity venues, we aim to provide automated support for reviewers of scientific contributions to improve the scalability of Artifact Evaluation. We introduce an Large Language Model-driven toolkit that analyzes paper texts and accompanying artifacts to (i) extract reproducibility indicators, (i) detect potential inconsistencies between claims and submitted artifacts, and (i) identify common pitfalls in experimental design and evaluation. By embedding these capabilities into the peer-review workflow, we aim to improve both the scalability and consistency of the Artifact Evaluation process. Contributions. We propose a three-step Large Language Model-driven toolkit to partially automate reproducibility assessments: â¶ Rate: Our Large Language Model-based method that conceptualizes reproducibility via concept vectors extracted from the modelâs hidden states achieves a recall of almost 95, allowing the automatic discarding of non-reproducible submissions. â¶ Prepare: Our Large Language Model-agent framework automatically sets up and runs submitted artifacts in sandboxed environments, preparing nearly 30 of manually reproducible submissions and offering supporting hints for all others. â¶ Assess: By repurposing the concept of the Rate stage, we reliably identify many design and evaluation pitfalls in security contributions, with an accuracy of >90. â¶ Integrated pipeline: A combination of these stages into a unified Artifact Evaluation workflow, which balances computational cost with assessment accuracy, correctly classifies more than 72 of the papers in our dataset regarding their reproducibility. Open Science. We have published our code on GitHub [8]. Organization. The remainder of this paper is structured as follows. SectionËI provides foundations and introduces recent works on reproducibility and relevant Artificial Intelligence techniques, particularly focusing on cybersecurity. SectionËI details our three-stage Large Language Model-driven pipeline. SectionËIV describes our implementation and empirical evaluation on a curated dataset of hundreds of security research papers. SectionËV discusses findings, lessons learned, and directions for future research before we conclude in SectionËVI. I Background and Related Work Artifact Evaluation at security conferences has become vital for ensuring the transparency and reliability of research, fostering a collaborative environment among researchers and experts. Cybersecurity presents unique challenges for Artifact Evaluation, often involving rapidly evolving threats, adversarial settings, and complex interactions. Manual Artifact Evaluation struggles to scale with growing submission volumes, complex software-hardware stacks, and deeper methodological flaws. Advancements in Artificial Intelligence, particularly concerning Large Language Models, offer promising solutions to automate and enhance certain tasks in this field. In this section, we review current Artifact Evaluation practices and their scalability challenges (SectionËI-A), common pitfalls in Artificial Intelligence-driven cybersecurity research (SectionËI-B), and emerging Artificial Intelligence-based automation techniques targeted toward Artificial Intelligence and cybersecurity (SectionËI-C). I-A Artifact Evaluation at Security Conferences Many top-tier cybersecurity conferences have introduced (currently non-mandatory) Artifact Evaluation into their peer-review process. Artifact Evaluation requires authors to submit the code, datasets, and documentation (typically including a Readme with setup and execution instructions) for independent reviewers to verify computational reproducibility. Consequently, the reviewers can award reproducibility badges if the code is available, runnable, or provides the claimed results [1, 26]. Several papers point out the importance of artifacts and their evaluation in computer science [33, 23, 15, 19, 12, 27, 35, 31, 39], and for cybersecurity in particular [28, 25, 34]. This process promotes transparency, encourages best practices in experiment reporting, and accelerates the adoption of the research code in the community and potentially in production environments [26]. Despite numerous attempts to formalize the requirements for experiment reporting and implementation description [20, 18, 33], many researchers emphasize various challenges with reproducing the results [4, 23, 26, 35, 7]. Despite these benefits, Artifact Evaluation often demands extensive manual effort and expertise, especially given the growing number of submissions to cybersecurity conferences [26] and conferences in general [9]. Reviewers must resolve complex dependencies and sometimes accommodate specialized hardware requirements [26]. Double-blind reviewing intensifies these challenges when anonymization forces the removal of identifying parts of the original code or documentation [4]. Further analyses reinforce these difficulties: Liu et al. [23] examine 21962196 papers and 14871487 corresponding artifacts submitted between 2017 and 2022 to software engineering venues and find no significant improvement in overall artifact quality, noting in particular that the provided Readme files often lack clear instructions and examples. Olszewski et al. [26] systematically inspect 744744 Artificial Intelligence-focused submissions at top security conferences and find that only 298298 include artifacts. Out of the available artifacts, only 57 provide setup instructions, and not all of these instructions lead to the successful execution [26]. Complementing the aforementioned study, we focus on exploring how Large Language Models can reduce the human workload of Artifact Evaluation by providing automated support for key steps and comparing our results against this manually established benchmark. Issue: Conventional Artifact Evaluation processes no longer scale with rising submission rates and the diversity in utilized software and hardware stacks. I-B Common Pitfalls in Cybersecurity Research A rigorous review of a research paper should not only reproduce results but also critically examine the underlying methodology for evaluation and design flaws, complementing Artifact Evaluation. Arp et al. [2] identify several recurring pitfalls that undermine the scientific validity of cybersecurity submissions. For example, Sampling Bias or Base Rate Fallacy may lead to overfitting on imbalanced data or inflated detection metrics due to unrealistically high attack rates in the evaluation data [24, 2, 28]. Lab-only evaluations restrict experiments to synthetic environments, failing to capture real-world operational networksâ diversity and adaptive strategies [2]. Conventional Artifact Evaluation, which focuses primarily on repeating an experiment by rerunning code, often misses these deeper, more foundational issues. However, they remain relevant from the artifact contribution to the community. In this work, we thus examine how Large Language Models can be used to detect textual indicators of these flaws and how to integrate the detection into a (semi-)automated Artifact Evaluation workflow. Issue: Detecting methodological flaws in a study is vital to determine its true contribution; however, these flaws are often hard to detect as part of standard reproducibility checks. I-C Artificial Intelligence-Induced Automation Improvements Large Language Models have demonstrated strong code understanding, generation, and document analysis capabilities [41]. In cybersecurity, they are already used for vulnerability detection [13], flagging anomalies or intrusions [40], and for guiding fuzzing campaigns and penetration tests [14]. Parallel efforts apply Large Language Models to peer review: Numerous authors [11, 6, 22, 42, 38, 17, 32, 16] introduce various techniques to support peer-review processes at academic conferences with Large Language Models. While their work provides a foundation for future research on automated academic peer-review systems, they note that some challenges such as susceptibility to adversarial inputs or biases must be resolved before the tools can be widely deployed. Regarding reproducibility, Bhaskar [5] introduces an Large Language Model-based tool to identify reproducibility indicators in Artificial Intelligence-related papers and their artifacts, achieving better agreement with human judgments when compared to keyword-based approaches. Despite these advances, a comprehensive automation of Artifact Evaluation, including execution environment provisioning, subsequent execution, and detection of methodological pitfalls, remains an open challenge. When complemented with an Large Language Model, such a system can substantially reduce Artifact Evaluation expertsâ manual workload and improve the consistency and reliability of the Artifact Evaluation process in cybersecurity research. Our intuition is that an Large Language Model-driven toolkit that integrates text-based reproducibility screening, automated setup and execution of artifacts, and the detection of common pitfalls may be marketable given the recent advances in Artificial Intelligence. Such a system can substantially reduce Artifact Evaluation expertsâ manual workload and improve the consistency and reliability of the Artifact Evaluation process in cybersecurity research. Issue: The utility of Artificial Intelligence for assessing the reproducibility of proposed concepts remains underexplored. I An Large Language Model-driven Pipeline to Automate Parts of Artifact Reproducibility Assessments Having the aforementioned issues in mind and employing recent Artificial Intelligence developments, we propose an Large Language Model-driven toolkit that provides automated support for three crucial stages of Artifact Evaluation: text-based reproducibility rating (Rate, cf. SectionËI-B), autonomous execution environment preparation (Prepare, cf. SectionËI-C), and methodological-pitfall assessment (Assess, cf. SectionËI-D). Before introducing the design details of the individual steps, we describe how they can be composed into a modular pipeline (SectionËI-A) to support the manual human peer review. I-A Design Overview Figure 1: The three pipeline stages require different inputs, and each stage utilizes an Large Language Model. Further, they can be used in combination as desired. Any available results can then be fed into the Artifact Evaluation (Complement stage). FigureË1 shows the workflow of our pipeline, consisting of three steps that address different parts of Artifact Evaluation processes. The steps can be combined as desired for the respective Artifact Evaluation process, or they can be used independently. Given this independence, the process can be interrupted at any point, and the generated results can be used or discarded according to the use-case-specific preferences (e.g., to exclude submissions with low reproducibility scores from the review). When using the pipeline in an Artifact Evaluation, the process could look as follows: First, the Rate stage checks how reproducible the contribution appears based on the paper and the Readme provided along with the source code. If the Large Language Model detects that reproducibility is likely impossible or very challenging, the subsequent stages could, if desired, be canceled. Second, the Prepare stage attempts to set up the entire research artifact in a fresh container environment to enable its execution using the provided documentation. The Large Language Model-based agent used in this stage iteratively issues shell commands to clone the repository, install dependencies, and compile and execute code, while parsing the commandâs outputs in a feedback loop. Suppose that the execution fails and the Large Language Model fails to identify further corrective actions. In that case, the resulting container and a detailed log of commands and errors are archived for further evaluation by an expert, providing them with first insights. Third, the Assess stage focuses on rating the methodological soundness of the submission: Based on the paper submission, it discovers pitfalls that are common in the design and evaluation of contributions in the field. The results could contain valuable insights and can improve the feedback on methodology that reviewers are returning eventually. Finally, the generated results of all stages, including any created runtime container (Prepare stage), can be forwarded to the Artifact Evaluation reviewers to serve as supplemental material for their âhumanâ expert review. The Complement stage is out of scope for this paper, since our goal is to support, streamline, and automate reviewersâ work using Artificial Intelligence rather than to replace their expert judgment. I-B Rate: Content-Based Reproducibility Ratings Figure 2: Mapping of concept vectors with a known concept âreproducible results.â When measuring a new vector âartifacts available,â it is mapped via a projection to the original vector to compute a score. Our toolkitâs first step, Rate, quantifies reproducibility as a semantic direction in an Large Language Modelâs hidden-state space. We adapt Yang et al.âs prior work [37] that extracts concept vectors from Large Language Modelsâ internal states. By projecting a new textâs embedding onto such a concept vector, they quantify how strongly that text represents the respective concept. The authors demonstrate that this approach yields consistent and valid measures for concepts in social science research contexts. In our case, we define the concept as reproducibility in cybersecurity research. We begin by crafting two descriptive prompts p+p^+ and pâp^- that define the opposite poles of our concept: one characterizing that a paper is âeasy to reproduceâ and the other describing that a paper is âdifficult to reproduce.â These prompts instruct the Large Language Model to attend to textual cues such as the clarity of methodological descriptions, the presence and quality of installation and execution instructions, as well as the completeness of supplementary materials. To extract a reproducibility concept vector, we randomly select a set of n probing texts ti,0â€i<nt_i,0†i<n and feed each twice into the Large Language Model, once under p+p^+ and once under pâp^-. We extract textual cues from each run in the form of embedding vectors from the final layer viv_i of the model, yielding pairs (vi+,viâ)(v_i^+,v_i^-). We then compute viÎŽ:=|vi+âviâ|v_i^ÎŽ:=|v_i^+-v_i^-| for each probe and apply Principal Component Analysis to the collection viÎŽ:0â€i<n\v_i^ÎŽ:0†i<n\. The first principal component serves as our distilled concept vector v v [37]. To evaluate the reproducibility of a new paper, we obtain its hidden-state embedding v under a neutral prompt and project it onto v v by computing a dot-product s:=vâ v^/âv^âs:=v· v/\| v\|. The resulting score s reflects how strongly the paperâs text aligns with the distilled reproducibility concept vector constructed from the training dataset. As the method relies only on hidden-state vectors and Principal Component Analysis, it is independent of the specific Large Language Model architecture and can be applied to any model that exposes final-layer embeddings. I-C Prepare: Autonomously Setting up Code Figure 3: The Large Language Model agent gets access to the paper, the relevant source code and data, as well as a Readme (if available). It then generates commands to execute the code and runs them in a terminal. The outputs are sent back to the agent to determine the next steps. In the Prepare stage, we deploy an Large Language Model-based agent to automate execution environment setup and code execution within a sandboxed environment, as we detail in FigureË3: Our agent has full access to a shell and is given (i) the paper, (i) the artifact codebase, and (i) existing documentation, such as a Readme file. We then prompt it to emit shell commands, which are executed sequentially in a container. As a first step, we instruct the agent to download any relevant code and datasets required to execute the artifact. After running a command, we capture the output and send it back to the Large Language Model, enabling it to diagnose errors such as missing dependencies, version mismatches, or compilation failures, and to generate follow-up commands to resolve errors. Optionally, the agent may be instructed to output natural-language explanations of each step for human review. This interactive feedback loop continues until the artifact runs or the agent indicates no further corrective actions are possible. By isolating each artifact in its own container, we ensure (i) reproducibility by starting from a clean system instance, (i) resource control, such as Graphics Processing Unit access, (i) clean teardown after the execution, and (iv) isolation from other processes running on the host that may otherwise interfere with the execution. The final deliverable of this stage is either a runnable container image ready for further analysis by an Artifact Evaluation expert or a structured error report that pinpoints issues the agent faced. In the latter case, an expert may manually try to fix the detected problems; nonetheless, the Large Language Model agent already completed large parts of trial-and-error setups beforehand. I-D Assess: Identifying Pitfalls in Contributions While the previous stages focus on computational reproducibility of the results of a research submission, this stage evaluates the scientific rigor of a submission. Most importantly, it may enhance the quality of reviews issued by Artifact Evaluation experts by supporting the detection of otherwise hard-to-notice flaws in the studyâs methodology and evaluation. We focus on Arp et al.âs [2] taxonomy of ten common pitfalls in Artificial Intelligence-driven cybersecurity research; however, our approach is conceptually independent of the specifically analyzed pitfalls. This stage works similarly to the Rate stage by independently extracting a concept vector from the underlying Large Language Model for each of the analyzed pitfalls. For each of the m analyzed pitfalls, we construct positive and negative prompts that characterize the opposite poles of the respective concept (i.e., pitfall present or pitfall not present). Using the procedure from the Rate stage, we derive a unique concept vector for each pitfall individually using a set of training papers. To assess a new paper, we compute scores si,0â€i<ms_i,0†i<m for each pitfall to obtain a feature vector s:=(s0,âŠ,smâ1)s:=(s_0,âŠ,s_m-1) which we input into a supervised classifier. The classifier outputs which pitfalls are most likely present. The report highlights potential design or evaluation flaws, providing reviewers with insights into the submissionâs potential methodological strengths and weaknesses. IV Evaluation To demonstrate the effectiveness of our toolkit on real paper submissions, we measure the accuracy and reliability of our individual steps on two expert-annotated datasets: We employ Olszewski et al.âs [26] dataset of several hundred Artificial Intelligence-based cybersecurity papers to benchmark Rate and Prepare, and Arp et al.âs [2] dataset of 3030 papers to assess Assess. We begin by introducing the datasets and our experimental setup in SectionsËIV-A and IV-B, respectively. We then present results for the combined pipeline and its individual components in SectionsËIV-C and IV-D. IV-A Datasets Reproducibility has no universally accepted quantitative benchmark. To evaluate our pipeline, we, therefore, rely on two expert-annotated datasets: first, Olszewski et al. [26] manually assessed the reproducibility of nearly 750750 Artificial Intelligence-based security research papers at top-tier conferences. Second, Arp et al. [2] compiled a dataset of 30 papers where they manually record the presence of ten common pitfalls found in studies in cybersecurity. Next, we introduce them and our experimental setup, including the configured Large Language Models. IV-A1 Olszewski-Study Olszewski et al. [26] invested over eight person-years to manually check the computational reproducibility of artifacts associated with papers on Artificial Intelligence in cybersecurity submitted to USENIX Security, ACM CCS, IEEE S&P, and NDSS between 2013 and 2022. They assign discrete reproducibility scores to each submission reflecting, e.g., the effort required to acquire its code and data, to execute its code, and to reproduce the correct results. Furthermore, they document the presence of metadata such as links to code repositories, hyperparameter settings, and dataset splitting. Most notably, the authors find that out of 744744 analyzed submissions, only 298298 include artifacts. Of those artifacts, roughly 57 include a Readme file that provides instructions for setting up and executing the corresponding code. The authors only manage to execute 46 of the provided artifacts, while only 20 of the tested code repositories produce the same results as advertised in the original papers. For our evaluation of Rate and Prepare, we rely on code repository availability, Readme presence, and manual execution success as ground-truth labels. We only consider the subset of papers where code is available for Prepare and where, additionally, a Readme is available for Rate. IV-A2 Arp-Study Arp et al. [2] manually reviewed 3030 papers in cybersecurity submitted to top-tier conferences (2011â2021) to identify ten recurring experimental and design pitfalls. They annotate each paper for the presence or partial presence of each pitfall and whether the authors discuss the flaw in their papers. Notably, they find that sampling bias affects 90 of the analyzed papers, 60 rely on an inappropriate threat model, and that other issues, such as base-rate fallacy and lab-only evaluation scenarios, affect a majority of papers. We rely on the dataset by Arp et al. [2] to evaluate the Assess step. While Arp et al. also track whether pitfalls are discussed by authors in the text, we only focus on detecting their presence. IV-B Experimental Setup We now introduce the hardware, models, and procedures used to implement and evaluate our pipeline and its individual stages. All Large Language Model-based components run with fixed prompts and thresholds for binary decisions. IV-B1 Rate and Assess For both Rate and Assess, we run a local instance of Llama-3.2-3B-Instruct111https://w.llama.com/docs/model-cards-and-prompt-formats/llama3_2/ on a machine equipped with an NVIDIA H-100 Tensor-Core Graphics Processing Unit. A prompt template informs the Large Language Model that it has access to the full paper text and, in the case of Rate, a Readme file associated with the submissionâs code artifact. To derive concept vectors, we fix a random sample of 1212 papers for Rate and 1010 papers for Assess from the respective datasets and run them through the Large Language Model under the positive and negative prompts. The remaining papers form the test set. We compute cutoff scores by optimizing for recall for Rate and via logistic regression for Assess. IV-B2 Prepare Our Large Language Model agent for Prepare uses OpenAIâs gpt-4o-mini222https://platform.openai.com/docs/models/o4-mini model and interacts with it through the respective web API. Initial experiments with Llama-3.2-3B-Instruct reveal that many of the generated commands to run the corresponding artifacts are invalid and that the model quickly runs out of ideas to fix any occurring issues. For each experiment, the agent spawns a Docker container based on Nvidiaâs cuda image, which in turn uses Ubuntu 22.04 as its base Linux distribution. We host the container on a machine equipped with two Intel Xeon Platinum 8160 CPUs and two NVIDIA Tesla V-100 Graphics Processing Units. Setting up artifacts that require graphical user interfaces or hardware emulations is, unfortunately, not possible in our setup, leading to failed executions of the corresponding code. IV-C Reproducibility Pipeline Evaluation Table I: Comparison of the output of the reproducibility pipeline with the Olszewski-Study. The pipeline correctly classifies almost three-quarters of the examined submissions, providing execution environments for more than 27 of all submissions marked as runnable in the ground-truth. Total 126 Olszewski-Study runs ÂŹ Pipeline runs 7.14%7.14\% 8.73%8.73\% 15.87%15.87\% ÂŹ 19.05%19.05\% 65.08%65.08\% 84.13%84.13\% 26.19%26.19\% 73.81%73.81\% Accuracy: 72.22â%72.22\% Precision: 45.00%45.00\% Recall: 27.27%27.27\% Table I: Comparison of the output of the Rate stage with the Olszewski-Study. The approach correctly classifies almost all submissions marked as runnable in the ground-truth. Total 130 Olszewski-Study runs ÂŹ Rate runs 40.77%40.77\% 54.62%54.62\% 95.38%95.38\% ÂŹ 2.31%2.31\% 2.31%2.31\% 4.62%4.62\% 43.08%43.08\% 56.92%56.92\% Accuracy: 43.08%43.08\% Precision: 42.74%42.74\% Recall: 94.64â%94.64\% Table I: Comparison of the output of the Prepare stage with the Olszewski-Study. The agent automatically sets up ready-to-use execution environments for almost 29 of all submissions marked as runnable in the ground-truth. Total 311 Olszewski-Study runs ÂŹ Prepare runs 7.40%7.40\% 14.79%14.79\% 22.19%22.19\% ÂŹ 18.97%18.97\% 58.84%58.84\% 77.81%77.81\% 26.37%26.37\% 73.63%73.63\% Accuracy: 66.24â%66.24\% Precision: 33.33%33.33\% Recall: 28.05%28.05\% TableËI illustrates the overall performance of our pipeline, i.e., the combination of the Rate and Prepare stages. We consider the intersection of papers from Olszewski et al.âs [26] dataset in both stages individually. Overall, our system correctly assesses whether an artifactâs code can be executed without major effort in more than 72 of cases. Although only about 7 of all attempted artifacts are fully containerized and executed by the pipeline, this performance corresponds to provisioning runnable environments for roughly 28 of the papers that Olszewski et al. [26] manage to execute out of the box, i.e., using only the instructions in the corresponding Readme files. Only about 7 of papers are misclassified as non-runnable when they, in fact, can run out of the box according to Olszewski et al. [26]. These false negatives are often induced by our Docker environment, which cannot emulate special hardware, or by the agent, which cannot fix them without external inputs, e.g., if the link to the sources from the dataset leads to an informative website instead of a Git repository. In any case, the agent provides a reason for failure that human Artifact Evaluation experts can use to try to fix the remaining issues manually. The true negative rate of the pipeline exceeds 85, meaning that our pipeline reliably filters out non-runnable submissions. These numbers underline that the proposed pipeline can indeed save reviewers from spending valuable time otherwise spent on setting up artifacts using trial-and-error, which is a process that is comparably easy to automate. By prepending the Assess stage, submissions deemed unlikely to be reproducible can even be discarded before entering the more costly Prepare stage, saving valuable computational resources. The results of the pipeline are shared with an Artifact Evaluation expert in the Complement stage, who can then decide on, e.g., whether to award a reproducibility badge to the given submission. IV-D Detailed Evaluation Results In this section, we give a more detailed overview of the classification results of the individual stages of our pipeline. We highlight how Rate reliably forwards nearly all runnable artifacts to the next stage, how Prepare autonomously provides numerous execution environments, and how Assess detects methodological flaws with high accuracy. IV-D1 Rate This stage aims to find papers whose code is likely not runnable and discard them early, before wasting computational resources and time setting up the code. For the evaluation of this stage, we consider only papers where the code as well as a Readme file are available. TableËI compares this stageâs classification results to the Olszewski-Study [26] dataset. The high recall of almost 95 indicates that almost all papers with runnable code are selected to move to the next stage (the false negative rate is just over 6). In fact, fewer than 3 of all analyzed papers are misclassified as not runnable. This result makes the step ideal as a first stage for our pipeline: if a paper is deemed not runnable, no computational resources need to be spent to try and execute the respective code. Instead, Artifact Evaluation experts are given the result of the stage. If they feel the results can be reproduced after all, the experts can still manually feed the respective code and paper into the Prepare stage. Given the small number of false negatives, only a few papers are discarded early in the pipeline and not automatically examined for execution. IV-D2 Prepare Unlike the previous stage, this stageâs goal is to automatically set up execution environments to enable experts to quickly run a paperâs code and manually assess the validity of the reproduced results. We consider 311 papers from the dataset since our agent does not explicitly require the presence of a Readme file; instead, it can autonomously analyze the code repository structure and, e.g., try to compile and execute relevant files. All papers that have been evaluated in the Rate stage are also analyzed in this stage. In TableËI, we summarize the results of the classification, which show that the agent yields a moderately high accuracy of more than 66. Notably, this stage alone reliably eliminates the need for experts to manually set up execution environments for papers whose code is not runnable for almost 60 of the analyzed papers. The false negatives are often induced by limitations of our execution environment (cf. SectionËIV-B): Even though our Large Language Model agent can effectively execute terminal commands, some artifacts require access to graphical desktop environments to, e.g., run Internet browsers, which is out of the scope of our experiments. IV-D3 Assess Given the relatively small size of the dataset from the Arp-Study [2], i.e., 3030 papers, we cannot analyze the pitfalls on sampling bias (P1) and data snooping (P3). This limitation is due to our training process requiring at least 55 papers per category âpitfall presentâ and âpitfall not present.â For the remaining pitfalls, our evaluation yields promising results. Except for the pitfall on biased parameters (P5), the classifier has an accuracy between 90 and 100. F1F_1 scores are between 0.920.92 and 11, and F2F_2 scores between 0.970.97 and 11, indicating an accurate response given by our approach. For (P5), our approach performs almost like a random predictor. However, Arp et al. [2] classify most papers as âunclear from textâ for this category. We presume that a larger and more representative dataset would fix this problem. The remaining seven pitfalls can be accurately detected using only a small human-annotated dataset. Overall, we conclude that Assess is well-suited for detecting common known pitfalls in security-related research papers on Artificial Intelligence. V Discussion and Future Work Our results show that an Large Language Model-driven toolkit can reliably filter out non-runnable submissions, autonomously provide execution environments for submitted artifacts, and accurately flag common methodological pitfalls in cybersecurity research. In the following, we discuss our findings in more detail while also addressing limitations of our design, implementation, and evaluation, and proposing directions for future research on the topic. Further, we complement this discussion of findings with a brief overview of lessons learned during our research activities in SectionËV-C. V-A Individual Findings Given the overall results of the toolkit, we reflect upon the individual componentsâ strengths and limitations. Furthermore, we outline targeted directions for future enhancements. V-A1 Rate This stage already yields promising results despite the Large Language Model used for this purpose not being fine-tuned to the given task. Instead, the training data is given to the Large Language Model as a prompt. Future work may evaluate whether fine-tuning an Large Language Model improves the quantification of the concept of artifact reproducibility within the model to generate more precise and consistent concept vectors. However, this change would require a large amount of training data, which is unavailable to us at the time of writing. This training data could, for example, be collected as part of a shadow Artifact Evaluation conducted to evaluate the pipeline further, as suggested in SectionËV-B. V-A2 Prepare While this stage automatically creates sandboxed execution environments for many paper artifacts, with the currently used execution environment, we are still unable to handle all submissions correctly. This situation is partly due to technical limitations, e.g., the lack of a desktop environment or specialized hardware required for some evaluations. The former could be solved by adding graphical user interface interaction support to the agent, e.g., using UI-TARS [29]. The latter can be solved by providing a more diverse hardware setup for the stage; this improvement, however, exceeds the scope of this paper, as our goal is to show the general feasibility of the approach. V-A3 Assess Our evaluation of this stage shows that the detection of pitfalls in cybersecurity papers on Artificial Intelligence performs very reliably. However, the small size of the evaluation dataset poses limitations to our evaluation. We propose to re-evaluate this stage on a more exhaustive dataset. The creation of such a dataset is, however, infeasible within the scope of this paper. V-B General Findings and Future Directions Our evaluation shows that our tools, when combined into a pipeline, can provide significant support in the Artifact Evaluation process conducted at security conferences. It provides a first step into automating this process by detecting submissions without reproducible artifacts and autonomously preparing their execution to enable Artifact Evaluation experts to more quickly assess the validity of the reproduced results. Hence, we provide means to significantly boost the scalability of the process, particularly as we facilitate the tedious task of setting up code environments for performing the evaluations. V-B1 Open Questions Despite the potential highlighted in our evaluation, we identify several open questions: (i) Better understanding in detail how âperfectâ prompts could look like for the different approaches in our toolkit. (i) Further comparing different underlying Large Language Models, as different models may be better in finding and understanding certain concepts or performing certain tasks, in particular, depending on the model size. (i) Assessing the security risks of applying our pipeline in practice, e.g., regarding the execution of arbitrary code in the artifacts, as well as better understanding the implications for intellectual property fed to closed-source commercial models. Concerning the last question, Prepare already provides execution environments that are sandboxed in individual Docker containers. However, access to hardware components such as Graphics Processing Units or other specialized devices may impose additional risks on the system. V-B2 Integration into Peer-Review After improving the techniques for automatically assessing artifact reproducibility proposed in this paper, future work may integrate them directly into the review process at cybersecurity conferences. Given more general training data, the process could also be integrated into conferences in other fields. Currently, artifact reproducibility checks are often only performed for accepted papers, i.e., after the review process is completed [3]. Automated reproducibility checks would allow checking a large number of submissions even before issuing an acceptance. While some authors might be concerned with participating in a potentially biased or low-quality Artifact Evaluations, an Artificial Intelligence-assisted pipeline may increase their trust in the process and, in turn, improve their willingness to participate. We believe that assessing the usability of the proposed pipeline workflow in the form of shadow Artifact Evaluation is a good next step for assessing its maturity. Integrating our tool into peer review requires addressing manipulation risks, e.g., prompt injection. The attack surface is limited in Rate and Assess, as outputs derive from concept vectors extracted from the Large Language Modelâs hidden states rather than direct generation from paper text. Reviewers read the paper before or in parallel to the partially automated Artifact Evaluation, aiding detection of any injections. In Prepare, injections may occur in Readme files or code comments; however, sandboxed execution and expert assessment of results render their impact negligible. We expect such malicious acts to be rare due to scientific integrity norms and penalized when discovered. V-B3 Future Evaluations Finally, future work may expand our research by applying the techniques to papers from different fields. We limit our evaluation to papers in these domains due to the availability of an exhaustive dataset. We believe the approach easily generalizes to other topic areas as it does not directly depend on the contents of the evaluated works. Most importantly, we show that Artificial Intelligence is a promising tool to employ in Artifact Evaluation with a great potential to complement the process to improve its quality, scalability, and thus sustainability. V-C Lessons Learned During the development of our pipeline, we noticed several unexpected behaviors across interactive handling, execution, verification, and contextual reasoning. For example, in the Prepare stage, the agent reports that it cannot engage with an interactive editor such as nano for one experiment. While it could have proposed an alternative non-interactive solution, e.g., using sed, the Large Language Model did not have this idea, revealing a gap in its problem-solving repertoire. We suggest evaluating this behavior with more powerful models to assess whether this problem can be resolved. Furthermore, in one experimental run during our implementation, the agent proposed to comment out an entire program to make it run successfullyâreturning in an inaccurate assessment. However, this change results in the program not performing any computations or providing any outputs. We worked around this issue by adapting the prompts given to the Large Language Model, highlighting the importance of carefully designed prompts and validation. In the Rate stage, we notice that the Large Language Modelâs grasp of the concept of reproducibility is less accurate than that of the different pitfalls analyzed in the Assess stage. This result may be induced by the training data of the utilized model: Reproducibility is a niche topic, and todayâs models are likely not trained on much input that covers this concept. Simultaneously, even on powerful hardware, local models execute much more slowly than commercial models like gpt-4o-mini, which are heavily optimized for mass use. In particular, to protect the privacy and confidentiality of submissions during the (confidential) Artifact Evaluation process, we propose that future work focuses on optimizing Large Language Models specifically for the use case of artifact reproducibility assessment. Overall, we have learned that Large Language Models constitute a powerful tool that has the potential to substantially complement and improve the Artifact Evaluation process at scientific conferences. They can be employed to automate tedious and repetitive tasks while simultaneously streamlining the whole process to help provide more consistent and high-quality feedback to authors. VI Conclusion Ensuring the reproducibility of research artifacts in cybersecurity is crucial in science to validate the potential for further use of given experimental and methodological results. It narrows the gap between experiments and simulations, as well as real-world deployments, since stakeholders can better assess the suitability of the approaches for their systems. Currently, some scientific conferences perform time-consuming manual artifact evaluations to assess whether the contributions of the submitted works are reproducible. We propose an Large Language Model-based toolkit that enhances the automation potential of otherwise manual and time-consuming artifact assessments. Our evaluation shows that, when combining the tools into a pipeline, a majority of submissions without runnable artifacts are automatically discarded. At the same time, execution environments are generated for many submissions with runnable code. We propose to integrate such a pipeline into the Artifact Evaluation (AE) process of conferences to incentivize researchers to deliver reproducible results. Furthermore, this change has potential to unburden reviewers by automating a time-consuming part of the review work, with the objective of improving the sustainability of the Artifact Evaluation process. Acknowledgments This work was funded by the Federal Ministry of Research, Technology and Space (BMFTR) in Germany under the grant number 16KIS2251 of the SUSTAINET-guardian project. The responsibility for the content of this publication lies with the authors. The authors thank Daniel Arp for supporting the pitfall evaluation (Assess stage), which builds upon the survey data by Arp et al. [2]. References [1] ACM (2020) Artifact Review and Badging - Current. Note: https://w.acm.org/publications/policies/artifact-review-and-badging-current Cited by: §I-A. [2] D. Arp, E. Quiring, F. Pendlebury, A. Warnecke, F. Pierazzi, C. Wressnegger, L. Cavallaro, and K. Rieck (2022) Dos and Donâts of Machine Learning in Computer Security. In Proceedings of the 31st USENIX Security Symposium (SEC â22), p. 3971â3988. External Links: ISBN 978-1-939133-31-1 Cited by: §I-B, §I-D, §IV-A2, §IV-A, §IV-D3, §IV, Acknowledgments. [3] M. Athanassoulis, P. Triantafillou, R. Appuswamy, R. Bordawekar, B. Chandramouli, X. Cheng, I. Manolescu, Y. Papakonstantinou, and N. Tatbul (2022) Artifacts Availability & Reproducibility (VLDB 2021 Round Table). ACM SIGMOD Record 51 (2), p. 74â77. External Links: Document, ISSN 0163-5808 Cited by: §I, §V-B2. [4] V. Bajpai, M. KĂŒhlewind, J. Ott, J. SchönwĂ€lder, A. Sperotto, and B. Trammell (2017) Challenges with Reproducibility. In Proceedings of the Reproducibility Workshop (Reproducibility â17), p. 1â4. External Links: Document, ISBN 978-1-4503-5060-0 Cited by: §I-A, §I-A. [5] A. Bhaskar and V. Stodden (2024) Reproscreener: Leveraging LLMs for Assessing Computational Reproducibility of Machine Learning Pipelines. In Proceedings of the 2nd ACM Conference on Reproducibility and Replicability (REP â24), p. 101â109. External Links: Document, ISBN 979-8-4007-0530-4 Cited by: §I-C. [6] L. Cao, L. You, and R&D Team (2025) CSPaper Review: Fast, Rubric-Faithful Conference Feedback. In Proceedings of the 18th International Natural Language Generation Conference: System Demonstrations (INLG â25), p. 3â7. Cited by: §I, §I-C. [7] C. Collberg, T. Proebsting, and A. M. Warren (2015) Repeatability and Benefaction in Computer Systems Research: A Study and a Modest Proposal. Technical report Technical Report TR 14-04, University of Arizona. Cited by: §I-A. [8] COMSYS Artifact Evaluation LLM Support. External Links: Link Cited by: §I. [9] C. Demetrescu, I. Finocchi, A. Ribichini, and M. Schaerf (2022) On computer science research and its temporal evolution. Scientometrics 127 (8), p. 4913â4938. External Links: Document, ISSN 0138-9130 Cited by: §I-A. [10] European Union Agency for Cybersecurity (2024) ENISA Threat Landscape 2024. Technical report European Union Agency for Cybersecurity. External Links: Document, ISBN 978-92-9204-675-0 Cited by: §I. [11] X. Gao, J. Ruan, Z. Zhang, J. Gao, T. Liu, and Y. Fu (2025) MMReview: A Multidisciplinary and Multimodal Benchmark for LLM-Based Peer Review Automation. Note: arXiv:2508.14146 External Links: Document Cited by: §I-C. [12] O. E. Gundersen and S. Kjensmo (2018) State of the art: Reproducibility in artificial intelligence. Proceedings of the AAAI Conference on Artificial Intelligence 32 (1). External Links: Document, ISSN 2159-5399 Cited by: §I-A. [13] Y. Guo, C. Patsakis, Q. Hu, Q. Tang, and F. Casino (2024) Outside the Comfort Zone: Analysing LLM Capabilities in Software Vulnerability Detection. In Proceedings of the 29th European Symposium on Research in Computer Security (ESORICS â24), Vol. 14982, p. 271â289. External Links: Document, ISBN 978-3-031-70878-7, ISSN 0302-9743 Cited by: §I-C. [14] A. Happe and J. Cito (2023) Getting pwnâd by AI: Penetration Testing with Large Language Models. In Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE â23), p. 2082â2086. External Links: Document, ISBN 979-8-4007-0327-0 Cited by: §I, §I-C. [15] B. Hermann, S. Winter, and J. Siegmund (2020) Community Expectations for Research Artifacts and Evaluation Processes. In Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE â20), p. 469â480. External Links: Document, ISBN 978-1-4503-7043-1 Cited by: §I-A. [16] S. Huang, Q. Wang, W. Lu, L. Liu, Z. Xu, and Y. Huang (2025) PaperEval: A universal, quantitative, and explainable paper evaluation method powered by a multi-agent system. Information Processing & Management 62 (6). External Links: Document, ISSN 0306-4573 Cited by: §I, §I-C. [17] M. Idahl and Z. Ahmadi (2025) OpenReviewer: A Specialized Large Language Model for Generating Critical Scientific Paper Reviews. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (System Demonstrations) (NAACL â25), p. 550â562. External Links: Document Cited by: §I, §I-C. [18] A. Jedlitschka and D. Pfahl (2005) Reporting Guidelines for Controlled Experiments in Software Engineering. In Proceedings of the 2005 International Symposium on Empirical Software Engineering, 2005 (ISESE â05), p. 95â104. External Links: Document, ISBN 978-0-7803-9507-7 Cited by: §I-A. [19] N. Juristo and S. Vegas (2011) The role of non-exact replications in software engineering experiments. Empirical Software Engineering 16 (3), p. 295â324. External Links: Document, ISSN 1382-3256 Cited by: §I-A. [20] B. A. Kitchenham, S. L. Pfleeger, L. M. Pickard, P. W. Jones, D. C. Hoaglin, K. El Emam, and J. Rosenberg (2002) Preliminary guidelines for empirical research in software engineering. IEEE Transactions on Software Engineering 28 (8), p. 721â734. External Links: Document, ISSN 0098-5589 Cited by: §I-A. [21] S. Krishnamurthi (2014 (accessed November 22, 2025)) About Artifact Evaluation. Note: https://artifact-eval.org/about.html Cited by: §I. [22] I. Kuznetsov, O. M. Afzal, K. Dercksen, N. Dycke, A. Goldberg, T. Hope, D. Hovy, J. K. Kummerfeld, A. Lauscher, K. Leyton-Brown, S. Lu, Mausam, M. Mieskes, A. NĂ©vĂ©ol, D. Pruthi, L. Qu, R. Schwartz, N. A. Smith, T. Solorio, J. Wang, X. Zhu, A. Rogers, N. B. Shah, and I. Gurevych (2024) What Can Natural Language Processing Do for Peer Review?. Note: arXiv:2405.06563 External Links: Document Cited by: §I, §I-C. [23] M. Liu, X. Huang, W. He, Y. Xie, J. M. Zhang, X. Jing, Z. Chen, and Y. Ma (2024) Research artifacts in software engineering publications: Status and trends. Journal of Systems and Software 213. External Links: Document, ISSN 0164-1212 Cited by: §I-A, §I-A. [24] M. A. Lones (2024) Avoiding Common Machine Learning Pitfalls. Patterns 5 (10). External Links: Document, ISSN 2666-3899 Cited by: §I-B. [25] R. H. Moulton, G. A. McCully, and J. D. Hastings (2024) Confronting the Reproducibility Crisis: A Case Study of Challenges in Cybersecurity AI. In Proceedings of the 2024 Cyber Awareness and Research Symposium (CARS â24), p. 1â6. External Links: Document, ISBN 979-8-3503-8641-7 Cited by: §I-A. [26] D. Olszewski, A. Lu, C. Stillman, K. Warren, C. Kitroser, A. Pascual, D. Ukirde, K. Butler, and P. Traynor (2023) "Get in Researchers; Weâre Measuring Reproducibility": A Reproducibility Study of Machine Learning Papers in Tier 1 Security Conferences. In Proceedings of the 2023 ACM SIGSAC Conference on Computer and Communications Security (CCS â23), p. 3433â3459. External Links: Document, ISBN 979-8-4007-0050-7 Cited by: §I, §I, §I-A, §I-A, §I-A, §I-A, §IV-A1, §IV-A, §IV-C, §IV-C, §IV-D1, §IV. [27] R. D. Peng (2011) Reproducible Research in Computational Science. Science 334 (6060), p. 1226â1227. External Links: Document, ISSN 0036-8075 Cited by: §I-A. [28] J. Pennekamp, E. Buchholz, M. Dahlmanns, I. Kunze, S. Braun, E. Wagner, M. Brockmann, K. Wehrle, and M. Henze (2021) Collaboration is not Evil: A Systematic Look at Security Research for Industrial Use. In Proceedings of the Workshop on Learning from Authoritative Security Experiment Results (LASER â20), External Links: Document, ISBN 978-1-891562-81-5 Cited by: §I-A, §I-B. [29] Y. Qin, Y. Ye, J. Fang, H. Wang, S. Liang, S. Tian, J. Zhang, J. Li, Y. Li, S. Huang, et al. (2025) UI-TARS: Pioneering Automated GUI Interaction with Native Agents. Note: arXiv:2501.12326 External Links: Document Cited by: §V-A2. [30] G. Sandoval, H. Pearce, T. Nys, R. Karri, S. Garg, and B. Dolan-Gavitt (2023) Lost at C: A User Study on the Security Implications of Large Language Model Code Assistants. In Proceedings of the 32nd USENIX Security Symposium (USENIX SEC â23), p. 2205â2222. External Links: ISBN 978-1-939133-37-3 Cited by: §I. [31] Q. Scheitle, M. WĂ€hlisch, O. Gasser, T. C. Schmidt, and G. Carle (2017) Towards an Ecosystem for Reproducible Research in Computer Networking. In Proceedings of the Reproducibility Workshop (Reproducibility â17), p. 5â8. External Links: Document, ISBN 978-1-4503-5060-0 Cited by: §I-A. [32] L. Sun, S. Tao, J. Hu, and S. P. Dow (2024) MetaWriter: Exploring the Potential and Perils of AI Writing Support in Scientific Peer Review. Proceedings of the ACM on Human-Computer Interaction 8 (CSCW1). External Links: Document, ISSN 2573-0142 Cited by: §I, §I-C. [33] C. S. Timperley, L. Herckis, C. Le Goues, and M. Hilton (2021) Understanding and Improving Artifact Sharing in Software Engineering Research. Empirical Software Engineering 26 (3). External Links: Document, ISSN 1382-3256 Cited by: §I-A. [34] R. Uetz, C. Hemminghaus, L. HacklĂ€nder, P. Schlipper, and M. Henze (2021) Reproducible and Adaptable Log Data Generation for Sound Cybersecurity Experiments. In Proceedings of the 37th Annual Computer Security Applications Conference (ACSAC â21), p. 690â705. External Links: Document, ISBN 978-1-4503-8579-4 Cited by: §I-A. [35] S. Winter, C. S. Timperley, B. Hermann, J. Cito, J. Bell, M. Hilton, and D. Beyer (2022) A Retrospective Study of One Decade of Artifact Evaluations. In Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE â22), p. 145â156. External Links: Document, ISBN 978-1-4503-9413-0 Cited by: §I-A. [36] World Economic Forum (2025) Global Cybersecurity Outlook 2025. Technical report World Economic Forum. Cited by: §I. [37] Y. Yang, H. Duan, J. Liu, and K. Y. Tam (2024) LLM-Measure: Generating Valid, Consistent, and Reproducible Text-Based Measures for Social Science Research. Note: arXiv:2409.12722 External Links: Document Cited by: §I-B, §I-B. [38] R. Ye, X. Pang, J. Chai, J. Chen, Z. Yin, Z. Xiang, X. Dong, J. Shao, and S. Chen (2024) Are We There Yet? Revealing the Risks of Utilizing Large Language Models in Scholarly Peer Review. Note: arXiv:2412.01708 External Links: Document Cited by: §I, §I-C. [39] B. Yildiz, H. Hung, J. H. Krijthe, C. C. S. Liem, M. Loog, G. Migut, F. A. Oliehoek, A. Panichella, P. PaweĆczak, S. Picek, M. de Weerdt, and J. van Gemert (2021) ReproducedPapers.org: Openly Teaching and Structuring Machine Learning Reproducibility. In Proceedings of the 3rd International Workshop on Reproducible Research in Pattern Recognition (RRPR â21), Vol. 12636, p. 3â11. External Links: Document, ISBN 978-3-030-76422-7, ISSN 0302-9743 Cited by: §I-A. [40] H. Zhang, A. Bin Sediq, A. Afana, and M. Erol-Kantarci (2024) Large Language Models in Wireless Application Design: In-Context Learning-enhanced Automatic Network Intrusion Detection. In Proceedings of the 2024 IEEE Global Communications Conference (GLOBECOM â24), p. 2479â2484. External Links: Document, ISBN 979-8-3503-5125-5, ISSN 2576-6813 Cited by: §I, §I-C. [41] J. Zhang, H. Bu, H. Wen, Y. Liu, H. Fei, R. Xi, L. Li, Y. Yang, H. Zhu, and D. Meng (2025) When LLMs meet cybersecurity: a systematic literature review. Cybersecurity 8. External Links: Document, ISSN 2523-3246 Cited by: §I-C. [42] M. Zhu, Y. Weng, L. Yang, and Y. Zhang (2025) DeepReview: Improving LLM-based Paper Review with Human-like Deep Thinking Process. Note: arXiv:2503.08569 External Links: Document Cited by: §I, §I-C.