Paper deep dive
Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recall--workload trade-offs and run-to-run consistency
Nikol Figalovå, Lynn Huestegge, Anne Böckler-Raettig
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 95%
Last extracted: 8/28/2026, 4:23:44 AM
Summary
This study evaluates and compares human and Large Language Model (LLM) screening workflows in a conceptually complex scoping review on gaze semantics. Using a benchmark dataset of 1,131 records, the authors assessed operational recall, retained workload, agreement, and run-to-run consistency across human reviewers (solo lead and distributed team) and seven LLM runs (GPT-5.4 and Gemini 3.1) with varying processing configurations. Results indicate that LLM performance is heavily dependent on workflow configuration rather than model identity alone. GPT-5.4 file-batch runs achieved 82.3-82.9% recall with 42.2-45.0% retention, while Gemini 3.1 achieved higher recall (83.9%) but with higher retention (56.7%). The study concludes that for high-recall tasks, LLMs are best suited for auditable, human-supervised workflows rather than autonomous exclusion, and emphasizes the importance of processing configuration and record-level consistency.
Entities (9)
Relation Signals (9)
Lynn Huestegge â affiliatedwith â Julius-Maximilians-UniversitĂ€t WĂŒrzburg
confidence 99% · Affiliation: Institute of Psychology, Julius-Maximilians-UniversitĂ€t WĂŒrzburg
Nikol FigalovĂĄ â affiliatedwith â Julius-Maximilians-UniversitĂ€t WĂŒrzburg
confidence 99% · Affiliation: Institute of Psychology, Julius-Maximilians-UniversitĂ€t WĂŒrzburg
Anne Böckler-Raettig â affiliatedwith â Julius-Maximilians-UniversitĂ€t WĂŒrzburg
confidence 99% · Affiliation: Institute of Psychology, Julius-Maximilians-UniversitĂ€t WĂŒrzburg
Gemini 3.1 â achievedrecall â 83.9%
confidence 95% · Gemini 3.1 file batches achieved the highest recall (83.9%)
GPT-5.4 â achievedrecall â 82.3-82.9%
confidence 95% · GPT-5.4 file-batch runs... achieving operational recall of 82.3â82.9%.
Nikol FigalovĂĄ â preregisteredstudyon â Open Science Framework
confidence 95% · The methodological study was preregistered on the Open Science Framework on February 24, 2026
GPT-5.4 â retainedworkload â 42.2-45.0%
confidence 95% · retaining 42.2â45.0% of records
Gemini 3.1 â retainedworkload â 56.7%
confidence 95% · retained 56.7% of records.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Background. Large language models (LLMs) are increasingly used for screening in evidence synthesis, where false negatives can remove relevant studies before full-text assessment. We compared human and LLM title-and-abstract screening workflows in a preregistered study embedded in a conceptually complex scoping review. Methods. After a conservative title-only screen, 1,131 records were screened by one review lead, four trained assistants screening non-overlapping subsets, and seven complete LLM runs using different models and processing configurations, including a nominally identical repeat run. We compared retained workload, operational recall against 316 verified eligible records, agreement, run-to-run consistency, and procedural burden. Because eligibility was verified only for records advanced and assessed in the parent review, recall estimates were operational. Results. No workflow recovered all verified eligible records. The human workflows and two GPT-5.4 file-batch runs retained 42.2-45.0% of records while achieving 82.3-82.9% recall. Gemini 3.1 file batches achieved the highest recall (83.9%) but retained 56.7% of records. All-at-once configurations recovered fewer eligible records than corresponding file-batch configurations. Two nominally identical GPT-5.4 file-batch runs agreed on 91.7% of records but differed on 94 records, including 29 verified eligible records retained by only one run. Discussion. LLM screening performance depended on the implemented workflow, not model identity alone. Processing configuration, workload, record-level variation, and human-LLM decision integration are therefore substantive properties of deployed systems. For high-recall tasks, LLMs are better suited to validated, auditable, human-supervised workflows than autonomous exclusion.
Tags
Links
- Source: https://arxiv.org/abs/2608.26885v1
- Canonical: https://arxiv.org/abs/2608.26885v1
Trouble viewing inline? Open PDF directly â
Full Text
109,179 characters extracted from source content.
Expand or collapse full text
Evaluating human and LLM screening workflows in a conceptually complex scoping review: Recallâworkload trade-offs and run-to-run consistency Nikol FigalovĂĄ Affiliation: Institute of Psychology, Julius-Maximilians-UniversitĂ€t WĂŒrzburgWĂŒrzburg, Germany*These authors contributed equallyCorresponding author:Nikol FigalovĂĄInstitute of Psychology, Julius-Maximilians-UniversitĂ€t WĂŒrzburgWĂŒrzburg, Germanynikol.figalova@uni-wuerzburg.de Lynn Huestegge Affiliation: Institute of Psychology, Julius-Maximilians-UniversitĂ€t WĂŒrzburgWĂŒrzburg, Germany*These authors contributed equallyCorresponding author:Nikol FigalovĂĄInstitute of Psychology, Julius-Maximilians-UniversitĂ€t WĂŒrzburgWĂŒrzburg, Germanynikol.figalova@uni-wuerzburg.de Anne Böckler-Raettig Affiliation: Institute of Psychology, Julius-Maximilians-UniversitĂ€t WĂŒrzburgWĂŒrzburg, Germany*These authors contributed equallyCorresponding author:Nikol FigalovĂĄInstitute of Psychology, Julius-Maximilians-UniversitĂ€t WĂŒrzburgWĂŒrzburg, Germanynikol.figalova@uni-wuerzburg.de Abstract Background. Large language models (LLMs) are increasingly used for document-classification tasks in evidence synthesis, where false-negative decisions can remove relevant studies before full-text assessment. We evaluated human and LLM-based title-and-abstract screening workflows in a preregistered methodological study embedded in a conceptually complex scoping review, treating the implemented workflow rather than the model alone as the primary unit of comparison. Methods. After a conservative preliminary title-only screen, 1,131 records were screened by one review lead, a distributed team of four trained assistants screening non-overlapping subsets, and seven complete LLM runs spanning different models and processing configurations, including a nominally identical repeat run. We compared retained workload, operational recall against 316 verified eligible records, conditional classification performance among 859 full-text-assessed records, agreement, combined recovery, run-to-run consistency, and procedural burden. Because eligibility was verified only for records advanced and assessed in the parent review, recall estimates were operational rather than complete-benchmark estimates. Results. No individual workflow recovered all verified eligible records. The two human workflows and two GPT-5.4 file-batch runs showed similar workloadârecall profiles, retaining 42.2â45.0% of records while achieving operational recall of 82.3â82.9%. Gemini 3.1 file batches achieved the highest recall (83.9%) but retained 56.7% of records. Both all-at-once configurations recovered fewer eligible records than their corresponding file-batch configurations. Two nominally identical GPT-5.4 file-batch runs agreed on 91.7% of records, yet differed on 94 records, including 29 verified eligible records retained by only one run. Discussion. LLM screening performance depended on the complete implemented workflow and could not be characterised by model identity or aggregate performance alone. Processing configuration, downstream workload, record-level variation, and the way model outputs are combined with human decisions are substantive properties of the deployed system. For high-recall screening tasks, LLMs are better suited to validated, auditable, human-supervised workflows than to autonomous exclusion. Keywords: large language models; artificial intelligence; humanâAI collaboration; document classification; evidence synthesis; literature screening; reproducibility; decision support Introduction Large language models (LLMs) are increasingly being used as natural-language classifiers within decision-support workflows, including tasks in which errors have asymmetric consequences. Title-and-abstract screening in evidence synthesis provides a demanding real-world example: large numbers of short documents must be classified against natural-language eligibility criteria, while false-negative decisions can permanently remove relevant evidence from subsequent assessment. At the same time, overly conservative classification increases the number of records requiring full-text retrieval and review. Screening therefore involves a practical trade-off between recovering relevant evidence and limiting downstream workload (Homiar et al., 2025; Landschaft et al., 2024; Shailendra et al., 2026). This trade-off is particularly consequential in scoping reviews. Such reviews often address broad or interdisciplinary questions, include heterogeneous study designs, and require judgements about conceptual relevance rather than application of narrowly specified eligibility criteria (Peters et al., 2020; Tricco et al., 2018). Conceptual and terminological heterogeneity may also create jingle and jangle problems, whereby distinct phenomena share a label or closely related phenomena are described using different terminology (Hanfstingl et al., 2024). Reviewers must consequently distinguish clearly relevant and irrelevant records from those for which the title and abstract provide insufficient information. Screening is commonly conservative in such cases, with ambiguous records retained for full-text assessment. LLMs offer a potentially useful means of supporting this process because eligibility criteria can be expressed directly in natural language and applied to titles and abstracts without review-specific model training. However, the relevant unit of evaluation is not necessarily the underlying model alone. An implemented LLM screening system also includes the prompt, uncertainty rule, record preparation, processing configuration, interaction procedure, output format, quality-control steps, and the rule by which model outputs are translated into downstream decisions. We therefore treat the screening workflow, rather than model identity in isolation, as the primary unit of comparison. LLM-based screening as a workflow-level problem Earlier approaches to automating evidence synthesis predominantly relied on conventional machine-learning and active-learning methods. Tools such as ASReview use reviewer-labelled examples to prioritise records for screening, reducing the number of irrelevant records humans must inspect while retaining human oversight (van de Schoot et al., 2021). These approaches generally depend on reviewer-generated labels, iterative learning, or corpus-specific calibration. LLMs enable a different form of automation. Through natural-language prompting, they can interpret eligibility criteria and return screening decisions without review-specific model training. LLMs are now being investigated across evidence-synthesis tasks including title-and-abstract screening, full-text assessment, data extraction, and synthesis (Harasgama et al., 2026; Lieberum et al., 2025; Njei et al., 2026; Cao et al., 2026; Laignelot et al., 2026). Recent reviews nevertheless recommend restricting their use to clearly defined tasks, transparently documenting the implemented procedure, validating performance in the intended context, and maintaining appropriate human supervision (Lieberum et al., 2025; Galli et al., 2025; Laignelot et al., 2026). Empirical evaluations show that LLM screening performance varies across models, prompts, datasets, review topics, decision thresholds, and reference standards. GPTscreenR, for example, used GPT-4 for title-and-abstract screening in scoping reviews and demonstrated useful but imperfect performance against consensus human decisions, with humanâGPT agreement below inter-human agreement (Wilkins, 2023). Other comparative studies have similarly reported substantial variation across model families, prompting strategies, and review datasets (Li et al., 2024; Syriani et al., 2024; Huotala et al., 2024; Krag et al., 2024; DennstĂ€dt et al., 2024). Even relatively small changes to instructions, response scales, or classification thresholds can alter the operating characteristics of the resulting screening procedure (DennstĂ€dt et al., 2024). Processing configuration may also be consequential. The number of records submitted within a prompt can affect both processing feasibility and classification performance, indicating that batch size may be a substantive workflow parameter rather than a neutral implementation choice (Fagerberg et al., 2026). More elaborate procedures do not necessarily improve performance. Akinseloyin et al. (Akinseloyin et al., 2026), for example, found that averaging independently generated relevance scores from several lower-cost LLMs provided robust and cost-effective prioritisation, whereas more complex debate and LLM-based adjudication procedures did not consistently add benefit. Repeated execution introduces a further source of variation. LLM outputs may differ even when the model, records, and instructions appear unchanged (Syriani et al., 2024). Similar aggregate recall or workload across repeated runs therefore does not imply that the same individual records were classified identically. From a workflow-evaluation perspective, reproducibility must consequently be assessed not only through aggregate performance but also through record-level agreement and the identities of discordantly classified records. These considerations make attribution of performance to a model name alone problematic. Reporting frameworks for LLM-assisted evidence synthesis accordingly emphasise documentation of the model and version, prompts, input preparation, processing configuration, decision rules, human involvement, validation procedures, and reference standards (Susnjak, 2023; Shailendra et al., 2026; Holst et al., 2025; Gallifant et al., 2025; Luo et al., 2025). Such information is necessary to characterise the implemented system and to understand which procedural choices may have contributed to observed performance. Evaluating screening performance No single metric fully characterises a screening workflow. Agreement with human title-and-abstract decisions, recovery of records ultimately judged eligible after full-text assessment, and reduction of downstream full-text workload capture related but distinct properties. An LLM may agree closely with human screeners while nevertheless missing records later found to be eligible. Conversely, high recovery may be achieved simply by retaining a large proportion of the screened dataset. Evaluations should therefore report recovered and missed evidence alongside the number of records retained for subsequent assessment rather than relying on a single summary measure (Homiar et al., 2025; Landschaft et al., 2024; Sciurti et al., 2026; Madeyski et al., 2025). Agreement provides complementary information but is not equivalent to screening performance. Because most records retrieved through review searches are typically ineligible, high overall agreement can be driven by concordant exclusions even when workflows differ on a smaller but consequential group of potentially eligible records. Conversely, workflows may use categories such as Include and Unclear differently while retaining largely overlapping sets of records. Agreement can therefore be examined both at the level of the original decision categories and at the operational retained/not-retained level, depending on the evaluation objective. Reference-based classification measures introduce an additional methodological constraint (Sokolova and Lapalme, 2009). Measures such as recall, precision, specificity, and the F1 score require verified reference outcomes for the records included in their calculation. In evidence synthesis, however, final eligibility is usually established only for records that pass title-and-abstract screening and proceed to full-text assessment. The reference set is therefore partly determined by the screening workflows themselves rather than by independent verification of the complete dataset. This creates a form of partial verification bias (Whiting et al., 2011). A record can receive a verified final eligibility outcome only if it is retained by at least one advancement workflow, its full text is successfully retrieved, and it undergoes full-text assessment. Records excluded by all advancement workflows, or those for which full texts cannot be retrieved, have unknown final eligibility and cannot automatically be treated as true negatives. Consequently, recall estimates based on subsequently verified eligible records are operational estimates rather than complete-benchmark measures, while other reference-based metrics apply only to the subset with verified full-text outcomes. The outputs of different workflows may also be complementary. Under a liberal combination rule, retaining a record when either member of a pair recommends advancement can reduce jointly missed eligible records, although generally at the cost of greater downstream workload. Large-scale evaluations have shown that humanâLLM and LLMâLLM combinations can improve sensitivity, with the benefit depending on the review and the chosen combination rule (Sanghera et al., 2025). More conservative approaches have instead accepted automated decisions only when multiple LLMs agree and referred discordant cases for human assessment (Hilkenmeier et al., 2026). Evaluation of an AI-assisted screening system must therefore consider not only individual classifiers but also how their outputs are integrated into the final decision process. Human decisions themselves are not an error-free reference standard (Wang et al., 2020). The number and organisation of reviewers can affect both reliability and resource requirements, and single-reviewer screening cannot be assumed to recover the same records as screening involving additional reviewers (Waffenschmidt et al., 2019). Human reviewers may differ in expertise, interpretation of eligibility criteria, calibration, and tolerance for uncertainty, while performance can also be influenced by task framing and fatigue (Belur et al., 2021; Wang et al., 2020). Human judgement under uncertainty is more generally susceptible to systematic heuristics and biases (Tversky and Kahneman, 1974). LLM errors arise through different mechanisms. A model may misunderstand an abstract, apply a criterion too strictly or too leniently, over-rely on salient topical terms, or fail to follow instructions for handling ambiguity. Such errors are consistent with broader evidence that LLM failures may reflect limitations in knowledge, reasoning, and instruction following (Wu et al., 2026). Comparing realistic human and LLM workflows is therefore more informative than treating either as an idealised or error-free decision maker. Despite the rapidly expanding literature on LLM-assisted evidence synthesis, several gaps remain. Existing evaluations are concentrated in biomedical or intervention-focused systematic reviews with comparatively structured eligibility criteria, leaving less evidence from conceptually complex and interdisciplinary screening tasks. Relatively few studies compare LLM procedures with both a solo review lead and a distributed trained human team using the same records, criteria, decision categories, and advancement rule. Processing configuration and repeated execution also remain insufficiently characterised under realistic web-interface conditions. More broadly, evaluations frequently report model-level performance without fully separating recovery, downstream workload, agreement, conditional classification performance, record-level reproducibility, and complementarity among different decision workflows. The present study The present study was a preregistered comparative methodological evaluation embedded within a PRISMA-ScR-aligned scoping review of gaze semantics in human social interaction (FigalovĂĄ et al., 2026). The parent review examined research in which gaze behaviours were treated as meaningful, interpretable, communicative, or socially consequential in human social contexts. This application provides a challenging natural-language classification task because eligibility depends on conceptual interpretation rather than simple topical matching. A study may mention gaze or eye tracking without examining the social meaning of gaze, whereas another may investigate social inference or communication without describing its contribution explicitly as gaze semantics. Screening therefore requires interpretation of how the phenomenon is conceptualised and whether the information provided in the title and abstract is sufficient for a confident decision. We compared screening by a solo review lead, a distributed team of four trained psychology student assistants, and seven complete LLM runs spanning multiple models and processing configurations. All workflows were evaluated on the same benchmark dataset and under a common operational retention rule. Rather than treating model identity as the sole unit of comparison, we evaluated the complete implemented workflows, including differences in how records were submitted to the LLM and whether nominally identical procedures produced stable record-level decisions across repeated execution. We use workflow to denote a complete screening procedure defined by the screening entity, model where applicable, input format, prompting strategy, and output-handling process. A run is one complete execution of a workflow on the benchmark records. Processing configuration describes how records were submitted to the LLM, for example as one complete file, separate file batches, or smaller sequential batches. We further distinguish advancement workflows, whose decisions contributed to the records retrieved for full-text assessment in the parent review, from comparative runs, which were evaluated retrospectively and did not influence retrieval. The evaluation deliberately separated several dimensions of performance: recovery of verified eligible records, retained full-text workload, agreement among workflows, conditional reference-based classification performance among records with verified full-text outcomes, and run-to-run consistency. Record-level analyses additionally examined overlap in missed eligible records, unique recovery, and the consequences of combining pairs of screening outputs under a liberal advancement rule. This design allowed us to examine not only whether LLM-based screening could approximate human performance, but also whether performance depended on processing configuration, whether repeated executions produced the same decisions, and whether human and LLM outputs provided complementary information. The study addressed the following research questions: 1. How did the human and LLM-based screening workflows differ in their recovery of verified eligible records and in the number of records retained for full-text assessment? 2. To what extent did the screening decisions produced by the different workflows agree? 3. How did the workflows perform on conditional reference-based classification measures among records assessed at full text? 4. How consistent were two LLM screening runs conducted under nominally identical conditions? 5. How did the workflows differ in the verified eligible records they failed to retain, and what did retrospective checks reveal about records retained only by a later comparative LLM run? Screening time, interaction burden, and other procedural characteristics were additionally summarised where corresponding data were available. Pairwise combined recovery was examined as an exploratory analysis. Materials and Methods Study design, preregistration, and open materials This preregistered comparative methodological study was embedded in a scoping review of gaze semantics in human social interaction (FigalovĂĄ et al., 2026). The methodological study was preregistered on the Open Science Framework on February 24, 2026, after database searching and deduplication but before the preliminary title-only screen and all title-and-abstract screening procedures evaluated here. The preregistration specified the principal human and LLM-based screening comparisons and analytical approach. Deviations from the preregistered plan and exploratory additions are documented in Supplementary Material S1. Additional open materials comprise detailed definitions of the outcome measures and statistical analyses (S2), the screening manual and version history (S3), input data (S4), complete LLM prompts (S5), screening outputs (S6), comprehensive analytical results (S7), the analysis script (S8), and the codebook (S9). These materials are publicly available on the Open Science Framework. The complete search strategies and PRISMA-ScR reporting for the parent review are provided in the companion manuscript (FigalovĂĄ et al., 2026). The present study concerns the comparative performance of the human and LLM-based title-and-abstract screening workflows and therefore reports only the review procedures required to define the screening task, benchmark dataset, and reference outcomes. Benchmark dataset and analytical sets Searches for the parent review were conducted in PubMed, PsycINFO, Web of Science Core Collection, and OSF Preprints on February 16, 2026, followed by a supplementary Web of Science search on February 23, 2026. After deduplication, the search corpus contained 5,291 records. The benchmark was constructed by the authors from bibliographic records retrieved for the parent scoping review and was not a pre-existing curated machine-learning or classification dataset. The review lead (NF) conducted a conservative preliminary title-only screen to remove records whose titles clearly indicated that they did not meet the eligibility criteria. Records were retained whenever eligibility could not be determined reliably from the title alone. This preliminary screen removed 4,160 records and was used only to construct the benchmark dataset; it was not evaluated as a screening workflow in the present study. The resulting benchmark comprised 1,131 records and is provided in Supplementary Material S4. The single-reviewer workflow, distributed-team workflow, and all seven LLM runs independently evaluated the same 1,131 records at the title-and-abstract level. Records identified subsequently through backward citation searching were considered in the parent review but were not added to the benchmark. All performance estimates therefore concern the 1,131 records remaining after the preliminary title-only screen, whose accuracy was not evaluated here. Of the 1,131 benchmark records, 922 were selected for full-text retrieval and 209 were not advanced. Full-text reports were successfully obtained and assessed for 859 records; 63 could not be retrieved. Among the 859 full-text-assessed records, 316 were judged eligible and 543 ineligible. Three analytical sets were therefore distinguished: âą the complete benchmark dataset (n=1,131n=1,131), used to quantify retained workload, agreement, run-to-run consistency, and record-level decision overlap; âą the full-text-assessed set (n=859n=859), used for conditional reference-based classification measures; and âą the verified eligible set (n=316n=316), used to estimate operational recall and missed eligible records. The relationship between outcome measures and analytical sets is summarised in Table 2. Data preprocessing The benchmark dataset was derived from bibliographic records retrieved for the parent scoping review. Database outputs were combined and deduplicated before the present study was preregistered. Following the conservative preliminary title-only screen described above, the remaining 1,131 records were assigned stable identifiers. For LLM screening, input files contained only the record identifier, title, and abstract. Records were divided into files or batches where required by the respective processing configuration, without substantive modification of the title or abstract text. No author-controlled tokenisation, embedding generation, feature engineering, model training, or fine-tuning was performed. After screening, LLM outputs were checked for structural problems such as missing or duplicated identifiers, missing decisions, and malformed rows. Outputs were linked by stable record identifier and formatting was standardised where necessary. No LLM screening decision was manually changed on substantive grounds. For the primary analyses, Include and Unclear were subsequently mapped to retained, whereas Exclude and Excludeâcitation seed were mapped to not retained. Screening task and decision structure An initial screening manual was developed before human calibration to operationalise the parent reviewâs eligibility criteria and provide a common decision framework for the human and LLM-based workflows. The manual was refined following human calibration and finalised before formal title-and-abstract screening began. The complete manual and its version history are provided in Supplementary Material S3. Records were considered potentially eligible when gaze or a related eye behaviour was substantively examined as carrying, expressing, signalling, cueing, shaping, or supporting the interpretation of meaning within a human social context. This included social interpretations of direct or averted gaze, eye contact, mutual gaze, gaze shifts, and gaze avoidance. Records were excluded when their primary focus concerned developmental or clinical populations, psychiatric classification, non-social or low-level visual attention, methods-only gaze research, or technology evaluation without a central focus on the human interpretation of gaze meaning. Studies involving robots, avatars, or virtual agents remained eligible when they examined human social interpretations of gaze. Each record received one of four mutually exclusive title-and-abstract decisions: 1. Include; 2. Unclear; 3. Exclude; or 4. Excludeâcitation seed. Unclear was assigned when the information available in the title and abstract was insufficient for a confident eligibility decision. Excludeâcitation seed identified an otherwise ineligible record that appeared potentially useful for supplementary backward citation searching. For the primary quantitative analyses, Include and Unclear were mapped to retained, whereas Exclude and Excludeâcitation seed were mapped to not retained. This binary mapping represented the operational decision of whether a record would proceed toward full-text assessment. Human reviewers and LLMs were also instructed to assign one primary exclusion reason to records classified as Exclude. Exclusion reasons were collected for transparency and descriptive inspection but were not included in the primary quantitative comparison. Human screening workflows Single-reviewer workflow The single-reviewer workflow was conducted by the review lead, NF, who screened all 1,131 benchmark records using the final screening manual and assigned one of the four predefined decisions to each record. Screening was conducted between March 2 and March 16, 2026. Distributed-team workflow The distributed-team workflow involved four psychology student assistants enrolled in a masterâs-level psychology programme. Each assistant screened a non-overlapping subset of the benchmark, and their decisions were subsequently combined to constitute one operational team workflow. Before formal screening, NF and all four assistants independently assessed a calibration set of 28 records. The decisions were discussed in a joint meeting to identify sources of disagreement and establish a shared interpretation of the eligibility criteria. The screening manual was then refined and finalised. No second calibration round was conducted. The calibration decisions themselves were not carried forward as final screening decisions. The 28 records remained in their original positions in the benchmark and were reassessed during formal screening by NF and by the assistant responsible for the corresponding file. No additional overlap subset, ongoing double-screening, or formal adjudication procedure was implemented during formal distributed-team screening. The benchmark was divided sequentially, in the original Zotero-export order, into 12 non-overlapping Excel files: 11 files containing 100 records and one containing 31 records. Neither records nor files were randomised. Files were placed in a shared cloud folder and selected by assistants on a first-come-first-served basis; allocation was tracked separately to prevent overlap. The four assistants screened 300, 331, 200, and 300 records, respectively. Because assistants assessed different non-random subsets, assistant-level differences could reflect both reviewer behaviour and file composition and were therefore interpreted descriptively rather than as direct comparisons of reviewer performance. The assistants were instructed not to use LLMs, automated summaries, AI-assisted classification tools, or other generative-AI applications. Distributed-team screening was conducted between March 2 and March 13, 2026. During formal screening, the assistants did not have access to NFâs decisions, and NF did not have access to the assistantsâ decisions. All human title-and-abstract screening was completed before the LLM-based runs began. Each assistant recorded their screening time. Aggregate distributed-team time represented summed person-time across assistants rather than elapsed calendar time. LLM-based screening workflows Models and run design The LLM-based procedures were conducted through the standard hosted web interfaces of ChatGPT 5.4 Thinking, Gemini 3 Thinking, and Gemini 3.1 Pro. Model names are reported exactly as displayed in the respective interfaces at the time of data collection. The providers did not expose the hardware, GPU configuration, inference software stack, model weights, or complete backend version used for individual runs; these infrastructure details were therefore unavailable to the authors. Seven complete LLM runs were conducted between March 17 and March 24, 2026. Each run constituted one complete pass through all 1,131 benchmark records. The implemented runs formed a purposive, non-factorial set of modelâconfiguration combinations. General-purpose reasoning-oriented LLMs available through standard hosted web interfaces were selected because the study aimed to evaluate screening procedures that could be implemented directly by review teams without review-specific model training, fine-tuning, or specialised computing infrastructure. Models from two provider ecosystems were included to avoid restricting the evaluation to a single hosted system. Processing configurations were selected to compare practically distinct ways of submitting the same screening task and to examine configuration effects and repeat-run consistency, rather than to estimate independent causal effects of model and processing configuration. Comparisons are therefore descriptive and refer to complete implemented workflows. The design was not intended to estimate independent causal effects of model identity and processing configuration. Accordingly, comparisons refer to complete implemented workflows rather than isolated model effects. The seven runs and their analytical roles are summarised in Table 1. These included file-batch and all-at-once configurations for ChatGPT 5.4 Thinking and Gemini 3 Thinking, a nominally identical repeat of the initial ChatGPT 5.4 file-batch procedure, a Gemini 3.1 Pro file-batch run, and an interactive ChatGPT 5.4 configuration processing records sequentially in groups of 10. Default interface settings were used. Memory was disabled, and user-adjustable sampling parameters were not modified. Input files contained only a stable record identifier, title, and abstract. Where records were divided across files, each file was processed in a separate conversation. Development of the LLM instructions The final human screening manual was translated into a structured LLM prompt specifying the eligibility criteria, decision categories, uncertainty rule, exclusion rules, and required output structure. The prompt was developed iteratively with assistance from ChatGPT 5.4 Thinking and was reviewed and refined by NF to ensure alignment with the human screening manual. The prompt was tested procedurally using the same 28 records used for human calibration. This testing was limited to execution-related properties, including whether all requested records were processed, identifiers were preserved, required variables were returned, and the requested output structure was followed. Human screening decisions and full-text eligibility outcomes were not supplied to the models during prompt development or testing. The 28 records remained in the benchmark and were reassessed in every formal LLM run. The calibration set was not used to select among prompt variants on the basis of agreement with human decisions, recall, or subsequent full-text eligibility. No systematic prompt-sensitivity analysis was conducted. Because ChatGPT 5.4 Thinking, one of the evaluated models, assisted with prompt development, the resulting estimates are interpreted as evaluations of the complete implemented workflows rather than model-independent estimates of screening capability. Complete prompts are archived in Supplementary Material S5. Processing configurations Approximately 100-record file batches. The benchmark was divided into the same 12 files used for distributed-team screening: 11 files containing 100 records and one containing 31 records. Each file was processed in a separate conversation. This configuration provided smaller independent outputs and allowed record coverage and output completeness to be inspected at the file level. The substantive screening criteria, decision categories, uncertainty rule, and requested output variables were held constant within the corresponding model-specific comparisons. All records at once. One Excel file containing all 1,131 benchmark records was uploaded in a single conversation. The model was instructed to assess every record from its title and abstract and return a structured output containing the record identifier, title, abstract, screening decision, and, for excluded records, one predefined exclusion reason. Interactive sequential groups of 10 records. The interactive configuration used ChatGPT 5.4 Thinking with the same screening manual, eligibility criteria, decision categories, and uncertainty rule. The model was instructed to process exactly 10 records per response, provide a decision and justification for each record, stop after each group, and wait for the instruction ânext.â The prompt and screening manual (Supplementary Materials S5 and S3) were uploaded at the beginning of each conversation, followed by one approximately 100-record input file. NF entered only ânextâ between successive groups and provided no substantive feedback or correction. Once all records in a file had been assessed, the model was asked to generate an Excel file reproducing the decisions displayed in the conversation. Spreadsheet generation was treated as transcription of the preceding screening output rather than as an independent classification step. Output handling, integrity checks, and safeguards Stable record identifiers were preserved throughout the input, model-output, merging, and analysis pipeline. During import and merging, outputs were inspected for apparent problems in record coverage, duplicate or missing identifiers, missing decisions, and malformed rows. These integrity checks were used to identify obvious structural failures but were not conducted under a prespecified or fully systematic validation protocol. In the interactive sequential run, the final generated spreadsheets were not exhaustively compared record by record with every decision displayed earlier in the conversations. No unresolved structural problem was apparent in the final merged dataset. Human title-and-abstract decisions, advancement status, full-text eligibility outcomes, and eventual inclusion in the parent review were absent from the files and prompts supplied to the models. LLM outputs were saved before they were linked to human decisions or full-text outcomes. Manual post-processing was restricted to identifier-based merging and formatting standardisation; no LLM screening decision was changed manually on substantive grounds. Additional safeguards were used to reduce outcome leakage, record loss, and post hoc intervention. Memory was disabled; file-batch inputs were processed in separate conversations; stable identifiers were retained throughout; and the instructions explicitly required ambiguous or insufficiently described records to receive an Unclear decision rather than a confident exclusion. These procedures were intended to reduce the risk of overconfident exclusion and contamination by study-specific reference outcomes. Because the evaluated systems were proprietary hosted services, their training corpora were unavailable and possible prior exposure to the underlying publications or bibliographic records could not be assessed. The investigators also had no control over undocumented model updates or other server-side changes. The study therefore does not constitute external validation of the underlying models and cannot isolate model-specific effects from all properties of the hosted interfaces. Its inferential target is the set of modelâpromptâprocessing-configuration workflows implemented under the conditions described here. Advancement to full-text assessment and reference outcomes Before the corresponding workflows were run and before their outputs were inspected, the parent review prospectively designated four advancement workflows: âą the single-reviewer workflow; âą the distributed-team workflow; âą Gemini 3 Thinking using approximately 100-record file batches; and âą the initial ChatGPT 5.4 Thinking approximately 100-record file-batch run. A record was selected for full-text retrieval when at least one of these four workflows assigned Include or Unclear. Thus, advancement followed a liberal union rule and did not require agreement, majority voting, or consensus adjudication. Records assigned Exclude or Excludeâcitation seed by all four advancement workflows were not selected for retrieval. The file-batch LLM configurations were prospectively selected as advancement workflows because their smaller, separate outputs were expected to permit more transparent inspection of record coverage and output completeness. The all-at-once configurations were designated as comparative runs. Their outputs were subsequently found to be structurally complete enough for inclusion in the methodological comparison, but they did not contribute to retrieval decisions. The retrieval set was fixed before the nominally identical ChatGPT file-batch repeat run, the Gemini 3.1 Pro run, and the interactive sequential run were conducted. Decisions from these later comparative runs therefore did not retrospectively alter which records underwent full-text retrieval. NF alone assessed the 859 retrieved full-text reports against the original eligibility criteria. Workflow- and run-specific title-and-abstract decisions were not displayed during full-text assessment. Because NF had previously completed the single-reviewer title-and-abstract screen, however, full-text assessment should not be interpreted as fully blinded to all prior screening experience. The operational reference standard therefore relied on a single full-text assessor. Full-text decisions and primary exclusion reasons were recorded using a standardised decision form and exclusion-reason scheme. Final eligibility was verified only for records selected by the union of the four advancement workflows and successfully retrieved. Eligible records missed by all four advancement workflows could therefore remain unidentified among the 209 non-advanced records, while the eligibility of the 63 unretrieved records remained unknown. Reference-based measures consequently quantify performance within the implemented review pathway rather than accuracy against a fully and independently verified benchmark. Comparisons involving later or comparative LLM runs are correspondingly not fully symmetric. Outcome measures and statistical analysis For the analyses, screening output denotes the decisions produced by one complete human workflow or LLM run. The nine screening outputs comprised two human workflows and seven LLM runs. Table 2 summarises the outcome measures, their interpretation, and the records contributing to each analysis. Detailed definitions, formulas, and additional interpretive limitations are provided in Supplementary Material S2. The primary screening trade-off was evaluated jointly using retained workload and operational recall. Retained workload was the number and proportion of the 1,131 benchmark records classified as Include or Unclear and therefore retained for possible full-text assessment. Operational recall quantified the proportion of the 316 verified eligible records retained by each workflow. Because final eligibility was not verified across the complete benchmark, this measure is an operational estimate rather than complete-benchmark recall. Conditional reference-based classification performance was evaluated among the 859 records with verified full-text outcomes. These analyses included precision, specificity, and F1F_1 alongside recall. Their interpretation is conditional on full-text verification and therefore does not extend to the 209 non-advanced or 63 unretrieved records with unknown final eligibility. Agreement was evaluated independently of final eligibility across the complete benchmark. All 36 pairwise comparisons among the nine screening outputs were analysed. Pairwise retained/not-retained agreement was summarised using overall agreement, Cohenâs Îș, Gwetâs AC1, and positive and negative agreement. AC1 was included alongside Îș because Îș can be strongly influenced by category prevalence and marginal decision distributions (Byrt et al., 1993; Gwet, 2008). Run-to-run consistency was examined specifically for the two ChatGPT 5.4 file-batch runs conducted under nominally identical conditions. In addition to aggregate performance, the analysis examined whether the same individual records were retained across runs and whether discordant decisions involved verified eligible records. Record-level analyses further examined the overlap of missed eligible records, records uniquely recovered by individual outputs, and complementarity between workflows. Pairwise liberal-union analyses treated a record as retained when either member of a pair retained it, allowing the gain in recovered eligible records to be considered jointly with the corresponding increase in retained workload. Ninety-five per cent confidence intervals were estimated using 10,000 resamples. For operational recall, precision, specificity, and F1F_1, resampling was stratified by full-text eligibility status within the full-text-assessed set. For agreement measures, confidence intervals were estimated through multinomial resampling of the corresponding observed 2Ă22Ă 2 table. Percentile intervals were obtained from the resulting empirical distributions. Because the complete constructed benchmark was analysed, these confidence intervals do not quantify uncertainty about the fixed observed values within this corpus. Instead, they describe resampling variability when the observed records are treated as an empirical distribution and may support cautious generalisation to comparable records. The intervals do not incorporate uncertainty about the eligibility of non-advanced or unretrieved records and do not account for potential clustering by assistant, source file, conversation, or multiple reports from the same underlying study. No paired hypothesis tests or between-workflow contrast intervals were calculated. Consequently, overlap or non-overlap between confidence intervals for separate workflows was not interpreted as statistical evidence of a difference between them. As a sensitivity analysis, the principal retained-workload, operational-recall, missed-record, conditional-classification, and agreement analyses were repeated after excluding the 28 records used for human calibration and procedural prompt testing, with all relevant denominators recalculated. No a priori power calculation was conducted because the complete available benchmark was analysed and the study was designed as a descriptive methodological comparison rather than a confirmatory hypothesis test. Software and computational reproducibility Analyses were conducted in R version 4.6.1 (2026-06-24 ucrt) using RStudio 2026.06.0 on Windows 11. Packages included tidyverse version 2.0.0, knitr version 1.51, and ggrepel version 0.9.8. The random seed used for resampling was 20260703. The analysis code, benchmark input data, LLM prompts, screening outputs, detailed statistical definitions, and comprehensive analytical results are archived in the accompanying open materials (Supplementary Materials S2 and S4âS8), allowing the reported comparisons to be reconstructed from the archived screening outputs. Re-execution of the proprietary hosted LLM workflows cannot be guaranteed to reproduce the original outputs because the underlying models and web interfaces are externally controlled and may change over time. Results Results are organised according to the five research questions. Exploratory analyses of pairwise workflow combinations, descriptive measures of screening effort and procedural burden, and the sensitivity analysis excluding calibration and prompt-testing records are reported subsequently. Detailed results, including complete pairwise comparisons, are provided in the Supplementary Materials. RQ1: Recovery of verified eligible records and retained workload The principal comparison considered operational recall jointly with retained workload (Table 3; Figure 1). No individual screening output recovered all 316 verified eligible records, and greater retention of benchmark records did not consistently correspond to greater recovery. Figure 1: Operational recall as a function of retained screening workload across human and LLM-based screening workflows. The two human workflows and the two GPT-5.4 file-batch runs occupied a closely situated region of the workloadârecall space. They retained 42.2â45.0% of the 1,131-record benchmark while recovering 82.3â82.9% of the verified eligible records. Within this group, the single reviewer retained the fewest benchmark records. The interactive GPT-5.4 batches-of-10 run had a similar retained workload (44.5%) but lower operational recall (79.7%). Gemini 3.1 file batches achieved the highest operational-recall point estimate (83.9%) but retained 56.7% of the benchmark. Gemini 3 file batches retained the largest proportion of records overall (63.1%) while recovering 77.2% of the verified eligible set. Thus, the most liberal workflow did not achieve the highest recovery. Within both model families for which file-batch and all-at-once configurations were evaluated, the file-batch runs recovered more verified eligible records. GPT-5.4 file batches achieved 82.6% operational recall compared with 66.8% for GPT-5.4 all at once. Gemini 3 file batches achieved 77.2% compared with 66.1% for Gemini 3 all at once. The associated workload patterns differed, however: GPT-5.4 all at once produced the lowest retained workload of all outputs (34.6%), whereas Gemini 3 all at once retained 59.0% of the benchmark. These results indicate substantial differences between processing configurations but, given the non-factorial design, do not isolate configuration as their causal source. RQ2: Agreement between screening workflows Binary retained/not-retained agreement varied considerably across workflow pairs (Table 4). Agreement between the single reviewer and distributed team was 72.8% (Îș=0.447Îș=0.447; AC1 = 0.464). The GPT-5.4 file-batch workflow showed a similar level of agreement with the two human workflows: 71.2% with the single reviewer and 74.0% with the distributed team. By contrast, Gemini 3 file batches agreed with the single reviewer and distributed team on 55.0% and 55.7% of records, respectively. Runs using the same displayed model but different processing configurations also produced different decisions. Agreement between GPT-5.4 file batches and GPT-5.4 all at once was 82.5%, whereas agreement between the two Gemini 3 configurations was 62.4%. Thus, sharing the same displayed model did not imply closely matching screening outputs across processing configurations. The highest agreement among the selected comparisons occurred between the two GPT-5.4 file-batch runs conducted under nominally identical conditions (91.7%). Their remaining record-level disagreement is examined under RQ4. Agreement did not map directly onto recovery of eligible records. Workflows with similar overall agreement could differ in which records they retained and in the consequences of their discordant decisions, supporting separate consideration of agreement and reference-based performance. RQ3: Conditional classification performance among full-text-assessed records Reference-based classification measures were calculated within the 859 records for which full-text eligibility was verified, comprising 316 eligible and 543 ineligible records (Table 5). Because all 316 verified eligible records belonged to this set, recall was numerically identical to the operational recall reported under RQ1. The single reviewer achieved the highest conditional F1F_1 score (0.693). The initial GPT-5.4 file-batch run (F1=0.670F_1=0.670), its rerun (F1=0.667F_1=0.667), and the distributed-team workflow (F1=0.663F_1=0.663) produced broadly similar conditional profiles. The interactive GPT-5.4 batches-of-10 run was slightly lower (F1=0.653F_1=0.653). GPT-5.4 all at once generated the fewest false positives and achieved the highest conditional specificity (0.718), but its lower recall resulted in an F1F_1 score of 0.621. Conversely, Gemini 3.1 file batches achieved the highest operational recall but retained more full-text-assessed ineligible records, yielding lower precision and specificity. The two Gemini 3 configurations had the lowest conditional F1F_1 scores overall. Gemini 3 file batches showed the lowest precision and specificity, whereas Gemini 3 all at once produced the lowest F1F_1 score. Accordingly, no single metric identified a uniformly superior workflow. Higher specificity could coincide with substantial loss of verified eligible records, while higher recall could require retaining substantially more ineligible records. These estimates are conditional on the 859 full-text-assessed records and should not be interpreted as complete-benchmark classification performance. RQ4: Run-to-run consistency of the repeated GPT-5.4 file-batch workflow The two GPT-5.4 file-batch runs conducted under nominally identical conditions produced closely similar aggregate performance but differed at the record level. Across the 1,131 benchmark records, the runs agreed on 1,037 records (91.7%) and disagreed on 94 (8.3%). They jointly retained 454 records and jointly did not retain 583; among the discordant records, 39 were retained only by the first run and 55 only by the rerun. Aggregate retained workload, operational recall, precision, specificity, and F1F_1 were closely aligned between the two runs (Tables 3 and 5). Their differences were therefore more apparent at the level of individual records than in the summary measures. Among the 316 verified eligible records, 247 were retained by both runs, 14 only by the first run, and 15 only by the rerun; 40 were missed by both. Thus, 29 verified eligible records received discordant advancement decisions despite the nominally identical procedure. Combining the two runs under a liberal union rule would have recovered 276 verified eligible records (87.3%) while retaining 548 benchmark records (48.5%). These results show that similar aggregate performance did not imply reproducible record-level classifications across repeated executions of the same hosted LLM workflow. RQ5: Missed verified eligible records and retrospective checks Overlap and unique recovery Missed eligible records were not distributed uniformly across screening outputs. Among the 316 verified eligible records, 84 (26.6%) were retained by all nine outputs and 91 (28.8%) were missed by only one. In total, 175 records (55.4%) were therefore retained by at least eight of the nine outputs. At the opposite extreme, seven verified eligible records were recovered by only one output. All seven of these uniquely recovered eligible records were retained by a human workflow: six only by the single reviewer and one only by the distributed team. No verified eligible record was uniquely recovered by an LLM output while being missed by both human workflows. By design, no verified eligible record could have been missed by all nine outputs, because final eligibility could only be established for records advanced by at least one of the four original advancement workflows. The observed distribution therefore characterises disagreement within the verified eligible set and cannot quantify potentially eligible records missed by every advancement workflow. Post hoc review of records retained only by the interactive LLM run Of the complete benchmark, 209 records (18.5%) were not retained by any of the four original advancement workflows and were consequently not selected for full-text retrieval. The later GPT-5.4 batches-of-10 run retained 13 of these records. Had this run contributed to the advancement rule, the retrieval set would therefore have increased from 922 to 935 records. Post hoc reassessment of these 13 titles and abstracts found that 11 were clearly outside the review scope, primarily because they concerned clinical or diagnosis-defined samples, developmental research, humanâcomputer interaction, or were not substantively focused on gaze. Two records could not be confidently excluded from the title and abstract alone and, retrospectively, would have warranted full-text retrieval. Because this assessment was conducted after the retrieval set and analytical denominators had been fixed, it did not alter the operational recall or conditional classification estimates. It nevertheless provides direct evidence that the set of non-advanced records could contain potentially relevant records whose final eligibility remained unverified. Exploratory analysis: pairwise combined recovery Combining screening outputs under a liberal union rule consistently increased recovery relative to the constituent workflows but also increased retained workload. The magnitude of this trade-off depended strongly on the specific pairing. The two human workflows together recovered 310 of 316 verified eligible records (98.1%) while retaining 647 benchmark records (57.2%). Of the recovered eligible records, 212 were retained by both workflows, 50 only by the single reviewer, and 48 only by the distributed team; six verified eligible records were missed by both. Combining the single reviewer with the initial GPT-5.4 file-batch workflow recovered 307 verified eligible records (97.2%) while retaining 648 benchmark records (57.3%), a nearly identical workload to the humanâhuman pair but with three fewer eligible records recovered. Combining the distributed team with GPT-5.4 file batches retained the same number of benchmark records but recovered 296 eligible records (93.7%). The single reviewer combined with Gemini 3.1 file batches recovered 310 verified eligible records, matching the humanâhuman pair, but retained 762 benchmark records (67.4%). This represented 115 additional retained records for the same recovery. Pairing either human workflow with Gemini 3 file batches resulted in still greater retained workloads. Among selected LLM-only combinations, Gemini 3 file batches combined with GPT-5.4 file batches recovered 295 verified eligible records while retaining 820 benchmark records. Combining the two GPT-5.4 file-batch runs recovered 276 eligible records while retaining 548 benchmark records. Thus, complementarity was present across several human and LLM combinations, but additional recovery could not be considered independently of the extra full-text workload generated. Among the pairs recovering 310 verified eligible records, the humanâhuman combination required the lowest retained workload. Screening effort and procedural burden Human screening effort The single reviewer screened all 1,131 records in 1,176 minutes of logged active screening time, corresponding to 1.04 minutes per record or 57.7 records per hour (Table 6). The distributed-team workflow required 1,628 minutes of summed person-time, corresponding to 1.44 minutes per record or 41.7 records per hour. Individual assistant screening rates ranged from 39.3 to 44.4 records per hour. These estimates refer to formal screening only. The single-reviewer measure excluded breaks and recovery time, while the distributed-team estimate excluded approximately three hours devoted to the introductory meeting, independent calibration screening, and subsequent calibration discussion. Distributed-team time represents summed person-time and is therefore not equivalent to elapsed calendar time. LLM procedural burden LLM timing measures describe interface processing and output generation rather than continuous human labour and are therefore not directly comparable with manual screening time. Active user interaction was not logged prospectively and is reported only as an approximate retrospective estimate. GPT-5.4 all at once required 46 minutes and 29 seconds of visible screening time and 5 minutes and 45 seconds for spreadsheet generation. Because this procedure involved a single conversation, active user interaction was estimated at approximately one to two minutes. The initial GPT-5.4 file-batch run used 12 conversations and required 66 minutes of visible screening time plus 8 minutes and 36 seconds for spreadsheet generation. The nominally identical rerun required 52 minutes plus 7 minutes and 18 seconds, respectively. Active user interaction for each file-batch run was estimated at approximately one to two minutes per conversation, or approximately 12â24 minutes in total. The interactive GPT-5.4 batches-of-10 run required 92 minutes and 54 seconds of visible screening time and approximately 49 minutes for spreadsheet generation. Because the model had to be prompted repeatedly to continue, active interaction was estimated at approximately five minutes per file. Combining batch outputs into the analytical dataset required approximately five additional minutes per run. Comparable timing data were unavailable for the Gemini workflows. Although the recorded GPT-5.4 procedures required substantially less active human interaction than manual screening, the comparison remains descriptive because human and LLM timing measures represented different forms of effort and were not collected using a common timing protocol. Sensitivity analysis excluding calibration and prompt-testing records Excluding the 28 records used for human calibration and procedural LLM prompt testing left 1,103 benchmark records, including 306 verified eligible records. The principal findings were unchanged. Across workflows, retained workload changed by between â0.7-0.7 and 0.20.2 percentage points, operational recall by between â0.003-0.003 and 0.0070.007, precision by between â0.003-0.003 and 0.0040.004, specificity by between â0.005-0.005 and 0.0070.007, and F1F_1 by between â0.003-0.003 and 0.0030.003. Workflow ordering on the principal measures remained unchanged: Gemini 3.1 file batches retained the highest operational-recall point estimate, the single reviewer retained the highest conditional F1F_1 score, and GPT-5.4 all at once retained the highest specificity while continuing to show substantially lower recall than the file-batch workflows. Agreement estimates were similarly stable. Agreement between the single reviewer and distributed team was 73.0% after exclusion, compared with 72.8% in the primary analysis. Agreement between the two GPT-5.4 file-batch runs remained 91.7%, and agreement between the distributed team and the initial GPT-5.4 file-batch run was 74.1%. Excluding the calibration and prompt-testing records therefore did not materially alter the principal performance or agreement patterns. Discussion Principal findings This preregistered methodological study compared two human workflows and seven LLM-based screening runs in a conceptually complex title-and-abstract classification task. The results show that screening performance was a property of the implemented workflow rather than of the displayed model alone. Across the evaluated procedures, recovery of verified eligible records, retained workload, agreement, conditional classification performance, complementarity, and run-to-run consistency captured distinct aspects of performance, and no individual output recovered all 316 verified eligible records. The two human workflows and the two GPT-5.4 file-batch runs occupied a similar region of the workloadârecall space, retaining approximately 42â45% of the benchmark while recovering approximately 82â83% of the verified eligible records. Gemini 3.1 file batches achieved the highest operational-recall point estimate, but at substantially greater retained workload. In contrast, both all-at-once configurations recovered only about two thirds of the verified eligible records. Greater retention therefore did not translate consistently into greater recovery, and the configuration with the lowest retained workload also missed a substantial proportion of the verified eligible set. Processing configuration was associated with marked differences within both model families for which file-batch and all-at-once procedures were compared. File-batch processing recovered more verified eligible records than all-at-once processing for both GPT-5.4 and Gemini 3, although the corresponding workload patterns differed. Because model, interface, date, interaction structure, and processing configuration were not varied independently, these comparisons do not establish a causal effect of batching. They nevertheless demonstrate that performance observed for one implementation of a model should not be assumed to generalise automatically to another implementation of the same displayed model. Independent screening outputs were also partly complementary. A liberal union of the two human workflows recovered 98.1% of the verified eligible records while retaining 57.2% of the benchmark. Some humanâLLM combinations approached or matched this level of recovery, but generally at greater retained workload. Repeating the GPT-5.4 file-batch workflow also increased combined recovery relative to either run alone, but the resulting operating point differed from that of the humanâhuman combination. The practical value of combining outputs therefore depended on both the specific pairing and the relative cost assigned to missed evidence versus additional downstream assessment. The repeated GPT-5.4 file-batch runs provide a further important result. Their aggregate workload, recall, and conditional classification measures were closely aligned, and they agreed on 91.7% of benchmark records. Nevertheless, they disagreed on 94 individual records, including 29 verified eligible records retained by only one run. Aggregate similarity therefore concealed meaningful item-level instability. Evaluations of LLM-based classifiers should consequently consider not only summary performance but also whether repeated executions identify the same individual cases. Implications for LLM workflow evaluation and deployment A central implication is that the model name is an insufficient description of an LLM-based decision system. The implemented workflow also comprises input preparation, prompt instructions, uncertainty handling, processing configuration, interaction structure, output generation, integrity checks, post-processing, and the rule by which classifications are translated into downstream actions. Differences in any of these components may alter the behaviour of the deployed system. This interpretation is consistent with evidence that prompt design, output format, decision thresholds, and batch size can affect LLM screening outcomes (DennstĂ€dt et al., 2024; Syriani et al., 2024; Fagerberg et al., 2026). Processing configuration should therefore be documented and evaluated as a substantive workflow characteristic rather than treated as an incidental implementation detail. In the present study, file batching also provided practical advantages for auditing record coverage and inspecting separate outputs. However, the observed differences cannot be attributed specifically to batch size. They could reflect context length, allocation of model attention across records, response-generation constraints, interface behaviour, or interactions among these factors. Likewise, the study does not establish superiority of one model family over another because model, interface, processing configuration, interaction structure, and run date were not manipulated factorially, and the common prompt was developed with assistance from ChatGPT 5.4 Thinking. The appropriate inferential target is therefore the complete modelâpromptâconfiguration workflow. The findings also illustrate why LLM performance should not be represented by a single accuracy-like measure. Operational recall quantified recovery of the verified eligible set, retained workload quantified the number of records that would proceed to later assessment, and agreement quantified similarity between classifiers without indicating whether their shared decisions preserved relevant evidence. Precision, specificity, and F1F_1 described classification only among records for which full-text outcomes were available. These measures favoured different workflows. Gemini 3.1 file batches achieved the highest operational-recall point estimate but retained substantially more records. GPT-5.4 all at once retained the fewest records and achieved comparatively high conditional specificity, yet missed approximately one third of the verified eligible set. In a high-recall task with asymmetric error costs, an apparently efficient reduction in workload may therefore result from an undesirable increase in false-negative decisions. Conversely, high recall can be achieved inefficiently by advancing many records. Evaluation should consequently report recovery and missed cases jointly with the workload required to achieve them rather than optimise either quantity in isolation (Madeyski et al., 2025; Sanghera et al., 2025). Agreement likewise requires careful interpretation. High overall agreement can be dominated by records on which both workflows agree not to retain, particularly when ineligible records are common. More importantly, the repeated-run analysis shows that even relatively high overall agreement may coexist with disagreement on cases that have disproportionate downstream consequences. For stochastic or otherwise nondeterministic LLM workflows, record-level consistency is therefore a separate property from aggregate predictive performance. This distinction has broader relevance for LLM-based decision support. If individual classifications determine which cases receive further assessment, two runs with similar aggregate metrics are not operationally interchangeable when they select meaningfully different cases. Reporting only average recall, precision, or agreement may obscure this instability. Repeated-run evaluations should therefore report the number and characteristics of discordant cases, including high-consequence cases classified differently across runs, rather than relying solely on differences between aggregate estimates. The complementarity analyses further show that combining classifiers changes the operating point rather than simply âimproving accuracy.â Liberal union rules increased recovery but also increased the number of records requiring downstream assessment. The two human workflows were particularly complementary: together they recovered substantially more eligible records than either alone while producing a more favourable workloadârecovery balance than the humanâLLM combination that reached the same recovery. LLM outputs nevertheless contributed additional decisions not shared with individual human workflows, consistent with previous evidence that humanâLLM and LLMâLLM combinations can increase sensitivity (Sanghera et al., 2025). Alternative combination rules would optimise different objectives. For example, systems that accept automated decisions only when multiple LLMs agree and refer discordant cases to humans place greater emphasis on reducing human workload while retaining human oversight of uncertain cases (Hilkenmeier et al., 2026). The appropriate architecture therefore depends on the relative costs of false negatives, false positives, and human review rather than on the assumption that one combination strategy is universally preferable. Among the 316 verified eligible records, all seven records recovered by only one of the nine outputs were uniquely retained by a human workflow. This result should not be interpreted as evidence that LLMs cannot identify eligible records missed by humans because final eligibility was not established for every non-advanced record. The later interactive GPT-5.4 run retained 13 records that had not been advanced by any of the four original workflows. Post hoc title-and-abstract reassessment suggested that 11 were clearly outside scope, whereas two could not be confidently excluded without full-text assessment. Their final eligibility remains unknown. This illustrates an important consequence of partially verified reference standards: records missed by all workflows contributing to verification can remain invisible to standard recall calculations. For deployment, the present results do not support treating an unvalidated LLM run as an autonomous exclusion component in a conceptually complex, high-recall screening task. A more defensible role is as one component of an auditable decision-support system, for example for supplementary screening, prioritisation, or identification of uncertain cases. Such systems require stable identifiers, structured machine-readable outputs, explicit uncertainty rules, checks for missing and duplicate records, transparent post-processing and merging procedures, predefined decision rules, and validation within the intended application context. Reporting should correspondingly describe enough of the complete pipeline to make the evaluated system interpretable: model and interface version, prompts, input preparation, processing configuration, interaction structure, human involvement, output handling, advancement or decision rules, recovery, missed cases, downstream workload, agreement, and repeated-run consistency (Susnjak, 2023; Shailendra et al., 2026; Holst et al., 2025; Gallifant et al., 2025; Luo et al., 2025). Reporting a model name and prompt alone may be insufficient when other workflow components materially affect the resulting decisions. No universally optimal humanâLLM arrangement is therefore implied by these findings. In applications where false negatives are particularly costly, a liberal combination rule may justify additional downstream workload. Where human resources are more constrained, prioritisation or selective human review of uncertain or discordant cases may be more appropriate. The optimal operating point depends on the application, error costs, conceptual complexity, available resources, reviewer expertise, and tolerance for missed cases. Strengths and limitations A principal strength of this study is that it was preregistered and embedded in an actual PRISMA-ScR-aligned review rather than evaluated solely on a retrospectively assembled or artificial benchmark. All nine screening outputs classified the same 1,131 records using the same substantive eligibility criteria and a common operational retained/not-retained rule. The comparison included a solo expert reviewer, a distributed trained human team, multiple practically implementable LLM workflows, different processing configurations, and a nominally repeated LLM workflow. A further strength is the multidimensional evaluation framework. Recovery, retained workload, agreement, conditional classification performance, pairwise complementarity, and run-to-run consistency were analysed separately rather than collapsed into one performance score. This distinction made it possible to identify cases in which superficially favourable performance on one measure corresponded to an unfavourable trade-off on another. The sensitivity analysis excluding the 28 calibration and prompt-testing records did not materially alter the principal findings. The study also used a conservative treatment of the reference outcomes. Records without verified full-text outcomes were not automatically classified as true negatives. This avoided producing apparently complete benchmark metrics by assuming that every record excluded before full-text assessment had been excluded correctly. The corresponding limitation is that the operational reference set was conditional on the parent reviewâs advancement and retrieval procedures. Final eligibility was not independently established for all 1,131 benchmark records. Potentially eligible records may therefore have remained among the 209 non-advanced records or the 63 records whose full texts could not be retrieved. Operational recall consequently represents recovery of the known verified eligible set rather than complete-benchmark sensitivity, while precision, specificity, and F1F_1 are conditional on the full-text-assessed subset. The four-workflow advancement union was selected prospectively before the relevant outputs were inspected, but this specific union was not specified in the preregistration. In addition, those four advancement workflows contributed to construction of the verified eligible set, whereas the later comparative LLM runs did not. Comparisons between advancement and comparative workflows are therefore partly asymmetric. Full-text eligibility was assessed by the review lead rather than by multiple independent assessors. Although workflow-specific title-and-abstract decisions were not displayed during full-text assessment, the review lead had previously completed the single-reviewer title-and-abstract screen. The verified eligible set should therefore be interpreted as an operational reference standard rather than an independent consensus gold standard. The human results are likewise specific to the implemented arrangements. The review lead had substantially greater familiarity with the conceptual scope, whereas the distributed workflow comprised four calibrated masterâs-level psychology assistants who screened non-overlapping subsets without ongoing double-screening or formal adjudication. Differences among assistants could also be confounded with the non-random record subsets they screened. Other training, expertise, overlap, or consensus procedures could produce different human performance. The LLM workflows were executed through proprietary hosted web interfaces. This increased ecological validity for researchers using readily available systems but reduced experimental control and strict reproducibility. Underlying model identifiers, weights, training corpora, server-side settings, and undocumented model updates were unavailable. The repeated GPT-5.4 workflow should consequently be understood as a nominal repetition of the observable procedure rather than a strict computational replication. The non-factorial design further limits causal interpretation. Model identity, processing configuration, interaction structure, interface, and date were not independently manipulated. The study can therefore establish that the evaluated workflows behaved differently, but not determine which individual component caused those differences. Similarly, because the common LLM prompt was developed with assistance from ChatGPT 5.4 Thinking, performance should not be interpreted as arising from a model-neutral prompt-development process. Output-integrity validation was also imperfect. Stable identifiers were maintained and outputs were inspected for missing records, duplicated identifiers, missing classifications, and malformed rows, but these checks were not conducted under a fully prespecified validation protocol. For the interactive batches-of-10 workflow, the final generated spreadsheets were not exhaustively compared record by record with every classification previously displayed in the conversations. Although no unresolved structural problem was apparent in the final analytical dataset, future evaluations should incorporate automated and prespecified completeness checks and verify exported outputs directly against the classifications generated during model interaction. Procedural burden was not measured using a common protocol across human and LLM workflows. Human measures represented active screening or summed person-time, whereas LLM measurements included visible interface processing, output generation, waiting, and subsequent data handling. Active LLM interaction was estimated retrospectively, and comparable timing was unavailable for the Gemini workflows. The study therefore supports only a descriptive comparison of procedural burden and cannot establish standardised time or economic efficiency. Generalisability is limited by the use of one interdisciplinary and conceptually complex scoping review. The benchmark was also constructed after a preliminary title-only screen removed 4,160 of the 5,291 deduplicated records. The resulting 1,131-record benchmark was therefore enriched for records that could not be confidently excluded from their titles alone. This may have affected difficulty, class prevalence, retained workload, and agreement and may also have advantaged the review lead through previous exposure to the complete search corpus. The study did not systematically vary prompts, languages, eligibility contexts, or review datasets and did not include external validation in an independent corpus. Possible prior exposure of the proprietary models to the underlying publications or bibliographic records could not be assessed because their training data were unavailable. Human screening decisions, advancement status, full-text outcomes, and eventual inclusion in the parent review were, however, not supplied to the models. The findings should therefore be interpreted as evidence about the specific workflows evaluated here rather than as general validation of the displayed models for literature screening or document classification. Future studies would benefit from independent datasets, prespecified factorial comparisons of workflow components, multiple prompt variants, systematic output-integrity checks, complete or independently sampled verification of non-advanced records, and repeated executions sufficient to characterise record-level variability. Finally, the findings reflect models and hosted interfaces available in March 2026. Rapid changes in proprietary LLM systems limit the durability of model-specific performance estimates. This strengthens the case for evaluating and documenting the complete workflow actually deployed rather than relying on performance reported previously for a model carrying the same or a similar name. Conclusion In this conceptually complex screening task, the two GPT-5.4 file-batch runs produced workloadârecall profiles close to those of the human workflows, whereas the all-at-once configurations recovered substantially fewer verified eligible records. No individual output recovered the complete verified eligible set, and greater recovery sometimes required substantially greater downstream workload. Recovery, workload, agreement, conditional classification performance, complementarity, and run-to-run consistency therefore represented distinct dimensions of screening performance. The AI approach evaluated here was the prompt-based application of general-purpose proprietary LLMs rather than a review-specific model trained or fine-tuned for screening. This approach was selected because it represents an immediately implementable form of LLM-assisted screening that does not require labelled training data, local model deployment, or specialised machine-learning infrastructure. The findings consequently concern the performance of complete implemented screening workflows rather than a newly developed model architecture or the intrinsic capabilities of the displayed models. The two human workflows were particularly complementary, jointly recovering 98.1% of verified eligible records while retaining fewer records than the humanâLLM combination that achieved the same recovery. LLM workflows also contributed complementary decisions, but their value depended on the workflow with which they were combined and the resulting workload. Moreover, similar aggregate results across nominally repeated LLM runs concealed meaningful differences in the individual records retained. These findings support using LLMs as documented, auditable, and human-supervised components of evidence-synthesis workflows rather than as autonomous replacements for human screening judgement. For conceptually complex reviews, the relevant question is therefore not simply which model is used, but how the complete workflow is implemented, validated, combined with human judgement, assessed for record-level variation, and transparently reported. The present findings are specific to the evaluated review, models, prompts, interfaces, and processing configurations and should not be interpreted as external validation of the displayed models across evidence-synthesis contexts. Future evaluations should test such workflows prospectively across independent review topics and model versions, with more complete reference verification and systematic variation of prompts and processing configurations. Competing interests The authors declare that they have no competing interests. Author contributions Nikol FigalovĂĄ: Conceptualization, Methodology, Investigation, Data curation, Formal analysis, Visualization, Project administration, Writing â original draft, and Writing â review and editing. Lynn Huestegge: Conceptualization, Methodology, Supervision, Funding acquisition, and Writing â review and editing. Anne Böckler-Raettig: Conceptualization, Methodology, Supervision, Funding acquisition, and Writing â review and editing. All authors reviewed and approved the final manuscript. Funding This work was funded by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under project number 562993814. Ethics statement This methodological study used bibliographic records and screening outputs produced during an evidence-synthesis project. No patient data or directly identifiable personal data were collected or analysed. Formal ethics approval was not required for this study. Data availability The preregistration for this study is available through the Open Science Framework at https://osf.io/cnsr2/overview?view_only=697b0ae9f0234f069bc18173626ba954. The screening manual, prompts, input data, screening outputs, detailed results, and analysis code supporting the findings are available through the Open Science Framework at https://osf.io/v2q4r. Generative AI use Generative AI tools were used both as part of the research methodology and during manuscript preparation. As described in the Methods, OpenAI ChatGPT 5.4 Thinking, Google Gemini 3 Thinking, and Google Gemini 3.1 Pro were used to perform the LLM-based screening workflows evaluated in this study. Full records of the prompts and methodological procedures used for these research workflows were retained and are provided in the associated supplementary materials and research repository. During manuscript preparation, OpenAI ChatGPT 5.4 and ChatGPT 5.6 were used for language editing and rephrasing, restructuring and shortening sections, assistance with LaTeX, checking analyses, and assistance with the development of R code. ChatGPT 5.4 also assisted with development of the R code used to generate Figure 1; the figure itself was generated from the study data using R and was not created using an AI image-generation model. The authors reviewed and verified all AI-assisted text, code, analyses, references, and other outputs for accuracy and originality. The terms of use of the AI tools were reviewed and considered suitable for the uses described above and for publication. The authors take full responsibility for the integrity and accuracy of the entire manuscript, including its references. Related work This manuscript contains original unpublished work and is not under consideration for publication elsewhere. The methodological study was embedded in the scoping review reported separately by FigalovĂĄ et al. The companion report presents the substantive findings of the scoping review, whereas the present manuscript evaluates the human and LLM-based screening workflows. The relationship between the two reports and any overlapping methods or data is disclosed in the manuscript. References Akinseloyin et al. (2026) O. Akinseloyin, X. Jiang, and V. Palade Large language model-based multiagent collaboration for abstract screening toward automated systematic reviews. Biology Methods and Protocols 11 (1), p. bpag006. External Links: Document Cited by: LLM-based screening as a workflow-level problem. Belur et al. (2021) J. Belur, L. Tompson, A. Thornton, and M. Simon Interrater reliability in systematic review methodology: exploring variation in coder decision-making. Sociological Methods & Research 50 (2), p. 837â865. External Links: Document Cited by: Evaluating screening performance. Byrt et al. (1993) T. Byrt, J. Bishop, and J. B. Carlin Bias, prevalence and kappa. Journal of Clinical Epidemiology 46 (5), p. 423â429. External Links: Document Cited by: Outcome measures and statistical analysis. Cao et al. (2026) C. Cao, R. Arora, P. Cento, A. Budak, K. Manta, E. Farahani, M. Cecere, A. Selemon, J. Sang, L. X. Gong, R. Kloosterman, S. Jiang, R. Saleh, D. Margalik, J. Lin, J. Jomy, J. Xie, D. Chen, J. Gorla, S. Lee, K. Zhang, J. Kuang, H. Ware, M. Whelan, B. Teja, A. A. Leung, R. K. Arora, J. Pillay, L. Hartling, M. Noetel, D. B. Emerson, A. S. Detsky, A. C. Tricco, G. M. Church, D. Moher, and N. Bobrovitz Automation of systematic reviews with large language models. Note: medRxiv preprintPreprint; latest version checked: version 4, May 4, 2026 External Links: Document Cited by: LLM-based screening as a workflow-level problem. DennstĂ€dt et al. (2024) F. DennstĂ€dt, J. Zink, P. M. Putora, J. Hastings, and N. Cihoric Title and abstract screening for literature reviews using large language models: an exploratory study in the biomedical domain. Systematic Reviews 13, p. 158. External Links: Document Cited by: LLM-based screening as a workflow-level problem, Implications for LLM workflow evaluation and deployment. Fagerberg et al. (2026) P. Fagerberg, O. Sallander, K. V. Patil, A. Berg, A. Nyman, N. Borg, and T. LindĂ©n Batch size effects on mid-2025 state-of-the-art large language model performance in automated title and abstract screening. Cochrane Evidence Synthesis and Methods 4 (3), p. e70082. External Links: Document Cited by: LLM-based screening as a workflow-level problem, Implications for LLM workflow evaluation and deployment. FigalovĂĄ et al. (2026) N. FigalovĂĄ, A. Böckler, and L. Huestegge Gaze semantics: a scoping review of how gaze conveys and constrains meaning in social contexts. PsyArXiv. Note: Preprint External Links: Document, Link Cited by: The present study, Study design, preregistration, and open materials, Study design, preregistration, and open materials. Galli et al. (2025) C. Galli, A. V. Gavrilova, and E. Calciolari Large language models in systematic review screening: opportunities, challenges, and methodological considerations. Information 16 (5), p. 378. External Links: Document Cited by: LLM-based screening as a workflow-level problem. Gallifant et al. (2025) J. Gallifant, M. Afshar, S. Ameen, Y. Aphinyanaphongs, S. Chen, G. Cacciamani, D. Demner-Fushman, D. Dligach, R. Daneshjou, C. Fernandes, L. H. Hansen, A. Landman, L. Lehmann, L. G. McCoy, T. Miller, A. Moreno, N. Munch, D. Restrepo, G. Savova, R. Umeton, J. W. Gichoya, G. S. Collins, K. G. M. Moons, L. A. Celi, and D. S. Bitterman The tripod-llm reporting guideline for studies using large language models. Nature Medicine 31 (1), p. 60â69. External Links: Document Cited by: LLM-based screening as a workflow-level problem, Implications for LLM workflow evaluation and deployment. Gwet (2008) K. L. Gwet Computing inter-rater reliability and its variance in the presence of high agreement. British Journal of Mathematical and Statistical Psychology 61 (1), p. 29â48. External Links: Document Cited by: Outcome measures and statistical analysis. Hanfstingl et al. (2024) B. Hanfstingl, S. Oberleiter, J. Pietschnig, U. S. Tran, and M. Voracek Detecting jingle and jangle fallacies by identifying consistencies and variabilities in study specificationsâa call for research. Frontiers in Psychology 15, p. 1404060. Cited by: Introduction. Harasgama et al. (2026) S. Harasgama, H. Pearce, C. Appel, L. Loftus, H. Painter, I. Kuhn, J. Karpusheff, A. Ceesay, and J. Ford Artificial intelligence tools for automating evidence synthesis: scoping review. Journal of Medical Internet Research 28, p. e81597. External Links: Document Cited by: LLM-based screening as a workflow-level problem. Hilkenmeier et al. (2026) F. Hilkenmeier, M. Stoltenberg, and C. Stierle Using full agreement across multiple large language models for title-and-abstract screening in systematic reviews: a proof-of-concept. Systematic Reviews 15, p. 191. External Links: Document Cited by: Evaluating screening performance, Implications for LLM workflow evaluation and deployment. Holst et al. (2025) D. Holst, K. Moenck, J. Koch, O. Schmedemann, and T. SchĂŒppstuhl Transparent reporting of ai in systematic literature reviews: development of the prisma-traice checklist. JMIR AI 4, p. e80247. External Links: Document Cited by: LLM-based screening as a workflow-level problem, Implications for LLM workflow evaluation and deployment. Homiar et al. (2025) A. Homiar, J. Thomas, E. G. Ostinelli, J. Kennett, C. Friedrich, P. Cuijpers, M. Harrer, S. Leucht, C. Miguel, A. Rodolico, Y. Kataoka, T. Takayama, K. Yoshimura, R. So, Y. Tsujimoto, Y. Yamagishi, S. Takagi, M. Sakata, Ä. BaĆĄiÄ, et al. Development and evaluation of prompts for a large language model to screen titles and abstracts in a living systematic review. BMJ Mental Health 28 (1), p. e301762. External Links: Document Cited by: Evaluating screening performance, Introduction. Huotala et al. (2024) A. Huotala, M. Kuutila, P. Ralph, and M. MĂ€ntylĂ€ The promise and challenges of using llms to accelerate the screening process of systematic reviews. In Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering, EASE 2024, New York, NY, USA, p. 273â283. External Links: Document Cited by: LLM-based screening as a workflow-level problem. Krag et al. (2024) C. H. Krag S. Balschmidt et al. Large language models for abstract screening in systematic- and scoping reviews: a diagnostic test accuracy study. Note: medRxiv preprintPreprint External Links: Document Cited by: LLM-based screening as a workflow-level problem. Laignelot et al. (2026) F. Laignelot, G. L. Martin, M. Ossman, O. Pingeon, A. Boubaker, E. Picovschi, J. Kim, X. Tannier, J. F. Cohen, and A. Dechartres Large language models show promising performance for some systematic review tasks but call for cautious implementation: a systematic review. Journal of Clinical Epidemiology 194, p. 112221. External Links: Document Cited by: LLM-based screening as a workflow-level problem. Landschaft et al. (2024) A. Landschaft, D. Antweiler, S. Mackay, S. Kugler, S. RĂŒping, S. Wrobel, T. Höres, and H. Allende-Cid Implementation and evaluation of an additional gpt-4-based reviewer in prisma-based medical systematic literature reviews. International Journal of Medical Informatics 189, p. 105531. External Links: Document Cited by: Evaluating screening performance, Introduction. Li et al. (2024) M. Li, J. Sun, and X. Tan Evaluating the effectiveness of large language models in abstract screening: a comparative analysis. Systematic Reviews 13 (1), p. 219. External Links: Document Cited by: LLM-based screening as a workflow-level problem. Lieberum et al. (2025) J. Lieberum, M. Töws, M. Metzendorf, F. Heilmeyer, W. Siemens, C. Haverkamp, D. Böhringer, J. J. Meerpohl, and A. Eisele-Metzger Large language models for conducting systematic reviews: on the rise, but not yet ready for useâa scoping review. Journal of Clinical Epidemiology 181, p. 111746. External Links: Document Cited by: LLM-based screening as a workflow-level problem. Luo et al. (2025) X. Luo, Y. C. Tham, M. GiuffrĂš, R. Ranisch, M. Daher, K. Lam, A. V. Eriksen, C. Hsu, A. Ozaki, F. Y. de Moraes, S. Khanna, K. Su, E. BegagiÄ, Z. Bian, Y. Chen, J. Estill, and G. W. Group Reporting guideline for the use of generative artificial intelligence tools in medical research: the gamer statement. BMJ Evidence-Based Medicine 30 (6), p. 390â400. External Links: Document Cited by: LLM-based screening as a workflow-level problem, Implications for LLM workflow evaluation and deployment. Madeyski et al. (2025) L. Madeyski, B. Kitchenham, and M. Shepperd LLM4SCREENLIT: recommendations on assessing the performance of large language models for screening literature in systematic reviews. Note: arXiv preprintPreprint External Links: 2511.12635, Document Cited by: Evaluating screening performance, Implications for LLM workflow evaluation and deployment. Njei et al. (2026) B. Njei, Y. A. Al-Ajlouni, U. Sidney Kanmounye, S. Boateng, G. Loic Nguefang, N. Njei, et al. Artificial intelligence agents in healthcare research: a scoping review. PLOS ONE 21 (2), p. e0342182. External Links: Document Cited by: LLM-based screening as a workflow-level problem. Peters et al. (2020) M. D. J. Peters, C. Marnie, A. C. Tricco, D. Pollock, Z. Munn, L. Alexander, P. McInerney, C. M. Godfrey, and H. Khalil Updated methodological guidance for the conduct of scoping reviews. JBI Evidence Synthesis 18 (10), p. 2119â2126. External Links: Document Cited by: Introduction. Sanghera et al. (2025) R. Sanghera, A. J. Thirunavukarasu, M. El Khoury, J. OâLogbon, Y. Chen, A. Watt, M. Mahmood, H. Butt, G. Nishimura, and A. A. S. Soltan High-performance automated abstract screening with large language model ensembles. Journal of the American Medical Informatics Association 32 (5), p. 893â904. External Links: Document Cited by: Evaluating screening performance, Implications for LLM workflow evaluation and deployment, Implications for LLM workflow evaluation and deployment. Sciurti et al. (2026) A. Sciurti, G. Migliara, L. M. Siena, C. Isonne, M. R. De Blasiis, A. Sinopoli, J. Iera, C. Marzuillo, C. De Vito, P. Villari, and V. Baccolini Compact large language models for title and abstract screening in systematic reviews: an assessment of feasibility, accuracy, and workload reduction. Research Synthesis Methods 17 (2), p. 332â347. Note: Published online 13 November 2025 External Links: Document Cited by: Evaluating screening performance. Shailendra et al. (2026) S. Shailendra, R. Kadel, A. Sharma, I. M. Tahidul, and U. R. Saxena L-prisma: an extension of prisma in the era of generative artificial intelligence (genai). Note: TechRxiv preprintPreprint; posted February 6, 2026 External Links: Document Cited by: LLM-based screening as a workflow-level problem, Introduction, Implications for LLM workflow evaluation and deployment. Sokolova and Lapalme (2009) M. Sokolova and G. Lapalme A systematic analysis of performance measures for classification tasks. Information Processing & Management 45 (4), p. 427â437. External Links: Document Cited by: Evaluating screening performance. Susnjak (2023) T. Susnjak PRISMA-dfllm: an extension of prisma for systematic literature reviews using domain-specific finetuned large language models. Note: arXiv preprintPreprint External Links: 2306.14905, Document Cited by: LLM-based screening as a workflow-level problem, Implications for LLM workflow evaluation and deployment. Syriani et al. (2024) E. Syriani, I. David, and G. Kumar Screening articles for systematic reviews with chatgpt. Journal of Computer Languages 80, p. 101287. External Links: Document Cited by: LLM-based screening as a workflow-level problem, LLM-based screening as a workflow-level problem, Implications for LLM workflow evaluation and deployment. Tricco et al. (2018) A. C. Tricco, E. Lillie, W. Zarin, K. K. OâBrien, H. Colquhoun, D. Levac, D. Moher, M. D. J. Peters, T. Horsley, L. Weeks, S. Hempel, E. A. Akl, C. Chang, J. McGowan, L. Stewart, L. Hartling, A. Aldcroft, M. G. Wilson, C. Garritty, S. Lewin, C. M. Godfrey, M. T. Macdonald, E. V. Langlois, K. Soares-Weiser, J. Moriarty, T. Clifford, Ă. Tunçalp, and S. E. Straus PRISMA extension for scoping reviews (prisma-scr): checklist and explanation. Annals of Internal Medicine 169 (7), p. 467â473. External Links: Document Cited by: Introduction. Tversky and Kahneman (1974) A. Tversky and D. Kahneman Judgment under uncertainty: heuristics and biases. Science 185 (4157), p. 1124â1131. External Links: Document Cited by: Evaluating screening performance. van de Schoot et al. (2021) R. van de Schoot, J. de Bruin, R. Schram, P. Zahedi, J. de Boer, F. Weijdema, B. Kramer, M. Huijts, M. Hoogerwerf, G. Ferdinands, A. Harkema, J. Willemsen, Y. Ma, Q. Fang, S. Hindriks, L. Tummers, and D. L. Oberski An open source machine learning framework for efficient and transparent systematic reviews. Nature Machine Intelligence 3 (2), p. 125â133. External Links: Document Cited by: LLM-based screening as a workflow-level problem. Waffenschmidt et al. (2019) S. Waffenschmidt, M. Knelangen, W. Sieben, S. BĂŒhn, and D. Pieper Single screening versus conventional double screening for study selection in systematic reviews: a methodological systematic review. BMC Medical Research Methodology 19, p. 132. External Links: Document Cited by: Evaluating screening performance. Wang et al. (2020) Z. Wang, T. Nayfeh, J. Tetzlaff, P. OâBlenis, and M. H. Murad Error rates of human reviewers during abstract screening in systematic reviews. PLOS ONE 15 (1), p. e0227742. External Links: Document Cited by: Evaluating screening performance. Whiting et al. (2011) P. F. Whiting, A. W. S. Rutjes, M. E. Westwood, S. Mallett, J. J. Deeks, J. B. Reitsma, M. M. G. Leeflang, J. A. C. Sterne, and P. M. M. Bossuyt QUADAS-2: a revised tool for the quality assessment of diagnostic accuracy studies. Annals of Internal Medicine 155 (8), p. 529â536. External Links: Document Cited by: Evaluating screening performance. Wilkins (2023) D. Wilkins Automated title and abstract screening for scoping reviews using the gpt-4 large language model. Note: arXiv preprintPreprint External Links: 2311.07918, Document Cited by: LLM-based screening as a workflow-level problem. Wu et al. (2026) Y. Wu, G. Wang, Y. Chen, J. Zhang, Y. Zhang, Y. Chen, J. Shang, G. Zhang, and Z. Liu PRISM: probing reasoning, instruction, and source memory in llm hallucinations. Note: arXiv preprintPreprint External Links: 2604.16909, Document Cited by: Evaluating screening performance. Table 1: LLM-based title-and-abstract screening workflows evaluated in the study. Each run comprised one complete pass through all 1,131 benchmark records. Model Date Processing configuration Role in study ChatGPT 5.4 Thinking March 17, 2026 Files of approx. 100 records Advancement workflow ChatGPT 5.4 Thinking March 17, 2026 All 1,131 records at once Comparative run ChatGPT 5.4 Thinking March 18, 2026 Files of approx. 100 records Nominally identical repeat of the March 17 file-batch workflow ChatGPT 5.4 Thinking March 19â24, 2026 Sequential groups of 10 records Exploratory comparative run Gemini 3 Thinking March 17, 2026 Files of approx. 100 records Advancement workflow Gemini 3 Thinking March 17, 2026 All 1,131 records at once Comparative run Gemini 3.1 Pro March 18, 2026 Files of approx. 100 records Comparative run Note. Model names are reported exactly as displayed in the hosted interfaces at the time of data collection. The March 18 ChatGPT file-batch run repeated the March 17 procedure using the same displayed model, prompt, record order, file structure, and interaction procedure; undocumented changes to the hosted system cannot be excluded. The sequential groups-of-10 configuration constituted one complete run conducted across several days. Table 2: Evaluation dimensions, outcome measures, and analytical sets used to address the research questions. RQ Evaluation dimension Outcome measures Analytical set RQ1 Recovery and retained workload Operational recall and missed eligible records quantified recovery of verified eligible records. Retained workload quantified the number of benchmark records retained for possible full-text assessment. 316 verified eligible records; all 1,131 benchmark records RQ2 Agreement between screening outputs Overall agreement, Cohenâs Îș, Gwetâs AC1, and positive and negative agreement compared binary retained/not-retained decisions. All 1,131 benchmark records RQ3 Conditional classification performance Precision, specificity, and F1F_1 described classification performance among records with verified full-text outcomes. 859 full-text-assessed records RQ4 Run-to-run consistency The two nominally identical GPT-5.4 file-batch runs were compared using record-level agreement, retained workload, operational recall, missed eligible records, and discordant decisions. All 1,131 benchmark records, including 316 verified eligible records RQ5 Missed-record patterns and retrospective checks Record-level analyses examined overlap in missed eligible records, unique recovery, and records that alternative procedures would have advanced. Thirteen non-advanced records retained by the later sequential run were reassessed post hoc. All 1,131 benchmark records, including 316 verified eligible records; 13 reassessed records Explor. Pairwise workflow complementarity A liberal union rule quantified recovery and retained workload when a record was retained if either member of a workflow pair retained it. All 1,131 benchmark records, including 316 verified eligible records Suppl. Procedural burden Human screening time and throughput and available LLM timing, interaction, and conversation-level measures were summarised descriptively. Available process logs Note. RQ = research question. Operational recall was calculated against the 316 verified eligible records and retained workload against the complete 1,131-record benchmark. Precision, specificity, and F1F_1 were conditional on the 859 records assessed at full text and should not be interpreted as complete-benchmark classification estimates. Table 3: Retained workload and recovery of verified eligible records for each screening workflow. Workflow Retained workload, n (%) Eligible retained, n (%) Eligible missed, n (%) Operational recall (95% CI) Single reviewer 477 (42.2%) 262 (82.9%) 54 (17.1%) 0.829 (0.785â0.870) Distributed team 509 (45.0%) 260 (82.3%) 56 (17.7%) 0.823 (0.778â0.864) Gemini 3: file batches 714 (63.1%) 244 (77.2%) 72 (22.8%) 0.772 (0.725â0.816) Gemini 3: all at once 667 (59.0%) 209 (66.1%) 107 (33.9%) 0.661 (0.608â0.715) GPT-5.4: file batches 493 (43.6%) 261 (82.6%) 55 (17.4%) 0.826 (0.785â0.867) GPT-5.4: all at once 391 (34.6%) 211 (66.8%) 105 (33.2%) 0.668 (0.614â0.718) Gemini 3.1: file batches 641 (56.7%) 265 (83.9%) 51 (16.1%) 0.839 (0.797â0.877) GPT-5.4: file-batch rerun 509 (45.0%) 262 (82.9%) 54 (17.1%) 0.829 (0.788â0.870) GPT-5.4: batches of 10 503 (44.5%) 252 (79.7%) 64 (20.3%) 0.797 (0.753â0.842) Note. Retained workload was calculated across all 1,131 benchmark records. Eligible retained, eligible missed, and operational recall were calculated against the 316 records judged eligible after full-text assessment. Operational recall is therefore an estimate against the verified eligible set rather than complete-benchmark recall. Confidence intervals are 95% percentile intervals from 10,000 stratified bootstrap resamples. Table 4: Selected pairwise agreement comparisons for binary retained/not-retained decisions across the 1,131-record benchmark. Workflow pair Agreement, % (95% CI) Positive agreement (95% CI) Negative agreement (95% CI) Cohenâs Îș (95% CI) Gwetâs AC1 (95% CI) Humanâhuman Single reviewer vs distributed team 72.8 (70.2â75.3) 0.688 (0.653â0.720) 0.759 (0.732â0.784) 0.447 (0.393â0.499) 0.464 (0.413â0.516) Single reviewer vs LLM Single reviewer vs Gemini 3 file batches 55.0 (52.1â57.8) 0.573 (0.539â0.605) 0.525 (0.488â0.560) 0.135 (0.083â0.187) 0.102 (0.045â0.160) Single reviewer vs GPT-5.4 file batches 71.2 (68.4â73.8) 0.664 (0.629â0.698) 0.748 (0.720â0.773) 0.412 (0.356â0.465) 0.435 (0.380â0.488) Distributed team vs LLM Distributed team vs Gemini 3 file batches 55.7 (52.7â58.6) 0.590 (0.556â0.623) 0.518 (0.479â0.554) 0.137 (0.081â0.190) 0.120 (0.060â0.180) Distributed team vs GPT-5.4 file batches 74.0 (71.4â76.5) 0.707 (0.673â0.737) 0.767 (0.740â0.791) 0.473 (0.420â0.523) 0.487 (0.434â0.537) Within-model processing-configuration comparisons Gemini 3 file batches vs Gemini 3 all at once 62.4 (59.5â65.3) 0.692 (0.663â0.720) 0.518 (0.476â0.557) 0.211 (0.153â0.268) 0.283 (0.224â0.342) GPT-5.4 file batches vs GPT-5.4 all at once 82.5 (80.3â84.7) 0.776 (0.745â0.805) 0.856 (0.836â0.876) 0.635 (0.590â0.680) 0.666 (0.622â0.709) Run-to-run consistency GPT-5.4 file batches vs GPT-5.4 file-batch rerun 91.7 (90.0â93.2) 0.906 (0.886â0.924) 0.925 (0.909â0.940) 0.832 (0.797â0.863) 0.836 (0.802â0.867) Note. Agreement was calculated for the binary retained/not-retained classification across all 1,131 benchmark records. Positive agreement quantifies concordance in retaining records; negative agreement quantifies concordance in not retaining records. Confidence intervals are 95% percentile intervals from 10,000 multinomial resamples of the observed 2Ă22Ă 2 agreement table for each workflow pair. Complete results for all 36 workflow pairs are provided in the Supplementary Materials. Table 5: Conditional classification performance among the 859 records with verified full-text outcomes. Workflow FP TN Precision (95% CI) Specificity (95% CI) F1F_1 (95% CI) Single reviewer 178 365 0.595 (0.565â0.627) 0.672 (0.634â0.711) 0.693 (0.662â0.723) Distributed team 208 335 0.556 (0.527â0.585) 0.617 (0.575â0.657) 0.663 (0.633â0.692) Gemini 3: file batches 427 116 0.364 (0.346â0.381) 0.214 (0.179â0.250) 0.494 (0.470â0.518) Gemini 3: all at once 335 208 0.384 (0.359â0.409) 0.383 (0.343â0.424) 0.486 (0.453â0.518) GPT-5.4: file batches 202 341 0.564 (0.535â0.594) 0.628 (0.587â0.669) 0.670 (0.640â0.699) GPT-5.4: all at once 153 390 0.580 (0.543â0.617) 0.718 (0.681â0.755) 0.621 (0.583â0.658) Gemini 3.1: file batches 317 226 0.455 (0.434â0.477) 0.416 (0.376â0.459) 0.590 (0.565â0.615) GPT-5.4: file-batch rerun 208 335 0.557 (0.529â0.587) 0.617 (0.576â0.657) 0.667 (0.638â0.696) GPT-5.4: batches of 10 204 339 0.553 (0.523â0.584) 0.624 (0.582â0.665) 0.653 (0.622â0.683) Note. Classification measures were calculated only among the 859 records assessed at full text, comprising 316 verified eligible and 543 ineligible records. FP = false positives, defined here as full-text-ineligible records retained during title-and-abstract screening; TN = true negatives, defined as full-text-ineligible records not retained during title-and-abstract screening. Recall is not repeated because it is numerically identical to the operational recall reported in Table 3. These estimates are conditional on full-text verification and are not complete-benchmark classification measures. Confidence intervals are 95% percentile intervals from 10,000 stratified bootstrap resamples. Table 6: Logged human screening time and throughput during formal title-and-abstract screening. Workflow or reviewer Records screened Time, min Min per record Records per hour Single reviewer 1,131 1,176 1.04 57.7 Distributed team, total 1,131 1,628 1.44 41.7 Assistant 1 300 418 1.39 43.1 Assistant 2 331 505 1.53 39.3 Assistant 3 200 300 1.50 40.0 Assistant 4 300 405 1.35 44.4 Note. Times refer to formal screening only. Calibration, training, breaks, and recovery time were excluded. Screening times for the single reviewer and assistants were self-reported in the screening sheets. Distributed-team time represents summed person-time across assistants rather than elapsed calendar time and is therefore not directly comparable with the elapsed processing times reported descriptively for the LLM workflows.