Paper deep dive
Artificial Intelligence Can Match Domain Experts in Evidence Extraction and Critical Appraisal of Microbial Oncogenesis Research Publications
Kaela Kokkas, Hairong Wang, Richard Klein, Nazir A. Ismail, Natalie Irwin, Mohammad Z. Moonsamy, Kubendran Naidoo, Jeremy Nel, Ekene E. Nweke, Raveen Parboosing, Emmanuel K. Sekyi, Rebecca T. van Dorsten, Bruce A. Bassett, Robert F. Breiman
Intelligence
Status: not_run | Model: - | Prompt: - | Confidence: 0%
Entities (0)
Relation Signals (0)
No relation signals yet.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying novel microbial oncogenicity could yield strategies that will reduce disease burdens. However, relevant evidence is dispersed and infeasible for humans to comprehensively synthesize. LLMs may enable scalable, expert-level systematic evidence synthesis to identify microbe-cancer pairs; however, such capabilities have not yet been demonstrated. Domain experts were recruited to create a dataset to benchmark LLM performance (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, GPT-5 Nano) on 24 research papers using MMTV-LV and breast cancer as a case study. We devised a structured template for evidence extraction and appraisal, consisting of MCQ, Likert-scale, multi-select, and free-text question types (77 items across 24 papers). Agreement between (1) experts and (2) experts and each LLM was determined per question instance using novel metrics. LLMs were assessed by comparing inter-expert and expert-LLM agreement distributions to determine whether LLMs behaved as additional experts by increasing or maintaining inter-expert agreement. Free-text responses were further evaluated qualitatively. Across all question types, LLM responses aligned closely with experts, with GPT-5 and GPT-5 Nano achieving score distributions indistinguishable from experts. Gemini models behaved similarly but were significantly more lenient in applying microbial oncogenesis criteria. Hallucinations were rare. Methodological appraisal and identification of contradictions within full-texts were the most persistent LLM vulnerabilities. GPT-5 and GPT-5 Nano were indistinguishable from experts on structured domain research paper evaluation tasks. This supports use of LLMs for automated systematic evidence synthesis. However, methodological appraisal tasks and contradiction identification in full-texts remain weaknesses requiring strengthening.
Tags
Links
- Source: https://arxiv.org/abs/2608.07250v1
- Canonical: https://arxiv.org/abs/2608.07250v1
Trouble viewing inline? Open PDF directly â
Full Text
162,539 characters extracted from source content.
Expand or collapse full text
ARTIFICIAL INTELLIGENCE CAN MATCH DOMAIN EXPERTS IN EVIDENCE EXTRACTION AND CRITICAL APPRAISAL OF MICROBIAL ONCOGENESIS RESEARCH PUBLICATIONS Kaela Kokkas 1 2423686@students.wits.ac.za Hairong Wang 4,15 Richard Klein 4,15 Nazir A. Ismail 2 Natalie Irwin 3,5 Mohammad Z. Moonsamy 4,5 Kubendran Naidoo 5,6,7,8,9,10 Jeremy Nel 5,11 Ekene E. Nweke 5,12 Raveen Parboosing 13 Emmanuel K. Sekyi 14 Rebecca T. van Dorsten 5,6,7 Bruce A. Bassett 4,5,15 Robert F. Breiman 5,15,16 1 Department of Clinical Microbiology and Infectious Diseases, Faculty of Health Sciences, University of the Witwater- srand, Johannesburg, South Africa 2 Department of Clinical Microbiology and Infectious Diseases, National Health Laboratory Service and Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa 3 Division of Medical Oncology, Department of Internal Medicine, University of the Witwatersrand, Johannesburg, South Africa 4 School of Computer Science and Applied Mathematics, University of the Witwatersrand, Johannesburg, South Africa 5 Infectious Diseases and Oncology Research Institute (IDORI), Faculty of Health Sciences, University of the Witwa- tersrand, Johannesburg, South Africa 6 South African Medical Research Council Vaccines and Infectious Diseases Analytics Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa 7 South African Medical Research Council Wits Antiviral Gene Therapy Research Unit, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa 8 National Health Laboratory Service, Johannesburg, South Africa 9 Department of Molecular Medicine and Haematology, School of Pathology, University of the Witwatersrand, Johan- nesburg, South Africa 10 Wits Research Institute for Malaria, Faculty of Health Sciences, National Health Laboratory Service, University of the Witwatersrand, Johannesburg, South Africa 11 Division of Infectious Diseases, School of Clinical Medicine, Faculty of Health Sciences, University of the Witwater- srand, Johannesburg, South Africa 12 Department of Surgery, Faculty of Health Sciences, University of the Witwatersrand, Johannesburg, South Africa 13 Division of Virology, University of the Witwatersrand & National Health Laboratory Service, Johannesburg, South Africa 14 OncoVectra, London, United Kingdom 15 Wits Machine Intelligence and Neural Discovery (MIND) Institute, University of the Witwatersrand, Johannesburg, South Africa 16 Department of Global Health, Rollins School of Public Health, Emory University, Atlanta, GA United States arXiv:2608.07250v1 [q-bio.QM] 7 Aug 2026 AI Can Match Domain Experts in Evidence Extraction and Appraisal ABSTRACT Background: Confirmed oncogenic microbes contribute significantly to cancer burden. Identifying and confirming novel microbial oncogenicity could yield strategies and tools that will reduce disease burdens. However, relevant evidence may be dispersed across a vast biomedical literature that is infeasible for humans to comprehensively synthesize. Large Language Models (LLMs) may enable scalable, expert-level systematic evidence synthesis to identify high priority microbe-cancer pairs; however, such capabilities have not yet been demonstrated. Methods: Domain experts were recruited to create a human-validated test dataset to benchmark the performance of LLMs (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, and GPT-5 Nano) on 24 original research papers using Mouse Mammary Tumor Virus-Like Virus and breast cancer as a case study. We devised a structured template for evidence extraction and appraisal of papers, consisting of multiple choice, Likert-scale, multi-select, and free-text question types (77 question items across 24 papers). Agreement between (1) experts, and (2) experts and each LLM, was determined per question instance using novel scoring metrics. LLMs were assessed by comparing inter-expert and expert-LLM agreement score distributions to determine whether LLMs behaved as additional experts by either increasing or maintaining inter-expert agreement. Free-text responses were further evaluated qualitatively. Results: Across all question types, LLM responses aligned closely with expert assessments, with two models (GPT-5, GPT-5 Nano) achieving score distributions statistically indistinguishable from those of experts. Gemini models behaved similarly for most tasks but were significantly more lenient in ap- plying microbial oncogenesis criteria, often over-attributing criteria fulfillment. Hallucinations were rare, although more frequent in smaller models (Gemini 2.5 Flash, GPT-5 Nano). Methodological appraisal and identification of contradictions within full-text papers were the most persistent areas of LLM vulnerability, however, the error rate could not be directly compared with experts. Conclusions: Two LLMs (GPT-5, GPT-5 Nano) were indistinguishable from domain experts on structured domain research paper evaluation tasks. This evidence supports use of LLMs for auto- mated systematic evidence synthesis. However, methodological appraisal tasks and contradiction identification in full-text papers remain weaknesses requiring further investigation, strengthening, and possibly multi-model strategies. Keywords large language models·artificial intelligence·critical appraisal·evidence extraction·evidence synthesis· evidence evaluation· biomedical literature· microbial oncogenesis 1 Introduction The identification of human papillomavirus (HPV) as the cause of cervical cancer [70] established cancer as a preventable disease through vaccination, screening, and treatment of infection [72]. Since then, 11 microbial Group 1 carcinogens (refers to those agents with sufficient evidence of carcinogenicity in humans) have been identified [29,27], with infectious agents linked to at least 20% of cancer cases worldwide and 30% of cases in sub-Saharan Africa [16,56]. The identification of additional oncogenic microbes could provide opportunities to significantly reduce cancer incidence, morbidity, and mortality via development and integration of targeted interventions into public health policies. However, identifying the most probable and impactful microbe-cancer pairs (MCPs) for further research is a significant challenge, with over 1400 possible microbial species [1], more than 100 cancer types [51,50], and variability in oncogenic mechanisms, the microbial types, and their molecular structures which could impact oncogenicity [47]. This has resulted in a division of efforts with multiple MCPs being investigated simultaneously [42,46,4,74,19]. While the existing literature exploring MCPs is not exhaustive, there are many proposed pairs with compelling evidence that warrant further investigation. Identifying a starting point requires a comprehensive, systematic evidence synthesis process to determine which MCPs offer both the greatest plausibility and public health benefit. However, the existing literature and one million articles published yearly in more than 13 000 active indexed biomedical science journals [20], renders a traditional âhuman-drivenâ literature search and evaluation infeasible, even after screening and filtering for relevance. Artificial Intelligence (AI) provides a potential for augmenting traditional literature search strategies with tools capable of rapidly evaluating and synthesizing available evidence. AI, especially in the form of Large Language Models (LLMs), is increasingly demonstrating value for strengthening and accelerating a variety of facets of medical research, from reviewing and evaluating medical data [36,7,6,22], drug discovery [75], predicting and designing protein folding structures [31,73], to multi-use research assistants [11,53,68,41,67,26], and specialized systemic review tools. Recent advancements in accessibility, affordability, and the capability of pre-trained LLMs, including the advent of reasoning LLMs capable of solving complex reasoning tasks, 2 AI Can Match Domain Experts in Evidence Extraction and Appraisal make analyzing the vast scientific literature potentially feasible. However, concerns regarding the trustworthiness of these tools necessitate validation [45]. These concerns include: (1) LLMsâ tendency to hallucinate and produce plausibly sounding yet factually incorrect outputs [34], and (2) limited assessments of the validity of outputs complicated by the âblack-box problemâ where neural networksâ rationale and critical thinking processes are indiscernible [9]. Thus, to allow for further integration of these systems into current research workflows, we need to ensure that modelsâ outputs are consistently trustworthy, accurate, and controllable when performing these tasks. We envision using AI to filter and analyze literature on MCPs, ultimately producing a ranked list based on a standardized scoring matrix for prioritizing investigations (Fig. 1). To begin validating this proposed method, we initially set up an automated search to retrieve papers from biomedical databases and assessed whether small LLMs could accurately screen these papers using nuanced classification criteria [15]. In this paper, we evaluate and compare pre-trained reasoning LLMsâ abilities to analyze and evaluate individual research papers using an extraction template (questionnaire), and benchmark four LLMsâ assessments (Gemini 2.5 Pro, Gemini 2.5 Flash, GPT-5, and GPT-5 Nano) with that of a panel of experts versed in biomedical science. One specific potential MCP was used for this validation: Human Mammary Tumor Virus (HMTV) / Mouse Mammary Tumor Virus-Like Virus (MMTV-LV) and breast cancer. Our objectives were to determine whether LLMs could (1) interpret nuanced biomedical language, (2) critically appraise research, (3) apply causal criteria to expert standards, and (4) produce accurate, trustworthy outputs with minimal hallucinations or omissions, and ultimately utilize our findings to highlight opportunities for improvement, with potential applicability across other biomedical research domains. 2 Methods 2.1 Extraction Template Development Rigorous evidence syntheses require transparent and reproducible methodology, including details of how studies are appraised and analyzed [23]. We developed a structured, standardized extraction template (questionnaire) to ensure repeatable and thorough information extraction. The template was designed to: (1) control and increase transparency of LLM outputs by documenting processes and reasoning, (2) improve reasoning via chain-of-thought (CoT) prompting [71], (3) ensure consideration of factors beyond each paperâs stated primary results, and (4) enable comparison between human and LLM responses. The template consisted of pre-prompt context describing the LLMâs role and six question sections addressing different aspects of the paper, namely âRelevancy", âPaper Bibliographic Details", âPaper Integrity and Reliability", âSummary of Paper Contents", âStrength of Evidence", and âMicrobial Oncogenesis Criteria" (see Table 1 for more details on each section). The template included different question types, with 22 multiple-choice questions (MCQs), four Likert scales, 10 multi-select questions, and 40 long-answer (free-text) questions. MCQs refer to questions where only one option can be selected, while multi-select questions refer to questions where any number of options can be selected concurrently. Completing all MCQs, Likert scales, and 14 specific long-answer questions was mandatory. Twenty-six long-answer questions and all 10 multi-select questions were conditional, based on the response to the previous question. Questions were iteratively refined prior to final experiments from team discussions and pilot sessions. The full extraction template can be found in the supplementary material. 2.2 Human-Validated Test Dataset Creation 2.2.1 Locating Papers The MCP of HMTV/MMTV-LV and breast cancer was chosen for validation of the system for four reasons. First, its carcinogenicity in humans is unconfirmed, which reflects the intended use-case for our pipeline. Second, there is conflicting evidence for this MCP, enabling us to see how the LLMs would interpret and rate positive, negative, and neutral findings. Third, limited evidence for this MCP in the African context, which is a common challenge across many MCPs under investigation, provides an opportunity to evaluate how the system interprets and applies findings from foreign studies to Africa. Finally, the combination of human and animal research (mouse models) for this MCP allowed us to determine whether LLMs could interpret and apply findings from animal models to humans. Twenty-four original research papers that could assist in determining whether HMTV/MMTV-LV causes breast cancer in humans, with differing methodologies and the type and strength of evidence provided, were selected from the PubMed database [32,49,17,65,60,48,37,10,59,52,57,18,38,25,64,24,33,30,8,21,28,44,55,35]. These 24 papers were drawn from the classification dataset (Relevant category) [15], consisting of an initial 20 that very clearly fulfilled the Relevant classification criteria, supplemented by four papers selected for their unique and nuanced study designs (i.e., where a Relevant article may initially seem unrelated to the topic, and require deeper biomedical 3 AI Can Match Domain Experts in Evidence Extraction and Appraisal Figure 1: Overview of proposed AI system. Initial keyword search is performed for microbe-cancer pairs (MCPs) on research databases, with web scraping and classification of titles and abstracts by small LLMs using pre-defined relevancy criteria. Papers classified as Somewhat relevant may be manually reviewed or included in next steps if limited papers retrieved. Full-text versions of Relevant papers are downloaded and analyzed by LLMs using extraction template. Potentially Relevant papers in references are cross-checked with paper storage database and searched for if absent. LLMs produce final MCP evaluation using individual paper summaries and data from infection and cancer registries to determine geographical impact and prevalence of both the cancer and the microbe (when results are available and accessible). This final MCP evaluation is then added to a dynamic list of MCPs, ranked based on a standardized scoring matrix including factors such as potential impact if targeted, confidence in causal relationship, and novelty, and updated with each new MCP evaluation. knowledge to discern its potential effect on the MCPâs plausibility), enabling evaluation of LLM performance under a variety of conditions. 2.2.2 Expert Recruitment The extraction template consists of two categories of analytical tasks: (1) questions requiring extensive domain-specific knowledge, and (2) extractive, rule-based, or statistical appraisal questions requiring formal research training but not 4 AI Can Match Domain Experts in Evidence Extraction and Appraisal Table 1: Overview of the contents of each question section in the extraction template. Extraction template question sections and their contents Question sectionContents RelevancyRelevancy classification criteria followed by questions asking for a classification of the paper (MCQ) and an explanation of why (long-answer). PaperBibliographic Details Questions for details to trace papers, including the title, publication date, journal of publication, study design, country of origin, and DOI. Paper Integrity and Reliability Questions to determine the reliability of the paper, including whether references were appropriate, and if there was any evidence of a conflict of interest, data or image manipulation, or any contradictions, with a final overall appraisal via a paper reliability score (Likert scale) from 1-5. Summary of Paper Contents Questions for more details of the paper, including the aim and hypothesis, sample characteristics, whether the studyâs findings are applicable to African populations, and important methodological factors of the study. Strength of EvidenceQuestions to determine the strength of the paperâs evidence, including how data was collected, the appropriateness of the sample size, whether statistical analysis was appropriate, if there were any errors in analysis or presentation of results, and whether there are any limitations to using the paperâs findings as evidence for or against the MCP, with a final question asking for an overall appraisal via a strength of evidence score (Likert scale) from 1-5. Microbial Oncogenesis Criteria Descriptions of each criterion followed by questions on whether the paper fulfilled, refuted, left the criterion uncertain, or did not examine the criterion, as well as why (e.g., âAppropriate methodology/tools", âInsufficient power"), and to extract the evidence for or against the criterion from the paper. The values for each of the criterions (1 for fulfilled, -1 for refuted, and 0 for either not investigated or left uncertain) are summed to produce the cumulative oncogenesis score (0-10, treated as a Likert scale), with a final question asking for a rating of the paperâs impact on the MCPâs plausibility (Likert scale) from 1-5 domain specialization. Category one questions were answered by domain experts, with three experts per paper from a group of seven experts (RD, NI, NAI, KN, JN, EEN, and RP), all of whom were recruited from the University of the Witwatersrand and included specialist clinicians and scientists in oncology, clinical microbiology, pathology, infectious diseases, and virology. Category two questions, which were objective and/or extractive in nature and therefore only required one evaluator, were answered by a clinician-scientist trainee with training in data analytics and biostatistics (K). These included extraction of bibliographic details of the paper (e.g., title, publication date, publishing journal), identifying contradictions (further evaluated during team discussions to ensure agreement), and bullet-point summaries of the study samples (see extraction template in supplementary material for questions in each category). 2.2.3 Expert Response Collection Author EKS created a web application for the purposes of this study (available atdata.oncovectra.com) that served as a united interface for extraction and appraisal, with dual display of the extraction template questions and the research paper. Papers were uploaded to the platform and assigned to experts based on training specialty and research interest suitability, while ensuring three experts per paper. Experts received a video tutorial on how to use the platform. Experts were aware that the study aimed to compare LLM and domain expert responses, but were blinded to LLM outputs during annotation. The human-validated test dataset is available in the supplementary material. 2.3 LLM Selection We chose four LLMs across two leading LLM vendors: Google (Gemini 2.5 Pro [13] and Gemini 2.5 Flash [13]) and OpenAI (GPT-5 [62] and GPT-5 Nano [62]). This included two frontier reasoning models (Gemini 2.5 Pro and GPT-5) selected for their competitive performance on various reasoning and medical benchmarks at the time of model selection and two lightweight and much cheaper models (Gemini 2.5 Flash and GPT-5 Nano). All four selected LLMs had context-windowsâ„128K to allow for input of the research paper and extraction template, and had multi-modal 5 AI Can Match Domain Experts in Evidence Extraction and Appraisal vision capability to facilitate interpretation of images and graphical data. Default parameters were used for all models, except for those related to thinking or reasoning effort, where the highest level was used. Initial outputs (using full versions of all 24 papers and the complete extraction template) were generated on 17 November 2025, using stable Gemini API versions gemini-2.5-pro and gemini-2.5-flash, and GPT-5 model API snapshots gpt-5-2025-08-07 and gpt-5-nano-2025-08-07. All subsequent experiments were done using the same APIs. 2.4 Generation of LLM Outputs LLMs evaluated the 24 papers separately, receiving the full extraction template and one of the 24 research papers in each API call. Text outputs were converted to JSON for analysis. The extraction template instructed models to only answer the extraction template if the paper was relevant to the research question (based on provided relevancy criteria). In cases where LLMs did not classify the paper as relevant, it was rerun with a modified prompt excluding this line to enable assessment of the LLMâs answers to the remaining questions. 2.5 Assessing LLM Outputs Krippendorffâs Alpha was initially used to measure inter-rater agreement, however, since this was calculated across three (inter-expert) or four (expert-LLM) raters per question instance, differences in a single response had a large impact on agreement, but did not reflect the broader trends in the data. For example, despite experts reaching consensus (either two or all three experts agreeing) in 92.26% (310/336) of MCQs (Fig. 2), Krippendorffâs Alpha for MCQ questions was 0.34, indicating poor expert agreement [43]. Therefore, question-specific metrics were designed. For clarity, question instances refer to individual questions and their responses (e.g., question 19 for paper one), while question items refer to the question as found in the extraction template (e.g., question 19 across all papers). To determine whether LLMs performed at an expert level, their impact on the inter-expert agreement score distribution was assessed. If LLMs responded to the extraction template at the level of domain experts (i.e., provided responses similar to that of experts), then integrating their responses into the inter-expert group produced a distribution that was either indistinguishable from or higher than the inter-expert distribution. If LLM responses did not resemble that of experts (i.e., behaved as an outlier), then integration of their responses would produce a distribution significantly lower than that of the inter-expert agreement distribution. LLM rater groups were created by including the specified LLMâs response with the expert responses, and are referred to by LLM name or collectively as the expert-LLM groups. 2.5.1 Multiple-Choice Questions (MCQ) For each MCQ, all inter-expert, expert-LLM, and inter-LLM response pairs were compared, producing a score of one if choices matched and zero if they did not. These scores were then averaged to produce a normalized score per question instance for each rater group, referred to as the MCQ match proportion. Additionally, frequency of question instances with full expert agreement (three matching responses), partial agreement (two matching responses), and no agreement (no matching responses) was determined. 2.5.2 Likert-Scale Questions Considering response pairs could differ by a single rating level and still have good agreement; match proportions were deemed inappropriate. Likert-scale questions were therefore assessed in two ways: 1.To determine the distance between response values, for each question instance the absolute differences of response pairs (e.g.,|LLM x Score â Expert y Score|) were averaged for each rater group to produce a normalized difference, referred to as the scale distance. 2.The scale distance is interpreted differently from other question type metrics, with lower values indicating greater agreement. Therefore to determine cumulative performance across all question instances from all question types, scale distances were inverted to enable assimilation (e.g., if the scale question was out of 5, and the scale distance was 1, the scale score would be 5 - 1 = 4). This is referred to as the inverted distance. Considering the four Likert-scale question items were high impact (providing scores for paper reliability, strength of evidence, cumulative oncogenesis causal criteria score, and the impact of the paper on the MCP) sub-analysis was done for each question item. 2.5.3 Multi-select Questions All multi-select questions were in the âMicrobial Oncogenesis Criteria" template section, providing potential expla- nations (e.g., âInsufficient power", âConcerns regarding paper reliability", and âAppropriate methodology/tools") for 6 AI Can Match Domain Experts in Evidence Extraction and Appraisal why a specific microbial oncogenesis criterion was supported, refuted, or remained uncertain. However, preliminary assessment of expert match rates revealed systematic expert disagreement (Fig. 2), with experts selecting opposing options (e.g., âAppropriate methodology/tools" and âInappropriate methodology/tools"). LLMs often returned either all provided options without indicating selection (Gemini Flash 2.5 and GPT-5-Nano), or provided a written explanation for each option. Therefore, no further analysis was performed, as these findings indicate confusion surrounding the question type (for both experts and LLMs) that precludes LLM assessment. 2.5.4 Long-Answer Questions Only question instances with three valid expert responses were used, with 75 question instances across eight question items assessed quantitatively and qualitatively (Table 2). Table 2: Long-answer question items included in LLM assessment and their question numbers. Long-answer question items with their corresponding question numbers Question Number Question 1.1 * Briefly explain why the paper is relevant, somewhat relevant, or irrelevant. 18List the most important factor(s) in the methodology that influence(s) the strength of the studyâs findings. 26.1 ** If there are limitations, what are they? 28.1.4State which kind of evidence is provided according to the tools to consider categories and identify the specific finding/s that fulfil or refute this criterion. (Epidemiologic Association) 28.2.4State which kind of evidence is provided according to the tools to consider categories and identify the specific finding/s that fulfil or refute this criterion. (Histopathologic Association) 28.4.4Identify the specific finding/s that fulfil or refute this criterion. (Experimental evidence of facilitation of oncogenesis) 28.5.4State which kind of evidence is provided according to the tools to consider categories and identify the specific finding/s that fulfil or refute this criterion. (Molecular and Multi-omics evidence for interaction) 28.9.4Identify the specific finding/s that fulfil or refute this criterion. (Plausibility) * Q1.1 follows Q1, which includes the detailed classification criteria. ** Q26.1 follows Q26: âAre there any limitations to using this research as evidence for or against a plausible causative relationship between HMTV/MMTV-like virus and breast cancer?â Quantitative Analysis of Long-Answers Long-answer responses were scored in two ways using an LLM judge (GPT-5, high-reasoning effort) to improve efficiency, supported by evidence that such models can approximate human judgment when assessing LLM outputs [14,58]. GPT-5 was selected due to its strong benchmark performance, and its scoring was validated on a subset of responses from each LLM to ensure reliability (see supplementary material Table S1 and Table S2). In both scoring methods, the LLM judge was blinded to the model being assessed when LLM responses were included. The potential impact of self-preference bias (whereby LLM judges favor their own responses) was considered negligible, as the LLM judge was instructed to score the overlap of response pairs rather than identify the superior response. The scoring methods were as follows: 1.All response pairs were compared and their overlap scored on a continuous scale from 0-4 up to 1 decimal place (see supplementary material for scoring rubric, Fig. S1). These scores were averaged for each rater group to produce a normalized score for each question instance, referred to as the pairwise overlap score. 2. Each LLMâs response was compared to a clearly indicated reference answer (created by combining all three expert responses) and scored on a discrete scale from 0-4 (see supplementary material for scoring rubric, Fig. S2). This score is referred to as the combined overlap score. Qualitative Analysis of Long-AnswersManual evaluation was performed to classify gaps between LLM responses and the combined expert reference answer to distinguish between three possibilities: (1) LLM omissions, (2) LLM hallucinations and/or distortions, and (3) paper-supported additions by the LLM that were omitted from the reference 7 AI Can Match Domain Experts in Evidence Extraction and Appraisal answer. Omissions were defined as information included in the expert response(s) that was absent from the LLM response. Hallucinations referred to instances where LLMs provided information not resembling the original research paper (fabrications), while distortions referred to misinterpretations of the text. Paper-supported additions were defined as information included in the LLM response that was absent from the expert responses but supported by the original paper. Hallucinations, distortions, and paper-supported additions were distinguished by confirming whether the additional information was present in the original research paper. 2.6 Data Analysis Specialized metrics (described in Sec. 2.5) were used to determine inter-rater agreement. Data was dependent and determined to be non-normal using Shapiro-Wilk tests; hence findings were reported as median (IQR) when applicable, with the exception of aggregate performance, where findings were reported as mean±SD. Per-paper scores were calculated by summing the agreement metric scores (MCQ match proportions, Likert-scale inverted distances, and long-answer pairwise overlap scores) per paper (isolated to question items with 3 expert responses). Aggregate performance was determined by averaging these per-paper scores. Friedman tests were used to determine whether there were differences in the distributions between groups (reported as test statisticF r andp-value), followed by post-hoc two-sided Wilcoxon Signed-Rank test with Bonferroni correction (reported as test statisticWandp-value) and calculation of the standardized effect size (r) when the group distributional difference (Friedman test) was determined to be significant (p < 0.05). Silvermanâs rule was used to calculate bandwidth for kernel density estimation (KDE) plots, and distributions were extrapolated. A cross-classified mixed-effect model was fitted to determine whether LLMs tended to match specific expert response patterns, with normalized individual pairwise score as the outcome (score range of 0-1, individual pairs e.g., LLM 1 to expert 1 , using MCQ match score, inverted Likert-scale distance, and long-answer pairwise overlap score), with LLM identity, expert identity, and their interaction as fixed effects, with question type as a covariate, and with question and paper numbers as random effects. All data was analyzed using Python 3.14. 3 Results 3.1 Inter-Expert Agreement 3.1.1 MCQ Agreement Experts displayed strong agreement across MCQs (Fig. 2), with full agreement in the majority of question instances (median MCQ match proportion of 1.00, IQR 0.33 - 1.00). No agreement (full expert disagreement) occurred in 7.7% (26/336) of MCQ instances and was confined to question items on applicability to African populations (5/336), sample size adequacy (6/336), and five microbial oncogenesis criteria (15/336). Most disagreements in the microbial oncogenesis criteria reflected differences in interpretation rather than opposing conclusions, while 1.2% (4/336) of instances involved directly contradictory assessments (e.g., one expert concluding that the findings supported microbial carcinogenicity and another concluding that they refuted it). 3.1.2 Likert-Scale Agreement Overall, expert Likert-scale ratings were closely clustered. For most question instances, experts had partial agreement with the outlier rating being one level away from consensus (Fig. 3). Considering that individuals may differ in their interpretation and application of Likert-scale ratings, deviations of one level indicate good agreement around a central value. No agreement (full expert disagreement) occurred in 19.8% (19/96) of question instances and was most common for cumulative oncogenesis score (7/96) and paper reliability score (5/96). In most instances (15.6%, 15/96), expert responses spanned three consecutive Likert-scale levels (e.g., 1, 2, and 3), while 4.2% (4/96) of instances contained an outlier rating two levels removed from the nearest score. Of the four Likert-scale question items, the strength of evidence and paper impact on MCP plausibility scores had the strongest expert agreement, followed by the paper reliability score, with score clustering and small scale distances. Cumulative oncogenesis score had the most disagreement with larger scale distances, reflecting moderate variability in interpretation and application of microbial oncogenesis criteria. 3.1.3 Long-Answer Question Agreement Expert responses captured similar key points, but differed on important nuances (Fig. 4A). Pairwise overlap scores differed between question items. Q1.1 (see Table 2 for full question items) had the highest scores, indicating strong agreement for reasoning on paper relevancy, while Q18 and Q26.1, which both focused on methodology, had the lowest pairwise overlap scores, indicating minor overlap with significant differences in responses. Manual evaluation found that 8 AI Can Match Domain Experts in Evidence Extraction and Appraisal Figure 2: Expert agreement levels for different question types. Expert consensus (consisting of full and partial agreement) was reached in the majority of MCQ and Likert-scale question instances, while most multi-select question instances had no agreement. Overall, experts had full or partial agreement for the majority (406/488, 83.2%) of question instances (see âCombined"). The frequency of question instances where all three expert responses matched (âFull agreement"), where two expert responses matched (âPartial agreement"), and where no expert responses matched (âNo agreement"), were determined for each question type. CombinedMCQLikert-scaleMulti-select Question Types 0 100 200 300 400 500 Number of Question Instances 210 187 23 196 123 54 19 82 26 19 37 Full agreement Partial agreement No agreement Expert agreement levels by question type Figure 3: Distribution of Likert-scale distances across rater groups. Gemini 2.5 Proâs Likert-scale ratings deviated substantially from experts, resulting in significantly increased scale distances when integrated, while Gemini 2.5 Flash, GPT-5, and GPT-5 Nano provided similar ratings to experts and did not alter the inter-expert scale distance distribution. Likert-scale distances were calculated per question instance (n=96) by averaging the differences between Likert-scale ratings for each rater pair (e.g., expert one and Gemini 2.5 Pro) in the rater group (e.g., Gemini 2.5 Pro). Outliers (circular outlines) were determined using Tukeyâs method (values outside of intervalQ 1 â 1.5Ă IQRand Q 3 + 1.5Ă IQR), whiskers indicate maximum and minimum values within interval. A Friedman test with post-hoc two-sided Wilcoxon-Signed-Rank tests were used to determine whether distributions differed. Inter-ExpertGemini 2.5 Pro Gemini 2.5 Flash GPT-5 GPT-5 Nano Inter-LLM Rater Groups 0 2 4 6 8 Scale Distance 0.67 0.67 1.33 0.00 2.00 1.00 0.33 1.67 0.00 3.67 0.83 0.33 1.33 0.00 2.33 0.67 0.33 1.00 0.00 2.00 0.67 0.33 1.33 0.00 2.33 1.00 0.50 1.17 0.00 2.17 Distribution of Likert-scale distances by rater group experts included non-contradicting points focusing on different aspects of the paper, thereby producing a comprehensive combined response. Therefore, assessment of LLM responses using the combined overlap score method would mitigate these differences. Question items 28.1.4, 28.2.4, 28.4.4, 28.5.4, and 28.9.4, were combined for analysis due to their small sample size (six combined) and conceptual similarity (involved applying and extracting evidence for or against 9 AI Can Match Domain Experts in Evidence Extraction and Appraisal the microbial oncogenesis criteria). Manual evaluation of responses for these question items determined differences to be due to varying response foci and detail levels. Similarly to Q18 and Q26.1, this indicates that experts focused on different aspects of the question item and the paperâs findings, collectively producing a well-rounded response. High levels of consensus and complementary responses among experts indicated that question items were clear and strengthened confidence in the human-validated test dataset as a means of benchmarking the LLMs. Figure 4: LLMs had higher combined overlap scores (B) than pairwise overlap scores (A), indicating that LLMs tended to include information from each of the experts rather than a single expert. (A) LLM long-answer pairwise overlap score distributions were indistinguishable from the inter-expert distribution, indicating that LLM responses were as similar to individual experts as experts were to each other. Inter-LLM distribution differed significantly from inter-expert and expert-LLM distributions (allp < 0.0001, r > 0.75), indicating that LLM responses had greater overlap with each other than with experts, and had greater overlap than experts had with each other. Long-answer response pairs for 75 question instances were scored on their overlap by an LLM-judge (GPT-5, blinded to LLM being assessed) on a continuous rubric scale from 0-4. Scores for each question instance were averaged for each rater group to produce the pairwise overlap score. (B) LLM long-answer responses captured most key points found across three expert responses, with LLM responses predominantly scoringâ„2. A Friedman test found no significant differences in combined overlap scores across the different LLMs (F r = 6.48, p = 0.0906), indicating that response completeness did not differ by model size. For 75 question instances, LLM responses were compared to a reference answer created by combining three separate expert responses. Each LLM response was scored on whether it covered the key and/or minor points in the reference answer by an LLM-judge (GPT-5, blinded to LLM being assessed) using a discrete rubric scale from 0-4 (see supplementary material Fig. S2 for scoring rubric). 0.00.51.01.52.02.53.03.5 Pairwise Overlap Score 0.0 0.1 0.2 0.3 0.4 0.5 0.6 Density 01234 Combined Overlap Score 0 5 10 15 20 25 30 35 Number of Question Instances 0 6 35 30 5 0 8 26 32 10 1 4 22 34 14 1 4 28 32 10 AB Inter-e xpertGemini 2.5 Pro Gemini 2.5 Flash GPT-5GPT-5 Nano Inter-LLM Rater Groups Pairwise overlap scores for rater group responses to all long-answer question instances Combined overlap scores for LLM rater group responses to all long-answer question instances 3.2 Overall LLM Performance Across Extraction Template 3.2.1 Aggregate Performance Across the 24 papers, GPT-5 (W = 139.5, p = 1.0) and GPT-5 Nano (W = 115.0, p = 1.0) consistently gave similar responses to experts, producing no overall change in inter-expert agreement when integrated (i.e., when LLM responses were included as one of the human experts for analysis of agreement). In contrast, the responses of Gemini 2.5 Pro (W = 17.0, p = 0.0004, r = 0.78) and Gemini 2.5 Flash (W = 39.0, p = 0.0127, r = 0.65) tended to deviate from experts, resulting in significantly decreased inter-expert agreement when integrated (Fig. 5). Inter-LLM agreement was significantly higher than inter-expert agreement (W = 45.0, p = 0.0267, r = 0.61), suggesting that LLM responses were more similar to each other than to expert responses, and exhibited greater inter-response similarity than that observed among experts. 10 AI Can Match Domain Experts in Evidence Extraction and Appraisal Overall, GPT-5 and GPT-5 Nano produced responses most similar to domain experts, resulting in indistinguishable score distributions, while Gemini 2.5 Pro and Gemini 2.5 Flash tended to produce responses that differed from domain experts, resulting in significantly lower score distributions. Figure 5: Average per-paper scores across rater groups. GPT-5 and GPT-5 Nano responses integrated seamlessly with expert responses across the 24 papers, producing no appreciable change in inter-expert agreement, while Gemini 2.5 Pro and Gemini 2.5 Flashâs responses were more inconsistent and tended to deviate from that of experts, significantly decreasing inter-expert agreement. Individual per-paper scores were determined by summing the MCQ match propor- tions, Likert-scale inverted distances, and long-answer pairwise overlap scores for each paper. Individual per-paper scores were averaged to produce the average per-paper score per rater group. Error bars represent±SD. 40.2 42.7 37.8 Inter-Expert Gemini 2.5 Pro Gemini 2.5 Flash GPT-5 GPT-5 Inter-LLM Rater Groups 0 5 10 15 20 25 30 35 40 45 Average P er -paper Sco re 37.9 40.9 35.0 33.9 38.3 29.4 34.9 39.3 30.6 37.8 39.9 35.7 37.1 39.9 34.4 Nano Average per-paper scores across rater groups 3.2.2 Performance by Question Type MCQ Performance All models provided similar MCQ answers to experts. GPT-5 (W = 3345.0,p = 0.0106, r = 0.76) and GPT-5 Nanoâs (W = 2601.5, p = 0.0005, r = 0.79) MCQ answers were consistently similar to experts, tending to agree with expert consensus (either partial or full), therefore significantly increasing expert agreement when integrated. Gemini 2.5 Pro (W = 6110.0, p = 1.0) and Gemini 2.5 Flash (W = 6007.0, p = 1.0) tended to agree with the expert outlier, with integration of their MCQ answers having no effect on expert agreement (Fig. 6). Likert-Scale PerformanceAcross the four Likert-scale questions, GPT-5 (W = 929.0, p = 1.00) and GPT-5 Nano (W = 1127.5, p = 1.00) rated the papers most similarly to experts, with integration of their ratings producing no overall change in expert scale distances. Gemini 2.5 Pro gave considerably higher ratings than experts, resulting in significantly greater expert scale distances (W = 639.0, p = 0.0013, r = 0.63). Similarly, while integration of Gemini 2.5 Flashâs ratings did not significantly alter expert scale distances (W = 769.0, p = 0.1981), these ratings were higher than expertsâ, resulting in larger scale distances (Fig. 3). For all question items, higher ratings equated to greater leniency (for example, greater paper reliability or strength of evidence), indicating a consistent positivity bias in Gemini 2.5 Proâs evaluations, with Gemini 2.5 Flash behaving similarly. Gemini 2.5 Pro and Gemini 2.5 Flash tended to agree with the expert outlier, while GPT-5 and GPT-5 Nano agreed with the expert consensus and outlier with similar frequency. Overall, GPT-5, GPT-5 Nano, and Gemini 2.5 Flash rated papers similarly to experts across the different Likert-scale question items, although Gemini 2.5 Flash had a tendency to rate papers slightly more positively than experts, while Gemini 2.5 Pro consistently rated papers more favorably than experts, indicating a positivity bias in its evaluations. Long-Answer Performance Quantitative Performance Pairwise Overlap LLM responses were as similar to individual expert responses as individual experts were to each other, with indistinguishable pairwise overlap score distributions between expert-LLM and inter-expert rater groups (Fig. 4A, see supplementary material Fig. S1 for scoring rubric). LLMs gave very similar 11 AI Can Match Domain Experts in Evidence Extraction and Appraisal Figure 6: Distribution of MCQ match proportions across rater groups. All LLMs provided MCQ answers similar to experts, with integration of their responses resulting in either no change in inter-expert agreement (Gemini 2.5 Pro, Gemini 2.5 Flash), or a significant increase in inter-expert agreement (GPT-5, GPT-5 Nano). MCQ match proportions were calculated per question instance (n=336) by averaging the binary match rates for each rater pair (e.g., expert one to Gemini 2.5 Pro) in the rater group (e.g., Gemini 2.5 Pro rater group). Friedman test with post-hoc two-sided Wilcoxon-Signed-Rank tests were used to determine whether distributions differed. 0.00.20.40.60.81.0 0.0 0.5 1.0 1.5 2.0 Density Inter-Expert Gemini 2.5 Pro Gemini 2.5 Flash GPT-5 GPT-5 Nano Inter-LLM Distribution of MCQ match proportions by rater group responses to one another, with significantly greater overlap between their responses compared to that of the expert-LLM and inter-expert groups (all pairwise tests p < 0.0001, r > 0.75). Combined Expert Response Overlap LLMs produced comprehensive responses that tended to capture key ideas dispersed across the three expert responses rather than aligning with a single expert, achieving greater combined overlap scores than pairwise overlap scores (Fig. 4B, see supplementary material Fig. S2 for scoring rubric). Response completeness did not differ by model complexity (F r = 6.48, p = 0.0906), with smaller models (Gemini 2.5 Flash, GPT-5 Nano) capturing the same amount of key points as larger models (Gemini 2.5 Pro, GPT-5) across all question instances. Overall, LLM long-answer responses were comprehensive, capturing key ideas from across the three expert responses. LLM responses were very similar, capturing the same amount of key points regardless of model size. Qualitative AnalysisOmissions Gemini 2.5 Flash and GPT-5 Nano had slightly less omissions than frontier LLMs Gemini 2.5 Pro and GPT-5 (Table 3). Per-question item analysis found that omissions for relevancy explanations (Q1.1) and responses involving extraction of findings supporting or refuting the microbial oncogenesis criteria (28.1.4, 28.2.4, 28.4.4, 28.5.4, 28.9.4) were generally minor and reflected differences in response focus between experts and LLMs, with singular exceptions for GPT-5 and GPT-5 Nano for the microbial oncogenesis criteria, while omissions for methodology-related question items (Q18 and Q26.1) were more significant. For methodology-related question items, LLMs tended to omit the same information. In some instances, experts included irrelevant information (for example, a description of results rather than methodological strengths or weaknesses), however, legitimate omissions included: (1) failure to comment on sample size or appropriateness, and (2) no mention of the country the study was conducted in, although models did occasionally mention the limited geographic generalizability of results. For Q26.1 specifically, LLMs were less likely to comment on potential errors or missing information in the papers than experts. Out of all omissions, the most severe occurred in the âMicrobial Oncogenesis Criteria" section for questions 28.2.4 (for âHistopathologic Association") and 28.4.4 (for âExperimental Evidence of Facilitation of Oncogenesis"), by GPT-5 Nano and GPT-5 respectively. For both question instances, all three experts agreed that the paper helped fulfilled these criteria, while these LLMs did not. While both modelsâ reasoning were valid and indicated stricter adherence to the criterionâs definition than experts, GPT-5 Nanoâs reasoning was inconsistent with its responses for this criterion in other 12 AI Can Match Domain Experts in Evidence Extraction and Appraisal Table 3: Of question instances where LLMs included information not in expert responses, almost all were paper-supported additions, with a few hallucinations and/or distortions found in smaller models (Gemini 2.5 Flash, GPT-5 Nano). While omissions were numerous, there were similar counts of paper-supported additions. Across 75 question instances, information gaps between experts and LLMs were manually assessed and categorized as either omissions, hallucinations and/or distortions, or paper-supported additions. Frequency of omissions, hallucinations and/or distortions, and paper-supported additions in LLM long-answer responses LLMOmissions 1 Information not in expert responses Hallucinations and/or distortions 2 Paper- supported additions 3 Gemini 2.5 Pro6851051 Gemini 2.5 Flash6464262 GPT-56762062 GPT-5 Nano6669762 1 Omissions refer to information LLMs excluded from their responses that experts included in theirs. 2 Hallucinations and/or distortions refer to information included in the LLM response that was inaccurate to the original paper, either fabricated or representing a misinterpretation of the text. 3 Paper-supported additions refer to information included in the LLM response, absent from expert responses, but faithful to or found within the original paper. papers, although these question instances were not included in the main analysis as they lacked three expert responses for comparison. Overall, smaller models had less omissions than larger models, with Gemini 2.5 Flash having the least and Gemini 2.5 Pro having the most, although models tended to omit the same information. For most question items omissions were minor, with the exception of methodological questions where LLMs often failed to consider factors surrounding sample size, country of study origin, and potential errors in the papers. Hallucinations and Distortions Frontier models (Gemini 2.5 Pro, GPT-5) had no hallucinations or distortions across the 75 question instances, while smaller models (Gemini 2.5 Flash, GPT-5 Nano) exhibited a limited number of such errors. Of the instances where Gemini 2.5 Flash (64) and GPT-5 Nano (69) introduced information not present in expert responses, 2 and 7 instances were classified as hallucinations and/or distortions (Table 3). These represented misinterpretations of nuanced, domain-specific methodological details rather than traditional fabrications, and occurred exclusively in methodology-focused questions (Q18 and Q26.1, see Table 2 for full questions). None of the errors overlapped between the two models, indicating distinct failure modes. Gemini 2.5 Flash Both errors were classified as distortions with minimal impact on evidence appraisal, and consisted of a misinterpretation of a PCR quality-control step and incomplete pooling of tissue sample type counts. GPT-5 Nano GPT-5 Nano repeated errors across the two methodology-related question items for three papers, resulting in four unique hallucinations/distortions out of the seven instances. GPT-5 Nano appeared more prone to text misinterpretation than Gemini 2.5 Flash, with its errors impacting evidence strength appraisal. These included a misunderstanding of standard PCR terminology, a misreading of a 2Ă2 contingency table, conflation of metastatic status with overall cancer prevalence, and falsely stating that appropriate controls were absent, which GPT-5 Nano characterized as a critical methodological flaw weakening reliability. Overall, frontier models Gemini 2.5 Pro and GPT-5 had no hallucinations and/or distortions in the 75 question instances assessed, while smaller models Gemini 2.5 Flash and GPT-5 Nano had minimal instances. Both smaller modelsâ errors represented misinterpretations of domain specific methodologies, however GPT-5 Nanoâs had slightly more errors which were more severe, with downstream effects on the evaluation of the paperâs strength of evidence. Paper-Supported Additions LLMs tended to include the same paper-supported additions as one another, even for methodology-related question items (Q18 and Q26.1) requiring applied reasoning over extraction. The majority of question instances for these question items had paper-supported additions. For Q18, LLMs tended to extract more detailed methodological information (e.g., the exact controls and processes used), while for Q26.1, LLMs tended to include additional limitations which required integration of paper-based information and domain knowledge (e.g., 13 AI Can Match Domain Experts in Evidence Extraction and Appraisal acknowledging that FFPE tissues have inherent nucleic acid degradation risks that may impact viral detection). Although LLMs tended to include similar limitations to one another, often repeating the same (yet contextually appropriate) set of limitations, there were instances where the LLMs noted pivotal points that the experts missed. 3.3 Performance by Template Section 3.3.1 Relevancy Gemini 2.5 Pro was maximally sensitive, classifying all 24 papers as Relevant. Similarly, Gemini 2.5 Flash classified one paper as Somewhat Relevant (aligned with expert consensus) and GPT-5 classified two papers as Somewhat Relevant, one aligning with consensus and one where experts agreed the paper was Irrelevant. GPT-5 Nano demonstrated the strictest classification threshold, labeling 54.17% (13/24) of papers as Somewhat Relevant. In most cases, this aligned with at least a single expert; however, in one instance all experts classified the paper as Relevant, and in another all agreed it was Irrelevant. LLM relevancy explanations consistently captured all major expert-identified points, with minor omissions (pairwise overlap scoresâ„2 and combined overlap scoresâ„3). Qualitative analysis of GPT-5 Nanoâs responses found that it consistently identified relevant mechanistic or associative findings in the papers, explaining the high pairwise overlap scores, but concluded that absence of temporality or direct causal proof warranted a Somewhat Relevant classification. Overall, LLMs demonstrated understanding of domain-specific nuances (both in the paper and in the classification criteria) and strong alignment with expert reasoning on paper relevancy. 3.3.2 Paper Integrity and Reliability Identifying Irrelevant References Only GPT-5 claimed to find irrelevant references, and this was for two papers. Manual re-evaluation found GPT-5 to be incorrect for the first paper, although this seemingly had no downstream effect on GPT-5âs evidence appraisal. However, for the second paper, GPT-5 identified three references that did appear to be irrelevant to the sections they were cited in, which initial human evaluation missed. Identifying Potential Conflicts of Interest GPT-5 Nano claimed that university funding constituted âa potential COI concern in some assessments", while Gemini 2.5 Flash claimed that the private laboratories the authors were affiliated with were contracted with the National Cancer Institute (NCI), which âcould introduce an indirect, potential commercial interest". In many instances, LLMs - most notably GPT-5 Nano - cited the paperâs own no conflict of interest declaration as reasoning for their response. Identifying Contradictions LLMs had difficulty consistently detecting human-identified contradictions in full-text papers, although these had minimal impact on the paperâs reliability. Gemini 2.5 Flash and GPT-5 were the only models to identify any human-identified contradictions, and this was for one of the 13 instances. Gemini 2.5 Pro and GPT-5 identified three and five contradictions, respectively, that the human evaluator missed. However, Gemini 2.5 Pro, GPT-5, and GPT-5 Nano had one, two, and four instances, respectively, where they either fabricated or misinterpreted a contradiction. Gemini 2.5 Proâs hallucination stated that values for a subgroup analysis did not add up to the correct value, yet in its response it demonstrated that they did, thereby fabricating a contradiction, while GPT-5 missed the authorsâ justification for methodological decisions. GPT-5 Nanoâs hallucinations were more severe, and included false inconsistencies in result values, sample sizes, and the counts in a 2Ă 2 contingency table. To determine whether failure arose from context window limitations or true inability, models were given the isolated sections of text, figures, and/or tables in which the human-identified contradictions were found. After repeated attempts, Gemini 2.5 Pro did not return any results, citing API rate limits despite request limit adherence. This isolation method improved performance considerably, with Gemini 2.5 Flash (11/13), GPT-5 (12/13), and GPT-5 Nano (9/13) detecting the majority of human-identified contradictions, indicating that prior failure was likely due to context window limitations rather than lack of ability. One of these contradictions required interpretation of an electrophoresis gel image, which Gemini 2.5 Flash and GPT-5 demonstrated, despite being generalist models without biomedical specialization. One contradiction, where the paper was inconsistent in which mouse models were used, was missed by all LLMs while the rest of the missed contractions varied between Gemini 2.5 Flash and GPT-5 Nano. In two of these isolated sections, Gemini 2.5 Flash and GPT-5 each identified two additional contradictions (a total of four each) that were missed by the human evaluator, with a total of six additional unique contradictions identified. Overall, LLMs were largely unable to find the human-identified contradictions within the full-text papers, although Gemini 2.5 Flash and GPT-5 did identify contradictions the human evaluator missed. Gemini 2.5 Pro, GPT-5, and GPT-5 14 AI Can Match Domain Experts in Evidence Extraction and Appraisal Nano all hallucinated some contradictions, although GPT-5 Nanoâs were the only severe ones. Isolation experiments improved performance considerably, indicating that prior failure was due to context window limitations rather than inability or lack of domain knowledge. Paper Reliability Score All models rated paper reliability similarly to experts, with no significant differences in the scale distances between the inter-expert and expert-LLM group pairs. Gemini 2.5 Pro and Gemini 2.5 Flash demonstrated the closest alignment with expert paper reliability ratings, followed by GPT-5, while GPT-5 Nano showed the greatest divergence, tending to rate paper reliability slightly lower than experts (Fig. 7). Figure 7: Distribution of paper reliability scores given by different rater groups (Likert-scale ratings, n=24). All three expert ratings for each question instance have been aggregated into the âExperts" group. Stacked bars represent the percentage of question instances assigned each score (1-5) by each rater group. Bars are aligned at 0% by the central rating value (3, dotted line), with lower scores (representing lower paper reliability) extending to the left, and higher scores (representing higher paper reliability) extending to the right. Gemini 2.5 Pro and Gemini 2.5 Flash gave very similar paper reliability scores to experts (predominantly 5s), while GPT-5 and GPT-5 Nano in particular gave lower paper reliability scores than experts (predominantly 3s and 4s), indicating more strict critique of paper reliability (although not significantly). Likert-scale rating meanings have been truncated for visualization (see supplementary material for extraction template with full meanings). 0% 10% 20% 30%10%20%30%40%50%60%70%80%90%100% Per centage of Question Instances GPT-5 Nano GPT-5 Gemini 2.5 Flash Gemini 2.5 Pro Experts 1 Unreliable due to extensive, major issues. 2 Concerns about bias due to major issues. 3 Integrity is reasonable but weakened by repeated minor issues. 4 Some noticeable but minor issues. 5 No issues detected. Paper reliability scores by rater group 3.3.3 Summary of Paper Contents Identifying Important Methodological Factors Influencing the Strength of Findings Expert responses for this question item tended to include different, non-contradicting focus points, resulting in low pairwise overlap scores for inter-expert and expert-LLM groups (Fig. 8A). However, LLMs tended to integrate key methodological considerations from multiple experts rather than aligning exclusively with a single expert, achieving higher combined overlap scores than pairwise overlap scores (Fig. 8B). Gemini 2.5 Flash and GPT-5 demonstrated slightly better agreement with the combined reference response compared to Gemini 2.5 Pro and GPT-5 Nano, although distributions were broadly similar across models (Fig. 8B). The majority of LLM responses had combined overlap scoresâ„2, indicating capture of some key points with omissions (see Sec. 3.2.2). Gemini 2.5 Flash and GPT-5 Nano had hallucinations and/or distortions in some of their responses (see Sec. 3.2.2), indicating difficulty interpreting nuanced biomedical research methodology. 3.3.4 Strength of Evidence Identifying Errors in Statistical Analysis or Result Presentation GPT-5 and Gemini 2.5 Flash alone identified one of the four human-identified errors in result presentation (percentage miscalculation), and Gemini 2.5 Pro, GPT-5, and GPT-5 Nano identified two, three, and one errors, respectively, that were missed by the human evaluator. The human-identified errors the LLMs missed included two instances of incorrectly presented Venn diagrams (difference values did not exclude intersection values and intersection values were duplicated), and another instance where a results table had an incorrect category label. Additionally, Gemini 2.5 Pro and GPT-5 Nano both repeated hallucinated and/or distorted errors they gave in their contradiction analysis (see Sec. 3.3.2 for more details). 15 AI Can Match Domain Experts in Evidence Extraction and Appraisal Figure 8: LLMs had higher combined overlap scores (B) than pairwise overlap scores (A) for Q18, indicating that LLMs tended to include important methodological factors from multiple experts rather than a single expert. Inter-LLM pairwise overlap scores were considerably greater than other rater groups (A), indicating that LLMs often included the same methodological factors as one another in their responses. (A) LLM long-answer pairwise overlap score distributions were near indistinguishable from the inter-expert distribution. Long-answer response pairs for 75 question instances were scored on their overlap by an LLM-judge (GPT-5, blinded to LLM being assessed) on a continuous rubric scale from 0-4 (see supplementary material, Fig. S1 for scoring rubric). Scores for each question instance were averaged for each rater group to produce the pairwise overlap score. (B) LLM long-answer responses captured most key points found across three expert responses, with LLM responses predominantly scoringâ„2. For 75 question instances, LLM responses were compared to a reference answer created by combining three separate expert responses. Each LLM response was scored on whether it covered the key and/or minor points in the reference answer by an LLM-judge (GPT-5, blinded to LLM being assessed) using a discrete rubric scale from 0 to 4 (see supplementary material Fig. S2 for rubric). 01234 Combined Overlap Score 0 2 4 6 8 10 12 14 Number of Question Instances 0 2 15 5 2 0 3 99 3 0 1 11 12 00 2 11 10 1 0.00.51.01.52.02.53.03.5 Pairwise Overlap Score 0.0 0.2 0.4 0.6 0.8 Density BA Inter-e xpertGemini 2.5 Pro Gemini 2.5 Flash GPT-5GPT-5 Nano Inter-LLM Rater Groups Pairwise overlap scores for rater group responses to methodological factors question instances (Q18) Combined overlap scores for LLM rater group responses to methodological factors question instances (Q18) Identifying Limitations LLMs tended to emphasize similar limitations to one another, while experts generally provided different (yet complementary) limitations, resulting in substantially higher inter-LLM pairwise overlap scores compared to the inter-expert and expert-LLM rater groups (Fig. 9A). LLMs tended to integrate limitations from multiple experts rather than aligning with a single expert, with higher combined overlap scores than pairwise overlap scores for all LLMs (Fig. 9B). Strength of Evidence ScoreAll models rated strength of evidence similarly to experts. Most LLM ratings were the same as or one level away from expert consensus (all LLM median scale distancesâ€1), with no significant differences in the scale distances across the rater groups (F r = 13.42,p = 0.0197). When ratings did deviate from experts, Gemini 2.5 Pro and Gemini 2.5 Flash tended to rate strength of evidence slightly higher than experts, indicating greater leniency, while GPT-5 tended to rate strength of evidence slightly lower than experts, indicating a slightly stricter interpretation of evidence strength (Fig. 10). 3.3.5 Microbial Oncogenesis Criteria Identifying Evidence Type and Findings Supporting/Refuting Microbial Oncogenesis Criteria Across the six questions analyzed, LLMs tended to include details from all three expert responses, with higher combined overlap scores than pairwise overlap scores. GPT-5 had the highest combined overlap scores (3.00, IQR 1.50 - 3.00), followed by Gemini 2.5 Pro (2.50, IQR 2.00 - 3.00) and Gemini 2.5 Flash (2.00, IQR 2.00 - 2.75), with GPT-5 Nano having the lowest scores (2.00, IQR 1.25 - 2.00), indicating that it tended to miss key points. Despite generally high performance, GPT-5, along with GPT-5 Nano, each had an instance where, contrary to expert consensus, they did not perceive a paper 16 AI Can Match Domain Experts in Evidence Extraction and Appraisal Figure 9: LLMs had higher combined overlap scores (B) than pairwise overlap scores (A) for Q26.1, indicating that LLMs tended to include limitations from multiple experts rather than a single expert. Higher inter-LLM pairwise overlap scores (A) indicated that LLMs included similar limitations to one another. (A) LLM long-answer pairwise overlap score distributions were similar to the inter-expert distribution. Long-answer response pairs for 75 question instances were scored on their overlap by an LLM-judge (GPT-5, blinded to LLM being assessed) on a continuous rubric scale from 0-4 (see supplementary material Fig. S1 for scoring rubric). Scores for each question instance were averaged for each rater group to produce the pairwise overlap score. (B) LLM long-answer responses captured most key points found across three expert responses, with LLM responses predominantly scoringâ„2. For 75 question instances, LLM responses were compared to a reference answer created by combining three separate expert responses. Each LLM response was scored on whether it covered the key and/or minor points in the reference answer by an LLM-judge (GPT-5, blinded to LLM being assessed) using a discrete rubric scale from 0 to 4 (see supplementary material Fig. S2 for scoring rubric). 1.01.52.02.53.03.5 Pairwise Overlap Score 0.0 0.2 0.4 0.6 0.8 1.0 1.2 Density 01234 Combined Overlap Score 0 2 4 6 8 10 12 Number of Question Instances 0 1 10 9 1 0 2 7 10 2 0 2 9 55 0 1 12 7 1 BA Inter-e xpertGemini 2.5 Pro Gemini 2.5 Flash GPT-5GPT-5 Nano Inter-LLM Rater Groups Pairwise overlap scores for rater group responses to limitations question instances (Q26.1) Combined overlap scores for LLM rater group responses to limitations question instances (Q26.1) as helping to fulfill or refute the specified criterion, possibly indicating that these LLMs were occasionally more strict than experts in applying microbial oncogenesis criteria. Neither Gemini 2.5 Pro nor Gemini 2.5 Flash had this issue. Cumulative Oncogenesis Score GPT-5 (W = 86.0,p = 1.0) and GPT-5 Nano (W = 60.0,p = 1.0) applied microbial oncogenesis criteria in a manner similar to experts, with inclusion of their cumulative oncogenesis scores leaving inter-expert agreement unchanged. Gemini 2.5 Pro (W = 3.0,p = 0.0004,r = 0.86) and Gemini 2.5 Flash (W = 21.0,p = 0.0091,r = 0.75) were far more lenient than experts in stating a paper helped fulfill various microbial oncogenesis criteria, giving higher Likert-scale ratings and therefore significantly decreasing expert agreement (Fig. 11A). Gemini 2.5 Pro and Gemini 2.5 Flash had a tendency to state papers helped fulfill a given microbial oncogenicity criterion in contexts where evidentiary support was indirect or absent. For example, for a review of publicly available ecological data correlating mouse population changes with breast cancer incidence [64], both Gemini models assigned cumulative oncogenesis scores of 8/10 and 10/10 respectively, indicating that they considered the paper to satisfy nearly all microbial oncogenesis criteria, including temporality, reproducibility/validation, dose-response relationship, molecular/multi-omics interaction, and even prevention. Notably, this behavior did not reflect misunderstanding of the papers, but rather an expansive interpretation of what constituted fulfillment of the oncogenesis criteria, with scoring justifications suggesting a susceptibility to ecological inference bias and false-cause reasoning. Overall, GPT-5 and GPT-5 Nano applied the microbial oncogenesis criteria most similar to experts, while Gemini 2.5 Pro and Gemini 2.5 Flash deviated from experts, applying the criteria much more leniently. Rating Impact of Paper on MCP LLMs rated each paperâs impact on the MCPâs plausibility similarly to experts, with no significant differences in the scale distances across the different rater groups (F r = 10.79,p = 0.0557), 17 AI Can Match Domain Experts in Evidence Extraction and Appraisal Figure 10: Distribution of strength of evidence scores given by different rater groups (Likert-scale ratings, n=24). All three expert ratings for each question instance have been aggregated into the âExperts" group. Stacked bars represent the percentage of question instances assigned each score (1-5) by each rater group. Bars are aligned at 0% by the central rating value (3, dotted line), with lower scores (representing lower strength of evidence) extending to the left, and higher scores (representing higher strength of evidence) extending to the right. Gemini 2.5 Pro and Gemini 2.5 Flash tended to rate strength of evidence slightly higher than experts (predominantly 3s and above), indicating greater leniency, while GPT-5 tended to rate strength of evidence slightly lower (predominantly 2s and below), indicating more strict evidence appraisal. GPT-5 Nano GPT-5 Gemini 2.5 Flash Gemini 2.5 Pro Experts 1 Severely weak evidence/no credible evidence. 2 Weak evidence. 3 Acceptable evidence. 4 Strong evidence. 5 Very strong evidence. 0% 25% 50% 75%25%50% Per centage of Question Instances Strength of evidence scores by rater group however, Gemini 2.5 Pro, Gemini 2.5 Flash, and GPT-5 all had a tendency towards rating the paperâs impact on the MCPâs plausibility higher, indicating slightly greater leniency in this regard than experts (Fig. 11B). 3.4 LLM-Expert Similarity Pattern Analysis A cross-classified mixed effects model investigating the pairwise scores of all question types (see Sec. 2.6) found no statistically significant evidence that any LLM preferentially matched specific experts (allp > 0.7), indicating that LLMs did not replicate the reasoning patterns of specific domain experts. 3.5 Stability of GPT-5 Outputs As GPT-5âs responses aligned most closely with experts, repeat experiments were performed to determine output stability. We ran 30 repeats on two papers, using a subset of six MCQs and two Likert-scale question items (see supplementary material Sec. 1.4 for included question items and the dataset for GPT-5âs outputs). GPT-5 demonstrated high response stability across both papers. For the first paper, six of eight question items received identical responses across all runs, resulting in an overall stability of 92.5% (222/240 responses). For the second paper, five of eight question items received identical responses across all runs, resulting in an overall stability of 92.1% (221/240 responses). Across both papers, overall response stability was 92.3% (443/480 responses). Where variation occurred, GPT-5âs responses generally remained closely aligned with expert assessments. For several MCQs, GPT-5 alternated between responses that matched either the expert consensus or the expert outlier. Similarly, variation in Likert-scale ratings was typically limited to a single rating level and remained close to expert evaluations. Overall, despite occasional deviations, GPT-5âs responses were consistent across repeated runs and generally remained within the range of expert interpretations, indicating stable expert-aligned reasoning across repeated evaluations. 4 Discussion Our results show that pre-trained reasoning LLMs are capable of analyzing and evaluating individual biomedical research papers at the level of domain experts on most tasks when using a structured extraction and appraisal template, with minimal hallucinations (see Fig. 12 for summary of results). 18 AI Can Match Domain Experts in Evidence Extraction and Appraisal Figure 11: (A) Distribution of cumulative oncogenesis scores given by different raters (Likert-scale ratings, n=24). All three expert ratings for each question instance have been aggregated into the âExperts" group. Stacked bars represent the percentage of question instances assigned each score (0-10) by each rater group. Bars are aligned at 0% by the central rating value (5, dotted line), with lower scores extending to the left, and higher scores extending to the right. Cumulative oncogenesis scores were determined by summing scores for each of the 10 microbial oncogenesis criteria. Higher scores approaching 10 indicated fulfillment of more microbial oncogenesis criteria. Lower scores approaching 1 indicated either: (1) fulfillment of less criteria, or (2) fulfillment of some criteria combined with evidence to refute other criteria, as combinations of positive scores for certain criteria (paper helps support specific criterion) and negative scores for other criteria (paper helps refute specific criterion) may cancel out to produce a low positive score or a score of 0 depending on the combination. GPT-5 and GPT-5 Nano had cumulative oncogenesis scores most similar to experts (between 0 and 4), while Gemini 2.5 Pro and Gemini 2.5 Flash tended to state that papers supported more microbial oncogenesis criteria (in some instancesâ„5 criteria for a single paper), and therefore gave higher cumulative oncogenesis scores than experts. (B) Comparison of paper impact on MCP plausibility scores (Likert-scale ratings, n=24) given by different raters. All three expert ratings for each question instance have been aggregated into the âExperts" group. Stacked bars represent the percentage of question instances assigned each score (1-5) by each rater group. Bars are aligned at 0% by the central rating value (3, dotted line), with lower scores (representing that the paper more strongly refuted the MCPâs plausibility) extending to the left, and higher scores (representing that the paper more strongly supported the MCPâs plausibility) extending to the right. Gemini 2.5 Pro, Gemini 2.5 Flash, and GPT-5 all tended to rate the paperâs impact on the MCPâs plausibility more highly, but Gemini 2.5 Pro and Gemini 2.5 Flash both gave more extreme ratings (1s for âStrongly refutes" and 5s for âStrongly supports") compared to experts, GPT-5, and GPT-5 Nano. 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% 100%10%20%30% Per centage of Question Instances GPT-5 Nano GPT-5 Gemini 2.5 Flash Gemini 2.5 Pro Experts 0 criteria fulfilled 1 criterion fulfilled 2 criteria fulfilled 3 criteria fulfilled 4 criteria fulfilled 5 criteria fulfilled 6 criteria fulfilled 7 criteria fulfilled 8 criteria fulfilled 9 criteria fulfilled 10 criteria fulfilled Cumulative oncogenesis scores by rater group 0% 10% 20% 30%10%20%30%40%50%60%70%80% Per centage of Question Instances GPT-5 Nano GPT-5 Gemini 2.5 Flash Gemini 2.5 Pro Experts 1 Strongly refutes (moderate to strong evidence) 2 Weakly refutes (weak to moderate evidence) 3 Does not impact (negligible or weak evidence) 4 Weakly supports (weak to moderate evidence) 5 Strongly supports (moderate to strong evidence) Paper impact on MCP plausibility scores by rater group BA The developed human-validated test dataset had strong inter-rater reliability, with high rates of full and partial agreement for MCQs (> 90%) and Likert-scale questions (> 80%, Fig. 2), and close clustering of Likert-scale ratings even when consensus was absent (Fig. 3). Long-answer responses demonstrated moderate overlap (Fig. 4A), with variability primarily reflecting complementary reasoning rather than substantive disagreement between experts, as expected with nuanced domain-specific tasks. Considering this variability, we assessed whether inclusion of LLM responses altered the overall distribution of expert agreement (i.e., whether LLMs behaved as an additional expert), rather than requiring exact matches to expert responses. This may be a more effective strategy for evaluating LLMs in domains characterized by complex, nuanced tasks with non-contradicting inter-expert variability. The exception to this pattern of strong inter-rater reliability was the multi-select (i.e., select all that apply) question type, which proved challenging for both experts and LLMs. Experts occasionally provided directly conflicting responses (e.g., one selecting âAppropriate methodology/tools" and another selecting âInappropriate methodology/tools"), potentially reflecting differences in domain expertise and interpretation of study strengths and limitations. LLMs also struggled with this format, frequently selecting all available options or generating long-answer responses rather than selections, potentially reflecting poor exposure to this question type in their training data. These question items were excluded from further analysis as expert disagreement precluded the establishment of a reliable reference standard. The fact that both experts and LLMs encountered difficulties suggests that the issue lay primarily with the question type rather than either evaluator group. Future iterations of the extraction template will therefore replace multi-select items with a series of binary (yes/no) statements for each microbial oncogenesis criterion, accompanied by a justification field. Based 19 AI Can Match Domain Experts in Evidence Extraction and Appraisal on these findings, multi-select question formats are not recommended for either human or LLM-based evaluation of biomedical literature. Across all 24 papers, GPT-5 and GPT-5 Nano performed most consistently and similarly to experts, while Gemini 2.5 Pro (W = 22.0,p < 0.0001,r = 0.75) and Gemini 2.5 Flash (W = 39.0,p = 0.0127,r = 0.65) behaved as outliers, significantly decreasing inter-expert agreement (Fig. 5). However, performance varied across question types and items, with methodological appraisal questions and identification of contradictions being the most persistent areas of vulnerability. Across MCQ items, LLMs performed similarly to experts - consistent with prior evidence that generalist LLMs (particularly GPT models) perform well on biomedical MCQ answering [61], with GPT-5 (W = 3345.0,p = 0.0106, r = 0.76) and GPT-5 Nano (W = 2601.5,p = 0.0005,r = 0.79) having significantly increased inter-expert agreement, agreeing most with expert consensus. Across Likert-scale questions, Gemini 2.5 Pro (W = 639.0,p = 0.0013, r = 0.63) had the greatest divergence from experts, significantly increasing scale distances by giving considerably higher Likert-scale ratings, indicating greater leniency in paper evaluation, while Gemini 2.5 Flash, GPT-5, and GPT-5 Nano gave similar scores to experts and therefore did not alter inter-expert distributions (Fig. 3). When Gemini 2.5 Pro and Gemini 2.5 Flashâs ratings did match an expertâs, it tended to be the outlier, while GPT-5 and GPT-5 Nano matched the expert consensus and outlier with similar frequency. Repeat evaluations with GPT-5 (30 runs on two papers for a subset of eight question items) demonstrated response consistency across MCQ and Likert-scale question types (overall average response stability of 92.3%), with responses consistently matching or remaining closely clustered around expert ratings. LLM long-answer responses were as similar to individual experts as experts were to each other, resulting in LLM long-answer pairwise overlap score distributions being indistinguishable from expert distributions (Fig. 4A). Despite operating at a disadvantage against three domain experts, LLMs captured the majority of key points identified collectively by the expert group, demonstrating greater alignment with the combined expert reference answers than with individual expert responses (Fig. 4B), suggesting integrative synthesis across perspectives and expertise. Despite differences in model size, complexity, and reported reasoning capacity, no significant superiority was observed across models in combined overlap scores, indicating functional convergence in synthesis capacity - however, limited instances of hallucinations and/or distortions by smaller models Gemini 2.5 Flash and GPT-5 Nano reduced their trustworthiness compared to frontier models Gemini 2.5 Pro and GPT-5 (see Sec. 3.2.2). Interestingly, LLMs exhibited greater agreement with one another than with experts, and greater agreement than was observed between experts themselves. However, this effect was driven entirely by long-answer questions; agreement patterns for MCQ and Likert-scale items were comparable between LLMs and experts. Among the model pairings, Gemini 2.5 Pro and Gemini 2.5 Flash showed the greatest overlap in long-answer responses, followed by Gemini 2.5 Flash and GPT-5, and then GPT-5 and GPT-5 Nano. High inter-LLM agreement may partially reflect similarities in training data or learned representations, although the strong agreement observed between Gemini 2.5 Flash and GPT-5 suggests that shared LLM developer lineage alone cannot fully explain the findings. Another likely contributor is response length: LLMs typically provided substantially more detailed answers than experts, often listing multiple observations where experts highlighted only one or two key points. Consequently, LLM responses had a greater opportunity to overlap with one another, whereas agreement between experts and LLMs was inherently constrained by the concise nature of expert responses. Given the inclusion of experts from multiple specialties, a cross-classified mixed-effects model was used to determine whether LLMs preferentially aligned with specific expert perspectives, thereby exhibiting specialty-specific reasoning patterns. However, there was no statistically significant evidence that any LLM consistently matched particular experts more closely than others (allp > 0.7). This finding suggests that the models did not replicate the reasoning patterns of individual domain specialists. One possible interpretation is that the LLMs integrated perspectives across multiple areas of expertise, producing responses that reflected a synthesis of expert viewpoints rather than alignment with any single specialist. This interpretation is consistent with the long-answer combined overlap scores, where LLMs achieved greater overlap with the combined expert responses than with any individual expert. An alternative explanation is that model responses were relatively random and thus did not systematically resemble any specific expert. However, the consistency of LLM performance across question types, together with their strong performance on long-answer questions, argues against largely random or inconsistent reasoning patterns and lends greater support to the integrative synthesis hypothesis. Nevertheless, the ability to detect expert-specific alignment may have been limited by the study design, as each paper was evaluated by only three experts and expert participation overlapped only partially across the dataset. Building on the test dataset reliability and overall LLM performance assessment, we return to the core objectives: whether generalist, pre-trained LLMs can 20 AI Can Match Domain Experts in Evidence Extraction and Appraisal 1. Interpret nuanced biomedical language, 2. Critically appraise biomedical research, 3. Apply microbial oncogenesis causal criteria to expert standards, and 4. Produce accurate, trustworthy outputs with minimal hallucinations or omissions 4.1 Interpreting Nuanced Biomedical Language While LLMs appeared to generally understand nuanced biomedical language, Gemini 2.5 Flash and GPT-5 Nano occasionally misinterpreted domain-specific methodological terminologies, leading to distortions and - in the case of GPT-5 Nano - downgrading of paper evidence (see Sec. 3.2.2). While smaller generalist models do not appear able to fully interpret nuanced biomedical language to the standard of domain experts, frontier models Gemini 2.5 Pro and GPT-5 performed sufficiently well, demonstrating understanding of complex domain-specific information, strengthening trust in their capabilities and enabling their use for systematic biomedical evidence synthesis processes such as ours. 4.2 Critically Appraising Biomedical Research Systematic differences in appraisal behavior were observed across models, suggesting architecture-dependent biases. Gemini 2.5 Pro significantly increased scale distances and demonstrated a consistent tendency toward more favorable appraisals of paper reliability (Fig. 7), strength of evidence (Fig. 10), and microbial oncogenesis plausibility (Fig. 11A). Gemini 2.5 Flash behaved similarly, however, it did not produce significant changes in inter-expert agreement. In contrast, GPT-based models exhibited relatively more conservative scoring behavior that predominantly matched that of experts. These findings converge with our prior work demonstrating increased leniency in Gemini models when applying classification criteria [15]. While leniency increased sensitivity in the literature screening stage, it presently reduced alignment with expert judgment. In evidence synthesis contexts, such differences could meaningfully influence cumulative grading, prioritization, and downstream decision-making. GPT-5 models demonstrated the highest alignment with expert critical appraisal. As all experts were recruited from a single institution, it is possible that some degree of shared institutional perspective influenced expert critical appraisal. However, several factors suggest that this is unlikely to explain the observed agreement. First, the expert panel was not educationally homogeneous, comprising individuals who received their undergraduate and postgraduate training from different institutions and who therefore entered the study with diverse academic backgrounds and methodological perspectives. Second, experts represented multiple specialties and departments, each of which place emphasis on different approaches to evidence evaluation, study design, and interpretation. This diversity was reflected in the fact that there was variance across the expert responses, and that they did not demonstrate complete consensus, with occasional disagreement in their assessments of paper reliability, strength of evidence, and microbial oncogenesis plausibility. Consequently, the expert evaluations cannot be viewed as the product of a single uniform institutional framework for critical appraisal. While it is conceivable that some shared perspectives could arise through a common institutional environment, given the vast and heterogeneous data sources used to train contemporary LLMs, it is unlikely that agreement with the expert panel resulted from exposure to appraisal patterns specific to a single South African institution. It is therefore more plausible that the observed alignment between GPT-5 models and experts reflects the modelsâ ability to identify broadly recognized indicators of paper reliability and strength of evidence than the replication of highly specific institutional appraisal norms. While GPT-5 models appeared to display evidence appraisal capabilities comparable to experts, calibration or multi- model strategies may still be necessary when deploying LLMs in evaluative biomedical workflows requiring evidence appraisal. GPT-5 models remained more stringent in aspects of critical evaluation (tended to provide slightly lower ratings for paper reliability score and strength of evidence score than experts), while simultaneously failing to identify issues detected by human evaluators, including contradictions and methodological concerns. Thus, similarity in overall appraisal scores did not necessarily correspond to human-level identification of specific weaknesses within a study. Future studies could combine models with complementary strengths, for example use a model specifically fine-tuned for research quality assessment or trustworthiness evaluation to identify potential limitations, sources of bias, and methodological concerns, and then incorporate these findings into the appraisal process of a more generalist model such as GPT-5. Such architectures could potentially leverage the human-like scoring behavior observed in frontier LLMs while improving the depth and consistency of evidence appraisal, resulting in evaluations that more closely approximate expert review. 21 AI Can Match Domain Experts in Evidence Extraction and Appraisal 4.3 Applying Microbial Oncogenesis Criteria While Gemini 2.5 Pro, Gemini 2.5 Flash, and GPT-5 all demonstrated superior ability in extracting evidence to support and/or refute microbial oncogenesis criteria (adapted from [69]), GPT-5 applied and scored these criteria most similarly to experts, thereby producing cumulative oncogenesis scores comparable to experts (Fig. 11A). GPT-5 had one instance of disagreement with expert consensus, but its reasoning was sound and indicated strict adherence to the provided microbial causal criteria definitions, increasing trust in its reasoning and consistency for future outputs. GPT-5 Nano appeared to apply and score the microbial oncogenesis criteria similarly to experts, yet had instances of inconsistent reasoning (see Sec. 3.2.2), in addition to less thorough evidence extraction than the other models (described in Sec. 3.2.2 and Sec. 4.4). Gemini 2.5 Pro and Gemini 2.5 Flash applied the criteria leniently, stating that criteria were fulfilled more often than experts did, thereby significantly decreasing inter-expert agreement on cumulative oncogenesis scores (Fig. 11A). Further analysis found this leniency to be highly inappropriate, with susceptibility to ecological inference bias and false-cause reasoning, in addition to claims that indirect associative evidence (e.g., reviews of publicly accessible prevalence data) fulfilled multiple criteria requiring experimental or clinical evidence (e.g., temporality, dose-response relationship, and prevention). Overall, Gemini 2.5 Pro, Gemini 2.5 Flash, and GPT-5 all demonstrated superior ability to extract evidence for nuanced microbial oncogenesis criteria to the level of domain experts, while GPT-5 alone applied the criteria most consistently and to expert standards. LLMs are therefore capable of applying and extracting evidence for nuanced criteria such as these, yet levels of competency and consistency in these tasks differ between models, requiring piloting and calibration of these tools before deployment. However, smaller generalist LLMs, such as GPT-5 Nano, may lack the domain knowledge and reasoning complexity to consistently apply and extract sufficient evidence for these criteria. 4.4 Producing Accurate, Trustworthy Outputs with Minimal Hallucinations or Omissions Contrary to widespread concerns regarding prevalence of hallucinations in biomedical applications [34,3], these were infrequent and largely restricted to smaller models, with most instances representing misinterpretations of the text (distortions) rather than typical fabrications (hallucinations). The majority of instances in which LLMs introduced information absent from expert responses represented paper-supported additions derived from the source paper rather than hallucinations or distortions (Table 3). LLMs more frequently had omissions in their responses, with most reflecting differences in emphasis or granularity rather than substantive misunderstanding. However, in two oncogenesis criteria question instances, GPT-5 and GPT-5 Nano failed to classify criteria as fulfilled where experts unanimously did so, suggesting potential rigidity in microbial oncogenicity criteria interpretation by these models (see Sec. 3.2.2). Conversely, LLMs frequently included paper-supported additions not identified by individual experts, suggesting differences in answer focus rather than inferential error. These low hallucination rates suggest that structured extraction templates, explicit criteria definitions, and constrained task framing may substantially mitigate hallucination risk. This is consistent with emerging evidence that structured prompting and reasoning scaffolds improve factual reliability [12,2], while other prompt refinements, including increased question specificity and minimum detail requirements, may reduce the frequency of omissions in future prompt iterations [63, 34]. Notably, hallucinations and distortions were isolated to two question groups: (1) methodology-focused questions, which required identification of study limitations and statistical interpretation, and (2) contradiction identification, which required understanding of nuanced biomedical language and interpretation of tables and figures. Methodological critique demands integration of contextual domain knowledge (which generalist LLMs - especially smaller models - may lack), statistical reasoning, and inferential judgment, all of which are areas in which LLM limitations have been previously observed [34,5]. Presently, when asked to identify potential limitations of studies, all LLMs tended to repeat the same set of limitations with minor changes to ensure contextual appropriateness, and were less likely than experts to comment on potential errors or information gaps. These failures illustrate the limits of domain-specific knowledge and skills in these generalist models, yet, there is clear potential to boost model performance in this regard, with instances where LLMs noted pivotal limitations that experts missed, and models tending to capture limitations from more than one expert in their responses. Overall, while integration of biomedical domain-specific models in a multi-model system may improve methodological appraisal ability, generalist model fine-tuning and prompt modifications may offer a similar performance boost while retaining the advanced reasoning abilities seen in these generalist models. Poor performance in identifying contradictions within the papers was unsurprising, considering prior evidence that LLMs have difficulty identifying contradictions in provided full-text papers [40] and isolated sentence pairs or text segments, with similar performance across models despite size differences in zero-shot settings [66]. This difficulty is not unique to LLMs, with studies indicating that humans struggle to identify in-text contradictions, from peer-reviewers failing to identify inconsistencies between paper abstracts and full report results [39], and studies demonstrating human 22 AI Can Match Domain Experts in Evidence Extraction and Appraisal difficulty in identifying contradictions in both full-text documents [40] and isolated paragraphs, even when provided text segments are short [54]. In contrast, when given the isolated paper sections containing the human-identified contradictions, LLMs performed considerably better, identifying close to all human-identified contradictions and detecting additional contradictions within these sections that the human evaluator missed. This indicates that while LLMs may have the ability to identify contradictions in research papers, current context window limitations likely preclude usability in this regard for full-text papers. This âcontext rot" poses an issue for our proposed pipeline, which relies on full-text papers. This requires the development of novel methods to circumvent large context size limitations and enable identification of errors for thorough evaluation of paper reliability and strength of evidence. Overall, frontier LLMsâ outputs appear to be sufficiently accurate to be trusted in evidence synthesis workflows, demonstrating comprehensiveness with minimal hallucinations. Although omissions were frequent, their impact was predominantly minor, with LLMs still tending to include information spanning across the three experts responses rather than aligning with singular experts. Smaller LLMsâ outputs were less trustworthy, with higher (although still limited) instances of hallucinations and/or distortions, likely due to their poorer ability to interpret nuanced biomedical language compared to frontier LLMs. Methodology-related questions and contradiction detection remain the greatest areas of vulnerability for generalist LLMs due to domain knowledge and context window limitations. 4.5 Limitations This study has several limitations. First, evaluation was restricted to a single MCP. This MCP was deliberately selected because of its unconfirmed status, conflicting evidence base, and predominance of associative rather than causal evidence, characteristics that are likely to resemble the intended real-world application of the system and many future MCPs of interest. Nevertheless, a single MCP cannot capture the full spectrum of evidence profiles encountered across potential MCPs. Although there is no obvious reason to expect the evaluated LLMs to be systematically biased towards or against this specific MCP, additional validation across MCPs with differing levels of evidential support would potentially strengthen confidence in the generalizability of these findings. In particular, future studies could evaluate MCPs approaching confirmed oncogenic status, including papers that challenge their plausibility, to assess whether model prior knowledge influences evaluation. Similarly, testing MCPs supported by sparse or emerging evidence could help determine whether model performance extends to more novel hypotheses. Consistent performance across such diverse MCPs could provide stronger evidence for the robustness and broader applicability of the proposed pipeline. Additionally, while the focus of this study is on applying this system to evaluating evidence on MCPs specifically, the choice of an MCP may limit generalizability of findings to other domains. Second, the dataset comprised 24 papers, reflecting the intensive nature of the expert annotation process. While the extraction template consisted of 77 questions, allowing for extensive points of analysis across this dataset, 24 papers remain a small sample. Completion of the extraction template typically required more than two hours per paper, making large-scale expert evaluation impractical. Although this constrained the dataset size, it mirrors the real-world challenge that motivated the present work: the volume of biomedical literature far exceeds the capacity for comprehensive manual review. The ability of LLMs to perform comparable evaluations within minutes suggests a potential route to scaling such analyses, with experts serving in oversight and validation roles rather than conducting every assessment manually. Third, the cohort of seven experts that provided responses for the human-validated test dataset were recruited from a single academic institution, and only three experts were assigned to each paper. While the study design of having independent review and extraction of information from papers was intentional to mimic current systematic evidence synthesis methods, a Delphi panel for consensus may have refined expert responses and further improved inter-expert agreement. Fourth, while the creation and use of a structured extraction template was intentional to standardize comparisons and enable control and transparency of LLM outputs, its use inevitably affects LLM reasoning patterns, which may improve model performance and reduce generalizability of findings to studies or domains that do not use such structured prompting. Finally, LLM developers make frequent changes to their models, including updates to existing, already released models (e.g., updates to the Gemini 2.5 Pro model without access to previous versions), and further releases of brand new models (e.g., the release of GPT-5.1, a distinct model from GPT-5). Updates to existing models can result in these LLMs having different abilities and competencies than when they were originally tested, including potential decreases in performance. The models included in this study are considered stable due to shifted focus to and release of newer model ranges by Google and OpenAI (e.g., Gemini 3 and GPT-5.5 model ranges), which reduces concerns regarding changes in the included modelsâ performance in future. However, as newer models are released, the potential for older ones to eventually be depreciated (i.e., no longer accessible) remains. Furthermore, the included older models may fail to represent the full spectrum of these newer frontier LLMsâ abilities in the various discussed domains. However, our aim was to determine whether current LLMs were generally competent and trustworthy enough to be used for automated systematic evidence synthesis purposes, which we have found to be true. Each newly released LLM tends to 23 AI Can Match Domain Experts in Evidence Extraction and Appraisal beat prior modelsâ performance, and if this trend continues, the concerns found in the present study may no longer be applicable to these new models. Figure 12: Summary of LLM performance by question type, overall, and on the core objectives. âGPT-5 models" refers to GPT-5 and GPT-5 Nano. âGemini models" refers to Gemini 2.5 Pro and Gemini 2.5 Flash. âLarger models" refers to Gemini 2.5 Pro and GPT-5, while âSmaller models" refers to Gemini 2.5 Flash and GPT-5 Nano. All models responded similarly to experts. GPT-5 models consistently responded most similarly to experts Multiple Choice Questions (MCQs) Likert-Scale Questions Long-Answer (Free-Text) Questions Across All Papers Interpreting Nuanced Biomedical Language Critically Appraising Biomedical Research Applying Microbial Oncogenesis Criteria Producing Accurate, Trustworthy Outputs with Minimal Hallucinations or Omissions PERFORMANCE OVERALL AND BY QUESTION TYPE PERFORMANCE ON CORE OBJECTIVES SectionMethod of LLM assessmentFinding(s) A 10 0 E1E2E3 B 0 1 1 Expert-LLM: ( 1 + 1 + 0 ) / 3 = 0.66 MCQ match proportions Inter-expert: ( 1 + 0 + 0 ) / 3 = 0.33 A E1E2E3 B Agreed most with consensus A GPT-5 models B Gemini models 12345 Gemini 2.5 Flash: slightly higher ratings Gemini 2.5 Pro: more lenient (paper appraisal and applying microbial oncogenesis criteria) GPT-5 models: most similar ratings Agreed most with outlier Expert consensus 533 01 1 E1E2E3 4 1 2 2 Response ResponseResponse 02 2 E1E2E3 and Response 3 1 1 E3 E2 E1 Response 4 Combined expert LLM Judge scored response overlap (0-4) using two methods: * (1) Pairwise overlap scores(2) Combined overlap scores Expert-LLM: ( 2 + 2 + 1 ) / 3 = 1 .66 Inter-expert: ( 0 + 1 + 1 ) / 3 = 0.66 Likert-scale distances All models produced comprehensive responses that captured key ideas across the three expert responses. response *Calculated in same way as MCQ and Likert-scale question types Response completeness did not differ by model size. Larger models Best Smaller models Misunderstood some of study methodology Worst GPT-5 models Most similar to experts Gemini models More lenient than experts Least similar to experts GPT-5 Most similar to experts Gemini models (extracted evidence similarly, lenient with criteria) GPT-5 Nano (applied criteria similarly, inconsistent reasoning). Least similar to experts Rare, majority by smaller models In methodology-related and contradiction- detection question items. Hallucinations Minor impact Exception: methodology- related question items. Omissions 4.6 Conclusion Overall, pre-trained reasoning LLMs (particularly GPT-5 models) were indistinguishable from experts across the structured biomedical evidence extraction and appraisal tasks. Hallucinations were rare in larger models (Gemini 2.5 Pro, GPT-5) under these constrained prompting conditions, but more frequent in smaller models (Gemini 2.5 Flash, GPT-5 Nano), although the rate of these errors could not be directly compared with human experts. These findings support integration of reasoning LLMs into AI evidence synthesis systems, particularly for evidence extraction and integrative summarization from individual biomedical research papers. However, methodological critique, scoring biases, and identification of contradictions remain key areas requiring further system refinement. Conflict of Interest Statement The authors declare that the research was conducted in the absence of any commercial or financial relationships that could be construed as a potential conflict of interest. 24 AI Can Match Domain Experts in Evidence Extraction and Appraisal Author Contributions K: Conceptualization, Methodology, Investigation, Data curation, Formal analysis, Writing - original draft. BAB, RFB, RK, MZM, EKS, and HW: Conceptualization, Methodology, Investigation, Software (development of the data collection application, by EKS), Supervision (BAB, RFB, and RK), and contribution to the development and refinement of the evaluation framework, including hallucination classification; Writing - review & editing. RD, NI, NAI, KN, JN, EEN, and RP: Investigation and Data curation (expert dataset generation); Writing - review & editing. All authors contributed to manuscript revision, read, and approved the submitted version. Funding The work reported herein was made possible through funding by the South African Medical Research Council (SAMRC) through its Division of Research Capacity Development under the SAMRC Clinician Researcher Development Pro- gramme with funding received from the National Department of Health. Additional support was provided by the Wits Health Consortium and the Infectious Diseases and Oncology Research Institute (IDORI). Article processing charges for this publication were supported by IDORI. Acknowledgments The authors thank Prof. Raquel Duarte and Ms Caryn McNamara for their advice and support on this project. 25 AI Can Match Domain Experts in Evidence Extraction and Appraisal References [1] Microbiology by numbers. Nature Reviews Microbiology, 9:628â628, 8 2011. [2]Dang Anh-Hoang, Vu Tran, and Le Minh Nguyen. Survey and analysis of hallucinations in large language models: attribution to prompting strategies or model behavior. Frontiers in Artificial Intelligence, 8:1622292, 9 2025. [3] Yaara Artsi, Vera Sorin, Benjamin S. Glicksberg, Panagiotis Korfiatis, Robert Freeman, Girish N. Nadkarni, and et al. Challenges of implementing llms in clinical practice: Perspectives. Journal of Clinical Medicine, 14:6169, 9 2025. [4]Giancarla Bernardo, Valentino Le Noci, Martina Di Modica, Elena Montanari, Tiziana Triulzi, Serenella M. Pupa, and et al. The emerging role of the microbiota in breast cancer progression. Cells, 12:1945, 8 2023. [5] Johan Boye and Birger Moell. Large language models and mathematical reasoning failures. 2 2025. [6] Francesco Brigo, Serena Broggi, Gionata Strigaro, Sasha Olivo, Valentina Tommasini, Magdalena Massar, and et al. Artificial intelligence (chatgpt 4.0) vs. human expertise for epileptic seizure and epilepsy diagnosis and classification in adults: An exploratory study. Epilepsy and Behavior, 166, 5 2025. [7]OgĂŒn BĂŒlbĂŒl, Hande Melike BĂŒlbĂŒl, and Esat Kaba. Assessing chatgptâs summarization of 68ga psma pet/ct reports for patients. Abdominal Radiology, 50:1467â1474, 3 2025. [8]Robert Callahan, Uma Mudunuri, Sharon Bargo, Ahmed Raafat, David Mccurdy, Corinne Boulanger, and et al. Genes affected by mouse mammary tumor virus (mmtv) proviral insertions in mouse mammary tumors are deregulated or mutated in primary human mammary tumors. Oncotarget, 3:1320, 2012. [9] Davide Castelvecchi. Can we open the black box of ai? Nature News, 538:20, 10 2016. [10] Alberto Cedro-Tanda, Alejandro CĂłrdova-Solis, Teresa JuĂĄrez-Cedillo, Emmanuel Pina-JimĂ©nez, Marta E. HernĂĄndez-Caballero, Christian Moctezuma-Meza, and et al. Prevalence of hmtv in breast carcinomas and unaffected tissue from mexican women. BMC cancer, 14:942, 12 2014. [11]Mary Chappell, Mary Edwards, Deborah Watkins, Christopher Marshall, and Sara Graziadio. Machine learning for accelerating screening in evidence reviews. Cochrane Evidence Synthesis and Methods, 1:e12021, 7 2023. [12]Jiahao Cheng, Tiancheng Su, Jia Yuan, Guoxiu He, Jiawei Liu, Xinqi Tao, Jingwen Xie, and Huaxia Li. Chain-of- thought prompting obscures hallucination cues in large language models: An empirical evaluation. In Christos Christodoulopoulos, Tanmoy Chakraborty, Carolyn Rose, and Violet Peng, editors, Findings of the Association for Computational Linguistics: EMNLP 2025, pages 1272â1305, Suzhou, China, November 2025. Association for Computational Linguistics. [13]Gheorghe Comanici, Eric Bieber, Mike Schaekermann, Ice Pasupat, Noveen Sachdeva, Inderjit Dhillon, and et al. Gemini 2.5: Pushing the frontier with advanced reasoning, multimodality, long context, and next generation agentic capabilities. 7 2025. [14] Emma Croxford, Yanjun Gao, Elliot First, Nicholas Pellegrino, Miranda Schnier, John Caskey, Madeline Oguss, Graham Wills, Guanhua Chen, Dmitriy Dligach, Matthew M. Churpek, Anoop Mayampurath, Frank Liao, Cherodeep Goswami, Karen K. Wong, Brian W. Patterson, and Majid Afshar. Evaluating clinical ai summaries with large language models as judges. npj Digital Medicine 2025 8:1, 8:640â, 11 2025. [15]Muhammed Muaaz Dawood, Mohammad Zaid Moonsamy, Kaela Kokkas, Hairong Wang, Robert F. Breiman, Richard Klein, and et al. Small language models can use nuanced reasoning for health science research classifica- tion: A microbial-oncogenesis case study. 12 2025. [16]Catherine de Martel, Damien Georges, Freddie Bray, Jacques Ferlay, and Gary M. Clifford. Global burden of cancer attributable to infections in 2018: a worldwide incidence analysis. The Lancet Global Health, 8:e180âe190, 2 2020. [17]NathĂĄlia de Sousa Pereira, Glauco Akelinghton Freire Vitiello, Bruna Karina Banin-Hirata, Glaura Scantam- burlo Alves Fernandes, Maria JosĂ© Sparça Salles, and et al. Mouse mammary tumor virus (mmtv)-like env sequence in brazilian breast cancer samples: Implications in clinicopathological parameters in molecular subtypes. International journal of environmental research and public health, 17:1â14, 12 2020. [18]Reem Al Dossary, Khaled R. Alkharsah, and Haitham Kussaibi. Prevalence of mouse mammary tumor virus (mmtv)-like sequences in human breast cancer tissues and adjacent normal breast tissues in saudi arabia. BMC cancer, 18:170, 2 2018. [19]Michael W. Dougherty and Christian Jobin. Intestinal bacteria and colorectal cancer: etiology and treatment. Gut Microbes, 15:2185028, 3 2023. 26 AI Can Match Domain Experts in Evidence Extraction and Appraisal [20]Asghar Ghasemi, Parvin Mirmiran, Khosrow Kashfi, and Zahra Bahadoran. Scientific publishing in biomedicine: A brief history of scientific journals. International Journal of Endocrinology and Metabolism, 21:e131812, 1 2022. [21] J. J. Goedert, C. S. Rabkin, and S. R. Ross. Prevalence of serologic reactivity against four strains of mouse mammary tumour virus among us women with breast cancer. British Journal of Cancer, 94:548â551, 2 2006. [22] Juraj Gottweis, Wei-Hung Weng, Alexander Daryin, Tao Tu, Petar Sirkovic, Artiom Myaskovsky, Grzegorz Glowaty, Felix Weissenberger, Alessio Orlandi, Dan Popovici, Anil Palepu, Keran Rong, Ryutaro Tanno, Khaled Saab, Fan Zhang, Jacob Blum, Andrew Carroll, Kavita Kulkarni, Nenad TomaĆĄev, Dina Zverinski, Ivor Rendulic, Elahe Vedadi, Florian Hasler, Luka Rimanic, Marina Boia, Ivan Budiselic, Ben Feinstein, Mathias Bellaiche, Tom Sheffer, Jan Freyberg, Jeremy Ratcliff, Ottavia Bertolli, Katherine Chou, Avinatan Hassidim, Burak Gokturk, Amin Vahdat, Yuan Guan, Vikram Dhillon, Eeshit Dhaval Vaishnav, Byron Lee, Tiago R. D. Costa, JosĂ© R. PenadĂ©s, Gary Peltz, Yossi Matias, James Manyika, Demis Hassabis, Yunhan Xu, Pushmeet Kohli, Annalisa Pawlosky, Alan Karthikesalingam, and Vivek Natarajan. Accelerating scientific discovery with co-scientist. Nature 2026 655:8122, 655:487â496, 5 2026. [23] David Gough, Phil Davies, Gro Jamtvedt, Etienne Langlois, Julia Littell, Tamara Lotfi, and et al. Evidence synthesis international (esi): Position statement. Systematic Reviews 2020 9:1, 9:155â, 7 2020. [24] Ishita Gupta, Reem Al-Sarraf, Hanan Farghaly, Semir Vranic, Ali A. Sultan, Hamda Al-Thawadi, and et al. Incidence of hpvs, ebv, and mmtv-like virus in breast cancer in qatar. Intervirology, 65:188â194, 10 2022. [25]Ishita Gupta, Monika Ulamec, Melita Peric-Balja, Snjezana Ramic, Ala Eddin Al Moustafa, Semir Vranic, and et al. Presence of high-risk hpvs, ebv, and mmtv in human triple-negative breast cancer. Human vaccines & immunotherapeutics, 17:4457â4466, 2021. [26]Lei Huang, Weijiang Yu, Weitao Ma, Weihong Zhong, Zhangyin Feng, Haotian Wang, and et al. A survey on hallucination in large language models: Principles, taxonomy, challenges, and open questions. ACM Trans. Inf. Syst., 43(2):42, January 2025. [27]IARC Working Group on the Evaluation of Carcinogenic Risks to Humans. General Remarks, volume 100B, page 35. International Agency for Research on Cancer (IARC), 2012. [28]Stanislav Indik, Walter H. GĂŒnzburg, Pavel Kulich, Brian Salmons, and Francoise Rouault. Rapid spread of mouse mammary tumor virus in cultured human breast cells. Retrovirology, 4:73, 10 2007. [29]International Agency for Research on Cancer (IARC). List of classifications: Agents classified by the iarc monographs, volumes 1-137, 2025. [30]Lisa M. James and Apostolos P. Georgopoulos. Breast cancer, viruses, and human leukocyte antigen (hla). Scientific reports, 14:16179, 12 2024. [31]John Jumper, Richard Evans, Alexander Pritzel, Tim Green, Michael Figurnov, Olaf Ronneberger, and et al. Highly accurate protein structure prediction with alphafold. Nature 2021 596:7873, 596:583â589, 7 2021. [32]Hafiz Fawad Khalid, Amjad Ali, Nida Fawad, Shazia Rafique, Inam Ullah, Gouhar Rehman, and et al. Mmtv-like virus and c-myc over-expression are associated with invasive breast cancer. Genetics and Evolution, 91:104827, 2021. [33] Hafiz Fawad Khalid, Sadia Bibi, Amjad Ali, Nida Fawad, Muhammad Usman Shams, Wafa Idrees, and et al. Decoding the mystery of mmtv-like virus and its relationship with breast cancer metastasis. Journal of infection and public health, 16:1396â1402, 9 2023. [34]Yubin Kim, Hyewon Jeong, Shan Chen, Shuyue Stella Li, Chanwoo Park, Mingyu Lu, and et al. Medical hallucinations in foundation models and their impact on healthcare. 11 2025. [35]Rodney P. Kincaid, Neena G. Panicker, Mary M. Lozano, Christopher S. Sullivan, Jaquelin P. Dudley, and Farah Mustafa. Mmtv does not encode viral micrornas but alters the levels of cancer-associated host micrornas. Virology, 513:180â187, 1 2018. [36] Matthew Chung Yi Koh, Jinghao Nicholas Ngiam, Jolene Ee Ling Oon, Lionel Hon Wai Lum, Nares Smitasin, and Sophia Archuleta. Using chatgpt for writing hospital inpatient discharge summaries â perspectives from an inpatient infectious diseases service. BMC Health Services Research, 25:221, 12 2025. [37]Bhushan B. Kulkarni, Shivaprakash V. Hiremath, Suyamindra S. Kulkarni, Umesh R. Hallikeri, Basavaraj R. Patil, and Pramod B. Gai. Genomic dna of mcf-7 breast cancer cells not an ideal choice as positive control for pcr amplification based detection of mouse mammary tumor virus-like sequences. Journal of Virological Methods, 193:304â307, 11 2013. 27 AI Can Match Domain Experts in Evidence Extraction and Appraisal [38]Francesca Lessi, Nicole Grandi, Chiara Maria Mazzanti, Prospero Civita, Cristian Scatena, Paolo Aretini, and et al. A human mmtv-like betaretrovirus linked to breast cancer has been present in humans at least since the copper age. Aging, 12:15978â15994, 8 2020. [39]Guowei Li, Luciana P.F. Abbade, Ikunna Nwosu, Yanling Jin, Alvin Leenus, Muhammad Maaz, and et al. A scoping review of comparisons between abstracts and full reports in primary biomedical research. BMC medical research methodology, 17:181, 12 2017. [40]Jierui Li, Vipul Raheja, and Dhruv Kumar. Contradoc: Understanding self-contradictions in documents with large language models. Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL 2024, 1:6509â6523, 2024. [41] Yilan Li, Tianshu Gu, Chengyuan Yang, Minghui Li, Congyi Wang, Lan Yao, and et al. Ai-assisted hypothesis generation to address challenges in cardiotoxicity research: Simulation study using chatgpt with gpt-4o. Journal of Medical Internet Research, 27:e66161, 2025. [42]Pauline Luczynski, Philip Poulin, Kamila Romanowski, and James C. Johnston. Tuberculosis and risk of cancer: A systematic review and meta-analysis. PLOS ONE, 17:e0278661, 12 2022. [43]Giacomo Marzi, Marco Balzano, and Davide Marchiori. K-alpha calculatorâkrippendorffâs alpha calculator: A user-friendly tool for computing krippendorffâs alpha inter-rater reliability coefficient. MethodsX, 12:102545, 6 2024. [44] Chiara Maria Mazzanti, Francesca Lessi, Ivana Armogida, Katia Zavaglia, Sara Franceschi, Mohammad Al Hamad, and et al. Human saliva as route of inter-human infection for mouse mammary tumor virus. Oncotarget, 6:18355â18363, 2015. [45]Benjamin R. McFadden, Mark Reynolds, and Timothy J.J. Inglis. Developing machine learning systems worthy of trust for infection science: a requirement for future implementation into clinical practice. Frontiers in Digital Health, 5:1260602, 9 2023. [46]Koji Miyabayashi, Hideaki Ijichi, and Mitsuhiro Fujishiro. The role of the microbiome in pancreatic cancer. Cancers, 14:4479, 9 2022. [47]Patrick S. Moore and Yuan Chang. Why do viruses cause cancer? highlights of the first century of human tumour virology. Nature reviews. Cancer, 10:878, 12 2010. [48] Abigail Morales-SĂĄnchez, TzindilĂș Molina-Muñoz, Juan L.E. MartĂnez-LĂłpez, Paulina HernĂĄndez-SancĂ©n, Alejandra Mantilla, Yelda A. Leal, and et al. No association between epstein-barr virus and mouse mammary tumor virus with breast cancer in mexican women. Scientific reports, 3, 10 2013. [49]Antonio Giuseppe Naccarato, Francesca Lessi, Katia Zavaglia, Cristian Scatena, Mohammad A. Al Hamad, Paolo Aretini, and et al. Mouse mammary tumor virus (mmtv) - like exogenous sequences are associated with sporadic but not hereditary human breast carcinoma. Aging, 11:7236â7241, 9 2019. [50] National Cancer Institute. A to z list of cancer types. [51] National Cancer Institute. What is cancer?, 10 2021. Accessed: 13/08/2025. [52]Wasifa Naushad, Orooj Surriya, and Hajra Sadia. Prevalence of ebv, hpv and mmtv in pakistani breast cancer patients: A possible etiological role of viruses in breast cancer. Infection, genetics and evolution : journal of molecular epidemiology and evolutionary genetics in infectious diseases, 54:230â237, 10 2017. [53] Kim Nordmann, Stefanie Sauter, Mirjam Stein, Johanna Aigner, Marie-Christin Redlich, Michael Schaller, and et al. Evaluating the performance of artificial intelligence in summarizing pre-coded text to support evidence synthesis: a comparison between chatbots and humans. BMC Medical Research Methodology, 25:150, 2025. [54] JosĂ© Otero and Walter Kintsch. Failures to detect contradictions in a text: What readers believe versus what they read. Psychological Science (0956-7976), 3:229â235, 7 1992. [55]Anne D. Otten, Michel M. Sanders, and G. Stanley McKnight. The mmtv ltr promoter is induced by progesterone and dihydrotestosterone but not by estrogen. Molecular endocrinology (Baltimore, Md.), 2:143â147, 1988. [56]Donald M. Parkin, Lucia HĂ€mmerl, Jacques Ferlay, and Eva J. Kantelhardt. Cancer in africa 2018: The role of infections. International Journal of Cancer, 146:2089â2103, 4 2020. [57]Raisa Perzova, Lynn Abbott, Patricia Benz, Steve Landas, Seema Khan, Jordan Glaser, and et al. Is mmtv associated with human breast cancer? maybe, but probably not. Virology journal, 14:196, 10 2017. [58] May Lynn Reese, Markela Zeneli, Mindy Ng, Jacob Haimes, Andreea Damien, and Elizabeth Stade. Using llm-as- a-judge/jury to advance scalable, clinically-validated safety evaluations of model responses to users demonstrating psychosis. Proceedings of IASEAI Conference, 2:610â624, 7 2026. 28 AI Can Match Domain Experts in Evidence Extraction and Appraisal [59]Malekpour Afshar Reza, Mollaie Hamid Reza, Lashkarizadeh Mahdiyeh, Fazlalipour Mehdi, and Zeinali Nejad Hamid. Evaluation frequency of merkel cell polyoma, epstein-barr and mouse mammary tumor viruses in patients with breast cancer in kerman, southeast of iran. Asian Pacific journal of cancer prevention : APJCP, 16:7351â7357, 2015. [60] Shervin Shariatpanahi, Najma Farahani, Ahmad Reza Salehi, and Rasoul Salehi. High prevalence of mouse mammary tumor virus-like gene sequences in breast cancer samples of iranian women. Nucleosides, nucleotides & nucleic acids, 36:621â630, 2017. [61]Md Kamrul Siam, Angel Varela, Md Jobair Hossain Faruk, Jerry Q. Cheng, Huanying Gu, Abdullah Al Maruf, and et al. Benchmarking large language models on the united states medical licensing examination for clinical reasoning and medical licensing scenarios. Scientific Reports 2025 16:1, 16:1387â, 12 2025. [62]Aaditya Singh, Adam Fry, Adam Perelman, Adam Tart, Adi Ganesh, Ahmed El-Kishky, and et al. Openai gpt-5 system card, 2026. [63]Karan Singhal, Shekoofeh Azizi, Tao Tu, S. Sara Mahdavi, Jason Wei, Hyung Won Chung, and et al. Large language models encode clinical knowledge. Nature 2023 620:7972, 620:172â180, 7 2023. [64]Alexandre F.R. Stewart and Hsiao Huei Chen. Revisiting the mmtv zoonotic hypothesis to account for geographic variation in breast cancer incidence. Viruses, 14:559, 3 2022. [65]Hedieh Moradi Tabriz, Kazem Zendehdel, Reza Shahsiah, Forouzandeh Fereidooni, Baharak Mehdipour, and Zahra Mostakhdemin Hosseini. Lack of detection of the mouse mammary tumor-like virus (mmtv) env gene in iranian women breast cancer using real time pcr. Asian Pacific journal of cancer prevention : APJCP, 14:2945â2948, 2013. [66]Derek Tam, Anisha Mascarenhas, Shiyue Zhang, Sarah Kwan, Mohit Bansal, and Colin Raffel. Evaluating the factual consistency of large language models through news summarization. Proceedings of the Annual Meeting of the Association for Computational Linguistics, Findings of the Association for Computational Linguistics: ACL 2023:5220â5255, 2023. [67]Guy Tsafnat, Paul Glasziou, Miew K. Choong, Adam Dunn, Filippo Galgani, and Enrico Coiera. Systematic review automation technologies. Systematic Reviews, 3:1â15, 4 2014. [68]Rens van de Schoot, Jonathan de Bruin, Raoul Schram, Parisa Zahedi, Jan de Boer, Felix Weijdema, and et al. An open source machine learning framework for efficient and transparent systematic reviews. Nature Machine Intelligence 2021 3:2, 3:125â133, 2 2021. [69]Rebecca Toumi van Dorsten and Robert F Breiman. A landscape review with novel criteria to evaluate microbial drivers for cancer: Priorities for innovative research targeting excessive cancer mortality in sub-saharan africa. Frontiers in Cellular and Infection Microbiology, 15, 7 2025. [70] Geoff Watts. Harald zur hausen. The Lancet, 402:20, 7 2023. [71]Jason Wei, Xuezhi Wang, Dale Schuurmans, Maarten Bosma, Brian Ichter, Fei Xia, and et al. Chain-of-thought prompting elicits reasoning in large language models. Advances in Neural Information Processing Systems, 35:24824â24837, 1 2022. [72] World Health Organisation (WHO). Cervical cancer elimination initiative, 2022. Accessed: 14/01/2025. [73]Vinicius Zambaldi, David La, Alexander E Chu, Harshnira Patani, Amy E Danson, Tristan O C Kwan, and et al. De novo design of high-affinity protein binders with alphaproteo. 9 2024. [74]Maroun Bou Zerdan, Joseph Kassab, Paul Meouchy, Elio Haroun, Rami Nehme, Morgan Bou Zerdan, and et al. The lung microbiota and lung cancer: A growing relationship. Cancers, 14:4813, 10 2022. [75]Kang Zhang, Xin Yang, Yifei Wang, Yunfang Yu, Niu Huang, Gen Li, and et al. Artificial intelligence in drug development. Nature Medicine, 31:45â59, 1 2025. 29 AI Can Match Domain Experts in Evidence Extraction and Appraisal A Supplementary Material A.1 Figures Figure S1: Assessment prompt for LLM judge (GPT-5, high reasoning effort) for pairwise long-answer comparisons (pairwise overlap scores). You are serving as an impartial medical research evaluator. Your task is to score how well two responses match each other in meaning and nuance. The topic is whether HMTV / MMTV-like virus can cause breast cancer in humans. SCORING RUBRIC: âą0: No overlap: Responses share no common points. âą1: Poor: Some minor overlap exists, but the responses differ substantially on key points or important nuances. âą2: Fair: Several key points overlap, but significant differences remain in multiple important nuances. âą3: Good: Most key points overlap, but the responses differ on one or two minor points. âą4: Excellent: The responses are fully comprehensive, capturing all the same points and nuances. DATA TO EVALUATE: Question: âquestionâ response 1 response 2 INSTRUCTIONS: Respond *only* with a JSON object containing: 1. âscore": A number from 0 to 4, expressed to one decimal place. 2. âjustification": A brief explanation for your score (2-4 sentences). If one response includes information not in the other response, begin your justification with: [EXTRA_INFO] If this extra information is also irrelevant to the question, begin your justification with: [EXTRA_INFO] [IRRELEVANT_INFO] Do not include any text before or after the JSON. 30 AI Can Match Domain Experts in Evidence Extraction and Appraisal Figure S2: Assessment prompt for LLM judge (GPT-5, high reasoning effort) for LLM to combined expert response comparison (combined overlap score). You are a medical researcher assessing the ability of LLMs to evaluate and summarise biomedical research papers. An LLM has been given a research paper on the topic of whether HMTV / MMTV-like virus can cause breast cancer in humans, along with a set of questions to answer based on the paper. Score the LLM answer using the scoring rubric below based on its completeness, accuracy, succinctness, and focus compared to the three provided expert reference answers. Scoring Rubric: âą0: Irrelevant or Incorrect: The generated answer is completely off-topic or factually wrong compared to the references. âą1: Poor: The answer includes a minor point but misses the main idea(s) or important nuances captured in the references. âą2: Fair: The answer captures some of the key ideas but is incomplete, misses some nuances, includes irrelevant information, or is more verbose than necessary. âą3: Good: The answer includes most of the key points from the references but either omits a minor point, includes some irrelevant information, or is slightly more verbose than necessary. âą4: Excellent: The answer is comprehensive, succinct, and does not contain irrelevant information. It is as good as, or better than, the expert reference answers or combination thereof. Here is the data to evaluate: Generated Answer: llm_response reference_block INSTRUCTIONS: Provide your evaluation in a JSON format with two keys: 1. âscore": An integer from 0 to 4. 2.âjustification": A brief explanation for your score. If the LLM answer includes any information not present in the references, begin your justification with the flag â[EXTRA_INFO]". Do not add any text before or after the JSON object. 31 AI Can Match Domain Experts in Evidence Extraction and Appraisal A.2 Tables Table S1: Manual scoring of pairwise overlap score was done for two question instances (all LLMs included, resulting in 30 scores for comparison). The table shows the number of question instances where the LLM-judge score was an exact match to the human-judge score, and where the LLM-judge score was higher or lower than the human-judge score. The human and LLM-judge scores, along with their respective justifications, can be found in the dataset. Initial human scores were discussed by the research team. The majority of LLM-judge scores were very close to the human-judge scores (â€1 point difference). In total, 10/30 scores were exact matches, 11/30 deviated byâ€0.5 points, 7/30 deviated by 1 point, and 2/30 deviated by 2 points. Comparison of human and LLM-judge pairwise overlap scores Evaluated groupExact MatchHigher than human judge Lower than human judge Experts15 * 0 Gemini 2.5 Pro31 ** 2 ** Gemini 2.5 Flash31 *** 2 ** GPT-532 *** 1 *** GPT-5 Nano05 *** 1 *** Overall10186 * Three deviated byâ€0.5 points, while two deviated by 2 points. ** Deviated from human score byâ€0.5 points. *** Deviated from human score byâ€1 point. Table S2: Manual scoring of combined overlap score was done for eight question instances (two per LLM). The table shows the number of question instances where the LLM-judge score was an exact match to the human-judge score, and where the LLM-judge score was higher or lower than the human-judge score. The human and LLM-judge scores, along with their respective justifications, can be found in the dataset. Initial human scores were further discussed by the research team. The majority of LLM- judge scores matched the human-judge scores exactly (6/8), with 2/8 scores deviating by 1 point. Comparison of human and LLM-judge combined overlap scores Evaluated LLMExact MatchHigher than human judge Lower than human judge Gemini 2.5 Pro11 * 0 Gemini 2.5 Flash101 * GPT-5200 GPT-5 Nano200 Overall61 * 1 * * Deviated from human score by 1 point. 32 AI Can Match Domain Experts in Evidence Extraction and Appraisal A.3 Extraction Template (Questionnaire) Category one questions (answered by domain experts) are colored black, while category two questions (Q2-11.1, Q13-16, Q19-21, and Q23-25.1, answered by clinician-scientist trainee) are colored blue. Note that âSAQ"s in the template are referred to as MCQs in the manuscript, while âMCQ"s in the template are referred to as multi-select questions in the manuscript. You are a medical researcher analysing the literature to evaluate the plausibility of HMTV/MMTV-like virus as a causative agent of breast cancer in humans. For each paper, you will generate a summary and evaluation of its findings by answering 77 predefined questions. These individual summaries will contribute to a final determination of whether HMTV/MMTV-like virus plays a causative role in human breast cancer and whether this warrants further research. To achieve this, carefully review the given paper and respond to questions 1 through 30, some of which are in SAQ or MCQ format, while others are long answer questions. Relevancy 1. How relevant is this article in determining whether it is plausible that HMTV/MMTV-like virus can cause breast cancer in humans (relevant, somewhat relevant, or irrelevant)? A relevant article would fulfil either of the following criteria: 1) gives causal evidence for whether or not HMTV/MMTV-like virus or its species/family of microbes cause breast cancer or other cancers either in humans or animal models, or 2) contains a theoretical model for how HMTV/MMTV-like virus might cause breast cancer, irrespective of whether there is experimental evidence. A somewhat relevant article would fulfil either of the following criteria: 1) mentions both HMTV/MMTV-like virus and breast cancer, but does not give any causal arguments or evidence for HMTV/MMTV-like virus causing breast cancer, or 2) contains evidence examining whether other microbes cause breast cancer, if the findings can plausibly generalise to HMTV/MMTV-like virus and breast cancer. An irrelevant article would be a paper that does not fulfil the criteria for the relevant or somewhat relevant classifications. a)Relevant b)Somewhat relevant c)Irrelevant 1.1 Briefly explain why the paper is relevant, somewhat relevant, or irrelevant. Only continue to answer the questions after this if the paper is relevant. Paper Bibliographic Details 2. Title of paper 3. Publication date 4. Journal of publication 5. Study design (e.g., randomised control trial, cohort study, case-control study, cross-sectional studies, laboratory study, etc) 6. What country was the research conducted in? 7. DOI Paper Integrity & Reliability 8. Are the references relevant to the topic of this paper and the sections of the text that they are cited in? a)Yes, all references are relevant. b)No, there were irrelevant references. 8.1 If any of the references were irrelevant, list them and state why they are irrelevant. 9. Was there any evidence of conflict of interest in this paper? This may include, but is not limited to, funding of research that could compromise the researchersâ objectivity, and patent ownership related to the research. a)Yes, there was evidence of a conflict of interest in the paper. b)No, there was no evidence of a conflict of interest in the paper. 33 AI Can Match Domain Experts in Evidence Extraction and Appraisal 9.1 If there was evidence of a conflict of interest, what was it? 10. Are there any contradictions within the paper? Examples of contradictions include inconsistencies in reported sample sizes and participant characteristics, methodological discrepancies between stated procedures and actual implementation, or variations in how results are reported across different sections of the paper. a)Yes, there were contradictions within the paper. b)No, there were no contradictions found within the paper. 10.1 If there were contradictions, list them. 11. Is there any evidence of image or data manipulation in this paper? E.g., using the same image to describe different results by re-labelling, rotating, changing the colours or contrast of the image, or cropping the image differently. a)Yes, there was evidence of image or data manipulation in the paper. b)No, there was no evidence of image or data manipulation in the paper. 11.1 If there was evidence of image or data manipulation, what specific alterations were identified? 12. Determine how reliable this paper is and produce a paper reliability score out of 5. 5)No issues detected. 4) Some noticeable issues, such as a few minor contradictions, or an identifiable (but not severe) conflict of interest. 3)Integrity is reasonable but weakened by repeated minor contradictions, a few irrelevant citations, or questionable choices in data presentation. 2)The paper contains major contradictions, numerous irrelevant references, and/or noticeable conflicts of interest that raise concerns about bias. 1)The paper is entirely unreliable due to extensive issues, including irrelevant citations, extreme bias from ma- jor conflicts of interest, multiple major contradictions, and/or strong indications of data or image manipulation. Summary of Paper Contents 13. What was the aim of the study? 14. What was the hypothesis of the study? 15. Provide a comprehensive, bullet-point summary detailing the demographic characteristics of the study sample, specifically the distribution of age, sex, ethnicity, and economic status. 16. Provide a comprehensive, bullet-point summary detailing the cancer types and subtypes included in the study sample, along with the quantity of each. Specify histological subtypes, grades, stages, receptor status, and any other relevant characteristics. 17. Based on the sample demographics and the country in which this study was performed, are the findings of this study applicable to African populations in general? a) Yes, the findings are applicable as populations across Africa share similar genetic, socioeconomic, and/or environmental conditions. b)Partially applicable, but caution should be exercised due to differences in genetic, socioeconomic, and/or environmental conditions. c)No, the findings cannot be applied to African populations due to significant contextual differences. d)Applicability cannot be determined without further information about the studyâs sample. 18. List the most important factor(s) in the methodology that influence(s) the strength of the studyâs findings. 19. According to the paper, do the relevant results of this research support any other paper/s published prior to it? a)Yes, this paperâs findings support previous research. b)No, this paperâs approach is novel and thus its results do not support or contradict previous findings. c)No, this paperâs findings contradict what has previous been found. 34 AI Can Match Domain Experts in Evidence Extraction and Appraisal d)Some of the findings support previous research, while other findings contradict previous research. 20. List all the papers referenced in this article that may assist in determining whether it is plausible that HMTV/MMTV- like virus can cause breast cancer in humans. Only include papers that would be classified as relevant as per the original classification criteria. Strength of Evidence 21. If applicable, was the data collected with probability sampling (random selection)? a)Yes, random selection was used. b)No, random selection was not used when it should have been. c)It is unclear whether random selection was used. d)N/A, random selection is not possible with this study design. 22. If applicable, was the sample size for cases - and, where relevant, controls - sufficient to support the paperâs conclusions? a)Yes, the sample size was sufficient. b)The sample size was smaller than ideal, yet sufficiently large that, combined with strong statistical findings, the conclusions remain credible. c)No, the sample size was insufficient to support the conclusions. d)N/A, no sample used in this study. 23. What is the statistical significance of the paperâs relevant findings? 24. Was the statistical analysis appropriate in this paper? a)Yes, all statistical analysis was appropriate. b)No, the statistical analysis was inappropriate. c) It is unclear what statistical analysis was performed, making it difficult to determine whether it was appropriate. 25. Are there any obvious errors in the statistical analysis or presentation of the paperâs results? a)Yes, there was an error in the paperâs statistical analysis or presentation of results. b)No, there were no errors in the paperâs statistical analysis or presentation of results. 25.1 If there was an error in the statistical analysis or presentation of the results, what was it? 26. Are there any limitations to using this research as evidence for or against a plausible causative relationship between HMTV/MMTV-like virus and breast cancer? a)Yes, there are limitations. b)No, there are no limitations. 26.1 If there are limitations, what are they? 27. Assess the strength of evidence provided by this paper and assign a score from 1 to 5, with 1 representing severely weak / no credible evidence and 5 representing very strong evidence. 5)Very strong evidence 4)Strong evidence 3)Acceptable evidence 2)Weak evidence 1)Severely weak evidence/no credible evidence 35 AI Can Match Domain Experts in Evidence Extraction and Appraisal Does this paper help fulfil or disprove the microbial oncogenesis criteria? 28. Determine if this paper helps fulfil the following listed 10 microbial oncogenicity criteria for the MMTV-like virus/HMTV and breast cancer pair. Give 1 point for each criterion the paper supports or clearly fulfils, 0 for a criterion that is either not examined by the paper or remains uncertain at the end of the paper, and -1 for any criterion that is disproved by this paper. Briefly explain the reasoning for each point or lack thereof given. Finally, determine the cumulative score from these criteria points, allocated as the causal criteria score. When answering these questions, consider the paper reliability score (Q12) and the strength of evidence score (Q30). 28.1 Epidemiologic Association: Consistent association of a microbe (or a combination of microbes) with a specific cancer type within a population (considering demographics and host factors) of humans (or across all populations), especially when compared to people within the same population(s) without cancer. a.Tools to consider: i.Case-control and Cohort studies i.Histopathology, immunochemistry, PCR i.Genomics and mass spectrometry iv. Strong association when consistent findings across different geographical regions and consistent risk ratios in case-control or cohort studies 28.1.1 Indicate whether the paper supports, refutes, does not investigate, or leaves this criterion uncertain. a)0, the paper did not examine this criterion. b) 0, the paper did investigate this criterion, but the findings are either inconclusive or uncertainty remains due to concerns about the studyâs reliability or strength of evidence. c)1, the paper provides evidence that supports this criterion. d)-1, the paper provides evidence that refutes this criterion. 28.1.2 If this criterion was supported, refuted, or remains uncertain, select all applicable explanations below that describe why. a)Insufficient power b)Inappropriate methodology/tools c)Concerns regarding paper reliability d)Error/s in the statistical analysis e)Sufficient power f)Appropriate methodology/tools g)Other (specify) 28.1.3 If you selected other, please specify. 28.1.4 State which kind of evidence is provided according to the tools to consider categories and identify the specific finding/s that fulfil or refute this criterion. 28.2 Histopathologic association: Consistent detection of the microbe (virus, bacterium, fungus, or parasite) in cancer tissues compared to healthy controls a.Tools to consider: i.In Vitro assays (PCR, FACS, imaging) i.Histopathology, immunochemistry, PCR i.Genomics and mass spectrometry 28.2.1 Indicate whether the paper supports, refutes, does not investigate, or leaves this criterion uncertain. a)0, the paper did not examine this criterion. b) 0, the paper did investigate this criterion, but the findings are either inconclusive or uncertainty remains due to concerns about the studyâs reliability or strength of evidence. 36 AI Can Match Domain Experts in Evidence Extraction and Appraisal c)1, the paper provides evidence that supports this criterion. d)-1, the paper provides evidence that refutes this criterion. 28.2.2 If this criterion was supported, refuted, or remains uncertain, select all applicable explanations below that describe why. a)Insufficient power b)Inappropriate methodology/tools c)Concerns regarding paper reliability d)Error/s in the statistical analysis e)Sufficient power f)Appropriate methodology/tools g)Other (specify) 28.2.3 If you selected other, please specify. 28.2.4 State which kind of evidence is provided according to the tools to consider categories and identify the specific finding/s that fulfil or refute this criterion. 28.3 Temporal association: Evidence that infection precedes onset of cancer a.Tools to consider: i.Longitudinal (long term) studies 28.3.1 Indicate whether the paper supports, refutes, does not investigate, or leaves this criterion uncertain. a)0, the paper did not examine this criterion. b)0, the paper did investigate this criterion, but the findings are either inconclusive or uncertainty remains due to concerns about the studyâs reliability or strength of evidence. c)1, the paper provides evidence that supports this criterion. d)-1, the paper provides evidence that refutes this criterion. 28.3.2 If this criterion was supported, refuted, or remains uncertain, select all applicable explanations below that describe why. a)Insufficient power b)Inappropriate methodology/tools c)Concerns regarding paper reliability d)Error/s in the statistical analysis e)Sufficient power f)Appropriate methodology/tools g)Other (specify) 28.3.3 If you selected other, please specify. 28.3.4 Identify the specific finding/s that fulfil or refute this criterion. 28.4 Experimental evidence of facilitation of oncogenesis: Induction of cancer/precancerous changes upon introduc- tion of the isolated microbe into appropriate models. Does oncogenic cellular transformation occur when a microbe is introduced into an animal (ideally a primate or other human-representative) model or tissue model? a.Tools to consider: i.3D biosystems (organoids) or animal models 28.4.1 Indicate whether the paper supports, refutes, does not investigate, or leaves this criterion uncertain. a)0, the paper did not examine this criterion. 37 AI Can Match Domain Experts in Evidence Extraction and Appraisal b)0, the paper did investigate this criterion, but the findings are either inconclusive or uncertainty remains due to concerns about the studyâs reliability or strength of evidence. c)1, the paper provides evidence that supports this criterion. d)-1, the paper provides evidence that refutes this criterion. 28.4.2 If this criterion was supported, refuted, or remains uncertain, select all applicable explanations below that describe why. a)Insufficient power b)Inappropriate methodology/tools c)Concerns regarding paper reliability d)Error/s in the statistical analysis e)Sufficient power f)Appropriate methodology/tools g)Other (specify) 28.4.3 If you selected other, please specify. 28.4.4 Identify the specific finding/s that fulfil or refute this criterion. 28.5 Molecular and Multi-omics evidence for interaction: i. Epigenetics (e.g. DNA methylation changes in H. pylori or HPV E6 oncogene resulting in P53 degradation), as well as mutational signatures i.Mutation effects (stimulating cellular mutations promoting cancer) i.Stimulating immune evasion mechanisms reducing immunosurveillance for emerging cancer cells iv.Integration of microbial DNA into host genome a.Tools to consider: i.in vitro assays, organoids, multi-omics assessments including: a.Molecular Evidence i.For bacteria: Identification of toxins, effector proteins, or metabolites that alter cellular signalling i.For viruses: Characterization of viral oncoproteins or insertional mutagenesis i.For fungi: Demonstration of mycotoxins or immune-modulating molecules with carcinogenic potential iv. For parasites: Identification of chronic inflammatory responses or direct tissue damage mechanisms v.Evidence of microbial interference with DNA repair, cell cycle regulation, apoptosis, or immune surveillance b.Genomic evidence: i.Microbial DNA sequences or integration sites in tumor genome; i. Host genetic susceptibility factors that enhance oncogenic potential of specific mi- crobes, including demonstration of how microbial exposure can alter (or promote) known host genetic risk factors for cancer. i.Demonstratable facilitation of cancer-associated mutational signatures. c.Transcriptomic evidence: i.Expression of microbial genes or altered host gene expression profiles d.Proteomic evidence: i.Detection of microbial proteins or altered host protein responses e.Metabolomic evidence: i.Microbial metabolites or altered host metabolic pathways f.Microbiome analyses: i.Consistent dysbiosis patterns associated with specific cancers 38 AI Can Match Domain Experts in Evidence Extraction and Appraisal 28.5.1 Indicate whether the paper supports, refutes, does not investigate, or leaves this criterion uncertain. a)0, the paper did not examine this criterion. b)0, the paper did investigate this criterion, but the findings are either inconclusive or uncertainty remains due to concerns about the studyâs reliability or strength of evidence. c)1, the paper provides evidence that supports this criterion. d)-1, the paper provides evidence that refutes this criterion. 28.5.2 If this criterion was supported, refuted, or remains uncertain, select all applicable explanations below that describe why. a)Insufficient power b)Inappropriate methodology/tools c)Concerns regarding paper reliability d)Error/s in the statistical analysis e)Sufficient power f)Appropriate methodology/tools g)Other (specify) 28.5.3 If you selected other, please specify. 28.5.4 State which kind of evidence is provided according to the tools to consider categories and identify the specific finding/s that fulfil or refute this criterion. 28.6 Prevention: Does preventing the infection (or removing exposure to the microbe) reduce cancer or evidence of oncogenesis a.Tools to consider: i.Clinical trials i.Relevant animal models i.For viruses: Reduced cancer incidence following vaccination (e.g., HPV, HBV) iv.For bacteria: Cancer prevention through antibiotic treatment or bacterial elimination v.For fungi: Antifungal intervention effects on precancerous lesions vi.For parasites: Impact of antiparasitic treatment on cancer development vii.Prevention trials showing reduced cancer incidence after targeting the microbe 28.6.1 Indicate whether the paper supports, refutes, does not investigate, or leaves this criterion uncertain. a)0, the paper did not examine this criterion. b) 0, the paper did investigate this criterion, but the findings are either inconclusive or uncertainty remains due to concerns about the studyâs reliability or strength of evidence. c)1, the paper provides evidence that supports this criterion. d)-1, the paper provides evidence that refutes this criterion. 28.6.2 If this criterion was supported, refuted, or remains uncertain, select all applicable explanations below that describe why. a)Insufficient power b)Inappropriate methodology/tools c)Concerns regarding paper reliability d)Error/s in the statistical analysis e)Sufficient power f)Appropriate methodology/tools g)Other (specify) 39 AI Can Match Domain Experts in Evidence Extraction and Appraisal 28.6.3 If you selected other, please specify. 28.6.4 State which kind of evidence is provided according to the tools to consider categories and identify the specific finding/s that fulfil or refute this criterion. 28.7 Reproducibility and Validation: Confirmation across independent laboratories, optimally using varied method- ologies or models. Replication in diverse human populations and geographic settings. Concordance between in vitro, animal, and human studies 28.7.1 Indicate whether the paper supports, refutes, does not investigate, or leaves this criterion uncertain. a)0, the paper did not examine this criterion. b) 0, the paper did investigate this criterion, but the findings are either inconclusive or uncertainty remains due to concerns about the studyâs reliability or strength of evidence. c)1, the paper provides evidence that supports this criterion. d)-1, the paper provides evidence that refutes this criterion. 28.7.2 If this criterion was supported, refuted, or remains uncertain, select all applicable explanations below that describe why. a)Insufficient power b)Inappropriate methodology/tools c)Concerns regarding paper reliability d)Error/s in the statistical analysis e)Sufficient power f)Appropriate methodology/tools g)Other (specify) 28.7.3 If you selected other, please specify. 28.7.4 Identify the specific finding/s that fulfil or refute this criterion. 28.8 Dose Response relationship: Relationship of severity or chronicity of the presumed offending infection with cancer initiation and severity in human longitudinal studies or in experimental models demonstrate the association of infectious dose with development of cancer a.Tools to consider: i.Organoids, animal models, or clinico-epidemiologic longitudinal studies 28.8.1 Indicate whether the paper supports, refutes, does not investigate, or leaves this criterion uncertain. a)0, the paper did not examine this criterion. b) 0, the paper did investigate this criterion, but the findings are either inconclusive or uncertainty remains due to concerns about the studyâs reliability or strength of evidence. c)1, the paper provides evidence that supports this criterion. d)-1, the paper provides evidence that refutes this criterion. 28.8.2 If this criterion was supported, refuted, or remains uncertain, select all applicable explanations below that describe why. a)Insufficient power b)Inappropriate methodology/tools c)Concerns regarding paper reliability d)Error/s in the statistical analysis e)Sufficient power f)Appropriate methodology/tools 40 AI Can Match Domain Experts in Evidence Extraction and Appraisal g)Other (specify) 28.8.3 If you selected other, please specify. 28.8.4 State which kind of evidence is provided according to the tools to consider categories and identify the specific finding/s that fulfil or refute this criterion. 28.9 Plausibility: Are there plausible mechanisms for considering a potential role for a microbe a.Tools to consider: i.Knowledge-based; i.e. understanding the physiology of microbial interaction with humans may suggest or support a role in carcinogenesis 28.9.1 Indicate whether the paper supports, refutes, does not discuss, or leaves this criterion uncertain. a)0, the paper does not suggest or discuss any plausible mechanisms for the MCP. b)0, the paper suggests a possible mechanism, but it is weakly justified, inconsistent with existing knowledge, or supported by insufficient evidence from the study. c) 1, the paper presents a plausible mechanism, either consistent with existing knowledge of cancer biology or microbial pathogenesis or supported by credible evidence produced by the present study. d)-1, the paper presents reasoning or mechanisms that argue against a plausible role for the microbe in carcinogenesis. 28.9.2 If this criterion was supported, refuted, or remains uncertain, select all applicable explanations below that describe why. a)Insufficient power b)Inappropriate methodology/tools c)Concerns regarding paper reliability d)Error/s in the statistical analysis e)Sufficient power f)Appropriate methodology/tools g)Other (specify) 28.9.3 If you selected other, please specify. 28.9.4 Identify the specific finding/s that fulfil or refute this criterion. 28.10 Impact of co-Factors: Microbial carcinogenesis occurs (or is accelerated) in an environment with promotive host factors, such as genetic risk, immunodeficiencies, nutritional status, and/or with environmental factors (including but not limited to known carcinogens) a.Tools to consider: i.Epidemiological studies in humans i.Laboratory studies i.Animal models 28.10.1 Indicate whether the paper supports, refutes, does not investigate, or leaves this criterion uncertain. a)0, the paper did not examine this criterion. b)0, the paper did investigate this criterion, but the findings are either inconclusive or uncertainty remains due to concerns about the studyâs reliability or strength of evidence. c)1, the paper provides evidence that supports this criterion. d)-1, the paper provides evidence that refutes this criterion. 28.10.2 If this criterion was supported, refuted, or remains uncertain, select all applicable explanations below that describe why. 41 AI Can Match Domain Experts in Evidence Extraction and Appraisal a)Insufficient power b)Inappropriate methodology/tools c)Concerns regarding paper reliability d)Error/s in the statistical analysis e)Sufficient power f)Appropriate methodology/tools g)Other (specify) 28.10.3 If you selected other, please specify. 28.10.4 State which kind of evidence is provided according to the tools to consider categories and identify the specific finding/s that fulfil or refute this criterion. 29 Calculate the cumulative score from these criteria points and provide the final causal criteria score. 30 Indicate whether the paper supports, refutes, or does not impact the plausibility of HMTV/MMTV-like virus causing breast cancer. 5)Strongly supports (moderate to strong evidence) 4)Weakly supports (weak to moderate evidence) 3)Does not impact (negligible or weak evidence) 2)Weakly refutes (weak to moderate evidence) 1)Strongly refutes (moderate to strong evidence) 42 AI Can Match Domain Experts in Evidence Extraction and Appraisal A.4 Abridged Extraction Template for Repeat Experiments This version of the extraction template was used for repeat experiments (30) with GPT-5. For the second paper, question item 4.4 (microbial oncogenesis criterion âImpact of co-factors") was replaced with the microbial oncogenesis criterion âPlausibility" question item (see question item 28.9 and 28.9.1 in the full extraction template). You are a medical researcher analysing the literature to evaluate the plausibility of HMTV/MMTV-like virus as a causative agent of breast cancer in humans. To achieve this, carefully review the given paper and respond to the questions that follow. Relevancy 1. How relevant is this article in determining whether it is plausible that HMTV/MMTV-like virus can cause breast cancer in humans (relevant, somewhat relevant, or irrelevant)? A relevant article would fulfil either of the following criteria: 1) gives causal evidence for whether or not HMTV/MMTV-like virus or its species/family of microbes cause breast cancer or other cancers either in humans or animal models, or 2) contains a theoretical model for how HMTV/MMTV-like virus might cause breast cancer, irrespective of whether there is experimental evidence. A somewhat relevant article would fulfil either of the following criteria: 1) mentions both HMTV/MMTV-like virus and breast cancer, but does not give any causal arguments or evidence for HMTV/MMTV-like virus causing breast cancer, or 2) contains evidence examining whether other microbes cause breast cancer, if the findings can plausibly generalise to HMTV/MMTV-like virus and breast cancer. An irrelevant article would be a paper that does not fulfil the criteria for the relevant or somewhat relevant classifications. a)Relevant b)Somewhat relevant c)Irrelevant Strength of Evidence 2.If applicable, was the sample size for cases - and, where relevant, controls - sufficient to support the paperâs conclusions? a)Yes, the sample size was sufficient. b) The sample size was smaller than ideal, yet sufficiently large that, combined with strong statistical findings, the conclusions remain credible. c)No, the sample size was insufficient to support the conclusions. d)N/A, no sample used in this study. 3. Assess the strength of evidence provided by this paper and assign a score from 1 to 5, with 1 representing severely weak / no credible evidence and 5 representing very strong evidence. 5)Very strong evidence 4)Strong evidence 3)Acceptable evidence 2)Weak evidence 1)Severely weak evidence/no credible evidence Does this paper help fulfil or disprove the microbial oncogenesis criteria? 4. Determine if this paper helps fulfil the following listed microbial oncogenicity criteria for the MMTV-like virus/HMTV and breast cancer pair. Give 1 point for each criterion the paper supports or clearly fulfils, 0 for a criterion that is either not examined by the paper or remains uncertain at the end of the paper, and -1 for any criterion that is disproved by this paper. When answering these questions, consider the strength of evidence score (Q3). 4.1 Epidemiologic Association: Consistent association of a microbe (or a combination of microbes) with a specific cancer type within a population (considering demographics and host factors) of humans (or across all populations), especially when compared to people within the same population(s) without cancer. a.Tools to consider: 43 AI Can Match Domain Experts in Evidence Extraction and Appraisal i.Case-control and Cohort studies i.Histopathology, immunochemistry, PCR i.Genomics and mass spectrometry iv. Strong association when consistent findings across different geographical regions and consistent risk ratios in case-control or cohort studies 4.1.1 Indicate whether the paper supports, refutes, does not investigate, or leaves this criterion uncertain. a)0, the paper did not examine this criterion. b)0, the paper did investigate this criterion, but the findings are either inconclusive or uncertainty remains due to concerns about the studyâs reliability or strength of evidence. c)1, the paper provides evidence that supports this criterion. d)-1, the paper provides evidence that refutes this criterion. 4.2 Histopathologic association: Consistent detection of the microbe (virus, bacterium, fungus, or parasite) in cancer tissues compared to healthy controls a.Tools to consider: i.In Vitro assays (PCR, FACS, imaging) i.Histopathology, immunochemistry, PCR i.Genomics and mass spectrometry 4.2.1 Indicate whether the paper supports, refutes, does not investigate, or leaves this criterion uncertain. a)0, the paper did not examine this criterion. b)0, the paper did investigate this criterion, but the findings are either inconclusive or uncertainty remains due to concerns about the studyâs reliability or strength of evidence. c)1, the paper provides evidence that supports this criterion. d)-1, the paper provides evidence that refutes this criterion. 4.3 Dose Response relationship: Relationship of severity or chronicity of the presumed offending infection with cancer initiation and severity in human longitudinal studies or in experimental models demonstrate the association of infectious dose with development of cancer a.Tools to consider: i.Organoids, animal models, or clinico-epidemiologic longitudinal studies 4.3.1 Indicate whether the paper supports, refutes, does not investigate, or leaves this criterion uncertain. a)0, the paper did not examine this criterion. b) 0, the paper did investigate this criterion, but the findings are either inconclusive or uncertainty remains due to concerns about the studyâs reliability or strength of evidence. c)1, the paper provides evidence that supports this criterion. d)-1, the paper provides evidence that refutes this criterion. 4.4 Impact of co-Factors: Microbial carcinogenesis occurs (or is accelerated) in an environment with promotive host factors, such as genetic risk, immunodeficiencies, nutritional status, and/or with environmental factors (including but not limited to known carcinogens) a.Tools to consider: i.Epidemiological studies in humans i.Laboratory studies i.Animal models 4.4.1 Indicate whether the paper supports, refutes, does not investigate, or leaves this criterion uncertain. 44 AI Can Match Domain Experts in Evidence Extraction and Appraisal a)0, the paper did not examine this criterion. b)0, the paper did investigate this criterion, but the findings are either inconclusive or uncertainty remains due to concerns about the studyâs reliability or strength of evidence. c)1, the paper provides evidence that supports this criterion. d)-1, the paper provides evidence that refutes this criterion. 5 Indicate whether the paper supports, refutes, or does not impact the plausibility of HMTV/MMTV-like virus causing breast cancer. 5)Strongly supports (moderate to strong evidence) 4)Weakly supports (weak to moderate evidence) 3)Does not impact (negligible or weak evidence) 2)Weakly refutes (weak to moderate evidence) 1)Strongly refutes (moderate to strong evidence) 45