Paper deep dive
Scientific Knowledge Discovery in the Age of Large Language Models
Eleni Adamidi, Serafeim Chatzopoulos, Thanasis Vergoulis
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 87%
Last extracted: 8/4/2026, 10:58:38 AM
Summary
This paper surveys 34 peer-reviewed studies on using generative Large Language Models (LLMs) for scientific knowledge discovery, specifically focusing on literature retrieval and screening. It analyzes methodologies, model types, evaluation metrics, and application domains, noting a concentration in biomedical research and a shift from keyword-based search to semantic, agentic, and RAG-based approaches.
Entities (8)
Relation Signals (6)
Large Language Models → supports → Literature Retrieval
confidence 95% · Generative large language models (LLMs) offer a more flexible alternative, supporting literature retrieval
Large Language Models → supports → Literature Screening
confidence 95% · supporting literature retrieval and the screening of candidate studies
BIP! Finder → uses → OpenAIRE Graph
confidence 90% · Relevant literature was identified through BIP! Finder... using the OpenAIRE Graph as its primary data source.
Chelli et al. (2024) → addresses → Literature Retrieval
confidence 85% · One paper, Chelli et al. (2024), appears in both categories [Literature Retrieval and Screening]
Chelli et al. (2024) → addresses → Literature Screening
confidence 85% · One paper, Chelli et al. (2024), appears in both categories
Liu et al. (2025b) → uses → Neo4j
confidence 80% · let GPT-4o generate Cypher queries against the resulting Neo4j graph
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The rapid growth of scholarly literature has made identifying relevant publications increasingly difficult, and conventional search systems still depend heavily on manually formulated queries and effortful manual inspection. Generative large language models (LLMs) offer a more flexible alternative, supporting literature retrieval and the screening of candidate studies against eligibility criteria. This chapter surveys 34 peer-reviewed papers applying generative LLMs to these two tasks, identified via a Boolean search over the OpenAIRE Graph (1,589 records screened to 34 inclusions). Reviewed studies are characterised by LLMs employed, model access and adaptation, prompting and architectural techniques, ground-truth sources, and evaluation metrics.
Tags
Links
- Source: https://arxiv.org/abs/2607.26670v1
- Canonical: https://arxiv.org/abs/2607.26670v1
Trouble viewing inline? Open PDF directly →
Full Text
65,967 characters extracted from source content.
Expand or collapse full text
Scientific Knowledge Discovery in the Age of Large Language Models Eleni Adamidi [0000−0001−9925−1560] , Serafeim Chatzopoulos [0000−0003−1714−5225] and Thanasis Vergoulis [0000−0003−0555−4128] Abstract The rapid growth of scholarly literature has made identifying relevant publi- cations increasingly difficult, and conventional search systems still depend heavily on manually formulated queries and effortful manual inspection. Generative large language models (LLMs) offer a more flexible alternative, supporting literature retrieval and the screening of candidate studies against eligibility criteria. This chapter surveys34 peer-reviewed papers applying generative LLMs to these two tasks, identified via a Boolean search over the OpenAIRE Graph (1, 589records screened to34inclusions). Reviewed studies are characterised by LLMs employed, model access and adaptation, prompting and architectural techniques, ground-truth sources, and evaluation metrics. Keywords Scientific Knowledge Discovery·LLMs·Literature retrieval·Literature screening 1 Introduction The volume of scientific literature continues to expand rapidly, complicating researchers’ efforts to identify, assess, and keep pace with relevant advances in their fields (Hanson et al., 2024). This growth shows no sign of slowing and may be accelerating further with AI writing assistants that enable faster article drafting and publication. At this scale, researchers risk overlooking relevant work, duplicating existing efforts, or incorporating evidence after substantial delay. These pressures intensify the need for more effective methods to support scientific knowledge discovery. This chapter surveys peer-reviewed work on the use of generative large language models (LLMs) 1 to support two core scientific knowledge-discovery tasks: literature retrieval, which concerns identifying, ranking, or recommending publications relevant to a query, seed paper, user profile, or research context, and literature screening, which Eleni Adamidi IMSI, ATHENA RC, e-mail: eleni.adamidi@athenarc.gr Serafeim Chatzopoulos IMSI, ATHENA RC, e-mail: schatz@athenarc.gr Thanasis Vergoulis IMSI, ATHENA RC, e-mail: vergoulis@athenarc.gr 1 From now on, we use the term LLM specifically for generative LLMs. 1 arXiv:2607.26670v1 [cs.DL] 29 Jul 2026 2Eleni Adamidi, Serafeim Chatzopoulos and Thanasis Vergoulis concerns assessing candidate publications against predefined eligibility criteria, typically in systematic reviews and other evidence-synthesis workflows. Traditionally, retrieval has relied on keyword matching (Gusenbauer and Haddaway, 2020), citation-based ranking (Bascur et al., 2023), and recommender techniques such as content-based and collaborative filtering (Beel et al., 2016), while screening has depended largely on manual review, with semi-automated tools (e.g., Abstrackr, EPPI-Reviewer) mainly helping to prioritize or classify records for human assessment (Marshall and Wallace, 2019). Despite these advances, both tasks remain labor-intensive. LLMs offer a more flexible alternative since they can interpret natural-language requests, capture semantic relationships, and produce structured judgments or explanations (Brown et al., 2020; Ouyang et al., 2022). The chapter reviews these approaches with the goal of providing a comprehensive overview of the landscape and identifying recurring patterns. 2 Survey Methodology and Literature Overview 2.1 Search and Screening Strategy Our survey was motivated by the need to characterize the rapidly evolving landscape of LLM-based methods for scientific literature retrieval and screening, aiming to assemble a corpus broad enough to cover the main methodological paradigms and application domains, yet focused enough to support meaningful synthesis. Relevant literature was identified through BIP! Finder (Vergoulis et al., 2019), 2 a scholarly search engine indexing250+ million research publications using the OpenAIRE Graph (Manghi et al., 2026), 3 as its primary data source. The search was restricted to publications from 2019 onward, covering the period in which LLMs became practically available in research-support contexts. The following Boolean query was applied: Search Query ("large language model" OR LLM OR agentic) AND ("literature exploration" OR "scientific exploration" OR "scientific knowledge discovery" OR "literature discovery" OR "academic search" OR "scholarly search" OR "paper recommendation" OR "citation recommendation" OR "study recommendation" OR "literature retrieval" OR "paper retrieval" OR "publication retrieval" OR "literature screening") The query was designed to capture work that explicitly addresses LLM-based methods for knowledge discovery tasks, with the focus on retrieval and screening. The search was conducted in May 2026; hence, the landscape presented in this survey reflects the state of the literature available at that time, based on BIP! DB-v20.1 4 (Vergoulis et al., 2021). 2 BIP! Finder: https://bip.athenarc.gr/ 3 OpenAIRE Graph: https://graph.openaire.eu/ 4 BIP! DB is the underlying database powering BIP! Finder. Scientific Knowledge Discovery in the Age of Large Language Models3 The search query returned1, 589candidate papers, which were subsequently dedu- plicated and screened for eligibility according to predefined inclusion and exclusion criteria (see below). Screening was performed in two stages: an initial assessment based on the title and abstract, followed by a full-text review for papers that met the initial screening criteria. Titles and abstracts were collected automatically by BIP! Finder, whereas full-text articles were manually retrieved and assessed by the authors, who served as domain experts throughout the screening process. To be included, a study had to have undergone peer review and satisfy all of the following criteria: (i) it addressed at least one of the target tasks; (i) it used a generative large language model as a primary component of the proposed system or method; (i) it reported a system, method, or evaluation with sufficient methodological detail to support comparison with related work; and (iv) the manuscript is written in English. Studies were excluded if they met any of the following criteria: (i) they were extended abstracts lacking sufficient methodological detail; (i) they had not undergone peer review 5 (e.g., preprints, technical reports); (i) the LLMs addressed tasks outside the scope of this survey (e.g., publication summarisation); (iv) the manuscript was not written in English; or (v) the full text could not be obtained. 2.2 The Final Corpus After deduplication, title-and-abstract screening, and full-text review, the final corpus comprised 34 papers published between 2024 and 2026. The temporal distribution (Figure 1, left) shows rapid growth, from9papers in 2024 to24in 2025, with1additional paper from early 2026 captured at the time of data collection. This pattern suggests the rapid uptake of LLMs in research workflows (the 2026 count is necessarily incomplete because the search was conducted early in this year). Citation patterns (Figure 1, right) reflect both recency and early-mover advantage. The 2024 papers have accumulated the most citations relative to their number, while the much larger 2025 cohort has received fewer citations so far, consistent with the citation lag phenomenon (Smith, 2009). The experts organised the papers into two thematic categories (Figure 2, left) by primary task: literature retrieval (푛= 21) and literature screening (푛= 14). One paper, Chelli et al. (2024), appears in both categories. Topic assignments from OpenAlex (Priem et al., 2022) (Figure 2, right) show that the corpus is concentrated in biomedical and clinical domains. The most frequent topic is “Biomedical Text Mining and Ontologies” (10papers), followed by “Topic Modeling” (7), and “AI in Healthcare and Education” and “Meta-analysis and Systematic Reviews” (6each). Overall, the distribution suggests that LLM-assisted literature workflows are being studied primarily in biomedical and evidence-synthesis settings. 5 For each preprint identified during screening, the authors manually checked for a peer-reviewed version not yet indexed by the search engine before final exclusion. 4Eleni Adamidi, Serafeim Chatzopoulos and Thanasis Vergoulis 2024 20252026 0 10 20 9 24 1 Year Number of papers 2024 20252026 0 100 200 236 54 0 Year Total citations Fig. 1: Temporal distribution of the reviewed corpus (푛= 34). Left: number of included papers per publication year. Right: Cumulative citation totals by publication year. 20 13 1 Chelli et al. (2024) Literature Retrieval Literature Screening 0 5 10 Biomedical Text Mining and Ontologies Topic Modeling AI in Healthcare and Education Meta-analysis and Systematic Reviews Radiomics and ML in Medical Imaging Semantic Web and Ontologies Computational Drug Discovery Methods Natural Language Processing Techniques Machine Learning in Healthcare Advanced Graph Neural Networks 10 7 6 6 3 3 3 3 2 2 Number of papers Fig. 2: Overview of the reviewed literature. Left: paper distribution across thematic categories. Right: top 10 research topics by paper count. 2.3 Presentation Approach The survey presents its two thematic areas in separate sections. Within each area, works are further organized into narrower groups defined by shared technical approaches or objectives. Each group is discussed using the same three-part structure: (1) an Overview summarizing the studies, their approaches, and the role of LLMs; (2) Models used, covering which LLMs are employed, how they are accessed (e.g., API, local deployment) and whether they are fine-tuned; and (3) Evaluation trends, covering how retrieval or screening quality is assessed, including the type of ground truth used, the metrics reported, and any indirect evidence from broader experiments. It is worth to mention two important conventions used. First, Author-Constructed ground truth refers to benchmarks, query sets, expert judgments, or synthetic data created specifically for the reviewed work; datasets introduced in earlier separate publications are treated as Existing Benchmarks, even if the author lists overlap. Second, only evaluations of LLM-assisted retrieval or screening components are reported: results for unrelated pipeline stages and results where LLMs appear only as baselines are excluded. Scientific Knowledge Discovery in the Age of Large Language Models5 2.4 Limitations Several potential sources of non-representativeness should be acknowledged. First, the search was conducted through a single scholarly search engine (BIP! Finder), which, while it has a broad coverage, may not index all venues uniformly. Second, the restriction to English-language publications excludes relevant work published in other languages. Third, the rapidly evolving nature of the field means that the corpus represents a snapshot: work published after the data collection date (see Section 2.1) is not included. Moreover, the survey is restricted to peer-reviewed publications, hence, preprints that had not yet undergone peer review at the time of screening (even if they are subsequently accepted and published) are not included. Finally, the inclusion criteria require explicit LLM usage as a primary method for retrieval or screening, meaning that important adjacent work is out of scope. Despite these limitations, the corpus provides a reasonably comprehensive view of the state of the field, capturing the primary methodological strands, the recent contributions, and the principal application domains in which LLM-based literature retrieval, recommendation, and screening have been deployed. 3 LLMs for Literature Retrieval Identifying relevant scientific literature underpins most research activities, yet established methods are increasingly strained by the growth described in Section 1. Effective querying requires anticipating shifting terminology, assessing relatedness requires semantic and conceptual understanding beyond surface overlap, and judging novelty or evidential relevance at scale exceeds what keyword matching or citation counts can reliably capture. Generative LLMs are therefore a natural fit for these challenges. The papers reviewed next fall into two broad groups: those where retrieval or recommendation is the primary contribution, and those where retrieval supports a downstream task such as automated peer review, clinical decision support, or survey writing. 3.1 Core Literature Search & Recommendation Overview. This is the largest and most technically diverse area identified in the survey’s literature-retrieval domain. It includes systems that retrieve scientific documents relevant either to an explicit natural-language or keyword-based query or to a stored profile of a user’s past interactions. We identify five distinct research strands within it. The first strand concerns semantic search. SemRank (Zhang et al., 2025b) addresses a limitation of dense retrievers: by encoding queries and passages holistically, they can miss the fine-grained scientific concepts that determine relevance. It uses an LLM to extract multi-granular concepts at indexing time and again at query time to identify the concepts the user is actually seeking, matching against a concept-based index 6Eleni Adamidi, Serafeim Chatzopoulos and Thanasis Vergoulis rather than raw embeddings alone. Pathfinder (Iyer et al., 2024) combines retrieval- augmented generation with time- and citation-based weighting over astronomy papers, allowing natural-language querying and embedding-space exploration; an LLM then performs consensus evaluation over the retrieved literature. Zheng et al. (2025) similarly treats relevance as a constrained optimization problem, using an LLM-conditioned encoder to embed queries and documents, then re-weighting their cosine similarity with constraints that penalize redundancy and reward semantic and logical consistency. BioMedSearch (Liu et al., 2025a) uses an LLM to decompose a biomedical question into sub-queries and keywords, builds a task graph, and fans out simultaneously across literature databases, protein databases, and general web search, applying a keyword-coverage threshold and embedding-similarity ranking to narrow the combined multi-source results before the LLM reasons over them. de Ara ́ujo Pessoa and Vogl (2025) builds a RAG pipeline for institutional publication analysis, where an LLM converts natural-language requests into full-text search queries and combines vector search with collaboration-network analysis. RefAI (Li et al., 2024b) similarly uses an LLM to turn user queries into PubMed keywords, then ranks results using citation and journal-impact signals alongside semantic relevance. Finally, Ding et al. (2025) reports using an LLM to embed patent titles, abstracts, and technical descriptions for retrieval. A second strand foregrounds the agentic dimension of search, using LLMs not just to encode meaning but to plan, route, and refine retrieval itself. TourSynbio-Search (Liu et al., 2024) use a classification layer with chain-of-thought (CoT) and few-shot prompting to decide, per query, whether to route to a paper-search agent or a protein-structure-search agent. LitChat (Huang et al., 2025) splits the problem across two cooperating agents: one that turns a conversational request into a Boolean query executable against Web of Science, Scopus, or Semantic Scholar (using in-context examples drawn from real systematic-review search strategies), and a second that mines the retrieved corpus, organized as a bibliographic knowledge graph, to answer follow-up questions. Finally, LITERAS (Gorenshtein et al., 2025) introduces a four-agent system: a Keyword Planner Agent generates PubMed-ready search strings, a Search Agent queries the PubMed API and deduplicates results, a Validation Agent filters and ranks the retrieved articles against relevance, recency, methodology, and quality criteria, and a Critic Agent loops back to refine the search if too few articles clear the threshold. A third strand grounds retrieval in structured knowledge rather than free-text similarity. GPTscholar (Pires et al., 2024) has the LLM generate a SPARQL query against the DBLP bibliographic knowledge graph before answering, explicitly adopting a “knowledge- aware inference” strategy to suppress hallucinated titles, authors, and DOIs. Liu et al. (2025b) takes a structurally similar approach in the opposite direction — it first builds the knowledge graph itself, via a novel ontology and an LLM-augmented multi-relation extraction model, then lets GPT-4o generate Cypher queries against the resulting Neo4j graph to answer content-level literature questions, arguing this substrate supports deeper exploration than keyword search or even a bare LLM. The fourth strand focuses not on retrieving relevant papers, but on the accuracy of citations that an LLM generates from its parametric memory without grounding them in a database query Chelli et al. (2024) conducted a controlled study in which ChatGPT and Bard/Gemini were asked to reproduce the reference lists of systematic reviews on rotator-cuff pathology. Scientific Knowledge Discovery in the Age of Large Language Models7 A fifth strand departs from query-driven search entirely, framing literature discovery as profile-driven literature recommendation. MCAP (Zhang et al., 2025a) represents users and papers as nodes in heterogeneous interaction graphs, and argues that the high-order message-passing propagation used by prior graph neural network (GNN) recommenders is sub-optimal, since long-distance propagation blurs useful signal with noise. Its fix (low-pass propagation with matrix completion) builds direct user–user and item–item relation graphs (from shared authorship/venue and content similarity) instead of relying on noisy multi-hop connections. LLMs (GPT-3.5-Turbo, GLM-4) play only a narrow, auxiliary role here: refining which papers count as “similar” for the item–item graph; the GNN recommender, trained from scratch, is the paper’s real contribution. Models used. In this area, many systems rely on general-purpose commercial models used off the shelf, most commonly via API, though access method is frequently left unstated. Pires et al. (2024) uses GPT-3.5-turbo variants, Li et al. (2024b) uses GPT-4- turbo, and Huang et al. (2025) and Liu et al. (2025b) use GPT-4o, all with API access explicitly confirmed. Zhang et al. (2025b) uses GPT-4.1-mini for concept extraction, although, as in Ding et al. (2025), the paper does not specify the access method. Liu et al. (2025a) compares five models side by side: GPT-4.1, DeepSeek-R1, Gemini-2.5, Llama-4, and Qwen3. de Ara ́ujo Pessoa and Vogl (2025), by contrast, uses the open- weight Llama-3.3-70B model through the UniGPT API service. Iyer et al. (2024) uses GPT-4o-mini for its post-retrieval consensus-checking module. Chelli et al. (2024) is the only Web UI-based study, testing GPT-3.5, GPT-4, and Bard/Gemini directly through their consumer chat interfaces rather than via API. Across these studies, the models are used without fine-tuning; Zhang et al. (2025a) follows the same pattern for its LLM components (GPT-3.5-Turbo, GLM-4), though these are used only to generate semantic side-information for a graph-based recommender rather than to perform retrieval directly. Domain-specific LLM fine-tuning appears in only one paper: TourSynbio-Search fine- tunes InternLM2-7B into TourSynbio-7B on curated protein literature and deploys it locally. Finally, Zheng et al. (2025) is the outlier that never names a concrete model. Columns 2-3 of Table 1 summarise the previous discussion. Evaluation trends. Many papers in this line of work evaluate literature retrieval with classic information-retrieval metrics (e.g., Recall, Precision, nDCG), reflecting that finding relevant documents is these papers’ primary focus. These metrics are calculated against existing open benchmarks (Zhang et al., 2025b,a; Zheng et al., 2025) or existing systematic reviews (Chelli et al., 2024). 6 Zhang et al. (2025a) specifically evaluates against established recommender-systems benchmarks, using standard top-푘ranking metrics (Recall@5, NDCG@5, Hit Ratio@5). Gorenshtein et al. (2025) and Pires et al. (2024) rely on manual quality and hallucination-frequency ratings on author-constructed prompt sets rather than ranking metrics. Li et al. (2024b) similarly evaluates against an author-constructed set of queries, using expert Likert-scale ratings of relevance and quality. Iyer et al. (2024) reports pipeline-level recall, nDCG, and MRR on an author-constructed ground truth, but the LLM component itself is not directly evaluated: its contribution, consensus evaluation of the retrieved set, is peripheral to these metrics 6 Chelli et al. (2024) does not involve genuine retrieval: ChatGPT and Bard generate citations from parametric memory without querying any database or corpus. Its Precision/Recall/F1 therefore measure overlap with a real reference list, not retrieval quality. 8Eleni Adamidi, Serafeim Chatzopoulos and Thanasis Vergoulis Table 1: Models, Access Modes, Retrieval-related Evaluation Ground Truth, and Metrics for Core Literature Search & Recommendation Approaches. N/R = Not reported PaperLLM ModelsAccessGr. TruthEvaluation Iyer et al. (2024) [Pathfinder] GPT-4o-miniN/R-No LLM-assisted retr. evaluation Zhang et al. (2025b) [SemRank] GPT-4.1-miniN/R Existing benchmark Recall@K, Time, Cost Zheng et al. (2025)N/RN/RExisting benchmark Precision@K,Recall@K, NDCG@K, MRR Liu et al. (2025a) [BioMedSearch] GPT-4.1, Llama-4, Gemini-2.5, DeepSeek-R1, Qwen3 API-No retr. evaluation de Ara ́ujo Pessoa and Vogl (2025) Llama-3.3-70BAPI-No retr. evaluation Li et al. (2024b) [Re- fAI] GPT-4-TurboAPIAuthor- constructed Expert satisfaction Ding et al. (2025)GPT-3.5-turbo, GPT-4 (as base- line comparison) N/RN/RAccuracy, Recall Liuetal.(2024) [TourSynbio-Search] TourSynbio-7B (fine-tuned from InternLM2-7B on curated Pro- teinLMDataset) Local-No retr. evaluation Huang et al. (2025) [LitChat] GPT-4oAPI-No retr. evaluation Gorenshteinetal. (2025) [LITERAS] GPT-4o-miniAPI Author- constructed Hallucination rate, Time (limited) Pires et al. (2024) [GPTscholar] GPT-3.5-turboAPIAuthor- constructed Correctness counts, Hallucination rate Liu et al. (2025b)GPT-4oAPIAuthor- constructed Expert satisfaction, Time (subjec- tive) Chelli et al. (2024)GPT-3.5/4, Gemini (Bard)Web UIExisting syst. reviews Precision, Recall, F1, Hallucination rate (based on simul. prompt) Zhang et al. (2025a) [MCAP] GPT-3.5-Turbo; GLM-4N/RExisting benchmark Recall@K, NDCG@K, Hit Ra- tio@K and is not separately assessed. Four papers (Liu et al., 2025a, 2024; Huang et al., 2025; de Ara ́ujo Pessoa and Vogl, 2025) report no formal quantitative evaluation of the core retrieval task at all, even though most of them do include other, non-retrieval-related experiments. Finally, cost and processing-time reporting is rare and mostly informal: Zhang et al. (2025b) stands out with a dedicated efficiency analysis (retriever/LLM calls per query, average LLM output length, running time, and actual dollar cost), while Gorenshtein et al. (2025) reports measured average processing time, and de Ara ́ujo Pessoa and Vogl (2025) and Liu et al. (2025b) offer only qualitative or subjective time/latency observations. Columns 4–5 of Table 1 summarise the previous discussion. 3.2 RAG-Style Literature Retrieval for Downstream Tasks Overview. All papers in this area adopt some form of RAG (Retrieval-Augmented Generation) combining literature retrieval with a downstream generation or reasoning stage conditioned on the retrieval material. However, only a subset explicitly use the term: Xu et al. (2025); Fu et al. (2025), and Bao et al. (2025) self-describe as RAG, while Peasley et al. (2025) and Zhu et al. (2025) use the term only in passing, and Zheng Scientific Knowledge Discovery in the Age of Large Language Models9 et al. (2026) and Balaskas et al. (2024) never use it at all despite following the identical retrieve-then-condition pattern. Consistent with this survey’s scope, we examine only the literature-retrieval component of these pipelines here; the quality of the downstream generation itself (e.g., review text, treatment plan, survey prose) is not covered. We identify two strands of works distinguished by the type of information retrieved. The related- and concurrent-work retrieval strand retrieves prior papers to contextualize a target paper’s novelty or limitations. DeepReview (Zhu et al., 2025) builds its retrieval stage on OpenScholar: Qwen-2.5-72B-Instruct generates key research questions about the target paper, Qwen-2.5-3B-Instruct converts these into search keywords used to pull roughly60candidate papers via the Semantic Scholar API, after which OpenScholar’s own (non-generative, cross-encoder) reranker narrows this to the top10most relevant papers and its QA model synthesizes the resulting novelty-analysis report. LIMITGEN (Xu et al., 2025) augments LLMs with RAG over the same Semantic Scholar API to ground generated limitations in prior findings. Its premise is that identifying weaknesses in a paper (e.g., missing baselines, limited datasets) requires literature awareness beyond that of a static LLM. However, the authors position their contribution primarily as a benchmark for assessing LLMs in this setting, rather than an optimized retrieval approach. PaperEval (Zheng et al., 2026) takes a more self-contained approach: ChatGPT generates representative topic keyphrases for a target paper, which are encoded via a CLIP text encoder. Cosine similarity against a corpus (restricted to concurrent, temporally-proximate publications) then retrieves a domain-aware reference set that a latent-reasoning module uses to assess the paper’s novelty and contribution. By contrast, the domain-literature evidence retrieval strand retrieves general support- ing evidence to inform a downstream generation or decision task. LITURAt (Peasley et al., 2025) is fundamentally a “plan-and-solve” scientific-data-analysis agent, built around Mixtral 8×7B, that dynamically constructs Entrez API queries to retrieve relevant PubMed abstracts alongside local datasets so its statistical summaries are contextual- ized by current literature. Fu et al. (2025) goes further, building an explicit “clinical data–literature knowledge–model decision-making” collaborative framework. The sys- tem combines RAG-retrieved oncology literature with structured patient data through a staged, fine-tuned 70B medical LLM. The system in (Balaskas et al., 2024) integrates a similar approach directly into an electronic health record interface: WizardVicuna-13B first generates a natural-language summary from a patient’s structured data; this is then embedded and used to query the literature. Dense vector retrieval returns100candidate documents, which a cross-encoder reranks to the top10. A relevance threshold then determines whether any results are shown to clinicians. The retrieval criteria explicitly prioritise patient applicability, treatment alignment, and source recency and credibility. SurveyGen/QUAL-SG (Bao et al., 2025) identifies literature retrieval as one of two key bottlenecks in existing RAG-based survey pipelines (the other being evaluation): standard retrieval selects papers by textual similarity alone and risks including low-impact or marginal work, so QUAL-SG expands the candidate pool via citation-graph analysis and re-ranks by three signals, including an LLM-judged (Claude-3.7-Sonnet) relevance score, before handing the retrieved literature to the generation stage. 10Eleni Adamidi, Serafeim Chatzopoulos and Thanasis Vergoulis Table 2: Models, Access Modes, Retrieval-related Evaluation Ground Truth, and Metrics for RAG-Style Lit. Retrieval for Downstream Tasks Approaches. N/R = Not reported PaperLLM ModelsAccess Gr. TruthEvaluation Zhu et al. (2025) [DeepReview] Qwen-2.5-3B-Instruct; Qwen-2.5- 72B-Instruct;Llama-3.1Open- Scholar-8B(fine-tunedfrom Llama3-8B) N/R- No retr. evaluation; Cost reporting (examining tradeoffs) Xu et al. (2025) [LIMITGEN] GPT-4o; GPT-4o-mini; Lla-Ma-3.3- 70B; Qwen2.5-72B API-No retr. evaluation (indirect signal from RAG-context ablation) Zheng et al. (2026) [PaperEval] GPTN/R-No retr. evaluation (indirect signal from ranking ablations) Peasley et al. (2025) [LITURAt] Mixtral-8x7BLocal- No retr. evaluation (indirect signal from the combined report stability experiment) Fu et al. (2025)Zhuomuniao-70B(undisclosed base architecture) N/R-Recall/Precision Proxies (e.g., per- centage of high-impact literature); Time (end-to-end proc. times) Balaskasetal. (2024) WizardVicuna-13B(fine-tuned from LlaMa-13B) Local-No retr. evaluation Bao et al. (2025) [QUAL-SG] Claude-3.7-SonnetN/R Author- constructed Precision, Recall, F1 Models used. The works in this area draw on a genuinely diverse set of models. Regarding LLM access, Xu et al. (2025) uses an API, 7 while Peasley et al. (2025) and Balaskas et al. (2024) instead use locally deployed open-weight models: Mixtral 8×7B and WizardVicuna-13B, respectively. The remaining studies—(Zhu et al., 2025), (Zheng et al., 2026), (Fu et al., 2025), and (Bao et al., 2025)-do not explicitly describe how their models are accessed or deployed. Regarding the models, Zhu et al. (2025) uses three generative LLMs in its retrieval pipeline: Qwen-2.5-72B-Instruct for research questions generation, Qwen-2.5-3B-Instruct for search keywords generation, and Llama- 3.1OpenScholar-8B (fine-tuned from Llama3-8B) for novelty-report synthesis. The pipeline also uses OpenScholar’s cross-encoder, a non-generative reranker fine-tuned from BAAI/bge-reranker-large, to reduce the retrieved candidate pool to the top 10. 8 (Xu et al., 2025) tests four frontier LLMs head-to-head, zero-shot: GPT-4o, GPT-4o-mini, Llama-3.3-70B, and Qwen2.5-72B. Zheng et al. (2026) names only ChatGPT (version unspecified) for keyphrase generation in its retrieval module 9 while the backbone of PaperEval’s own reasoning module is never named. Fu et al. (2025)’s 70B “Zhuomuniao” model has an undisclosed base architecture, fine-tuned in three stages (LoRA domain adaptation, contrastive subtype discrimination, RLHF-based contraindication reduction). Bao et al. (2025)’s retrieval pipeline is anchored by Claude-3.7-Sonnet, which performs the LLM-judged topical-relevance scoring inside QUAL-SG itself and is also the paper’s best-performing and most extensively used model overall, spanning all three evaluation tasks. Columns 2-3 of Table 2 summarise the previous discussion. 7 Based on the information in the author’s code repository. 8 The work additionally uses other LLMs for tasks that do not touch literature retrieval, including DeepReviewer-14B (paper’s full-parameter fine-tune of Phi-4 14B). 9 GPT-4o and GPT-4o-mini appearing elsewhere in that paper belong to separate baseline methods, not to PaperEval’s own pipeline Scientific Knowledge Discovery in the Age of Large Language Models11 Evaluation trends. The dominant pattern is that literature-retrieval quality is rarely evaluated against a dedicated ground truth: six of the seven papers report no retrieval- specific ground truth, relying instead on indirect signals or ground truth scoped to the downstream task. This likely reflects where these papers place their contribution: retrieval is a means to an end (e.g., grounding a review, a treatment plan, a patient summary) so evaluation effort naturally concentrates on the quality of that downstream output rather than on the retrieval step that feeds it. Zhu et al. (2025) reports no retrieval-relevant signal of any kind; even its LLM-as-judge review-quality scores are holistic ratings of the final review text, with no rubric item or ablation attributing any part of the judgment to retrieval specifically. Xu et al. (2025) and Zheng et al. (2026) each provide an indirect signal via ranking-position ablations: Xu et al. (2025) varies which subset of already-retrieved papers is used downstream (Top 5 vs. Top 3 vs. Last 5 of 18), and Zheng et al. (2026) both removes its retrieval module entirely (“w/o Ret.”) and sweeps the number of retrieved papers푘; however, neither compares retrieved content against a labeled relevant-document set, so both remain indirect signals rather than genuine retrieval evaluations. Peasley et al. (2025) provides a similar indirect signal. The authors limit their only ground-truth accuracy evaluation to the local-data-analysis module, noting that they could not establish ground truth for the literature review. Hence, insights come from an internal/external consistency experiment (Cohen’s휅, self-judged by Mixtral) on the combined report. Fu et al. (2025) reports what we term recall/precision proxies. Compared with a no-RAG baseline, its RAG-enabled system achieves a 42.3% increase in citations to recent literature, serving as a proxy for recall or coverage, and cites a higher proportion of high-impact literature, serving as a proxy for precision or quality. Both measures are computed from the system’s own outputs rather than evaluated against an external ground truth of relevant documents. Balaskas et al. (2024) is a clean no-evaluation case, reinforced by the authors’ own admission that retrieval-quality assessment is left to future work. Bao et al. (2025) is the only paper with genuine, validated retrieval ground truth: it builds a dedicated author-constructed dataset (4, 200+ human-written surveys with real citations) and reports citation Precision/Recall/F1 against human-selected references. A four-component ablation further isolates each ranking signal’s individual contribution: removing the LLM-judged relevance component (Claude-3.7-Sonnet) drops citation F1 by4.44points (from16.73to12.29). QUAL-SG’s retrieval quality is also validated against dedicated re-ranking baselines (RankGPT, UPR). Cost/time reporting is rare but not entirely absent: no paper reports a retrieval-specific cost or timing figure, but Zhu et al. (2025) reports a total token-generation/inference-cost tradeoff and Fu et al. (2025) reports end-to-end system processing times. Columns 4-5 of Table 2 summarise the previous discussion. 4 LLMs for Literature Screening Literature screening, the classification of candidate publications against predefined eligibility criteria, is widely seen as the most labor-intensive stage of systematic reviews and related evidence-synthesis workflows. The LLM-based papers reviewed here split into two broad lines of work: some treat screening as a single model call per paper, with 12Eleni Adamidi, Serafeim Chatzopoulos and Thanasis Vergoulis the main technical variation lying in prompt design, while others use multi-component systems intended to move beyond one-shot classification toward more practical screening tools. The following sections present these two areas in turn. 4.1 Single-Pass LLM Screening Overview. This group treats the screening decision as the output of a single, independent model call per candidate paper: one model, one prompt, one classification or, in Chelli et al. (2024)’s case, one generative response. Several papers in the group involve more than one LLM, but the models are compared, not combined. There is no chaining, no retrieval infrastructure, and no inter-model arbitration deciding the final answer. This architectural simplicity means these papers function largely as benchmarking studies, establishing how well an off-the-shelf LLM call performs at screening. The real technical variation across this group lies entirely in how the prompt and evaluation protocol are constructed, and the spread is considerable. At the simplest end, Chelli et al. (2024) and Delgado-Chaves et al. (2025) use fixed, static criteria-statement prompts: Chelli et al. (2024) tests only whether specifying a minimum number of papers to retrieve changes the model’s output, while Delgado-Chaves et al. (2025) holds the prompting technique constant across18models and instead demonstrates that the formulation of the inclusion/exclusion criteria themselves is a major performance driver, independent of model choice. Ronquillo et al. (2024) moves one step further by treating the prompt itself as the experimental variable: the same three models are each run through four escalating prompt-optimization variants, directly measuring how much structure added to the prompt (rather than which model receives it) moves accuracy. Li et al. (2024a) pushes prompt engineering further still, using few-shot examples paired with explicit reasoning explanations and running a systematic ablation across two axes: shot count (0-shot through 10-shot) and the source of the reasoning explanations attached to each example, comparing human-expert-written versus LLM-generated reasoning as a variable in its own right. Tao et al. (2025) develops this same idea into the most elaborate prompt-engineering process in the group: although the deployed screening call is still a single model (GPT-4o, Kimi, or DeepSeek, tested separately), the prompt behind that call is built offline through an iterative, three-phase process: a small seed set of PICOS-derived few-shot examples, progressive expansion with wording refinement, and finally folding misclassified cases directly back in as new few-shot examples, an active-learning-style loop that shapes the prompt without ever changing the underlying single-call architecture. Ruan et al. (2025) takes a further, distinctive detour from hand-designed prompting altogether, using meta-prompting (asking the LLMs themselves to draft candidate screening prompts), iteratively refining the machine-suggested prompts before applying them uniformly across all four models being compared. Oami et al. (2025) use a more conventional fixed role/PICOS-structured prompt but add a methodological check: the identical prompt is run three times to test reproducibility, and a post hoc ensemble rule (include if flagged by either of two models) is explored as a secondary analysis. Scientific Knowledge Discovery in the Age of Large Language Models13 Table 3: Models, Access Modes, Screening Evaluation Ground Truth, and Metrics for Single-Pass LLM Screening Approaches. N/R = Not reported PaperLLM ModelsAccessGr. TruthEvaluation Chelli et al. (2024) GPT-3.5/4; Gemini (Bard)Web UIExisting syst. reviews Precision, Recall, F1, Hallucination rate (based on simul. prompt) Delgado- Chaves et al. (2025) GPT-3.5-turbo, GPT-4o, GPT-4o-miniAPIExisting syst. reviews Precision, Recall, Specificity, F1, MCC, PABAK Gemma-7B; Gemma2-9B/27B; Llama3- 8B/-8B-instruct-fp16; Llama3.1-8B/70B; Llama3.2-3B;Reflection-70B(Llama- based); Athene-70B (Llama-based); Mistral- v0.2/v0.3-7B; Mistral-Nemo-12B; Mixtral- 8x22B; Qwen2.5-7B Local Li et al. (2024a) GPT-3.5; GPT-4; Claude-2API Existing benchmark Recall, Specificity, Precision, F1, Con- sistency (of multiple runs) Ronquillo etal. (2024) GPT-3.5 Turbo; GPT-4; Claude-2APIAuthor- constructed Accuracy (chi-square comparisons) Ruan et al. (2025) GPT-4; Claude-3.5; Gemini-1.5; DeepSeek-V3 N/R Author- constructed Recall- and Specificity-based metrics (FNF, FPF, PLR, Youden’s Index, NNS), Risk difference, Risk ratio, Re- dundancy number, Screening time Oami et al. (2025) GPT-4o; Claude-3.5-Sonnet; Gemini-1.5- Pro; Llama-3.3-70B APIExisting syst. reviews Recall, Specificity, Screening time, Cost, Consistency (of multiple runs) Tao et al. (2025) GPT-4o;Kimi(Moonshot-v1-128k); DeepSeek-Chat 2.5 API Author- constructed Accuracy, Precision, Recall, Exclusion- reason concordance rate, Time Models used. Most papers in this group use commercial chat-model families: OpenAI’s GPT-3.5/GPT-4/GPT-4o, Anthropic’s Claude (2 and 3.5), and Google’s Gemini/Bard. Delgado-Chaves et al. (2025) stands out as the exception: alongside GPT-3.5-turbo, GPT-4o, and GPT-4o-mini, they run a genuinely broad sweep of15open-source models spanning Google’s Gemma/Gemma2 family, Meta’s Llama3/3.1/3.2 line, Mistral AI’s Mistral/Mistral-Nemo/Mixtral models, and Qwen2.5-7B, 18 models in total, making it the widest single-paper model comparison in the whole set. Two of these open-source entries, Reflection-70B and Athene-70B, are themselves third-party fine-tunes of Llama (by Reflection AI and Nexusflow, respectively); the authors though do not fine-tune any model, and this holds across the whole group: all papers use off-the-shelf models, with “tuning” limited to prompt-level techniques rather than gradient-based adaptation of the underlying model. Access is mostly API-based, except for Delgado-Chaves et al. (2025), which runs open-source models locally via Ollama, and Chelli et al. (2024), which uses the web chat UI instead of an API. Columns 2-3 of Table 3 summarise the previous discussion. Evaluation trends. Ground truth sources in this group include existing systematic reviews used as a known answer key (Chelli et al., 2024; Delgado-Chaves et al., 2025; Oami et al., 2025), author-constructed ground truth built specifically for the study (Ronquillo et al., 2024; Ruan et al., 2025; Tao et al., 2025) and one pre-existing, independently curated benchmark dataset Li et al. (2024a). Regarding the used metrics, this group leans heavily on the classic diagnostic-accuracy toolkit: precision, recall/sensitivity, specificity, and F1 appear in nearly every paper, supplemented in (Delgado-Chaves et al., 2025) by agreement statistics that correct for class imbalance (MCC and PABAK). Beyond this common core, several papers introduce task-specific or efficiency-oriented metrics 14Eleni Adamidi, Serafeim Chatzopoulos and Thanasis Vergoulis that reflect their particular angle: Chelli et al. (2024)’s hallucination rate (since its task is generative rather than classificatory); Li et al. (2024a)’s and Oami et al. (2025)’s consistency metrics, both measuring agreement between repeated runs (computed by varying the few-shot examples between runs in the former and by simply rerunning the identical prompt in the latter); Ruan et al. (2025)’s false-negative and false-positive fractions (mathematically equivalent to 1 minus sensitivity and 1 minus specificity, respectively), reported alongside a set of diagnostic-performance measures borrowed from the clinical diagnostic-test tradition (PLR, Youden’s Index, NNS, risk difference, risk ratio) and a paper-specific redundancy number tracking the manual re-review burden left after screening (computed for a non-LLM comparator, RobotSearch, as well as for the four LLMs); and Oami et al. (2025)’s and Tao et al. (2025) shared emphasis on cost and processing-time efficiency alongside accuracy. Columns 4-5 of Table 3 summarise the previous discussion. 4.2 Engineered Multi-Component Pipelines Overview. This group treats the screening decision itself as the output of a designed system with multiple interacting components: pre-filtering stages, retrieval infrastructure, inter-model arbitration, or chained/tiered prompts. Critically, in every paper here the multi-component architecture is built specifically to produce the include/exclude decision, not some other downstream task; where a paper also performs data extraction or synthesis afterward, that later stage runs as a separate process on the already-screened subset and is excluded from this discussion. The shared ambition across this group is to move beyond a one-shot classification benchmark toward something closer to a production-usable screening tool, generally by reproducing the checks and balances a human review team would apply (e.g., a second reviewer, a tie-breaking adjudicator, a retrieval step) rather than relying on a single model’s raw judgment. The lightest form of this engineering is a pre-filter-then-classify cascade. Ji et al. (2024)’s ModelDB pipeline narrows a large corpus using SPECTER2 embeddings and k-nearest-neighbours distance before GPT-3.5/GPT-4 confirm inclusion criteria on the reduced set, with temperature and chain-of-thought (CoT) reasoning further tuned on top of the confirmation step; the two stages work together purely to produce the relevance decision, with any subsequent metadata extraction handled as a separate step on the papers that already passed screening. A step up in complexity is explicit inter-model arbitration, where multiple models screen independently but a designated component resolves their disagreements rather than leaving that to a human or a simple majority vote. Dai et al. (2025) has two models (ChatGPT-4o, Claude-3.5 Sonnet) screen in parallel, with a third model, Gemini-1.5 Pro, explicitly arbitrating any disagreements between them, a genuine consensus mechanism, further refined via a post hoc, error-driven CoT prompt revision cycle targeting identified false-negative patterns. Pan et al. (2025) builds a heavier version of the same idea: a tripartite framework (Doubao-1.5-pro-32k, DeepSeek-v3, DeepSeek-R1-Distill-Qwen-7B) in which the third model is purpose-built as an adjudicator, resolving conflicting inclusion/exclusion decisions between the other two and producing supporting reasoning for its ruling. Scientific Knowledge Discovery in the Age of Large Language Models15 A second architectural pattern is retrieval-augmented screening, where the system doesn’t just prompt a static context but actively retrieves information to answer the screening question. Trad et al. (2025) pairs a prompt-engineered, role-based title/abstract stage (with a conservative yes/no/unsure decision rule that retains uncertain articles) with a full-text screening stage built on retrieval-augmented generation (full-text PDFs are indexed into a vector store, and GPT-4 retrieves relevant passages to answer eligibility questions rather than relying on a fixed prompt alone). This is architecturally distinct from the arbitration papers above: rather than combining multiple models’ judgments, it’s a single model whose screening decision is scaffolded by a dedicated retrieval component. Lin et al. (2025)’s decision-tree pipeline for precision-oncology curation follows a related but distinct chaining logic: prompts are organized into sequential tiers, each tier’s output feeding the next, with each stage referencing external regulatory databases (ARTG, PBS, FDA) to assign a tiered evidence classification. The most elaborate systems in this group are explicitly agentic and iterative, with the multi-agent mechanism dedicated entirely to the screening questions. Hu et al. (2025)’s GREP-Agent integrates four LLMs with a dedicated fine-tuning phase, in which human reviewers provide feedback on a small labelled subset to iteratively refine the screening prompts before an operational phase applies workload-reduction thresholds to decide when human review of a screening decision is still needed. The entire architecture, from fine-tuning to operation, is scoped to the three inclusion/exclusion questions the system answers. Mourtzinis et al. (2025) similarly requires consensus across four LLMs on each of six screening criteria, with human experts stepping in specifically to break ties. Models used. The works in this group combine commercial and open-source models within each system rather than benchmarking many models against one another. GPT- 4/GPT-4o/GPT-4o-mini and Claude 3.5 Sonnet appear in the arbitration- and retrieval- based systems of (Dai et al., 2025; Trad et al., 2025; Hu et al., 2025). Lin et al. (2025) uses a single fully local open-source model, Mistral-Nemo-12B, while Pan et al. (2025); Mourtzinis et al. (2025) favor newer open-weight families. Hu et al. (2025)’s GREP-Agent mixes GPT-4o/GPT-4o-mini with Llama-3.1 and Phi-4. Similarly to the approaches presented in the previous section, none of these papers fine-tunes model weights. Access is confirmed as API-based for (Ji et al., 2024), (Trad et al., 2025), and (Pan et al., 2025); Lin et al. (2025) instead uses a local deployment, while Dai et al. (2025) uses a Web UI. Access is not explicitly reported for (Hu et al., 2025) or (Mourtzinis et al., 2025), despite their automated, multi-model pipelines making API access plausible. Columns 2-3 of Table 4 summarise the previous discussion. Evaluation trends. Regarding the ground truth, existing systematic reviews or ongoing systematic-review-style processes are used by three papers (Trad et al., 2025; Hu et al., 2025; Pan et al., 2025). The remaining four papers rely on author-constructed ground truth: Ji et al. (2024)’s inter-annotator-agreement gold standard, Dai et al. (2025)’s own manual screening as reference standard, Lin et al. (2025)’s single-expert review, and Mourtzinis et al. (2025)’s 20%-sample expert-labeled subset. Regarding the metrics used, this group again uses the standard precision/recall/F1/specificity core (Ji et al., 2024; Dai et al., 2025; Lin et al., 2025; Hu et al., 2025; Pan et al., 2025) but several papers layer on agreement- or efficiency-oriented metrics reflecting their engineered, deployment-oriented framing. Pan et al. (2025) reports Cohen’s휅and PABAK alongside cost and accuracy; Trad et al. (2025) use a distinctive set tailored to measuring workload 16Eleni Adamidi, Serafeim Chatzopoulos and Thanasis Vergoulis Table 4: Models, Access Modes, Screening Evaluation Ground Truth, and Metrics for Engineered Multi-Component Pipelines. N/R = Not reported PaperLLM ModelsAccessGr. TruthEvaluation Ji et al. (2024) [ModelDB] GPT-3.5; GPT-4APIAuthor- constructed Accuracy, Cohen’s휅 , F1 Dai et al. (2025)GPT-4o; Claude-3.5-Sonnet; Gemini- 1.5-Pro Web UI Author- constructed Recall, Specificity, AUC, SROC Trad et al. (2025)GPT-4APIExisting syste- matic reviews AER, FNR, Specificity, PPV, NPV, Efficiency Hu et al. (2025) [GREP-Agent] GPT-4o; GPT-4o-mini; Llama-3.1; Phi-4 N/RExisting syste- matic reviews Accuracy, Precision, Recall, F1 Lin et al. (2025) Mistral-NeMo-12BLocalAuthor- constructed Accuracy, Recall, Specificity, F1, Precision Pan et al. (2025)DeepSeek-V3; Doubao-1.5-pro; Deep- Seek-R1-Distill-Qwen-7B APIExisting syste- matic reviews Accuracy, Precision, Recall, F1, 휅 , PABAK, Efficiency, Cost Mourtzinis et al. (2025) GPT-4.1-mini;Gemini-2.5-Flash, DeepSeek-R1-70B; Llama-3.3-70B N/RAuthor- constructed F1 reduction specifically; and Ji et al. (2024) reports Cohen’s휅as an inter-rater agreement measure between the LLMs and experts, alongside standard accuracy and F1. Dai et al. (2025) add AUC and SROC curve analysis on top of pooled sensitivity/specificity. Finally, Mourtzinis et al. (2025) reports only F1. Columns 4-5 of Table 4 summarise the previous discussion. 5 Discussion Across the reviewed papers, the survey reveals a rapidly expanding body of work applying generative LLMs to both literature retrieval and screening, with the two areas showing distinct methodological profiles. Retrieval is technically diverse, spanning semantic search, agentic query planning, knowledge-graph grounding, citation auditing, and RAG pipelines that support downstream tasks (e.g., automated review, clinical decision support, survey generation). In many systems, retrieval serves as a supporting component rather than the paper’s main contribution, so authors are often more likely to construct or identify ground truth for the downstream task than for retrieval itself. As a result, focused evaluation of the LLM-based retrieval component remains uncommon. In screening, by contrast, the methodological range is narrower, centered mainly on single-pass prompting and more engineered multi-component pipelines. Evaluation is also markedly more rigorous and consistent than in retrieval: studies are almost always benchmarked against real systematic-review decisions or carefully constructed expert-labeled ground truth, typically using an established diagnostic-accuracy toolkit. Screening therefore emerges as the more methodologically standardized of the two areas. The survey also identifies several cross-cutting trends. First, both tasks are shifting from single LLM calls to composed, often agentic, multi-component systems. In retrieval, this appears in multi-agent pipelines for query planning, search, validation, and routing. In screening, this more often takes the form of multi-model consensus or arbitration, or escalation of uncertain cases to humans, designs that mirror how a human review team Scientific Knowledge Discovery in the Age of Large Language Models17 operates. This design serves not only to improve screening accuracy but also to calibrate when human involvement remains necessary. Second, the corpus suggests that open-weight models are gaining ground alongside still-dominant closed, API-based frontier models. OpenAI’s GPT-3.5/4/4o family remains the most common choice, with Claude and Gemini also appearing regularly. But 2025 papers increasingly use open-weight models, either alone, as in local Mistral-NeMo-12B deployment for precision-oncology curation, or alongside closed models in mixed pipelines. Where authors give reasons, the main drivers are privacy-preserving local deployment and cost control in workflows requiring many repeated or parallel calls. Direct open-versus-closed comparisons, however, remain too limited for firm conclusions. Third, access and adaptation are unevenly reported in both areas. API access to commercial frontier models dominates, but retrieval papers often leave the access mode unstated, whereas it is confirmed for nearly every screening paper. At the same time, domain-specific fine-tuning is rare. Fourth, the corpus is heavily concentrated in biomedical and clinical applications. The most common topics are biomedical (see Section 2.2), and most presented engineered screening pipelines, along with several RAG-based retrieval systems, target medical evidence synthesis or clinical decision support. This likely reflects both the maturity of systematic-review methodology in biomedicine, which provides established ground truth and review conventions, and the field’s strong need for scalable evidence synthesis. As a result, the patterns identified here are best understood as characteristic primarily of LLM-assisted biomedical literature work and may not transfer directly to domains without similar benchmarks or review practices. 6 Conclusion This chapter has surveyed34peer-reviewed papers published between 2024 and 2026 that apply generative LLMs to two core scientific knowledge-discovery tasks, literature retrieval and literature screening, motivated by the growing strain that expanding publication volumes place on traditional, largely manual search and review workflows. The picture that emerges is of an active and rapidly growing but highly concentrated body of work. The great majority of the reviewed systems target biomedical and clinical applications, particularly systematic reviews and other evidence-synthesis workflows. This is unsurprising since biomedicine offers mature review methodology, ready-made ground truth, and clear evaluation conventions, while also facing a particularly acute need to keep pace with a fast-growing evidence base. Outside this setting, comparable benchmarks and conventions are largely absent, so the transferability of these approaches to other domains remains unclear. Taken together, the surveyed work suggests that LLM-assisted literature retrieval and screening is moving from isolated demonstrations toward a more established, if still domain-bound, area of practice. 18Eleni Adamidi, Serafeim Chatzopoulos and Thanasis Vergoulis References de Ara ́ujo Pessoa LF, Vogl R (2025) Applications of llm and nlp in the retrieval and analysis of institutional publications. In: Proc. of EUNIS 2025, p 60–49 Balaskas G, Papadopoulos H, Korakis A (2024) Leveraging large language models for simplified patient summary generation, literature retrieval and medical information summarization: A health cascade study. In: Proc. of the Annual Hawaii International Conference on System Sciences, p 3255–3264 Bao T, Nayeem MT, Rafiei D, Zhang C (2025) Surveygen: Quality-aware scientific survey generation with large language models. In: Proc. of the 2025 Conference on Empirical Methods in Natural Language Processing, p 2712–2736 Bascur JP, Verberne S, van Eck NJ, Waltman L (2023) Academic information retrieval using citation clusters: in-depth evaluation based on systematic reviews. Scientometrics 128(5):2895–2921 Beel J, Gipp B, Langer S, Breitinger C (2016) Research-paper recommender systems: A literature survey. International Journal on Digital Libraries 17:305–338 Brown TB, et al. (2020) Language models are few-shot learners. In: Advances in Neural Information Processing Systems, vol 33, p 1877–1901 Chelli M, Descamps J, Lavou ́ e V, Trojani C, Azar M, Deckert M, Raynier JL, Clowez G, Boileau P, Ruetsch-Chelli C (2024) Hallucination rates and reference accuracy of chatgpt and bard for systematic reviews: Comparative analysis. Journal of Medical Internet Research 26:e53164 Dai ZY, Wang FQ, Shen C, Ji YL, Li ZY, Wang Y, Pu Q (2025) Accuracy of large language models for literature screening in thoracic surgery: Diagnostic study. Journal of Medical Internet Research p e67488 Delgado-Chaves FM, Jennings MJ, Atalaia A, Wolff J, Horvath R, Mamdouh ZM, Baumbach J, Baumbach L (2025) Transforming literature screening: The emerging role of large language models in systematic reviews. Proc of the National Academy of Sciences 122:e2411962122 Ding Y, Wu Y, Ding Z (2025) An automatic patent literature retrieval system based on llm-rag. Journal of Technology Innovation and Engineering 1:2025 Fu H, Liang R, Lin S, Li X (2025) Study on the generation of tumor individualized treatment schemes based on large language models and literature retrieval. Frontiers in Computing and Intelligent Systems 13:93–97 Gorenshtein A, Shihada K, Sorka M, Aran D, Shelly S (2025) Literas: Biomedical literature review and citation retrieval agents. Computers in Biology and Medicine 192:110363 Gusenbauer M, Haddaway NR (2020) Which academic search systems are suitable for systematic reviews or meta-analyses? evaluating retrieval qualities of google scholar, pubmed, and 26 other resources. Research synthesis methods 11(2):181–217 Hanson MA, G ́ omez Barreiro P, Crosetto P, Brockington D (2024) The strain on scientific publishing. Quantitative Science Studies 5(4):823–843 Hu B, Tomini E, Corrin T, Pussegoda K, Sandner E, Henriques A, Simniceanu A, Fontana L, Wagner A, Brazeau S, Waddell L (2025) Enhancing evidence synthesis Scientific Knowledge Discovery in the Age of Large Language Models19 efficiency: Leveraging large language models and agentic workflows for optimized literature screening. Cochrane Evidence Synthesis and Methods 3:e70042 Huang M, Zhou S, Chen Y, Li K (2025) Conversational exploration of literature landscape with litchat. In: Proc. of the 34th International Joint Conference on Artificial Intelligence, vol 12, p 11058–11061 Iyer KG, Yunus M, O’Neill C, Ye C, Hyk A, McCormick K, Ciuc ̆ a I, Wu JF, Accomazzi A, Astarita S, Chakrabarty R, Cranney J, Field A, Ghosal T, Ginolfi M, Huertas-Company M, Jab lo ́ nska M, Kruk S, Liu H, Marchidan G, Mistry R, Naiman JP, Peek JEG, Polimera M, M ́ endez SJR, Schawinski K, Sharma S, Smith MJ, Ting YS, Walmsley M (2024) pathfinder: A semantic framework for literature review and knowledge discovery in astronomy. The Astrophysical Journal Supplement Series 275:38 Ji Z, Guo S, Qiao Y, McDougal RA (2024) Automating literature screening and curation with applications to computational neuroscience. Journal of the American Medical Informatics Association 31:1463–1470 Li D, Wu L, Zhang M, Shpyleva S, Lin YC, Huang HY, Li T, Xu J (2024a) Assessing the performance of large language models in literature screening for pharmacovigilance: a comparative study. Frontiers in Drug Safety and Regulation 4:1379260 Li Y, Zhao J, Li M, Dang Y, Yu E, Li J, Sun Z, Hussein U, Wen J, Abdelhameed AM, Mai J, Li S, Yu Y, Hu X, Yang D, Feng J, Li Z, He J, Tao W, Duan T, Lou Y, Li F, Tao C (2024b) Refai: a gpt-powered retrieval-augmented generative tool for biomedical literature recommendation and summarization. Journal of the American Medical Informatics Association 31:2030–2039 Lin FP, Tran M, Thavaneswaran S, Grady JP, Kansara M, Chan J, Ballinger ML, Simes J, Thomas DM (2025) Automating precision oncology literature curation: A decision tree approach using large language models. Studies in Health Technology and Informatics 329:248–252 Liu C, Wei X, Liu P, Shen Y, Mao Y, Cui T (2025a) Biomedsearch: A multi-source biomedical retrieval framework based on llms. In: 2025 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), p 2516–2521 Liu Q, Wang Y, Mu L, Li J (2025b) Research on the construction and application of problem-method-oriented academic graph empowered by llm. Discover Computing 28:168 Liu Y, Chen Z, Wang YG, Shen Y (2024) Toursynbio-search: A large language model driven agent framework for unified search method for protein engineering. In: 2024 IEEE International Conference on Bioinformatics and Biomedicine (BIBM), p 5395–5400 Manghi P, Atzori C, Bardi A, Baglioni M, Schirrwagen J, Dimitropoulos H, La Bruzzo S, Foufoulas I, Mannocci A, Horst M, Czerniak A, Iatropoulou K, Kokogiannaki A, De Bo- nis M, Artini M, Lempesis A, Ioannidis A, Manola N, Principe P, Vergoulis T, Chat- zopoulos S, Pierrakos D (2026) Openaire graph dump. DOI 10.5281/zenodo.7488618 Marshall IJ, Wallace BC (2019) Toward systematic review automation: A practical guide to using machine learning tools in research synthesis. Systematic Reviews 8(1):163 Mourtzinis S, Silva TS, Lo JCM, Smith DL, Conley SP (2025) A human in the loop approach to applying large language models for farm management insight. Scientific Reports 16:1273 20Eleni Adamidi, Serafeim Chatzopoulos and Thanasis Vergoulis Oami T, Okada Y, aki Nakada T (2025) Optimal large language models to screen citations for systematic reviews. Research Synthesis Methods 16:859–875 Ouyang L, et al. (2022) Training language models to follow instructions with human feedback. In: Advances in Neural Information Processing Systems, vol 35, p 27730– 27744 Pan C, Lu W, Chen B, Zhang G, Yang Z, Hao J (2025) Automated literature screening for hepatocellular carcinoma treatment through integration of 3 large language models: Methodological study. JMIR Medical Informatics 13:e76252 Peasley D, Kuplicki R, Sen S, Paulus M (2025) Leveraging large language models and agent-based systems for scientific data analysis: Validation study. JMIR Mental Health 12:e68135–e68135 Pires C, Correia P, Silva P, Ferreira L (2024) Enhancing llms with knowledge graphs for academic literature retrieval. In: Proc. of the 16th International Joint Conference on Knowledge Discovery, Knowledge Engineering and Knowledge Management, vol 1, p 343–347 Priem J, Piwowar H, Orr R (2022) Openalex: A fully-open index of scholarly works, authors, venues, institutions, and concepts. arXiv preprint arXiv:220501833 Ronquillo JG, Ye J, Gorman D, Lemeshow AR, Watt SJ (2024) Practical aspects of using large language models to screen abstracts for cardiovascular drug development: Cross-sectional study. JMIR Medical Informatics 12:e64143–e64143 Ruan M, Fan J, Liu M, Meng Z, Zhang X, Zhang C (2025) Artificial intelligence for the science of evidence synthesis: how good are ai-powered tools for automatic literature screening? BMC Medical Research Methodology 25:199 Smith DR (2009) A 30-year citation analysis of bibliometric trends at the archives of environmental health, 1975–2004. Archives of environmental & occupational health 64(sup1):43–54 Tao Y, Li X, Yisha Z, Yang S, Zhan S, Sun F (2025) Litautoscreener: Development and validation of an automated literature screening tool in evidence-based medicine driven by large language models. Health Data Science 5 Trad F, Yammine R, Charafeddine J, Chakhtoura M, Rahme M, Fuleihan GEH, Chehab A (2025) Streamlining systematic reviews with large language models using prompt engineering and retrieval augmented generation. BMC Medical Research Methodology 25:130 Vergoulis T, Chatzopoulos S, Kanellos I, Deligiannis P, Tryfonopoulos C, Dalamagas T (2019) BIP! Finder: Facilitating scientific literature search by exploiting impact-based ranking. CIKM Vergoulis T, Kanellos I, Atzori C, Mannocci A, Chatzopoulos S, Bruzzo SL, Manola N, Manghi P (2021) BIP! DB: A dataset of impact measures for scientific publications. In: Companion Proc. of the Web Conf., p 456–460 Xu Z, Zhao Y, Patwardhan M, Vig L, Cohan A (2025) Can llms identify critical limitations within scientific research? a systematic evaluation on ai research papers. In: Proc. of the 63rd Annual Meeting of the Association for Comp. Linguistics, vol 1, p 20652–20706 Zhang D, Zheng S, Zhu Y, Yuan H, Gong J, Tang J (2025a) Mcap: Low-pass gnns with matrix completion for academic recommendations. ACM Transactions on Inf Systems 43:1–29 Scientific Knowledge Discovery in the Age of Large Language Models21 Zhang Y, Yang R, Jiao S, Kang S, Han J (2025b) Scientific paper retrieval with llm- guided semantic-based ranking. In: Findings of the Association for Comp. Linguistics: EMNLP 2025, p 2049–2060 Zheng J, Chen Y, Zhou Z, Peng C, Deng H, Yin S (2025) Information-constrained retrieval for scientific literature via large language model agents. In: 2025 6th International Conference on Big Data, Artificial Intelligence and Internet of Things Engineering (ICBAIE), IEEE, p 366–370 Zheng W, Xu Y, Lin X, Gao C, Wang W, Feng F (2026) Navigating through paper flood: Advancing llm-based paper evaluation through domain-aware retrieval and latent reasoning. Proc of the AAAI Conference on Artificial Intelligence 40:35041–35049 Zhu M, Weng Y, Yang L, Zhang Y (2025) Deepreview: Improving llm-based paper review with human-like deep thinking process. In: Proc. of the 63rd Annual Meeting of the Association for Comp. Linguistics, vol 1, p 29330–29355