Paper deep dive
Most biomedical publications show signs of LLM-assisted writing
Lena Holzwarth, Rita GonzĂĄlez-MĂĄrquez, Dmitry Kobak
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 90%
Last extracted: 8/16/2026, 3:11:41 AM
Summary
This study introduces and validates an unbiased method to estimate the prevalence of Large Language Model (LLM) usage in biomedical publications by analyzing changes in word frequencies of LLM-associated marker words. Applying this method to 1,194,287 open-access papers from PubMed Central (PMC), the authors found that by the end of 2025, 89% of papers exhibited signs of LLM-assisted writing. Usage varied significantly by section, with the Discussion section showing the highest prevalence (68%) and the Methods section the lowest (32%) when normalized for length. The study highlights concerns regarding academic integrity, hallucinations, and linguistic homogenization, urging for updated policy guidelines.
Entities (10)
Relation Signals (6)
PubMed Central â contains â biomedical publications
confidence 99% ¡ We apply our method to the full texts of open-access biomedical papers from Pubmed Central
LLM â usedin â academic writing
confidence 95% ¡ LLM-powered chatbots and agents have become widely used as a tool for academic writing
Discussion â hashigherllmusage â Methods
confidence 90% ¡ LLMs are twice as likely to be used when writing a paragraph in the Discussion section (68%) compared to a paragraph in the Methods section (32%)
South Korea â hashighestllmusage â PubMed Central
confidence 85% ¡ We found the highest β^ in 2025 in South Korea (0.85)
LLM â causes â Hallucination
confidence 80% ¡ there are concerns about hallucinations... that can negatively impact academic publishing
LLM â causes â Fraud
confidence 80% ¡ concerns about misconduct and fraud
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decisions, it is necessary to monitor the prevalence of LLM-altered texts in scholarly publications. Despite some recent progress in this direction, no existing method can produce reliable estimates. Here we suggest and validate a new unbiased approach to estimate LLM usage in a corpus of texts based on changing word frequencies. We apply our method to the full texts of open-access biomedical papers from Pubmed Central, and show that by the end of 2025, 89% of papers show excess of LLM-associated vocabulary. We also find that LLMs are twice as likely to be used when writing a paragraph in the Discussion section (68%) compared to a paragraph in the Methods section (32%), but even inside the Methods section, the overall prevalence of LLM usage is over 50%. We believe that our estimates are crucial to shape future guidelines and policies.
Tags
Links
- Source: https://arxiv.org/abs/2608.10715v1
- Canonical: https://arxiv.org/abs/2608.10715v1
Trouble viewing inline? Open PDF directly â
Full Text
44,800 characters extracted from source content.
Expand or collapse full text
Most biomedical publications show signs of LLM-assisted writing Lena Holzwarth Hertie Institute for AI in Brain Health, University of TĂźbingen, Germany Rita GonzĂĄlez-MĂĄrquez Hertie Institute for AI in Brain Health, University of TĂźbingen, Germany Dmitry Kobak Hertie Institute for AI in Brain Health, University of TĂźbingen, Germany Department of Mathematics, Computer Science, and Statistics, Ghent University, Belgium VIB Center for AI and Computational Biology, VIB, Ghent, Belgium Abstract Over the past several years, LLM-powered chatbots and agents have become widely used as a tool for academic writing. LLM-assisted writing can be valuable by removing language barriers but at the same time causes concerns about misconduct and fraud. To inform policy decisions, it is necessary to monitor the prevalence of LLM-altered texts in scholarly publications. Despite some recent progress in this direction, no existing method can produce reliable estimates. Here we suggest and validate a new unbiased approach to estimate LLM usage in a corpus of texts based on changing word frequencies. We apply our method to the full texts of open-access biomedical papers from Pubmed Central, and show that by the end of 2025, 89% of papers show excess of LLM-associated vocabulary. We also find that LLMs are twice as likely to be used when writing a paragraph in the Discussion section (68%) compared to a paragraph in the Methods section (32%), but even inside the Methods section, the overall prevalence of LLM usage is over 50%. We believe that our estimates are crucial to shape future guidelines and policies. 1 Introduction Figure 1: Overview of the estimation and validation process. (a) Usage frequencies of 379 LLM-associated marker words in PMC abstracts before (2022) and after (2025) widespread LLM adoption. (bâd) Observed (q) and expected (p^human p_human) frequencies for individual words in PMC abstracts. Linear regression is fit to the 60 months in 2018â2022. (e) The same for a set of 200 rarest marker words. (f) Observed and expected frequencies in 2025 PMC abstracts for different sets of marker words, each containing all marker words with frequency below a given threshold. Here p^human p_human is computed based on yearly data with regression based on five years 2018â2022. (g) Lower bounds on the LLM usage derived from (f). Thresholds yielding standard error of β^LB β_LB above 0.025 are shown in gray and discarded. The maximum remaining value of β^LB β_LB is chosen as the LLM usage estimate β β. (h) Validation experiment with ground-truth values of LLM usage β correctly retrieved by β β, and strongly underestimated by the maximal frequency gap maxâĄ(Î) ( ). Since the release of OpenAIâs ChatGPT in November 2022, chatbots and agents based on large language models (LLMs) have become increasingly helpful in any workplace dealing with writing, such as academia. While LLMs can increase inclusivity and accessibility for non-native English speakers, there are concerns about hallucinations (Walters and Wilder, 2023; Topaz et al., 2026; Zhao et al., 2026) and fraud (Kendall and Silva, 2024) that can negatively impact academic publishing. To inform any policy decisions, it is important to estimate the prevalence of actual LLM usage. While LLMs produce text that can be difficult to distinguish from human-written, they introduce effects visible on the corpus level. For example, certain style words appear with higher frequency in LLM-generated texts, affecting word frequencies in a corpus of documents, as observed in publications from, e.g., linguistics (Botes et al., 2025), astronomy (Astarita et al., 2024), medicine (Matsui, 2024), and dentistry (Uribe and Maldupa, 2024). Comparing the pre-LLM and post-LLM frequencies of LLM-associated marker words, either pre-selected (Gray, 2024, 2025; Kousha and Thelwall, 2025), or optimized (Kobak et al., 2025; Siler, 2026) can yield an estimate of the overall LLM usage. However, all existing methods in this group can only provide a lower bound on the LLM usage. Another approach is to measure word frequencies in LLM-generated texts and then employ a mixture model to estimate LLM usage in a corpus (Liang et al., 2024a, 2025; Geng and Trotta, 2024). The inherent limitation here is that the results can strongly depend on the exact prompts and models used to generate LLM texts. In this work, we go beyond prior literature in two respects. First, we suggest and validate a simple method to estimate LLM usage in an assumption-free way based on a large set of marker words, going beyond the lower bounds. Second, we apply this method to analyze full texts of open-access biomedical papers from PubMed Central (PMC), which allows us to compare LLM usage across different sections of the papers, such as Methods, Results, and Discussion. Our estimate of the prevalence of LLM-assisted writing in PMC in the end of 2025 is 89%, which is higher than previously reported estimates. 2 Results We based our analysis on PubMed Central (PMC), a digital repository of open-access scholarly articles published in biomedical and life-science journals. We accessed PMC in early 2026 and restricted our analysis to 1,194,287 papers written in English, published in 2017â2025, and containing sections called Introduction, Methods, Results, and Discussion (or equivalent names; see Methods). In our prior work (Kobak et al., 2025), we identified 379 non-content words that showed markedly increased usage in PubMed abstracts in 2024 compared to earlier years, and attributed this excess to LLM-assisted writing or editing. These LLM markers are words that can be used in any research context, such as these, potential, or delves. We confirmed that in our PMC dataset almost all of these words saw increased usage in abstracts in 2025 (Figure 1a), so we based our subsequent analysis on this list. For any single of the marker words, we obtained its usage frequency qâ(t)q(t), defined as the fraction of documents containing at least one occurrence of the marker word, as a function of time (Figure 1bâd). We fit a linear regression to the five-year period (2018â2022) before the widespread adoption of LLMs (ChatGPT was released in November 2022), and extrapolated the regression line to 2023â2025 to obtain p^humanâ(t) p_human(t), the counterfactual usage frequency of this word in human-written texts in the absence of LLMs. The frequency gap Î=qâp^human =q- p_human gives a lower bound on LLM usage (Kobak et al., 2025): for example, for the word these in abstracts in December 2025, we have q=0.50q=0.50 and p^human=0.33 p_human=0.33 (Figure 1b), meaning that the excess 17% usage must be due to LLM editing. A tighter lower bound is also possible. If pLLMp_LLM is the (unobservable) usage frequency in LLM-edited or LLM-written texts, and β is the proportion of papers written with some help of an LLM, then q=(1âβ)âphuman+βâpLLM.q=(1-β)p_human+β p_LLM. From here we obtain β=qâphumanpLLMâphumanâĽqâphuman1âphumanâĽqâphuman.β= q-p_humanp_LLM-p_human⼠q-p_human1-p_human⼠q-p_human. To calculate β, we would need access to phumanp_human and pLLMp_LLM. For the former, we can plug in our estimate p^human p_human, but the latter cannot be inferred directly from the data. By replacing it with its upper bound pLLMâ¤1p_LLM⤠1, we arrive at a lower bound that is stronger than Î : βâĽÎ˛^LB=qâp^human1âp^humanâĽÎ.β⼠β_LB= q- p_human1- p_human⼠. For the same word these in PMC abstracts in December 2025, we get β^LB=0.17/0.67=0.25 β_LB=0.17/0.67=0.25, implying at least 25% LLM usage. 2023 2024 2025 Dec 2025 Entire sections Abstract .09 (2) .31 (2) .53 (1) .68 (2) Introduction .11 (2) .34 (2) .50 (2) .63 (2) Methods .11 (2) .29 (2) .46 (2) .54 (2) Results .11 (3) .30 (3) .47 (2) .58 (2) Discussion .14 (2) .43 (2) .68 (1) .78 (2) Full paper .19 (5) .52 (4) .77 (2) .89 (3) Random 255-word crops Abstract .10 (2) .32 (2) .54 (2) .67 (2) Introduction .10 (2) .30 (1) .48 (1) .59 (2) Methods .03 (2) .14 (2) .28 (2) .32 (4) Results .08 (1) .20 (1) .35 (1) .46 (2) Discussion .10 (2) .35 (2) .57 (2) .68 (2) Full paper .07 (2) .21 (2) .36 (1) .40 (3) Table 1: Estimated frequencies of LLM usage for each manuscript section. Each row used its own threshold T, which was optimized for 2025 (and then used for all years). Top part: entire sections. Bottom part: random 255-word crops. The row with the highest values in each part is highlighted in bold. âFull paperâ refers to Introduction, Methods, Results, and Discussion concatenated together. Standard errors for the last digit are shown in parentheses. As each of the marker words is associated with LLM usage, using a set of marker words G to compute the fraction of documents qâ(t)q(t) containing at least one occurrence of at least one of the words from G can give higher bounds on β. However, making G too large can lead to p^humanâ1 p_humanâ 1, leading to unstable estimates with large standard errors (see Methods), so some selection has to be made. Following Kobak et al. (2025), we sorted all marker words by their frequency in 2024, and selected all words with q<Tq<T into a set Gâ(T)G(T), parameterized by the threshold T. As an example, selecting T to yield 200 words in PMC abstracts (Figure 1e), resulted in β^LB=0.11 β_LB=0.11 in December 2025. To find an optimal value of T, we used a grid of T values. For each T value, we found q and p^human p_human for 2025 using regression on yearly 2018â2022 frequencies (Figure 1f), and selected T with the highest β^LB β_LB (Figure 1g), after discarding all T values yielding standard error above 0.0250.025 (see Methods). As this maximum was typically achieved at q close to 1 (and hence pLLMp_LLM close to 1, which is the approximation used in computing the lower bound), we assume that our lower bound is acceptably tight and use it as the estimate of LLM usage frequency β^=maxTâĄ(β^LB). β= _T( β_LB). We used a simulation to validate our estimation procedure and could retrieve the ground-truth β with high accuracy (see Methods for details). The simulation assumed 100,000 documents per year, which was close to the amount of papers in our PMC sample, and realistic ranges for phumanp_human and pLLMp_LLM (see Methods). We obtained |β^âβ|<0.02| β-β|<0.02 for all simulated values βâ[0,1]βâ[0,1] (Figure 1h), confirming the reliability of our estimation procedure. We applied this procedure to estimate the frequency of LLM usage with yearly (Table 1) and monthly (Figure 2a) resolution for each of the PMC manuscript sections. Threshold T was optimized for each section separately using the 2025 data, and then used for other years and months as well. All sections showed continually increasing estimated LLM usage starting in early 2023. For the PMC abstracts, we obtained β^=0.53 β=0.53 in all 2025 and 0.68 in December 2025. For the full paper (Introduction, Methods, Results, and Discussion concatenated together), we obtained β^=0.77 β=0.77 in all 2025 and 0.89 in December 2025. The interpretation is that by the end of 2025, 9 out of 10 papers exhibited signs of LLM-assisted writing or editing. As expected, the estimates for the individual sections were all lower than the estimate for the full paper text, as our approach detects LLM usage in any part of the paper. Similarly, comparing different sections with each other is confounded by their average length, because a long section presents more opportunity for authors to employ LLM editing and for an LLM to use its marker words. Therefore, we repeated our estimation procedure using random 255-word crops from each section (corresponding to the median abstract length). Here, the Discussion had the highest β β values in each year, reaching 0.68 in December 2025, closely followed by the abstracts reaching 0.67 (Table 1). We obtained the lowest LLM usage in the Methods section, reaching only 0.32 in December 2025. The Methods also had the biggest gap between the estimate based on the entire and on the cropped sections (0.54 vs. 0.32), indicating that LLM usage was not uniform over the length of this section. Figure 2: Estimated LLM usage for individual sections. (a) Estimated LLM usage over time using entire sections. All estimates were computed with the section-specific optimal threshold for the year 2025. The standard errors are shown as shaded areas. (b) The same for random 255-word crops from each section. Finally, we sorted all PMC papers in our sample by country of affiliation of the first author and estimated LLM usage for each of the top 20 countries in terms of the number of papers in our dataset. We found the highest β β in 2025 in South Korea (0.85), followed by China (0.82) and Taiwan (0.80), while the lowest value was observed in the UK (0.28), in agreement with prior studies reporting large differences in estimated LLM usage between affiliation countries (Liao et al., 2024; Lin et al., 2025; Prakash et al., 2025; Kobak et al., 2025; Siler, 2026). Note that in this analysis we used a fixed set of marker words and did not optimize it for individual countries, so some of our estimates here can be an underestimation. The estimated LLM usage was negatively correlated with the frequency of LLM marker words before LLMs, i.e. qâ(2022)q(2022) (Figure 3a), which can be taken as a rough proxy for the mastery of English language, as many marker words are uncommon vocabulary (e.g. exacerbating or invaluable). When grouping together all papers from countries with predominantly native English speakers (Australia, Canada, Ireland, New Zealand, South Africa, UK, USA) and the rest of the world, we obtained β β values in 2025 of 0.37 and 0.72 respectively (Figure 3b). The pre-2023 frequency of LLM marker words was much lower in non-natively English-speaking countries, but it became very close to the frequency in natively English-speaking countries by 2025 (Figure 3c), indicating strong linguistic convergence driven by the LLM usage (Lin et al., 2025; Prakash et al., 2025). Figure 3: Estimated LLM usage for individual countries (full papers). (a) Estimated LLM usage in 2025 for some of the top 20 countries by paper count, plotted against the pre-LLM frequency of the set of selected marker words. Estimates with a standard error âĽ0.07⼠0.07 are not shown here but can be found in Table S3. The same set of 251 marker words was used for all countries (see Methods), so some of the values can be an underestimation. Countries are colored based on whether they have a majority of native English speakers. (bâc) Estimated LLM usage and marker word frequencies for all countries in PMC with a majority of native English speakers, and all other countries. Dashed lines show the expected frequencies. 3 Discussion Comparison to existing estimates We found that by the end of December 2025, almost 90% of PMC papers showed signs of LLM-assisted writing or editing. Our estimates for 2023â2025 are higher than most existing estimates reported in the literature (Table S1), but are consistent with the usage ranges in existing surveys (Table S2). Studies relying on word frequency gaps (Kobak et al., 2025; Gray, 2025; Kousha and Thelwall, 2025) reported estimates for 2024 ranging from 12% to 16% based on abstracts and full texts of papers from PubMed, PMC, and Dimensions databases, which is much smaller than our 2024 estimate of 31% (abstracts) and 52% (full papers). As we showed, this is because frequency gaps systematically underestimate the true fraction of LLM-assisted articles. A recent study based on the distribution of marker-word usage rates per 1000 words (Siler, 2026) found 57% in 2025. This is also an underestimation, because this method is based on the overlap of the usage rate distributions pre- and post-2022 and implicitly assumes non-overlapping rates for human and LLM writing, which in reality is not the case. Studies based on measuring pLLMp_LLM via prompting popular LLMs (Geng and Trotta, 2024; Liang et al., 2024b) produced very different results (35% and 18% respectively) on the same data (computer-science arXiv abstracts in early 2024), illustrating strong dependence on chosen prompts and models (Dou et al., 2024). Existing studies based on LLM detectors (Akram, 2024; Liu and Bu, 2024; Picazo-Sanchez and Ortiz-Martin, 2024) estimated LLM usage in 2023 when it was much lower. Furthermore, various LLM detectors have been repeatedly found to be inaccurate (Weber-Wulff et al., 2023; Lazebnik and Rosenfeld, 2024). Several surveys of scientistsâ use of LLMs for academic writing have been conducted by publishing houses (Elsevier, 2024; Kwon, 2025; Nordling, 2023; Noorden and Perkel, 2023; Wiley, 2025, 2026) and independent researchers (Eppler et al., 2024; Liao et al., 2024; Mishra et al., 2024; Mohammadi et al., 2026). Many of them reported much higher usage (Table S2) than the estimation efforts reviewed above (Table S1). For example, Wiley (2026) found over 70% of survey respondents in mid-2025 using LLM for writing assistance. This is in good agreement with our estimates, especially taking into the account a possible underreporting bias in surveys. People are likely using LLMs more often than they are admitting when surveyed. Also, surveys typically ask about specific use cases, so actual LLM usage can be higher. Limitations Our estimate hinges on the central assumption that p^human p_human is a faithful estimate of phumanp_human in 2025. While linear extrapolation remains reasonable for some time after the release of ChatGPT, it becomes more tenuous as we move away from this date. As people become more and more exposed to the writing style of LLMs, they may start consciously or subconsciously copying it, and indeed there is some evidence for that in spoken communication (Yakura et al., 2024). On the other hand, there is also an opposite effect of people consciously avoiding LLM-associated marker words (Geng and Trotta, 2025) (Figure 1d), and it is unclear which effect is stronger. We believe that any effects arising from changes to human vocabulary are smaller and slower than the effects arising from direct LLM-assisted editing. Reassuringly, our estimates are in broad agreement with existing surveys of LLM usage (see above). Note that the same limitation applies to all other approaches estimating LLM usage, including those based on frequency gaps (Kobak et al., 2025), mixture models (Liang et al., 2025), and distribution overlaps (Siler, 2026). Policy implications LLMs can be a powerful tool for increasing equity by overcoming language barriers (Berdejo-Espinola and Amano, 2023), and indeed surveyed researchers often report using LLMs for translating manuscript drafts (Kwon, 2025; Mohammadi et al., 2026; Wiley, 2025). It has also been argued that LLMs can make academic writing more efficient, allowing scientists to devote more time to research (Filimonovic et al., 2025). There is evidence that the widespread use of LLMs led to an increasing number of submissions and faster reviewing (Qi et al., 2025). On the other hand, LLM-assisted writing also carries a number of dangers. First, LLMs can hallucinate and misattribute their statements. An evidence of increased rate of hallucinated citations (Walters and Wilder, 2023; Topaz et al., 2026; Zhao et al., 2026) implies that scientists may sometimes fail to critically review LLM outputs, and suggests that LLMs can also be used for outright scientific fraud (Kendall and Silva, 2024). Second, the increase in writing efficiency can reduce the overall quality of published research by putting even stronger strains on the review process (Gray, 2026) and amplifying publish-or-perish biases (Kendall and Silva, 2024). Third, the global dependency on LLMs can have a homogenizing effect not only on language but also on reasoning, threatening diversity as a source of innovation (Sourati et al., 2026; Zhou et al., 2026). It does not help that many academic authors neglect already existing guidelines for ethical LLM usage (Kim, 2024) and often do not disclose it (Kousha, 2024; Kwon, 2025). All these developments could reduce societal trust in the process of science as a whole (Gray, 2026). Beyond that, LLM-assisted writing can carry individual dangers as well. The process of writing is closely connected to the process of thinking, and delegating writing to an LLM and only passively proofreading their outputs can lead to less reflection and sloppier arguments staying unnoticed (Messeri and Crockett, 2024; Tang, 2025). Relying on LLMs for academic writing can also cause a loss of individual creativity (Zhou et al., 2026). In summary, while LLM-assisted writing does have its advantages such as closing the language barrier, widespread reliance on LLMs can pose serious challenges to academia. These need to be addressed, both on an individual level by critical reflection, and on an institutional level by developing guidelines and increased safeguards against fraud and misconduct. 4 Methods PubMed Central Pubmed Central (PMC) is a collection of open-access biomedical literature curated by the U.S. National Institutes of Healthâs (NIH) National Library of Medicine. In this work we used the PMC Open Access Subset (https://pmc.ncbi.nlm.nih.gov/tools/openftlist/), downloading articles with licences allowing commercial as well as non-commercial use. We used the semiannual baseline published on January 23, 2026 containing papers dated up to December 31, 2025. In our subset of 1,194,287 papers (see below), the top three journals by the number of papers were Scientific Reports (8.1%), Nature Communications (4.7%) and PLoS ONE (3.9%), all of which are multidisciplinary journals covering a wide range of topics in natural sciences. In our subset, 1.8% of papers were preprints from bioRxiv, medRxiv, arXiv, and Research Square, that had funding requirement to be uploaded to PMC. If such a preprint gets published in a journal that is part of PMC, the preprint version is not replaced, so that two versions of the same paper can exist in PMC. Constructing the dataset From the downloaded XML files, each corresponding to one paper, we extracted the article type, publication date, language, affiliation country, PMC ID, Pubmed ID, journal name, title, abstract, and also titles and texts of each section. If the language was not specified, we used Polyglotâs detector module (https://polyglot.readthedocs.io/). All non-English papers were discarded. Some publication dates only specified the month (e.g. âFeb 2025â) or the season (e.g. âFall 2025â), in which case we replaced them with the first day of the month/season. All papers without date information were discarded. For each section, we extracted its title and text, but sometimes some of the sections were untitled. If the first section was untitled or if there was some text in the body tag but outside of a section tag, we assumed it to be the Introduction. All other untitled sections were discarded. To standardize section titles, we renamed âExperimental resultsâ, âFindingsâ, âMain resultsâ, etc., into âResultsâ, and similarly âDataâ, âExperimental methodsâ, etc., into âMethodsâ. If afterwards a paper had multiple sections with the same standardized title, they were concatenated. We only kept papers containing sections titled Introduction, Methods, Results, and Discussion (this was the most common paper structure in the data). Papers with any of these sections or the abstract below 250 characters were discarded. An overview of the section lengths can be found in Table S4. These steps reduced the dataset from 7,125,722 to 1,602,404 papers, 1,194,287 of which were dated 2017 or later. As the affiliation country we used the country of the first affiliation of the first author, which was available for 80.2% of the papers in our sample. Standard errors Our estimator (qâp^human)/(1âp^human)(q- p_human)/(1- p_human) depends on p^human p_human and on q. We determined their standard errors separately and then propagated them to obtain the error for β β. The value q is a fraction of papers containing a given word (or any of the set of words), and so it has binomial variance VarâĄ[q]=qâ(1âq)/nVar[q]=q(1-q)/n, where n is the number of papers in the current month or year. The estimate p^human p_human is obtained by extrapolating the regression line. If X is the design matrix (consisting only of pre-ChatGPT time points on which the regression was fit and a column of ones), XnewX_new is the predictor matrix with post-ChatGPT time points, and y the response vector containing pre-ChatGPT q values, then VarâĄ(p^human)=Ď^2âXnewâ(XTâX)â1âXnewT+Ď^2.Var( p_human)= Ď^2X_new(X^TX)^-1X_new^T+ Ď^2. Here Ď^2=âyâXâb^â2/(kâ2) Ď^2=\|y-X b\|^2/(k-2) is the unbiased estimate of the noise variance, where k is the number of time-points in the regression (60 for monthly data and 5 for yearly data), and b^=(XTâX)â1âXTây b=(X^TX)^-1X^Ty are the estimated regression coefficients. The first term gives the uncertainty of the mean prediction and the second term captures expected deviations from this mean. To compute the variance of β β, we used the delta method assuming independence of the two variance terms: VarâĄ[β^] [ β] =âβ^âqâVarâĄ[q]+âβ^âp^humanâVarâĄ[p^human] = â βâ qVar[q]+ â βâ p_humanVar[ p_human] =qâ(1âq)/n1âp^human+(qâ1)âVarâĄ[p^human](1âp^human)2. = q(1-q)/n1- p_human+ (q-1)Var[ p_human](1- p_human)^2. When optimizing the threshold T, we excluded all estimates with standard error (square root of the variance) above 0.025. We also excluded all thresholds leading to p^humanâĽ0.999 p_human⼠0.999. Note that when the optimized threshold was applied to other years or affiliation countries, resulting standard errors sometimes got higher than 0.025 (Table 1). Simulation experiment To validate our estimation procedure, we performed the same analysis on simulated data with known pLLMp_LLM and phumanp_human and varying β. We simulated data for n=100,000n=100,000 texts, which is close to the typical number of papers per year in our dataset, and for m=500m=500 words, all of which were treated as LLM marker words. For each word, its phumanp_human was sampled from a gamma distribution with shape parameter 22 and scale parameter 0.020.02. Its pLLMp_LLM was set (1+δ)(1+δ) times larger, with δâźUâ(0.5,5)δ U(0.5,5). If needed, δ was repeatedly sampled until the resulting probability was below 1. We then sampled two binary matrices, MpreM_pre and MpostM_post, each with dimensions nĂmnĂ m, corresponding to data before and after the adoption of LLMs. For MpreM_pre, each entry is a Bernoulli draw with corresponding phumanp_human (corresponding to the m-th column). For MpostM_post, (1âβ)ân(1-β)n rows were filled in the same way, and the remaining βânβ n rows used the corresponding pLLMp_LLM. We then carried out our estimation procedure, treating word frequencies in MpreM_pre as p^human p_human, and those in MpostM_post as the observed frequencies q. All words were ordered by frequency in MpreM_pre, word sets were formed for various thresholds T, and we found β^=maxTâĄ(β^LB) β= _T( β_LB) as the maximum over thresholds T, after discarding thresholds yielding p^human>0.999 p_human>0.999 or standard error above 0.025. As we did not need to use regression here, the variance of p^human p_human was computed as p^humanâ(1âp^human)/n p_human(1- p_human)/n. Threshold grids When optimizing the threshold T, we used a grid with 12 values spaced evenly on a log-scale between 0 and 1 and added several additional values in the high-interest region between 0.03 and 0.7, resulting in 19 thresholds. When doing the analysis by country, we did not optimize the threshold T for each country separately, but used a fixed set of marker words G because otherwise the comparison in Figure 3c would not be possible. To generate the fixed set of marker words, we used the full papers in our entire dataset, but employed the threshold two steps lower in our grid compared to the threshold optimized for the full dataset. This decrease was necessary to avoid unreliable estimates with large standard errors. Note that as the result, some of our estimates for individual countries should be treated as lower bounds. Cropping-based analysis To create section segments of comparable length, every sample was cropped to a randomly placed 255-word long segment. If a sample was shorter than that, it was used as is. The length of 255 words corresponds to the median abstract length in a subset of the PMC data that we used for initial exploration and is slightly larger than the median abstract length in the final dataset (248 words; Table S4). 4.1 Data availability All data are openly available at the PubMed Central website (https://pmc.ncbi.nlm.nih.gov/tools/openftlist/). 4.2 Code availability All code is available on Github at https://github.com/kobaklab/llm-usage-in-pmc. Acknowledgments The authors would like to thank Jan Lause for discussions. Author contributions All authors contributed to conceptualization and design, as well as editing the paper. L.H. also performed all statistical analyses and drafted the paper. Competing Interests The authors declare no competing interests. References A. Akram (2024) Quantitative analysis of AI-generated texts in academic research: a study of AI presence in arxiv submissions using AI detection tool. arXiv preprint arXiv:2403.13812. Cited by: §3, Table S1. S. Astarita, S. Kruk, J. Reerink, and P. GĂłmez (2024) Delving into the utilisation of ChatGPT in scientific publications in astronomy. arXiv preprint arXiv:2406.17324. Cited by: §1. H. Bao, M. Sun, and M. Teplitskiy (2025) Where thereâs a will thereâs a way: ChatGPT is used more for science in countries where it is prohibited. Quantitative Science Studies 6, p. 716â731. Cited by: Table S1, Table S1. V. Berdejo-Espinola and T. Amano (2023) AI tools can improve equity in science. Science 379, p. 991. Cited by: §3. E. Botes, J. Dewaele, J. Colling, and Z. Teuber (2025) Initial indications of generative AI writing in linguistics research publications. PsyArXiv. https://doi. org/10.31234/osf. io/4yvbp_v1. Cited by: §1. Z. Dou, Y. Guo, C. Chang, H. H. Nguyen, and I. Echizen (2024) Enhancing robustness of LLM-synthetic text detectors for academic writing: a comprehensive analysis. In International Conference on Advanced Information Networking and Applications, p. 266â277. Cited by: §3. Elsevier (2024) Insights 2024 â attitudes toward AI â Elsevier. External Links: Link Cited by: §3. M. Eppler, C. Ganjavi, L. S. Ramacciotti, P. Piazza, S. Rodler, E. Checcucci, J. G. Rivas, K. F. Kowalewski, I. R. BelenchĂłn, S. Puliatti, M. Taratkin, A. Veccia, L. Baekelandt, J. Y.C. Teoh, B. K. Somani, M. Wroclawski, A. Abreu, F. Porpiglia, I. S. Gill, D. G. Murphy, D. Canes, and G. E. Cacciamani (2024) Awareness and use of ChatGPT and Large Language Models: a prospective cross-sectional global survey in urology. European Urology 85, p. 146â153. Cited by: §3, Table S2. D. Filimonovic, C. Rutzer, and C. Wunsch (2025) Can GenAI improve academic performance? evidence from the social and behavioral sciences. arXiv preprint arXiv:2510.02408. Cited by: §3. M. Geng and R. Trotta (2024) Is ChatGPT transforming academicsâ writing style?. arXiv preprint arXiv:2404.08627. Cited by: §1, §3, Table S1. M. Geng and R. Trotta (2025) Human-LLM coevolution: evidence from academic writing. In Findings of the Association for Computational Linguistics: ACL 2025, p. 12689â12696. Cited by: §3. A. Gray (2024) ChatGPT âcontaminationâ: estimating the prevalence of LLMs in the scholarly literature. arXiv preprint arXiv:2403.16887. Cited by: §1, Table S1. A. Gray (2025) Estimating the prevalence of LLM-assisted text in scholarly writing. arXiv preprint arXiv:2512.01560. Cited by: §1, §3, Table S1. A. Gray (2026) How AI use in scholarly publishing threatens research integrity, lessens trust, and invites misinformation. Taylor & Francis 82, p. 85â88. Cited by: §3. G. Kendall and J. A. T. D. Silva (2024) Risks of abuse of large language models, like ChatGPT, in scientific publishing: authorship, predatory publishing, and paper mills. Learned Publishing 37. Cited by: §1, §3. S. Kim (2024) Research ethics and issues regarding the use of ChatGPT-like artificial intelligence platforms by authors and reviewers: a narrative review. Science Editing 11 (2), p. 96â106. Cited by: §3. D. Kobak, R. GonzĂĄlez-MĂĄrquez, E. HorvĂĄt, and J. Lause (2025) Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances 11 (27), p. eadt3813. Cited by: §1, §2, §2, §2, §2, §3, §3, Table S1. K. Kousha and M. Thelwall (2025) How much are LLMs changing the language of academic papers after ChatGPT? a multi-database and full text analysis. Scientometrics 2026, p. 1â21. Cited by: §1, §3, Table S1. K. Kousha (2024) How is ChatGPT acknowledged in academic publications?. Scientometrics 129, p. 7959â7969. Cited by: §3. D. Kwon (2025) Is it OK for AI to write science papers? nature survey shows researchers are split. Nature 641, p. 574â578. Cited by: §3, §3, §3, Table S2. T. Lazebnik and A. Rosenfeld (2024) Detecting LLM-assisted writing in scientific communication: are we there yet?. JDIS Journal of Data and Information Science. Cited by: §3. W. Liang, Z. Izzo, Y. Zhang, H. Lepp, H. Cao, X. Zhao, L. Chen, H. Ye, S. Liu, Z. Huang, et al. (2024a) Monitoring AI-modified content at scale: a case study on the impact of ChatGPT on AI conference peer reviews. In Proceedings of the 41st International Conference on Machine Learning, p. 29575â29620. Cited by: §1. W. Liang, Y. Zhang, Z. Wu, H. Lepp, W. Ji, X. Zhao, H. Cao, S. Liu, S. He, Z. Huang, et al. (2024b) Mapping the increasing use of LLMs in scientific papers. arXiv preprint arXiv:2404.01268. Cited by: §3, Table S1. W. Liang, Y. Zhang, Z. Wu, H. Lepp, W. Ji, X. Zhao, H. Cao, S. Liu, S. He, Z. Huang, D. Yang, C. Potts, C. D. Manning, and J. Zou (2025) Quantifying large language model usage in scientific papers. Nature Human Behaviour 9, p. 2599â2609. Cited by: §1, §3, Table S1. Z. Liao, M. Antoniak, I. Cheong, E. Y. Cheng, A. Lee, K. Lo, J. C. Chang, and A. X. Zhang (2024) Llms as research tools: a large scale survey of researchersâ usage and perceptions. arXiv preprint arXiv:2411.05025. Cited by: §2, §3, Table S2. D. Lin, N. Zhao, D. Tian, and J. Li (2025) ChatGPT as linguistic equalizer? quantifying LLM-driven lexical shifts in academic writing. arXiv preprint arXiv:2504.12317. Cited by: §2. J. Liu and Y. Bu (2024) Towards the relationship between AIGC in manuscript writing and author profiles: evidence from preprints in LLMs. arXiv preprint arXiv:2404.15799. Cited by: §3, Table S1, Table S1. K. Matsui (2024) Delving into PubMed records: some terms in medical writing have drastically changed after the arrival of ChatGPT. MedRxiv, p. 2024â05. Cited by: §1. L. Messeri and M. J. Crockett (2024) Artificial intelligence and illusions of understanding in scientific research. Nature 627 (8002), p. 49â58. Cited by: §3. T. Mishra, E. Sutanto, R. Rossanti, N. Pant, A. Ashraf, A. Raut, G. Uwabareze, A. Oluwatomiwa, and B. Zeeshan (2024) Use of large language models as artificial intelligence tools in academic research and publishing among global clinical researchers. Scientific Reports 14, p. 31672. Cited by: §3, Table S2. E. Mohammadi, M. Thelwall, Y. Cai, T. Collier, I. Tahamtan, and A. Eftekhar (2026) Is generative AI reshaping academic practices worldwide? a survey of adoption, benefits, and concerns. Information Processing & Management 63, p. 104350. Cited by: §3, §3, Table S2. J. Y. Ng, S. G. Maduranayagam, N. Suthakar, A. Li, C. Lokker, A. Iorio, R. B. Haynes, and D. Moher (2025) Attitudes and perceptions of medical researchers towards the use of artificial intelligence chatbots in the scientific process: an international cross-sectional survey. The Lancet Digital Health 7, p. e94âe102. Cited by: Table S2. R. V. Noorden and J. M. Perkel (2023) AI and science: what 1,600 researchers think. Nature 621, p. 672â675. External Links: Document, ISSN 14764687 Cited by: §3, Table S2. L. Nordling (2023) How ChatGPT is transforming the postdoc experience. Nature 622, p. 655â657. Cited by: §3, Table S2. P. Picazo-Sanchez and L. Ortiz-Martin (2024) Analysing the impact of ChatGPT in research. Applied Intelligence 54, p. 4172â4188. Cited by: §3, Table S1. A. Prakash, S. Aggarwal, J. J. Varghese, and J. J. Varghese (2025) Writing without borders: AI and cross-cultural convergence in academic writing quality. Humanities and Social Sciences Communications 2025 12:1 12, p. 1058â. Cited by: §2. M. Qi, Z. Cao, Q. Wang, N. Li, and T. Zhu (2025) Does genai rewrite how we write? an empirical study on two-million preprints. arXiv preprint arXiv:2510.17882. Cited by: §3. K. Siler (2026) The diffusion of large language models in published academic articles. Proceedings of the National Academy of Sciences 123 (22), p. e2605754123. Cited by: §1, §2, §3, §3, Table S1. Z. Sourati, A. S. Ziabari, and M. Dehghani (2026) The homogenizing effect of large language models on human expression and thought. Trends in Cognitive Sciences. Cited by: §3. B. L. Tang (2025) The epistemic downside of using LLM-based generative AI in academic writing. Publications 2025 13, p. 63. Cited by: §3. M. Topaz, N. Roguin, P. Gupta, Z. Zhang, and L. Peltonen (2026) Fabricated citations: an audit across 2¡ 5 million biomedical papers. The Lancet 407 (10541), p. 1779â1781. Cited by: §1, §3. S. E. Uribe and I. Maldupa (2024) Estimating the use of ChatGPT in dental research publications. Journal of Dentistry 149, p. 105275. Cited by: §1. W. H. Walters and E. I. Wilder (2023) Fabrication and errors in the bibliographic citations generated by ChatGPT. Scientific Reports 13, p. 14045â. Cited by: §1, §3. D. Weber-Wulff, A. Anohina-Naumeca, S. Bjelobaba, T. FoltĂ˝nek, J. Guerrero-Dib, O. Popoola, P. Ĺ igut, and L. Waddington (2023) Testing of detection tools for AI-generated text. International Journal for Educational Integrity 19. Cited by: §3. Wiley (2025) ExplanAItions: an AI study by wiley. External Links: Link Cited by: §3, §3, Table S2. Wiley (2026) ExplanAItions 2025-2026: the evolution of AI in research.. External Links: Link Cited by: §3, Table S2. H. Yakura, E. Lopez-Lopez, L. Brinkmann, I. Serna, P. Gupta, I. Soraperra, and I. Rahwan (2024) Empirical evidence of large language modelâs influence on human spoken communication. arXiv preprint arXiv:2409.01754. Cited by: §3. Z. Zhao, Y. Wang, T. Stuart, M. De Vaan, P. Ginsparg, and Y. Yin (2026) LLM hallucinations in the wild: large-scale evidence from non-existent citations. arXiv preprint arXiv:2605.07723. Cited by: §1, §3. Y. Zhou, Q. Liu, J. Huang, and G. Li (2026) Creative scar without generative AI: individual creativity fails to sustain while homogeneity keeps climbing. Technology in Society 84, p. 103087. Cited by: §3, §3. Supplementary Tables Period Section Dataset Estimate Method Ref. 12/2022â02/2023 abstract Various journals 0.10 Various LLM detectors 2024 08/2023 abstract arXiv, bioRxiv 0.13 LLM detector (custom) 2025 11/2023 full paper arXiv (CS) 0.07 LLM detector (Originality.AI) 2024 2023 full paper Dimensions 0.02 frequency gap 2024 2023 abstract arXiv NLP papers 0.07 LLM detector (GPTKIT) 2024 0.04 LLM detector (Smodin) 01/2024 abstract arXiv (CS) 0.35 mixture model 2024 02/2024 abstract arXiv (CS) 0.18 mixture model 2024b introduction arxiv (CS) 0.14 abstract Nature portfolio 0.06 introduction Nature portfolio 0.04 09/2024 abstract arXiv (CS) 0.23 mixture model 2025 introduction arxiv (CS) 0.20 abstract Nature portfolio 0.09 introduction Nature portfolio 0.09 2024 abstract PubMed 0.14 frequency gap 2025 2024 full paper Dimensions 0.12 frequency gap 2025 2024 full paper Dim., OpenAlex, PMC 0.16 frequency gap 2025 2023 full paper Elsevier, MDPI, Frontiers, PLoS 0.12 distribution overlap 2026 2024 0.36 2025 0.57 Table S1: Literature overview: Estimates of LLM usage in academia. Dimensions and OpenAlex are databases containing academic papers from diverse disciplines. All studies on LLM detectors used existing detectors, except for Bao et al. (2025), who trained their own. We excluded one of the detectors used by Liu and Bu (2024) (Sapling) because it returned high LLM usage for pre-2022 papers. Period Sample size Sample LLMs used for Usage Ref. 04â05/2023 456 urologists Summarizing text 27% 2024 06â07/2023 3,838 postdocs Refining text 20% 2023 07â08/2023 2,165 PubMed authors Writing or editing manuscripts 22% 2025 ?â09/2023 1,659 recent authors To help write research manuscripts 21% 2023 11/2023â04/2024 816 researchers (Semantic Scholar) Fix grammar or rephrase 50% 2024 Rewrite for another style 24% Shorten or summarize 23% Draft paragraphs from ideas 21% Direct writing (overall) 43% 01â03/2024 2,021 researchers (Scopus) Synthesizing drafts 16% 2026 Proofreading drafts 17% Translating text 18% 03â04/2024 1,043 researchers Help with translation 40% 2025 Proofreading and editing 38% 03â05/2024 5,229 researchers To edit your paper 28% 2025 To translate your paper 8% To summarize other papers 8% To write the first draft 8% 04â06/2024 226 medical researchers Writing 8% 2024 Revision and editing 8% 07â08/2025 341â557 academia, healthcare, etc. Writing assistance (copyediting, translation, etc.) 71% 2026 Table S2: Literature overview: Surveys on LLM usage in academia. We only included surveys that explicitly mention LLM use-cases that are likely to have an effect on the published manuscript text, such as translation, editing, etc. We excluded questions about LLM usage in general, or about use-cases with limited influence on the text, such as grammar checks. The use-cases are copied verbatim from the surveys, and sometimes abbreviated. Country 2022 frequency 2025 frequency LLM usage Sample size Entire PMC corpus 0.83 0.95 0.68 (2) 1,194,287 Majority native English 0.92 0.96 0.37 (3) 333,161 Other countries 0.79 0.95 0.72 (2) 719,959 China 0.75 0.97 0.82 (2) 246,000 USA 0.93 0.97 0.39 (5) 208,178 UK 0.91 0.95 0.28 (7) 64,347 Japan 0.68 0.86 0.55 (4) 59,863 Germany 0.90 0.96 0.53 (4) 54,827 Italy 0.85 0.95 0.61 (5) 28,716 Canada 0.92 0.96 0.46 (14) 28,649 Netherlands 0.86 0.95 0.49 (5) 27,538 South Korea 0.72 0.96 0.85 (2) 25,537 France 0.88 0.93 0.33 (11) 24,577 Australia 0.91 0.95 0.31 (11) 22,956 Spain 0.86 0.97 0.70 (7) 22,657 Brazil 0.82 0.94 0.63 (4) 19,670 Sweden 0.87 0.93 0.47 (11) 16,015 Switzerland 0.90 0.97 0.67 (7) 15,389 India 0.86 0.96 0.69 (5) 14,173 Taiwan 0.77 0.96 0.80 (4) 11,001 Turkey 0.64 0.90 0.67 (5) 10,287 Iran 0.68 0.94 0.78 (3) 9,399 Poland 0.83 0.94 0.48 (8) 8,210 Table S3: Estimated LLM usage in 2025 for individual countries (full paper). Top 20 countries by the number of papers are sorted by sample size. The results for countries with majority native English speakers include all PMC countries with a majority of native English speakers, not only those from the top 20. The reported frequencies correspond to the fixed set of marker words identified using the entire PMC corpus, but using a lower threshold value compared to what we used in Table 1 (see Methods). This explains the difference between the first row in this table and Table 1. Section 5th percentile Median 95th percentile Abstract 148 248 449 Introduction 246 595 1,314 Methods 433 1,270 3,452 Results 487 1,852 5,354 Discussion 510 1,179 2,284 Full paper 2,527 5,156 10,684 Table S4: Section lengths in words.