Paper deep dive
Impact of large language models on peer review opinions from a fine-grained perspective: Evidence from top conference proceedings in AI
Wenqing Wu, Chengzhi Zhang, Yi Zhao, Tong Bao
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 97%
Last extracted: 4/26/2026, 11:36:23 PM
Summary
This study investigates the impact of Large Language Models (LLMs) on the peer review process in top AI conferences (ICLR and NeurIPS). By analyzing review reports at a fine-grained level (word, sentence, and aspect), the researchers found that LLM-assisted reviews tend to be longer, more fluent, and more standardized, with an increased emphasis on surface-level aspects like 'summary' and 'clarity' at the expense of deeper evaluative dimensions like 'originality' and 'replicability'. The study also utilizes a maximum likelihood estimation method to detect LLM-assisted content and examines the relationship between these linguistic shifts and reviewer recommendations.
Entities (9)
Relation Signals (3)
Originality â isadimensionof â Peer Review
confidence 100% · The primary function of peer review is improving the quality of academic manuscripts, such as clarity, originality and other evaluation aspects.
ICLR â issourceof â Peer Review Data
confidence 100% · this study selects two top-tier conferences in the field of computer science, ICLR and NeurIPS, using OpenReview 2 as the data sources.
Large Language Models â affects â Peer Review
confidence 95% · The results indicate that following the emergence of LLMs, peer review texts have become longer and more fluent...
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:With the rapid advancement of Large Language Models (LLMs), the academic community has faced unprecedented disruptions, particularly in the realm of academic communication. The primary function of peer review is improving the quality of academic manuscripts, such as clarity, originality and other evaluation aspects. Although prior studies suggest that LLMs are beginning to influence peer review, it remains unclear whether they are altering its core evaluative functions. Moreover, the extent to which LLMs affect the linguistic form, evaluative focus, and recommendation-related signals of peer-review reports has yet to be systematically examined. In this study, we examine the changes in peer review reports for academic articles following the emergence of LLMs, emphasizing variations at fine-grained level. Specifically, we investigate linguistic features such as the length and complexity of words and sentences in review comments, while also automatically annotating the evaluation aspects of individual review sentences. We also use a maximum likelihood estimation method, previously established, to identify review reports that potentially have modified or generated by LLMs. Finally, we assess the impact of evaluation aspects mentioned in LLM-assisted review reports on the informativeness of recommendation for paper decision-making. The results indicate that following the emergence of LLMs, peer review texts have become longer and more fluent, with increased emphasis on summaries and surface-level clarity, as well as more standardized linguistic patterns, particularly reviewers with lower confidence score. At the same time, attention to deeper evaluative dimensions, such as originality, replicability, and nuanced critical reasoning, has declined.
Tags
Links
- Source: https://arxiv.org/abs/2604.19578v1
- Canonical: https://arxiv.org/abs/2604.19578v1
Trouble viewing inline? Open PDF directly â
Full Text
91,379 characters extracted from source content.
Expand or collapse full text
Impact of large language models on peer review opinions from a fine-grained perspective: Evidence from top conference proceedings in AI Wenqing Wu 1 , Chengzhi Zhang 1* , Yi Zhao 2 , Tong Bao 1 1 Department of Information Management, Nanjing Universityof Science and Technology, Nanjing, 210094, Jiangsu, China. 2 School of Management, Anhui University, Hefei, 230601, Anhui, China. *Corresponding author(s). E-mail(s): zhangcz@njust.edu.cn; Contributing authors: winchywwq@njust.edu.cn;yizhao93@ahu.edu.cn; tbao@njust.edu.cn; Abstract With the rapid advancement of Large Language Models (LLMs),the academic community has faced unprecedented disruptions, particularly in the realm of academic communication. The primary function of peer review is improving the quality of academic manuscripts, such as clarity, originality and other evaluation aspects. Although prior studies suggest that LLMs are beginning to influence peer review, it remains unclear whether they are altering its core evaluative func- tions. Moreover, the extent to which LLMs affect the linguistic form, evaluative focus, and recommendation-related signals of peer-reviewreports has yet to be systematically examined. In this study, we examine the changes in peer review reports for academic articles following the emergence of LLMs, emphasizing vari- ations at fine-grained level. Specifically, we investigate linguistic features such as the length and complexity of words and sentences in review comments, while also automatically annotating the evaluation aspects of individual review sentences. We also use a maximum likelihood estimation method, previously established, to identify review reports that potentially have modified orgenerated by LLMs. Finally, we assess the impact of evaluation aspects mentioned in LLM-assisted review reports on the informativeness of recommendation for paper decision- making. The results indicate that following the emergence of LLMs, peer review texts have become longer and more fluent, with increased emphasis on sum- maries and surface-level clarity, as well as more standardized linguistic patterns, particularly reviewers with lower confidence score. At the same time, attention to deeper evaluative dimensions, such as originality, replicability, and nuanced 1 arXiv:2604.19578v1 [cs.CL] 21 Apr 2026 critical reasoning, has declined. These phenomena are moreobvious when com- paring LLM-assisted and non-LLM-assisted reviews, and theaspects mentioned in LLM-assisted reports have a modest positive influence on informativeness of the recommendations. Keywords:Peer review, Large language model, Academic communication, LLM-assisted text detection, Fine-grained analysis 1 Introduction Peer review is a critical quality control mechanism in the academic research and publi- cation process [1]. Its primary purpose is to ensure the rigor and credibility of academic research, assist authors in improving their work, and identify potential errors and shortcomings [2,3]. However, in recent years, the peer review mechanism has faced widespread criticism due to the surge in paper submissions and the shortage of domain experts qualified to serve as reviewers [4,5], particularly at top artificial intelligence (AI) conferences [6]. Current peer review processes face several challenges, including bias [7], variability in review quality [7], unclear reviewer motivations [8], and imper- fect review mechanism [9]. As submission volumes continue to rise, these issues are becoming increasingly pronounced. Some researchers have sought to mitigate these problems by enhancing fairness [8], reducing biases among novice reviewers [7], cali- brating noisy peer review ratings [10], and improving mechanisms for matching papers with reviewersâ expertise [11,12]. Other studies [13â19] have explored the use of natu- ral language processing techniques to support or refine the peerreview process. These studies introduce the possibility of leveraging artificial intelligence toassist overbur- dened scientists in the peer review process [20,21]. While these technologies may aid reviewers to some extent, their impact on the peer review processstill requires further study. In recent years, the impressive capabilities demonstrated by largelanguage mod- els (LLMs) [22] have sparked extensive research and discussion within the academic community. At the same time, concerns [23,24] have emerged within the academic community about the potential erosion of peer review by LLMs. Researchers have also begun to study and analyze the application and impact of LLMs in the peer review process. For example, Liang et al. [25] not only evaluated the effectiveness of GPT-4 in generating scientific feedback but also proposed [26] a method to estimate the extent of LLMs usage in peer review texts. They found that some reviews from recent AI conferences may have been modified by LLMs. Latona etal. [27] investi- gated the prevalence and impact of LLM-assisted peer reviews at the ICLR 2024, finding that LLM-assisted reviews significantly influence review scores and submission acceptance rates. While these preliminary studies indicate that LLMs have begun to affect peer review, whether they are altering the core functions of peer review remains underexplored. As LLMs become increasingly integrated into scholarly workflows, it is therefore important to analyse their impacts on peer review frommultiple perspec- tives, including linguistic patterns, evaluation aspects, and the influence of reviewer 2 recommendations. Understanding these effects can help refine peer review practices, ensure fairness and transparency, and provide actionable insights for adapting to the evolving academic landscape. Notably, major conferences such asNeurIPS 1 have not yet established explicit policies on whether reviewers may use LLMs toassist in writ- ing their reports. This policy ambiguity underscores the importanceof examining how LLM assistance may already be shaping the linguistic and evaluative characteristics of peer review texts. This study is guided by the following three research questions: RQ1:How has the emergence of LLMs affected the linguistic complexity and aspect-level content expression of peer review texts? RQ2:In comparison to non-LLM-assisted reviews, which evaluation aspects are more prominently emphasized in LLM-assisted reviews? RQ3:How do the evaluation aspects emphasized in LLM-assisted reviews relate to reviewersâ scoring and confidence levels? In summary, we conduct a fine-grained analysis of peer review reports before and after the emergence of LLMs, aiming to address the three research questions outlined above. To answer Research Question 1 (RQ1), we examine the linguistic characteris- tics of review texts, such as average word and sentence length, lexical sophistication, syntactic complexity, and the distribution of evaluation aspects, to determine whether and how linguistic expression has evolved with the introduction of LLMassistance. To address Research Question 2 (RQ2), we automatically annotatereview sentences with predefined evaluation aspects (e.g., clarity, summary, and soundness) and com- pare their distributions between LLM-assisted and non LLM-assisted reviews, in order to identify which aspects are more prominently emphasized by LLM-assisted review- ers. For Research Question 3 (RQ3), we analyse the relationship between reviewers? recommendations and the identified evaluation aspects, assessinghow LLM assistance influences the connection between reviewersâ focal aspects andtheir recommendations. Finally, to complement these analyses, we apply a maximum likelihood estimation model trained on both expert-authored and AI-generated reference texts to estimate review texts that may have been substantially modified or generated by LLMs. The main contributions of this paper are as follows: First, this paper conducts a fine-grained statistical analysis of the review report texts from two artificial intelligence conferences. We find that theproliferation of LLMs has had the most pronounced impact on ICLR review reports,as evidenced by an observable increase in word and sentence length, along with a significant rise in the use of nominal subjects. Additionally, the average length of reviews from reviewers with a confidence score of 1 has shown an upward trend in recent years. Second, we analyzed the proportion of evaluation aspects in the review reports of LLM-assisted and non-LLM-assisted reviewers. We found thatcompared to non- LLM-assisted reviewers, LLM-assisted reviewers significantly increased the proportion of summaries, while the proportion of evaluations focusing on originality decreased. Finally, we explored the relationship between the evaluation aspectsmentioned in the review reports of LLM-assisted reviewers and the scores they assigned. The results 1 https://neurips.c/Conferences/2024/ReviewerGuidelines. Since 2018, NIPS has been renamed to NeurIPS. Therefore, this paper uniformly uses NeurIPS to refer to the conference. 3 indicate that the impact of various evaluation aspects on scoring is relatively weak, and the correlation with reviewersâ confidence scores is generally low. The code and dataset for this paper can be accessed athttps://github.com /njust-winchy/LLMimpact 2 Related work 2.1 Pre-LLM Approaches to Automated Peer Review Automated Scholarly Paper Review (ASPR) [ 28] refers to the process in which computers or intelligent machines independently evaluate the content of a scholarly paper and generate a review report automatically. Currently, research on ASPR is still in its early stages, with machines only serving as an aid to human reviewers. For ASPR, the availability of large datasets is crucial. As pioneers in this field, Kang et al. [29] introduced PeerRead, a publicly available dataset for research purposes, consisting of 14.7K paper drafts and corresponding peer review reports from top venues, including ACL, NeurIPS, and ICLR. On this basis, Wang et al.[30] presents ReviewRobot, a novel tool designed to assist human reviewers by automatically assigning review scores and generating constructive comments across multiple cate- gories, demonstrating high accuracy and effectiveness in enhancing the peer review process. Li et al. [31] introduced a multi-task shared structure encoding approach for predicting peer-review aspect scores of academic papers, demonstrating improved performance over single-task and naive multi-task methods by effectively leveraging auxiliary task information. Furthermore, Yuan et al. [ 14] proposed ASAP-Review, a dataset annotated with evaluation aspects of review content, and trained a targeted summarization model for generating peer reviews. Later, Yuan and Liu [15] proposed an end-to-end knowledge-guided review generation framework for scientific papers, introducing an oracle pre-training strategy to enhance the modelâsunderstanding and coverage of review aspects. These studies represent early-stage auxiliary methods for automated academic peer review that laid the groundwork for further advances prior to the emergence of LLMs. 2.2 LLMs for Scientific Peer Review: Applications and Impacts The emergence of LLMs has sparked a global wave of research, including investigations into their application in peer review and their impact on the peer reviewprocess [ 32â 34]. Liang et al. [25] first evaluated GPT-4âs utility in generating scientific feedback, revealing that GPT-4âs feedback overlaps significantly with human peer reviews and is perceived as helpful by researchers, suggesting its potential toenhance the peer-review process despite noted limitations. Robertson [35] conducted a pilot study on the use of GPT-4 in the peer review process, demonstrating that LLM-generated reviews can be as useful as human reviews and highlighting the potential applications for address- ing resource constraints in peer review. Liu and Shah [ 36] demonstrated promising results in using LLMs to identify errors and verify checklist tasks in peer reviews, 4 although limitations were noted in abstract quality comparisons. Furthermore, Thel- wall [37] evaluated GPT-4âs capabilities in journal article evaluations and found its accuracy insufficient for full automation, recommending editorial oversight. Zhou et al. [38] assessed GPT-3.5 and GPT-4 via score prediction and review generation, highlighting valuable insights but also significant limitations. Du et al. [39] analyzed LLM-generated reviews compared to human-written reviews, noting the LLMsâ poten- tial to identify deficiencies. Amidst the flourishing development of LLMs, research has begun focusing on optimization strategies for LLM-generated peer reviews. Jin et al.[19] introduce an LLM-based peer review simulation framework AGENTREVIEW that addresses the complexities and privacy concerns of traditional peer review analysis, revealing signif- icant biases in paper decisions. Gao et al. [40] presented REVIEWER2, a two-stage review generation framework that explicitly models review aspect distributions, sup- ported by a dataset of 27,000 papers and 99,000 annotated reviews. Tan et al. [41] reformulated peer review as a multi-turn dialogue among authors, reviewers, and decision-makers, with novel evaluation metrics and a large-scale dataset. DâArcy et al. [42] proposed MARG, a feedback generation framework using multiple LLM instances for internal discussion to enhance feedback specificity. Yu et al. [43] introduced SEA, an automated reviewing framework aimed at improving the quality andconsistency of LLM-generated reviews. Wu et al. [44] explored the potential of LLMs to enhance post- publication peer review, demonstrating that fine-tuned models caneffectively identify high-quality articles but still face challenges in providing consistent and context- sensitive evaluations. Beyond generation and optimization, recent studies have also investigated the real-world impact of LLM-assisted reviewing. Recently, Liang et al. [26] proposed a method to estimate the proportion of LLM-generated or modifiedtexts in peer review corpora, applying it to AI conference reviews post-ChatGPT release. Similarly, Latona et al. [27] found that at least 15.8% of ICLR 2024 reviews were LLM-assisted and that these reviews tended to assign higher scores, with such papers more likely to be accepted. Inspired by these findings, we propose three research questions to conduct a fine-grained comparative analysis of peer-review reports before and after LLMs entered the public sphere (i.e., prior to and following the releaseof Chat- GPT), with a particular focus on changes in linguistic patterns, evaluative aspects, and recommendation-related signals. 2.3 Detection of Large Language Model-Generated Content LLMs possess enough linguistic capabilities for academic writing [ 45] and other scholarly tasks [46,47], which has inevitably raised concerns within the academic community [24]. Consequently, efforts to detect the use of LLMs have also begun. Mitchell et al. [48] proposed DetectGPT, which detects text generated by LLMs by analyzing the negative curvature regions of the modelâs log probability function. Yang et al. [49] proposed a novel training-free detection method, Divergent N-Gram Analysis (DNA-GPT), which identifies machine-generated text by analyzing the dif- ferences between original and regenerated text segments. Furthermore, Li et al. [ 50] developed a comprehensive testbed by collecting texts from diverse human writings 5 and LLM-generated deepfake texts from various LLMs. Liang et al. [26] proposed a simple and effective method for estimating the proportion of text within a large corpus that has been significantly altered or generated by LLM. 3 Methodology In contrast to the previous studies, we adopted a more refined analytical approach, focusing on the characteristics at the word, sentence, and aspect levels, while Liang et al. [ 26] and Latona et al. [27] primarily conducted analysis at the overall text level. This methodological difference enables our research to provide a deeper understanding and more detailed analysis of the text. It is important to note that this study detects review comments that may have been assisted by LLMs, rather than those entirely generated by LLMs. Figure1presents the research framework for the fine-grained analysis of peer review texts, which primarily includes peer review data collection and processing, and fine-grained analysis of review texts. RQ1 RQ2 RQ3 ReviewsReviews Data collection and processing NLTK Lexicon by LLMs Lexical sophistication Syntactic sophistication and complexity LLM-assisted detection Review Aspect Identification Word level Sentence level Aspect level Fine-grained analysis of review texts Fig. 1: Framework of this study. Colored blocks on the right denote the corresponding research questions. 3.1 Data Collection and Processing Considering the availability of large-scale data, this study selects two top-tier con- ferences in the field of computer science, ICLR and NeurIPS, usingOpenReview 2 as the data sources. It is important to note, however, that while thisstudy uses these two conferences as examples, the proposed methodology is generalizable and can be applied and analyzed in other domains, if data in a similar format is available. The 2 https://openreview.net/ 6 rationale for selecting this field and these two conferences is that they provide years of continuous and publicly available peer review data, which serves asa valuable ref- erence. The detailed data statistics are shown in Table1. It is important to note that since we rely on publicly available data, and neither OpenReview nor thereleased datasets include NeurIPS 2020, the data for that year are therefore missing. These reviews include textual evaluations, confidence scores, and overall score rat- Table 1: Number of ICLR and NeurIPS reviews and papers per year. VenueYear# Papers# Reviews# Accept# Reject ICLR 20174891,498245244 20189112,748425486 20191,5794,7645021,077 20202,2136,7216871,526 20212,97211,0588602,112 20223,32512,7777332,592 20234,91518,5601,5753,340 20245,74922,2452,2613,488 202510,49742,5713,7036,794 NeurIPS 20165693,1835690 20176791,9416790 20181,0093,0141,0090 20191,4284,2531,4280 20212,76610,7292,630136 20222,82410,3302,671153 20233,39415,1713,218176 20244,23416,6354,033201 Total49,613188,19827,22822,325 ings. The confidence score is measured on a consistent scale from 1to 5, reflecting the reviewersâ confidence in the comments they have provided. The overall score is based on a scale ranging from 1 to 10. The descriptive statistics for the two types of scores are provided in AppendixA, TablesA1andA2. Each score is accompanied by a detailed definition to clarify its meaning and provide guidance, the detail can be found in the Reviewer Guidelines 3 , which outline the expectations and criteria for reviewing papers. These guidelines include instructions for evaluating various aspectsof the submitted content, such as clarity, originality, soundness, and significance,and require scoring for certain aspects. Then, we use the Natural Language Toolkit (NLTK 4 ) to segment the text from peer review reports into words and sentences, facilitating subsequent aspect identification and LLM-assisted detection. 3.2 Review Sentence Aspect Identification To conduct a fine-grained analysis of peer review texts, it is necessary to identify evaluation aspects within the segmented peer review sentences. Based on the review 3 https://iclr.c/Conferences/2024/ReviewerGuide, https://neurips.c/Conferences/2024/ReviewerGuidelines 4 https://w.nltk.org 7 guidelines and the study by Yuan et al. [14], we can define eight review aspects: summary, motivation, originality, soundness, substance, replicability, meaningful comparison, and clarity. Descriptions of these evaluation aspectscan be found at https://github.com/neulab/ReviewAdvisor/blob/main/materials/AnnotationGuideli ne.pdf. For sentence-level evaluation aspect identification, we adopted the pre-trained aspect identification model from the study by Yuan et al. [14] and further optimized it. Their model was trained on manually annotated data with the addition of heuris- tic rules, and its performance was ultimately evaluated through human assessment, achieving an accuracy of 92.75%. Additionally, the model also identifies the sentiment polarity of the sentenceâs aspect. After processing the review text through a series of steps, we can analyze peer review texts from recent years at afine-grained level, including word, sentence, and aspect-based perspectives. In the method proposed by Yuan et al. [14], aspect identification is treated as a sequence labeling task. Specifically, given a review sentence consisting ofnwords S=w 1 , ..., w n , the objective is to convey appropriate aspect information in the input sequence through a mapping function. First, BERT is used to represent the sentence Sas a sequence containingntokens (e 1 , e 2 , ..., e n ), converting it into a context- enriched representation. This sequence is then fed into the trained mapping function to obtain feature classifications for labeling: p i =sof tmax(W e i +b)(1) WhereWandbare trained parameters of the multilayer perceptron.p i is a vector that represents the probability of tokenibeing assigned to different aspects. Specifically, given a review text, we first split it into sentences and identify the aspect mentioned in each sentence. For all aspects except Summary, senti- ment polarity (positive or negative) is also annotated. In Section 4.1.3and4.1.4, for each year, we compute the average number of aspect mentions and the corre- sponding sentiment distributions across review texts. For example, consider a review text containing 10 sentences, in which Summary is mentioned twice, and Original- ity and Clarity are each mentioned once, with Originality expressed positively and Clarity negatively. The corresponding dictionary used for aggregation can be repre- sented as:summary: 2, originalitypositive: 1, originalitynegative: 0, claritypositive: 0, claritynegative: 1, replicabilitypositive: 0, replicabilitynegative: 0, sound- nesspositive: 0, soundnessnegative: 0, motivationpositive: 0, motivationnegative: 0, substancepositive:0, substancenegative: 0, meaningfulcomparisonpositive: 0, mean- ingfulcomparisonnegative: 0. Each review text is converted into such a dictionary, which is then used for subsequent analyses. 3.3 LLM-Assisted Peer Review Text Detection To identify which review reports are LLM-assisted, we need to detect whether the review texts have been modified or generated by LLMs. In contrast to the previous studies, we adopted a more refined analytical approach, focusingon the characteristics at the word, sentence, and aspect levels, while Liang et al. [ 26] and Latona et al. [27] 8 primarily conducted analysis at the overall text level. This methodological difference enables our research to provide a deeper understanding and moredetailed analysis of the text. It is important to note that this study detects reviewcomments that may have been assisted by LLMs, rather than those entirely generated by LLMs. We employ the maximum likelihood model designed by Liang et al. [26] to detect text that may have been assisted by artificial intelligence. We chose this model because our dataset has a very high similarity to the data used in the original study, ensuring that its performance remains valid for our case. This model leverages expert-authored and LLM-generated reference texts to accurately and efficiently assess corpus-level real- world usage of LLMs. And the average prediction error of the model is less than 5%. The objective of this model is used maximum likelihood estimation (MLE)L(α) to estimate theαscore for LLM-generated content: L(α) = n â i=1 log((1âα)P(x i ) +αQ(x i ))(2) Wherexto refer to a corpus,PandQdenote the probability distribution of documents written by scientists and generated by LLM in Liang et al. [ 26]âs model, respectively. We replacexinQwith our data to calculate the value ofα, thereby determining whether the input text is LLM-assisted generated. In addition to utilizing the model for detection, we also employed a lexicon of commonly and primarily used terms by LLMs, as developed in the studies by Liang et al. [26] and Latona et al. [27], to aid in detection. The specific lexicon is provided in Appendix TableB1. Prior to applying the detection model described above, we first perform a filtering step using a predefined terminology dictionary. If the review text does not contain any terms from this dictionary, we assume it is unlikely to have been LLM-assisted and exclude it from further analysis. Only the texts that pass this initial filter are subsequently input into the detection model for evaluation. 3.4 Lexical Sophistication and Syntactic Sophistication and Complexity for Review Text In addition to segmenting the review texts into word- and sentence-level units, we also calculated the lexical and syntactic complexity of the review texts for each year. Lexical complexity was measured using TAALES [ 51], a tool that evaluates over 400 classic and novel indices of lexical complexity, including indices relatedto various sub- structures. Syntactic sophistication and complexity were assessed using the advanced syntactic analysis tool TAASSC [52], which captures a wide range of metrics asso- ciated with syntactic development. These two tools do not calculatea single value to measure syntax but instead compute multiple distinct metrics forcomparison. We selected several key metrics for our study, as shown in Table2and3. The reason for selecting these three metrics to measure lexical complexity is that they assess the frequency of a wordâs occurrence across different documents, reflecting the breadth of vocabulary usage. Additionally, analyzingthe frequency of common phrases (bigrams and trigrams) reveals common expression structures and word collocations in the text. These metrics enable the evaluation ofa textâs lexical 9 Table 2: The lexical sophistication metrics used for peer review text. MetricsDescriptionTypes of Words COCA Academic Range AW Mean Range (number of documents that a word occurs in) score All words COCA Academic Bigram Frequency Mean bigram frequency scoreAll words COCA Academic Trigram Frequency Mean trigram frequency scoreAll words Note:COCA is corpus of contemporary American English. AW is all words. Table 3: The syntactic sophistication and complexity metrics used for peerreview. MetricsDescriptionType of Sentences advclpercl Number of adverbial clauses per clause Clause Complexity nsubj percl Number of nominal subjects per clause Clause Complexity mark percl Number of subordinating conjunctions per clause Clause Complexity aux percl Number of auxilliary verbs per clause Clause Complexity dobj percl Number of direct objects per clause Clause Complexity diversity, the richness of language expression, and sentence complexity. The five met- rics chosen to calculate syntactic sophistication and complexity were selected because adverbial clauses, noun clauses as subjects, and subordinating conjunctions indicate sentence complexity and logical relationships. Auxiliary verbs, on the other hand, help to understand the reviewerâs stance and tone, particularly in speculative or definitive judgments. Furthermore, the frequency of direct objects reflects the reviewerâs atten- tion to specific entities and the concreteness of the text. These metrics allow for a deeper analysis of the structural diversity, subjectivity, and rigor of argumentation in the review texts. 4 Result In this section, we first present the trends in review text length and aspect mentions in ICLR and NeurIPS before and after the advent of LLMs. We thenexamine aspect evaluations by reviewers using LLM assistance versus those without, and finally, we analyze the relationship between LLM-assisted review texts and assigned scores. 4.1 Impact of LLMs on Linguistic Complexity and Content-Level Expression This section addresses RQ1, which investigates how the emergenceof LLMs has influ- enced the linguistic and content-level properties of peer review texts. The analysis 10 (a) Average sentence count for ICLR After LLM (a) Average sentence count for ICLR After LLM (b) Average word count for ICLR After LLM (b) Average word count for ICLR After LLM (c) Average sentence count for NeurIPS After LLM (c) Average sentence count for NeurIPS After LLM (d) Average word count for NeurIPS After LLM (d) Average word count for NeurIPS After LLM (a) Average sentence count for ICLR After LLM (a) Average sentence count for ICLR After LLM (b) Average word count for ICLR After LLM (b) Average word count for ICLR After LLM (c) Average sentence count for NeurIPS After LLM (c) Average sentence count for NeurIPS After LLM (d) Average word count for NeurIPS After LLM (d) Average word count for NeurIPS After LLM Fig. 2: The trends in sentence and word lengths over the years for the ICLR and NeurIPS conferences. examines both linguistic complexity (e.g., length, lexical and syntacticsophistication) and aspect-level expression (e.g., number, length, and sentimentof aspect mentions). 4.1.1 Overall Text Length Patterns Before and After the Emergence of LLMs We first compare the length of reviews across pre- and post-LLM periods 5 . The trends in the length of review texts for ICLR and NeurIPS are illustrated in Figure2. The subgraph (a), (b), (c), and (d) in Figure2represent the average sentence count and average word count for ICLR, and the average sentence count and average word count for NeurIPS, respectively. The average sentence and word counts in ICLR reviews exhibit a rise-fall-rise pattern over time. This suggests that when LLMs first became available, only a small number of reviewers adopted them, resulting ina continuation 5 The review release date for ICLR 2023 was November 5, 2022 (https://iclr.c/Conferences/2023/Dates), and the rebuttal start date for NeurIPS 2022 was July 26, 2022(https://nips.c/Conferences/2022/Dates). In contrast, the rebuttal start date for NeurIPS 2023 was August 2, 2023 (https://nips.c/Conferences/2023/Dates). Since ChatGPT was released on November 30, 2022, we con- sider the review processes of ICLR 2023 and NeurIPS 2022 to beunaffected by it, whereas reviews for subsequent ICLR and NeurIPS conferences may have been influenced by LLMs. 11 Fig. 3: Average length of LLM assisted and non LLM assisted review textsin ICLR (2024-2025) and NeurIPS (2023-2024). of earlier trends. In contrast, ICLR 2025 6 explicitly mentioned the use of LLMs, indi- cating that reviewers assisted by LLMs were able to provide longer and more detailed feedback. Similarly, NeurIPS shows a comparable trend, but with noticeable differ- ences in the average word count after the emergence of LLMs. A plausible explanation is that NeurIPS reviews more frequently correspond to acceptedpapers, which tend to include more concise yet targeted evaluations. To further examinethe reasons behind the change in review length following the introduction of LLMs, we compared review texts from ICLR 2024-2025 and NeurIPS 2023-2024, distinguishing between LLM- assisted and non-assisted reviews, as illustrated in Figure 3. The results show that reviews potentially assisted by LLMs tend to be longer than those without LLM assis- tance, which aligns with the intuition that LLM-generated content typically increases text length. We think that the decline in review quality may be a contributing factor to the observed reduction in review length. Recent studies [53â56] have indicated a gradual decrease in the quality of peer reviews over the years, and the trend in text length may serve as a tangible reflection of this decline in quality. 4.1.2 Linguistic Complexity To capture the fine-grained changes in language use, we analyze lexical and syntactic complexity. Lexical richness reflects vocabulary diversity and sophistication, while syntactic indicators measure grammatical elaboration. 6 https://blog.iclr.c/2025/04/15/leveraging-llm-feedback-to-enhance-review-quality/ 12 (a) Lexical Sophistication (b) 6 6RSKLVWLFDWLRQDQG&RPSOH[LW\ (c) Lexical Sophistication (d) 6 6RSKLVWLFDWLRQDQG&RPSOH[LW\ ICLR NeurIPS After LLM After LLM After LLMAfter LLM Fig. 4: Lexical sophistication and syntactic sophistication and complexityof review texts in ICLR and NeurIPS. As shown in Figure 4, the lexical complexity of reviews in both conferences remains relatively stable over time, though it exhibits a slight decline following the emergence of LLMs. In terms of syntactic complexity, both conferences display similar overall trends; however, a notable exception is the significant increase in the number of nominal subjects in ICLR reviews after the introductionof LLMs. We hypothesize that the decrease in lexical complexity observed in ICLR 2024 may be associated with the adoption of LLMs. LLMs are capable of generating concise, clear, and readable text while generally avoiding overly complex terminology,which may contribute to a reduction in lexical complexity within review reports.Recent studies [45,57] have examined the applications and impacts of LLMs in academic writing. Their findings indicate that since the release of ChatGPT, the use ofChatGPT or other LLMs in abstracts has been steadily increasing, and by February 2024, approximately 35% of arXiv abstracts may have been modified using ChatGPT or other LLMs. Peer review reports from ICLR 2024 and NeurIPS 2023 have also been influenced, to varying degrees, by the use of language learning models [26,27]. The increase in nominal subject usage within ICLR reviews suggests that reviewers are becoming more inclined to employ noun phrases as subjects rather than personal pronouns (e.g., âIâ or âweâ), thereby enhancing the objectivity and professionalism of their writing. Additionally, we observe a downward trend in the useof auxiliary 13 verbs across both conferences. A possible explanation for this trend is that LLMs tend to generate concise and straightforward sentences, leading to a lower frequency of auxiliary verb usage. Consequently, reviewers assisted by language learning models may have contributed to this overall decline. 4.1.3 Aspect-Level Expression Patterns We then examine how content expression has shifted at the aspectlevel. Specifically, we analyze: The average length of aspect-related contents to measure elaboration depth (Figure 5). The average number of aspect mentions (e.g., Clarity, Soundness, Originality) to assess topic coverage (Figure6). The average sentiment polarity associated with each aspect to explore tonal tendencies (Figure7). The detailed computation procedures are provided in Section3.2. (c) Average sentence count for aspect mentions in NeurIPS (a) Average sentence count for aspect mentions in ICLR (b) Average word count for aspect mentions in ICLR (d) Average word count for aspect mentions in NeurIPS After LLM After LLM After LLM After LLM (c) Average sentence count for aspect mentions in NeurIPS (a) Average sentence count for aspect mentions in ICLR (b) Average word count for aspect mentions in ICLR (d) Average word count for aspect mentions in NeurIPS After LLM After LLM After LLM After LLM Fig. 5: The average length of sentences and words related to the reviewaspects in ICLR 2017-2025 and NeurIPS 2016-2024 (except 2020). 14 As shown in Figure5, after 2018, the average sentence and word lengths of aspect- related mentions in ICLR reviews exhibit an overall declining trend, followed by a noticeable increase after the emergence of LLMs. In contrast, the variation in NeurIPS is less pronounced, but a similar upward trend in sentence count canbe observed after the emergence of LLMs. This suggests that the advent of LLMs may have influenced the way reviewers compose their reports and may even have encouraged some reviewers to employ LLMs as assistance in conducting aspect-level evaluations of papers. In addition, the results in the figure reveal a sharp increase from 2017 to 2018. We attribute this phenomenon to the emergence of the Attention mechanism [58] in 2017, which triggered a rapid expansion of research activity in the deep learning community, thereby leading to the observed surge during this period. (a) (b) Fig. 6: The distribution of aspect mentions and their sentiment in ICLR 2017-2025 and NeurIPS 2016-2024. Comparison indicates a meaningful comparison. 15 Then, we analyzed the overall aspect mentions in ICLR and NeurIPSpeer review texts before and after the advent of LLMs. As shown in Figure6, we present the annual distribution of reviewersâ focus across various aspects. Overall,both venues consis- tently emphasize summary, which remains the dominant linguistic component across years. The recent rise in the frequency of summary mentions may be attributed to the growing complexity of academic papers, prompting reviewers to produce longer and more detailed summaries. For ICLR, a pronounced peak is observedin clarity, sound- ness, and summary between 2022 and 2023. This sharp increase coincides with the early diffusion of LLM-assisted reviewing, suggesting that automated orsemi-automated tools enhanced linguistic fluency and logical coherence. However, after 2023, both con- ferences show a decline in the frequency of originality-related evaluations, which we hypothesize may also be influenced by the use of LLMs. In contrast, NeurIPS exhibits a relatively stable pattern with only slight increases, indicating a morestandardized and less volatile reviewing style. A possible explanation for this discrepancy is that, in our dataset, the number of rejected papers from ICLR is substantially higher than that from NeurIPS. Overall, the results suggest that the emergence of LLMs has had a positive impact on the clarity and soundness dimensions of reviews, yet it has not sub- stantially improved meaningful comparison or replicability. This findingimplies that automated tools primarily enhance expressive fluency rather thanreviewersâ critical reasoning capabilities. 16 Fig. 7: The sentiment distribution of aspect mentions in ICLR 2017-2025 and NeurIPS 2016-2024 (except 2020). Comparison indicates a meaningful comparison. Finally, the sentiment polarity distribution for the above aspects is illustrated in Figure7. A notable shift in sentiment can be observed in ICLR between 2022 and 2024. This suggests that as submission volumes surged, reviewersexhibited changes in emotional tone. With the emergence of LLMs, while submission volumes continued to rise, reviewers gained access to tools potentially enhancing theirefficiency, likely contributing to increased emotional stability. The sentiment distribution for NeurIPS demonstrates relatively stable trends, which aligns with the steadyaspect mention patterns. 17 4.1.4 Interaction Between Reviewer Confidence and Linguistic Features Then, we test whether the above linguistic and content-level changes vary by review- ersâ self-reported confidence scores. This analysis reveals whether LLMs amplify or mitigate stylistic and expressive differences between confident and less-confident reviewers. (a) ICLR (b) NeurIPS Sentence Word Sentence Word Fig. 8: Number of words and sentences by confidence scores (1-5) in ICLR (2017-2025, except 2020) and NeurIPS (2021-2024). Firstly, we analyze the average length variation of review reports with different confidence scores. As shown in Figure 8(a), we observe that ICLR reviewers with confidence scores between 2 and 5 exhibit a trend where average review length initially increases and then decreases. Meanwhile, reviewers with a confidence score of 1 show an upward trend after 2023, indicating a differing trend from otherconfidence scores. In Figure8(b), we see that NeurIPS reviewers with varying confidence scores generally follow similar trends, except for those with a confidence score of 1,whose patterns resemble those observed for ICLR. From these observations, weinfer that while higher- confidence reviewers tend to write longer reports, the length of these reports has been decreasing in recent years. This indicates that high confidence score reviewers become increasingly familiar with the field over time, especially after the emergence of LLMs. High confidence score reviewers can write more accurate and concise evaluations that directly address key issues, resulting in shorter overall lengths. The observed upward trend among reviewers with a confidence score of 1 indicates that the emergence of LLMs can help reviewers with low confidence scores learn more about areas and knowledge they are not familiar with. This enables them to provide relatively informed 18 reviews instead of submitting low value evaluations due to unfamiliaritywith the topic, as in the past. Fig. 9: Yearly trends of syntactic sophistication and complexity across reviewer confidence levels in ICLR (2017-2025, except 2020). 19 Fig. 10: Yearly trends of syntactic sophistication and complexity across reviewer confidence levels in NeurIPS (2021-2024). We then analyze the lexical and syntactic complexity of review textswith differ- ent levels of reviewer confidence. Figure 9and10present the variation of syntactic sophistication and complexity across reviewer confidence levels forICLR (2017-2025, except 2020) and NeurIPS (2021-2024). Overall, reviews with higher confidence tend to exhibit greater syntactic complexity, characterized by a higher proportion of auxiliary verbs (auxpercl) and adverbial clauses (advclpercl), reflecting more elaborated and logically structured sentences. In contrast, features such as direct objects (dobjpercl) and clause markers (markpercl) show more fluctuation, indicating stylistic diversity among reviewers. After the emergence of LLMs (notably ICLR 2024-2025 and NeurIPS 2024), the syntactic patterns become more stable and consistent across confidence lev- els. Specifically, the increased use of auxiliaries and adverbial clauses suggests that LLM-assisted reviews employ more standardized and cohesive syntactic constructions, 20 while the slight decline in direct object and marker frequencies implies reduced syn- tactic depth and less structural variety. Compared with NeurIPS, ICLR demonstrates a more pronounced increase in syntactic regularity, highlighting a stronger influence of LLM assistance on the linguistic formulation of review texts. These findings indicate that LLMs may contribute to grammatically refined yet more homogeneous sen- tence structures, reinforcing fluency but potentially diminishing individual variation in writing style. Fig. 11: Yearly trends of lexical sophistication across reviewer confidence levels in ICLR (2017-2025, except 2020). 21 Fig. 12: Yearly trends of lexical sophistication across reviewer confidence levels in NeurIPS (2021-2024). Figure 11and12present the variation in academic lexical frequencies across reviewer confidence levels for ICLR (2017-2025, except 2020) and NeurIPS (2021- 2024). In both venues, reviews with higher confidence generally show lower bigram and trigram frequencies, indicating that more confident reviewerstend to employ less formulaic and more flexible academic language. However, after the introduction of LLMs (notably in ICLR 2024-2025 and NeurIPS 2024), these differences become attenuated: academic word frequency increases while bigram/trigram frequency declines, suggesting that LLM-assisted reviews adopt more standardized yet less lexically diverse expressions. Compared with NeurIPS, ICLR demonstrates a clearer rise in academic word frequency in the LLM period, implying a strongerinfluence of model-assisted writing on linguistic normalization. Finally, we analyzed the aspect mentions and sentiment polarity distribution for reviewers with different confidence scores, as shown in Figures13and14. For these two figures, we first briefly describe the criteria used for groupingby year and aspect. For example, Figure13(a) shows the number of review texts from ICLR 2021 with a confidence score of 1 (the number of reviews with a confidence score of 1 in that year is 92, as reported in TableA2). We identify the number of mentions of each aspect in every review text and the corresponding sentiment polarity, and compute their aver- ages; accordingly, each value shown for an aspect in the figure represents the average number of mentions. In addition, for sentiment polarity (Figure13(f)), we sum the sentiment polarity of all aspects within each review text for that year (positive and negative are calculated separately), and then compute the average sentiment across 22 all review texts with different confidence scores in that year. ÎŽaΔΎaΔ ÎŽdΔ ÎŽcΔ ÎŽeΔΎeΔ ÎŽbΔΎbΔ ÎŽfΔ ÎŽaΔΎaΔ ÎŽdΔ ÎŽcΔ ÎŽeΔΎeΔ ÎŽbΔΎbΔ ÎŽfΔ Fig. 13: Distribution of aspect mentions and sentiment polarity across different confidence scores in ICLR (2017-2025, except 2020). Figures (a)-(e) represent the distribution of aspect mentions for confidence scores ranging from 1 to 5, while figure (f) shows the sentiment polarity distribution. 23 ÎŽaΔ ÎŽbΔ ÎŽcΔ ÎŽdΔ ÎŽeΔ ÎŽfΔ ÎŽaΔ ÎŽbΔ ÎŽcΔ ÎŽdΔ ÎŽeΔ ÎŽfΔ Fig. 14: Distribution of aspect mentions and sentiment polarity across different confidence cores in NeurIPS (2021-2024). Figures (a)-(e) represent the distribution of aspect mentions for confidence scores ranging from 1 to 5, while figure (f) shows the sentiment polarity distribution. Firstly, in Figures 13(a)-(e), we present the average number of aspects mentioned by reviewers with confidence scores ranging from 1 to 5 in ICLR. Theresults indicate that, aside from the aspect of summary, the trends in mentions ofother aspects correlate with the changes in average length. Additionally, we observe that regardless of the confidence score, mentions of the summary have shown an upward trend after 24 2023, particularly pronounced among reviewers with a confidence score of 1. We think that reviewers with a confidence score of 1 may be less familiar with the review domain, thus relying more on the abstract section of the paper fortheir evaluation. Furthermore, after 2023, LLM-assisted review may have placed greater emphasis on the abstract, or the generated review reports may have highlighted the summary more prominently, facilitating easier reference for reviewers, particularly those who are less knowledgeable about the field. This also suggests that reviewers utilizing LLM assistance may primarily focus on the evaluation of the summaryrather than assessing core aspects of the paper, such as originality. We also present the average number of aspects mentioned by reviewers with confidence scoresranging from 1 to 5 in NeurIPS in Figures14(a)-(e). From the results in the figure, we can observe an upward trend in the number of mentions of the summary by reviewers across all confidence scores in 2023. However, the average length in NeurIPS that year was able to accommodate this increasing trend. Therefore, we infer that NeurIPS reviewers may have been less affected by LLM assistance. We also present the sentiment polarity changes for reviewers with different con- fidence scores at ICLR and NeurIPS, as shown in Figures13(f) and14(f). From the figures, we observe that reviewers with higher confidence scores(scores of 4 and 5) exhibit more pronounced sentiment fluctuations, particularly in negative sentiment. LLM may assist these high-confidence reviewers in conducting moredetailed and in- depth assessments, enabling them to identify more issues within thepapers, thereby increasing the expression of negative sentiment. These reviewersmay use LLM for efficient critical evaluation. On the other hand, reviewers with lowerconfidence scores (scores of 1 and 2) display significant sentiment variation, which could reflect their useness on LLM during the review process. Given their potential lack of deep understanding in the field, LLM may partially compensate for their knowledge gaps, assisting them in analyzing the content of the papers and resulting innoticeable shifts in sentiment polarity. The above analysis assumes that reviewers are utilizing LLM-assisted tools during the review process. Previous studies byLiang et al. [26] and Latona et al. [27] have indicated that at least 15% of review reports were LLM- assisted. 25 ÎŽaΔICLR LLM-assisted ÎŽbΔICLR non-LLM-assisted ÎŽcΔNeurIPS LLM-assisted ÎŽdΔNeurIPS non-LLM-assisted Fig. 15: The distribution of aspect mentions between LLM-assisted reviewers and non-LLM-assisted reviewers. Comparison indicates a meaningful comparison. 4.2 Which Evaluation Aspects Are More Prominently Reflected in LLM-assisted Reviews To address RQ2, we conducted LLM-assistance detection on the review reports from ICLR 2024-2025 and NeurIPS 2023-2024, obtaining 9775 and 5884samples, respec- tively. We then performed aspect recognition on the reports identified as LLM-assisted. Similarly, we carried out aspect recognition on an equivalent number of reports that were not identified as LLM-assisted, and compared the aspect mention distributions between the two, as shown in Figure 15. The non-LLM-assisted peer review texts were sampled from ICLR 2023 and NeurIPS 2022, with the number of samples matched to those of ICLR 2024-2025 and NeurIPS 2023-2024, respectively. As shown in Figure15, in both ICLR and NeurIPS, the summary section occu- pies the largest proportion of content in both LLM-assisted and non-LLM-assisted review reports, highlighting its central role in reviewersâ evaluationprocesses. However, in ICLR reviews, LLM-assisted reports tend to allocate a larger share to the sum- mary compared to non-LLM-assisted ones, whereas the oppositepattern is observed in NeurIPS, although the difference there is less pronounced than in ICLR. We attribute this discrepancy to differences in data composition: in our sample, the vast majority of NeurIPS reviews correspond to accepted papers, while ICLR includes a substantial 26 number of rejected submissions, which may lead to different distributions of emphasis on summaries. Overall, these results indicate that LLM assistance does not systemati- cally increase reviewersâ emphasis on the summary section. The motivation aspect also receives slightly greater emphasis in LLM-assisted reviews, indicating that reviewers may rely on generative models to elaborate on the contextual significance of a study. A similar pattern is observed for NeurIPS, where both substance and originality receive comparable levels of attention across review types, though originality tends to be marginally higher in non-LLM-assisted reports. Conversely, LLM-assisted reports generally devote less focus to clarity, particularly in ICLR, implying that the lin- guistic fluency provided by LLMs may reduce reviewersâ perceived need to explicitly comment on the readability or organization of the paper. Furthermore, non-LLM- assisted reports place relatively greater emphasis on replicability and soundness than LLM-assisted ones, suggesting that human reviewers continue toengage more with methodological robustness and reproducibility concerns dimensions that are less fre- quently discussed in LLM-assisted writing. Notably, while NeurIPS exhibits limited distributional change across LLM-assisted and non-LLM-assisted groups, ICLR demonstrates more pronounced shifts. Specif- ically, LLM-assisted reviewers in ICLR show heightened attention tosummary, substance, and replicability, while their mentions of originality and clarity decrease considerably. Non-LLM-assisted reviewers, by contrast, focusprimarily on summary, clarity, soundness, and originality, whereas LLM-assisted reviewers emphasize sum- mary, soundness, substance, and clarity. This shift suggests that, with the aid of LLM, reviewers tend to generate broader and more structured overviews, potentially improving their comprehension of a paperâs content. However, such assistance appears to come at the costof reduced engagement with critical and creative dimensions of evaluation, such as originalityand replicabil- ity. The declining attention to originality, in particular, raises concerns about whether LLM-assisted reviews may inadvertently promote uniformity and reduce the diversity of evaluative perspectives within peer review discourse. 4.3 Effects of LLM Assistance on Reviewersâ Scoring and Confidence To address RQ3, we calculated the correlation between aspect mentions by LLM- assisted reviewers and the scores they assigned, including their confidence scores, using the Spearman correlation coefficient. 4.3.1 Correlation Analysis Between Aspect Mentions and Reviewer-Assigned Scores First, we investigate the relationship between the aspects mentioned in review texts and the scores assigned by reviewers. Specifically, we examine whether the frequency with which different review aspects are discussed is associated with reviewersâ overall scores. To this end, we conduct a correlation analysis and provide visualizations in the form of scatter plots to illustrate the distributional patterns between aspect mentions and reviewer-assigned scores across ICLR and NeurIPS. 27 (a) ICLR (b) NeurIPS Fig. 16: Scatter plots of overall score versus aspects mentioned in ICLRand NeurIPS.Note:The solid lines represent simple linear regression fits and the shaded areas show the associated confidence intervals. The figure is intended as a visual aid for qualitative comparison rather than as the basis for formal regression analysis. 28 Table 4 : The results of correlation analysis between aspect mentions and o verall scores in ICLR and NeurIPS. Summ represents Summary, Moti represents Motivation, Rep represents Replicability, Ori rep resents Originality, Meaningful represents Meaningful compariso n. Venue Summ Moti Substance Rep Ori Soundness Clarity Meaning ful ICLR score 0.04*** 0.05*** 0.04*** -0.02** 0.01 0.01 0.00 -0.05*** NeurIPS score 0.01 0.06*** -0.01 -0.01 0.02 0.00 0.03** -0.04*** Note: ***p < 0 . 001; **p < 0 . 05. 29 As shown in Figure16(a) and (b), We have plotted scatter diagrams of the review- ersâ scores against the mentioned aspects in both ICLR and NeurIPS. ICLR and NeurIPS show similar overall results. From the figure, we can observe that, the men- tion of summary has a slightly positive influence on the assigned score, though the effect is very minimal. Motivation follows a similar pattern, showing a weak positive correlation with the score, but the scattered distribution of datapoints indicates that the impact is not significant. The regression lines for substance, soundness, and clarity are nearly horizontal, suggesting no significant relationship with thescore. Mentions of replicability and meaningful comparison may have a weak negative impact on the score. In contrast, originality shows a very slight positive correlation, implying that this aspect could contribute to a higher score. We note that the regression lines are included for visualization purposes only and are not used for formalstatistical infer- ence. Furthermore, we calculated the correlation between the aspectsmentioned and the scores assigned by reviewers in ICLR and NeurIPS, as shown in Table4. As shown in the results from the table, consistent with the scatter plot findings, aspects such as summary, motivation, and originality are positively correlated withthe assigned scores, while aspects like replicability and meaningful comparison exhibit a negative correlation with the scores. These results indicate that the number of mentions of var- ious aspects has a weak influence on the assigned scores, with the correlations being not statistically significant. This suggests that even with LLM-assisted reviews, the use of LLM does not significantly impact the reviewersâ assigned scores. 4.3.2 Correlation Analysis Between Aspect Mentions and Reviewer-Assigned Confidence Scores We think that the use of LLM assistance may also have some influenceon the confi- dence scores. Therefore, we also plotted scatter diagrams of the reviewersâ confidence scores against the mentioned aspects in both ICLR and NeurIPS, shown in Figure 17(a) and (b). ICLR and NeurIPS show similar overall results. 30 (a) ICLR (b) NeurIPS Fig. 17: Scatter plots of confidence score versus aspects mentioned in ICLR and NeurIPS.Note:The solid lines represent simple linear regression fits and the shaded areas show the associated confidence intervals. The figure is intended as a visual aid for qualitative comparison rather than as the basis for formal regression analysis. From the figure, we can observe that, the relationship between aspect mention and reviewer confidence scores is generally weak in both ICLR and NeurIPS. While some aspects (e.g., abstract, substance, originality, and meaningful comparisons) show a slight positive correlation with confidence scores, these trends are accompanied by significant dispersion of data points. In contrast, clarity consistently shows a weak negative correlation with confidence scores. Overall, the scatterplots do not reveal a clear or robust linear relationship between aspect mention and confidence scores. We 31 note that the regression lines are included for visualization purposes only and are not used for formal statistical inference. To quantitatively assess these observations, we calculated the Spearman correla- tion coefficient between aspect mention and confidence scores, asshown in Table5. The results indicate that most correlations are weak in both conferences. Although some coefficients reached statistical significance, their values remain low, suggesting limited actual correlations. In particular, clarity shows a persistent negative correla- tion with confidence scores in both ICLR and NeurIPS, while most other arguments show a weak positive correlation. Overall, the association between the frequency of aspect mentions and review- ersâ confidence scores is limited, indicating that reviewersâ confidence under LLM assistance is not necessarily driven primarily by the relative emphasisthey place on individual review aspects. 32 Table 5 : The results of correlation analysis between aspect mentions and c onfidence scores in ICLR and NeurIPS. Summ represents Summary, Moti represents Motivation, Rep represents Replicabilit y, Ori represents Originality, Meaningful represents Meaningful comparison. Venue Summ Moti Substance Rep Ori Soundness Clarity Meaning ful ICLR Confidence score 0.01 -0.02** 0.04*** 0.01 0.02** 0.00 -0.02** 0.03*** NeurIPS Confidence score 0.03** 0.01*** 0.03** 0.02** 0.03** -0.01 -0.02** 0.02** Note: ***p < 0.001; **p < 0.05. 33 5 Discussion In this section, we will discuss the implication of our study on theoretical and practical, and limitation of our study. 5.1 Implication (1) Theoretical Implication This study used peer review reports from two artificial intelligence conferences as data sources, collecting peer review report data from ICLR 2017-2025 and NeurIPS 2016-2024 (excluding 2020). We conducted aspect identification at the sentence level and performed word-level, sentence-level, and aspect-level analyses of these reports, both holistically and across different confidence scores. Furthermore, we also identi- fied instances of LLM-assisted writing in the peer review reports for ICLR 2024-2025 and NeurIPS 2023-2024, and conducted a fine-grained analysis based on these find- ings. Based on the above results, we present the following insights. First, from a more fine-grained perspective, our results indicate that the number of words and sentences in ICLR and NeurIPS reviews have been affected to varying degrees following the emergence of LLMs. Similarly, with the exception of summary, the frequency of aspect mentions aligns with the observed trendsin text length. More- over, our findings reveal a decline in lexical complexity accompanied by an increase in the use of nominal subjects after the introduction of LLMs. These observations are consistent with the findings reported by Liang et al. [ 26] and Latona et al. [27], their results indicate that at least 15% of review reports were producedwith the assistance of artificial intelligence, and we believe that this proportion will continue to increase. Therefore, we argue that the rapid advancement of artificial intelligence in recent years has influenced the way peer review reports are written. Secondly, we performed aspect identification on LLM-assisted andnon-LLM- assisted review reports from ICLR 2024-2025 and NeurIPS 2023-2024, comparing the distribution of aspect mentions between the two groups. The results show that in ICLR, aside from the summary, the top three aspects emphasized in non-LLM- assisted reports were clarity, soundness, and originality, while in LLM-assisted reports, the focus shifted to clarity, soundness, and substance, accompanied by a noticeable reduction in originality-related mentions. In contrast, NeurIPS exhibits no significant changes across these aspects in our dataset, likely because the reviews primarily from accepted papers. This suggests that reviewers using LLM assistance may lack a clear framework for assessing originality, leading to fewer mentions of this aspect and plac- ing more emphasis on others. Finally, we examine the relationship between aspect mentions in LLM-assisted review reports and reviewersâ assigned scores. Both the scatter plots and Spear- man correlation results indicate that aspect mentions exhibit only weak and largely insignificant correlations with overall scores, suggesting that LLMassistance does not substantially influence the final evaluations assigned to papers. This finding is encour- aging, as it implies that reviewers primarily leverage LLMs to support the writing and organization of review reports, rather than to directly determineevaluative judgments. We further analyze the correlation between aspect mentions and reviewersâ confidence 34 scores. While most aspects show weak positive correlations with confidence, clarity consistently exhibits a negative correlation across both conferences. Although the cor- relations are weak, this result allows us to speculate that LLM assistance may not necessarily improve reviewersâ understanding of the paper content. On the contrary, when clarity is low, reviewers may experience greater uncertainty even with LLM sup- port, which can in turn reduce their confidence in their evaluations. In conclusion, these findings contribute to a nuanced theoreticalunderstanding of how LLM assistance can reshape traditional peer review processes by altering the structure and focus of review content without directly modifying evaluation outcomes. These insights underscore the need for further research in developing LLM that not only improve efficiency but also enhance comprehension in expert review settings. It is essential to clarify that we do not intend to suggest that using LLM in review writ- ing is inherently beneficial or detrimental. Nor do we claim (nor do we believe) that many reviewers are using LLM to compose entire reviews directly. Rather, we posit that current peer review processes are unlikely to be replaced by LLM but can instead be meaningfully supported by it. (2) Practical Implication Artificial intelligence is not a looming threat, and we need to approachits impact on academia with a rational perspective. While some studies [23,24] have raised concerns that LLMs (such as ChatGPT) may hinder the peer review process,the ultimate con- trol still lies in human hands. At present, LLM is merely a tool for assistance. Recent studies [35â38] have also shown that LLMs have the potential to support peer review. We believe that appropriate guidelines or regulations should be established to restrict the use of LLMs in peer review, without completely prohibiting their use. Instead, journals or conferences should provide peer reviewers with LLM-assisted tools, rather than leaving the choice entirely to the reviewers. We are confident that the integra- tion of LLM will significantly alleviate the pressure caused by the increasing volume of submissions. Recently, ICLR 2025 introducing 7 a review feedback agent that iden- tifies potential issues in reviews and provides feedback to reviewers for improvements. The goal of this system is to help make reviews more constructive and actionable for authors. From our results, the implementation of this system appears to have yielded a certain positive effect. Based on our findings, LLMs have not yet had a significant impact on the peer review process, particularly when examined at a fine-grained level. Reviewers using LLM assistance tend to employ it primarily for summarization or improving clar- ity rather than for evaluating specific aspects of a paper, such assoundness. While the study by Latona et al. [27] suggests that LLM-assisted reviews tend to assign higher scores, our fine-grained analysis reveals both positive and negative correlations between the use of LLM and the scores assigned to papers. This means that LLM assistance may lead to higher or lower scores, though overall, positive correlations are more common, aligning with Latona et al.âs findings. Therefore, we believe that the judicious use of LLM in tasks such as language polishing or summarizingpaper content has a relatively minor impact on the review process. As previously discussed, certain 7 https://blog.iclr.c/2024/10/09/iclr2025-assisting-reviewers/ 35 LLM functionalities should be restricted until the technology has fully matured, which remains a long-term objective. 5.2 Discussion on the Impact of LLMs on Review Report While our analysis primarily focuses on linguistic and structural shiftsin review texts, an important open question remains: whether the involvement of LLMs leads to more accurate, insightful, or useful reviews. Quantifying such dimensions of review quality is inherently challenging, as they involve subjective and context-dependent judgments that go beyond surface-level linguistic features. Nevertheless,our findings offer indi- rect evidence that may inform this broader discussion. The increase in clarity and soundness mentions among LLM-assistedreviews sug- gests that reviewers aided by LLMs are able to produce more coherent and logically consistent feedback. This linguistic improvement could enhance thereadability and interpretability of review reports, potentially making them more accessible and action- able for authors. However, we also observe a decrease in the frequency of originality and meaningful comparison mentions-dimensions that are crucial for deep, evaluative reasoning. This may imply that while LLMs help reviewers articulate their ideas more fluently, they might also encourage reliance on generic or template-like expressions, thereby reducing the depth and specificity of critical engagement. In other words, LLM assistance appears to improve the form and presentation of reviews rather than their analytical rigor. Future work could investigate this issue more directly by combining linguistic analyses with expert evaluations of review helpful- ness, accuracy, and constructiveness. Such a multi-dimensionalassessment-integrating human judgments, author feedback, and content-based scoring-would provide a more definitive understanding of whether LLMs enhance not only how reviewers write but also how well they evaluate. 5.3 Limitation Our study has several limitations worth highlighting. First, our datais primarily drawn from top conferences in the fields of machine learning and deep learning. To generalize these findings to other fields, corresponding data from those areas would be required. The NeurIPS data we used primarily comprises accepted papers, while the ICLR data contains more rejected papers than accepted ones. This difference may influence cer- tain results, such as average length, aspect mentions, and potential LLM assistance usage. Additionally, our research conducts fine-grained analysis mainly through word and sentence counts as well as lexical and syntactic complexity. Wedid not further examine detailed word usage or conduct more in-depth syntactic analyses, such as changes in syntactic dependencies. Secondly, the LLM-assisted detection method we used identifies reports that may have LLM assistance but does notprovide absolute certainty, as we enhance detection through likely LLM-assisted lexicons. Finally, our study is based on analyses of existing correlations or observed patterns rather than aiming to infer causal relationships through experimental or quasi-experimental meth- ods. Finally, this study focuses on identifying correlations and observed patterns rather than establishing causation. Although our analyses highlight associations between 36 review text, review aspect, sentiment, overall scores, confidence scores, and LLM- assisted, they do not imply causal relationships due to the observational nature of our study. Experimental or quasi-experimental approaches, such as controlled studies of LLMâs impact on review writing, would be necessary to draw causal inferences. 6 Conclusion and Future Works In this paper, we conducted a fine-grained analysis of open peer review reports from AI conferences in recent years. First, we examined the overall changes in the length of review texts, the distribution of aspect mentions, and the sentiment polarity of different aspects over time. We then performed the same analysis on review reports with varying levels of confidence scores. Additionally, we analyzed the lexical and syn- tactic complexity of peer review texts in recent years. Then, we analyzed the aspect mentions in LLM-assisted and non-LLM-assisted review reports. Finally, we calcu- lated the correlation between the assigned scores and aspect mentions in LLM-assisted review reports. Overall, our findings indicate that the emergence of LLMs is associated with longer and more fluent peer-review texts, increased emphasison summaries and surface-level clarity, and increasingly standardized linguistic patterns effects that are particularly pronounced among reviewers with lower self-reportedconfidence. At the same time, attention to deeper evaluative dimensions, including originality, replicabil- ity, and fine-grained critical reasoning, has declined. These shiftsbecome more evident when comparing LLM-assisted and non-LLM-assisted reviews. Although LLM-assisted reports show a modest improvement in the informativeness of recommendations, this trend raises concerns about a potential trade-off between linguistic fluency and eval- uative depth in peer review. We believe that our approach possesses a degree of generalizability across different domains and is not limited to the specific fields asso- ciated with the data discussed in this study. In future work, we plan to explore the causal relationship betweenLLM usage and changes in peer review comments by employing experimental or quasi-experimental methodologies. Additionally, we intend to broaden our scope by analyzing peer review reports from leading journals such as PLOS ONE and Nature Communications, to identify trends and patterns in the review process across different fields and publi- cation standards. Finally, we aim to adopt more sophisticated tools and techniques for detecting the use of LLMs in peer review comments. These toolscould include advanced text analysis methods and LLM-based detection systems, which will help us identify and quantify the extent to which LLMs influence scholarly reviews and ensure transparency in the review process. Acknowledgements.This study is supported by the National Natural Science Foundation of China (Grant No. 72074113). Declarations The author(s) declared no potential conlicts of interest with respect to the research, author- ship, and/or publication of this article. 37 Appendix A Distribution of Review Scores and Confidence Scores Table A1andA2are distribution of review scores and confidence scores in ICLR and NeurIPS. Table A1: Distribution of review scores in ICLR 2017-2025 and NeurIPS 2016-2024. Year12345678910 ICLR 2017114251258338331138 930 366 ICLR 2018560224536560621523 170 472 ICLR 201915 91418983 1063 1067 833 222 666 ICLR 2020924 -2562--2393-842-- ICLR 202114 167 905 2289 2685 2832 1999 495 112 10 ICLR 2022256 -2769-3224 2972-3508 -48 ICLR 2023465 -4704-5121 5324-2845 - 101 ICLR 2024353 -4246-6153 7632-3759 - 102 ICLR 2025941 - 10506-11547 13457-5950 - 170 NeurIPS 2016 941 - 10506-11547 13457-5950 - 170 NeurIPS 2017 941 - 10506-11547 13457-5950 - 170 NeurIPS 2018 941 - 10506-11547 13457-5950 - 170 NeurIPS 2019 941 - 10506-11547 13457-5950 - 170 NeurIPS 2021 218140513 1259 3939 3788 905 156 9 NeurIPS 2022 437320836 1935 3445 2913 767 694 NeurIPS 2023 850477 1197 3468 5075 3933 887 66 10 NeurIPS 2024 845506 1316 4163 5433 4140 930 80 14 Table A2: Distribution of confidence scores in ICLR 2017-2025 and NeurIPS 2021-2024. Year12345 ICLR 2017751387804249 ICLR 201834 1306921391 501 ICLR 201940 275 1239 2378 832 ICLR 2020----- ICLR 202192 641 3327 5614 1834 ICLR 202217 652 4223 6521 1364 ICLR 202387 1364 6077 8922 2110 ICLR 2024113 1736 7639 10226 2531 ICLR 2025149 2787 13842 20181 5612 NeurIPS 2021 100 674 3578 5182 1195 NeurIPS 2022 166 855 3582 4625 1102 NeurIPS 2023 271 1315 5231 6669 1685 NeurIPS 2024 201 1326 5391 7643 2074 38 Appendix B Lexicon of Frequently and Predominantly Used Terms by LLMs Table B1is the lexicon of frequently and predominantly used terms by LLMs. Table B1: Lexicon of frequently and predominantly used terms by LLMs meticulouslyreportedlylucidly innovativelyaptlymethodically excellentlycompellinglyimpressively undoubtedlyscholarlystrategically intriguinglycompetentlyintelligently hithertothoughtfullyprofoundly undeniablyadmirablycreatively logicallymarkedlythereby contextuallydistinctlyjudiciously cleverlyinvariablysuccessfully chieflyrefreshinglyconstructively inadvertentlyeffectivelyintellectually rightlyconvincinglycomprehensively seamlesslypredominantly coherently evidentlynotablyprofessionally subtlysynergisticallyproductively purportedlyremarkablytraditionally starklypromptlyrichly nonethelesselegantlysmartly solidlyinadequatelyeffortlessly forthfirmlyautonomously dulycriticallyimmensely beautifullymaliciouslyfinely succinctlyfurtherrobustly decidedlyconclusivelydiversely exceptionallyconcurrentlyappreciably methodologically universallythoroughly soundlyparticularlyelaborately uniquelyneatlydefinitively substantivelyusefullyadversely primarilyprincipallydiscriminatively efficientlyscientificallyalike hereinadditionallysubsequently potentiallycommendableinnovative meticulousintricatenotable versatilenoteworthyinvaluable pivotalpotentfresh ingeniouscogentongoing 39 tangibleprofoundmethodical laudablelucidappreciable fascinatingadaptableadmirable refreshingproficientintriguing thoughtfulcredibleexceptional digestibleprevalentinterpretative remarkableseamlesseconomical proactiveinterdisciplinary sustainable optimizablecomprehensive vital pragmaticcomprehensible unique fullerauthenticfoundational distinctivepertinentvaluable invasivespeedyinherent considerableholisticinsightful operationalsubstantialcompelling technologicalbeneficialexcellent keenculturalunauthorized strategicexpansiveprospective vividconsequentialmanageable unprecedentedinclusiveasymmetrical cohesivereplicablequicker defensivewiderimaginative traditionalcompetentcontentious widespreadenvironmentalinstrumental substantivecreativeacademic sizeableextantdemonstrable prudentpracticablesignatory continentalunnoticedautomotive minimalisticintelligentunderscores necessitatingdelvesadaptability delveddelveelucidated underscorecredibilityadvancements elucidationunderpinningsequitable perplexingexcelsintricacies persuasivenessdelineationelucidate provisionbolsterdiscourse meticulousendeavorstangible commendableshowcasingimperative encompassingoffering References [1] Siler, K., Lee, K., Bero, L.: Measuring the effectiveness of scientific gatekeep- ing. Proceedings of the National Academy of Sciences112(2), 360â365 (2015) https://doi.org/10.1073/pnas.1418218112 40 [2] Huisman, J., Smits, J.: Duration and quality of the peer review pro- cess: the authorâs perspective. Scientometrics113(1), 633â650 (2017) https://doi.org/10.1007/s11192-017-2310-5 [3] Yu, H., Liang, Y., Xie, Y.: Understanding the sustainability of supplyâdemand in peer review system: an analysis based on scholarsâ research and review activities. Scientometrics130(3), 1547â1569 (2025) https://doi.org/10.1007/s11192-025-05264-8 [4] Tennant, J.P.: The state of the art in peer review. FEMS Microbiology Letters 365(19), 204 (2018)https://doi.org/10.1093/femsle/fny204 [5] Russo, A.: Some Ethical Issues in the Review Process of Machine Learning Conferences (2021).https://arxiv.org/abs/2106.00810 [6] Tran, D., Valtchanov, A.V., Ganapathy, K.R., Feng, R., Slud, E.V., Gold- blum, M., Goldstein, T.: An Open Review of OpenReview: A Criti- cal Analysis of the Machine Learning Conference Review Process (2021). https://openreview.net/forum?id=Cn706AbJaKW [7] Stelmakh, I., Shah, N.B., Singh, A., Daum Ìe, H.: Prior and prejudice: The novice reviewersâ bias against resubmissions in conference peer review. Proc. ACM Hum.- Comput. Interact.5(CSCW1) (2021) https://doi.org/10.1145/3449149 [8] Zhang, J., Zhang, H., Deng, Z., Roth, D.: Investigating Fairness Dis- parities in Peer Review: A Language Model Enhanced Approach (2022). https://arxiv.org/abs/2211.06398 [9] Fox, C.W., Meyer, J., Aim Ìe, E.: Double-blind peer review affects reviewer ratings and editor decisions at an ecology journal. Functional Ecology37(5), 1144â1157 (2023)https://doi.org/10.1111/1365-2435.14259 [10] Lu, Y., Kong, Y.: Calibrating âcheap signalsâ in peer review withouta prior. In: Proceedings of the 37th International Conference on Neural Information Pro- cessing Systems. NIPS â23. Curran Associates Inc., Red Hook, NY,USA (2024). https://doi.org/10.5555/3666122.3667028 [11] Xu, Y.E., Jecmen, S., Song, Z., Fang, F.: A one-size-fits-all approach to improving randomness in paper assignment. In: Proceedings of the 37th International Con- ference on Neural Information Processing Systems. NIPS â23. Curran Associates Inc., Red Hook, NY, USA (2024) [12] Liu, Y., Yang, K., Liu, Y., Drew, M.G.B.: The Shackles of Peer Review:Unveiling the Flaws in the Ivory Tower (2023). https://arxiv.org/abs/2310.05966 [13] Wang, Q., Zeng, Q., Huang, L., Knight, K., Ji, H., Rajani, N.F.: ReviewRobot: Explainable paper review generation based on knowledge synthesis.In: 41 Davis, B., Graham, Y., Kelleher, J., Sripada, Y. (eds.) Proceedings of the 13th International Conference on Natural Language Generation, p. 384â397. Association for Computational Linguistics, Dublin, Ireland (2020). https://doi.org/10.18653/v1/2020.inlg-1.44 [14] Yuan, W., Liu, P., Neubig, G.: Can we automate scientific reviewing?J. Artif. Int. Res.75(2022)https://doi.org/10.1613/jair.1.12862 [15] Yuan, W., Liu, P.: Kid-review: Knowledge-guided scientific review generation with oracle pre-training. Proceedings of the AAAI Conference on Artificial Intelligence 36(10), 11639â11647 (2022)https://doi.org/10.1609/aaai.v36i10.21418 [16] Gao, Z., Brantley, K., Joachims, T.: Reviewer2: Optimizing Review Generation Through Prompt Generation (2024).https://arxiv.org/abs/2402.10886 [17] DâArcy, M., Hope, T., Birnbaum, L., Downey, D.: MARG: Multi-AgentReview Generation for Scientific Papers (2024).https://arxiv.org/abs/2401.04259 [18] Yu, J., Ding, Z., Tan, J., Luo, K., Weng, Z., Gong, C., Zeng, L., Cui, R., Han, C., Sun, Q., Wu, Z., Lan, Y., Li, X.: Automated peer reviewing in paper SEA:Stan- dardization, evaluation, and analysis. In: Al-Onaizan, Y., Bansal, M.,Chen, Y.-N. (eds.) Findings of the Association for Computational Linguistics: EMNLP 2024, p. 10164â10184. Association for Computational Linguistics, Miami, Florida, USA (2024).https://aclanthology.org/2024.findings-emnlp.595 [19] Jin, Y., Zhao, Q., Wang, Y., Chen, H., Zhu, K., Xiao, Y., Wang, J.: AgentReview: Exploring peer review dynamics with LLM agents. In: Al-Onaizan, Y., Bansal, M., Chen, Y.-N. (eds.) Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 1208â1226. Association forComputational Linguistics, Miami, Florida, USA (2024).https://aclanthology.org/2024.emnlp- main.70 [20] Barnett, A., Mewburn, I., Schroter, S.: Working 9 to 5, not theway to make an academic living: observational analysis of manuscript and peer review submissions over time. BMJ367(2019) https://doi.org/10.1136/bmj.l6460 [21] Geng, Y., Cao, R., Han, X., Tian, W., Zhang, G., Wang, X.: Scientistsare working overtime: when do scientists download scientific papers? Scientometrics127(11), 6413â6429 (2022)https://doi.org/10.1007/s11192-022-04524-1 [22] OpenAI: GPT-4 Technical Report (2024).https://arxiv.org/abs/2303.08774 [23] Donker, T.: The dangers of using large language models for peer review. The Lancet Infectious Diseases23(7), 781 (2023) https://doi.org/10.1016/S1473-3099(23)00290-6 [24] Chawla, D.S.: Is chatgpt corrupting peer review? telltale words hint at ai use. 42 Nature628(8008), 483â484 (2024)https://doi.org/10.1038/d41586-024-01051-2 [25] Liang, W., Zhang, Y., Cao, H., Wang, B., Ding, D.Y., Yang, X., Vodrahalli, K., He, S., Smith, D.S., Yin, Y., McFarland, D.A., Zou, J.: Can large language mod- els provide useful feedback on research papers? a large-scale empirical analysis. NEJM AI1(8), 2400196 (2024)https://doi.org/10.1056/AIoa2400196 [26] Liang, W., Izzo, Z., Zhang, Y., Lepp, H., Cao, H., Zhao, X., Chen, L., Ye, H., Liu, S., Huang, Z., McFarland, D., Zou, J.Y.: Monitoring AI-modified content at scale: A case study on the impact of ChatGPT on AI con- ference peer reviews. In: Salakhutdinov, R., Kolter, Z., Heller, K., Weller, A., Oliver, N., Scarlett, J., Berkenkamp, F. (eds.) Proceedings of the 41st International Conference on Machine Learning. Proceedings of Machine Learn- ing Research, vol. 235, p. 29575â29620. PMLR, Vienna, Austria (2024). https://proceedings.mlr.press/v235/liang24a.html [27] Latona, G.R., Ribeiro, M.H., Davidson, T.R., Veselovsky, V., West, R.: The AI Review Lottery: Widespread AI-Assisted Peer Reviews Boost Paper Scores and Acceptance Rates (2024). https://arxiv.org/abs/2405.02150 [28] Lin, J., Song, J., Zhou, Z., Chen, Y., Shi, X.: Automated scholarly paper review: Concepts, technologies, and challenges. Information Fusion98, 101830 (2023) https://doi.org/10.1016/j.inffus.2023.101830 [29] Kang, D., Ammar, W., Dalvi, B., Zuylen, M., Kohlmeier, S., Hovy, E., Schwartz, R.: A dataset of peer reviews (PeerRead): Collection, insights and NLP applica- tions. In: Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, Volume 1 (Long Papers), p. 1647â1661. Association for Computational Linguis- tics, New Orleans, Louisiana (2018).https://doi.org/10.18653/v1/N18-1149 [30] Wang, Q., Zeng, Q., Huang, L., Knight, K., Ji, H., Rajani, N.F.: ReviewRobot: Explainable paper review generation based on knowledge synthesis.In: Davis, B., Graham, Y., Kelleher, J., Sripada, Y. (eds.) Proceedings of the 13th International Conference on Natural Language Generation, p. 384â397. Association for Computational Linguistics, Dublin, Ireland (2020). https://doi.org/10.18653/v1/2020.inlg-1.44 [31] Li, J., Sato, A., Shimura, K., Fukumoto, F.: Multi-task peer-review score predic- tion. In: Chandrasekaran, M.K., Waard, A., Feigenblat, G., Freitag,D., Ghosal, T., Hovy, E., Knoth, P., Konopnicki, D., Mayr, P., Patton, R.M., Shmueli-Scheuer, M. (eds.) Proceedings of the First Workshop on Scholarly DocumentProcess- ing, p. 121â126. Association for Computational Linguistics, Online(2020). https://doi.org/10.18653/v1/2020.sdp-1.14 [32] Yu, S., Luo, M., Madusu, A., Lal, V., Howard, P.: Is Your Paper Being Reviewed by an LLM? Benchmarking AI Text Detection in Peer Review (2025). 43 https://arxiv.org/abs/2502.19614 [33] Garg, M.K., Prasad, T., Singhal, T., Kirtani, C., Mandal, M., Kumar, D.: ReviewEval: An Evaluation Framework for AI-Generated Reviews (2025). https://arxiv.org/abs/2502.11736 [34] Zhuang, Z., Chen, J., Xu, H., Jiang, Y., Lin, J.: Large language models for auto- mated scholarly paper review: A survey. Information Fusion124, 103332 (2025) https://doi.org/10.1016/j.inffus.2025.103332 [35] Robertson, Z.: GPT4 is Slightly Helpful for Peer-Review Assistance: A Pilot Study (2023).https://arxiv.org/abs/2307.05492 [36] Liu, R., Shah, N.B.: ReviewerGPT? An Exploratory Study on Using Large Language Models for Paper Reviewing (2023).https://arxiv.org/abs/2306.00622 [37] Thelwall, M.: Can chatgpt evaluate research quality? Journal ofData and Information Science9(2), 1â21 (2024)https://doi.org/10.2478/jdis-2024-0013 [38] Zhou, R., Chen, L., Yu, K.: Is LLM a reliable reviewer? a comprehensive evalua- tion of LLM on automatic paper reviewing tasks. In: Calzolari, N., Kan, M.-Y., Hoste, V., Lenci, A., Sakti, S., Xue, N. (eds.) Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation (LREC-COLING 2024), p. 9340â9351. ELRA and ICCL,Torino, Italia (2024).https://aclanthology.org/2024.lrec-main.816 [39] Du, J., Wang, Y., Zhao, W., Deng, Z., Liu, S., Lou, R., Zou, H.P., Narayanan Venkit, P., Zhang, N., Srinath, M., Zhang, H.R., Gupta, V.,Li, Y., Li, T., Wang, F., Liu, Q., Liu, T., Gao, P., Xia, C., Xing, C., Jiayang, C., Wang, Z., Su, Y., Shah, R.S., Guo, R., Gu, J., Li, H., Wei, K., Wang, Z., Cheng, L., Ranathunga, S., Fang, M., Fu, J., Liu, F., Huang, R., Blanco, E., Cao, Y., Zhang, R., Yu, P.S., Yin, W.: LLMs assist NLP researchers: Critique paper (meta- )reviewing. In: Al-Onaizan, Y., Bansal, M., Chen, Y.-N. (eds.) Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, p. 5081â5099. Association for Computational Linguistics, Miami, Florida, USA (2024).https://aclanthology.org/2024.emnlp-main.292 [40] Gao, Z., Brantley, K., Joachims, T.: Reviewer2: Optimizing Review Generation Through Prompt Generation (2024). https://arxiv.org/abs/2402.10886 [41] Tan, C., Lyu, D., Li, S., Gao, Z., Wei, J., Ma, S., Liu, Z., Li, S.Z.: Peer Review as A Multi-Turn and Long-Context Dialogue with Role-Based Interactions (2024). https://arxiv.org/abs/2406.05688 [42] DâArcy, M., Hope, T., Birnbaum, L., Downey, D.: MARG: Multi-AgentReview Generation for Scientific Papers (2024).https://arxiv.org/abs/2401.04259 44 [43] Yu, J., Ding, Z., Tan, J., Luo, K., Weng, Z., Gong, C., Zeng, L., Cui, R., Han, C., Sun, Q., Wu, Z., Lan, Y., Li, X.: Automated peer reviewing in paper SEA:Stan- dardization, evaluation, and analysis. In: Al-Onaizan, Y., Bansal, M.,Chen, Y.-N. (eds.) Findings of the Association for Computational Linguistics: EMNLP 2024, p. 10164â10184. Association for Computational Linguistics, Miami, Florida, USA (2024).https://aclanthology.org/2024.findings-emnlp.595 [44] Wu, W., Zhang, Y., Haunschild, R., Bornmann, L.: Leveraging largelanguage models for post-publication peer review: Potential and limitations. In: Editors: Shushanik Sargsyan, Wolfgang Gl Ìanzel, Giovanni Abramo: 20th International Conference On Scientometrics & Informetrics, p. 23â27 (2025) [45] GENG, M., Trotta, R.: Is chatGPT transforming academicsâ writing style? In: ICML 2024 Next Generation of AI Safety Workshop (2024). https://openreview.net/forum?id=0CfKpG94co [46] Evans, J., DâSouza, J., Auer, S.: Large language models as evaluators for scien- tific synthesis. In: Araujo, P.H., Baumann, A., Gromann, D., Krenn,B., Roth, B., Wiegand, M. (eds.) Proceedings of the 20th Conference on Natural Language Pro- cessing (KONVENS 2024), p. 1â22. Association for Computational Linguistics, Vienna, Austria (2024).https://aclanthology.org/2024.konvens-main.1 [47] Ma, Y., Liu, J., Yi, F., Cheng, Q., Huang, Y., Lu, W., Liu, X.: AI vs. Human â Differentiation Analysis of Scientific Content Generation (2023). https://arxiv.org/abs/2301.10416 [48] Mitchell, E., Lee, Y., Khazatsky, A., Manning, C.D., Finn, C.: Detectgpt: Zero- shot machine-generated text detection using probability curvature. In: Krause, A., Brunskill, E., Cho, K., Engelhardt, B., Sabato, S., Scarlett, J. (eds.)International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA. Proceedings of Machine Learning Research, vol. 202, p. 24950â24962. PMLR, ??? (2023).https://proceedings.mlr.press/v202/mitchell23a.html [49] Yang, X., Cheng, W., Wu, Y., Petzold, L.R., Wang, W.Y., Chen, H.: DNA-GPT: Divergent n-gram analysis for training-free detection of GPT-generated text. In: The Twelfth International Conference on Learning Representations (2024). https://openreview.net/forum?id=Xlayxj2fWp [50] Li, Y., Li, Q., Cui, L., Bi, W., Wang, Z., Wang, L., Yang, L., Shi, S., Zhang, Y.: MAGE: Machine-generated text detection in the wild. In: Ku, L.-W., Mar- tins, A., Srikumar, V. (eds.) Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 36â53. Association for Computational Linguistics, Bangkok, Thailand (2024). https://doi.org/10.18653/v1/2024.acl-long.3 45 [51] KYLE, K., CROSSLEY, S.A.: Measuring syntactic complexity in l2 writ- ing using fine-grained clausal and phrasal indices. The Modern Lan- guage Journal102(2), 333â349 (2018)https://doi.org/10.1111/modl.12468 https://onlinelibrary.wiley.com/doi/pdf/10.1111/modl.12468 [52] Kyle, K., Crossley, S., Berger, C.: The tool for the automatic analysis of lexical sophistication (taales): version 2.0. Behavior Research Methods50(3), 1030â1046 (2018)https://doi.org/10.3758/s13428-017-0924-4 [53] Maddi, A., Miotti, L.: On the peer review reports: does size matter? Scientomet- rics129(10), 5893â5913 (2024)https://doi.org/10.1007/s11192-024-04977-6 [54] Guo, Y., Shang, G., Rennard, V., Vazirgiannis, M., Clavel, C.: Automatic analysis of substantiation in scientific peer reviews. In: Bouamor, H., Pino, J., Bali, K. (eds.) Findings of the Association for Computational Linguistics: EMNLP 2023, p. 10198â10216. Association for Computational Linguistics, Singapore (2023). https://doi.org/10.18653/v1/2023.findings-emnlp.684 [55] Oviedo-Garc Ìıa, M. Ì A.: The review mills, not just (self-)plagiarism in review reports, but a step further. Scientometrics129(9), 5805â5813 (2024) https://doi.org/10.1007/s11192-024-05125-w [56] Cooke, S.J., Young, N., Peiman, K.S., Roche, D.G., Clements, J.C., Kadykalo, A.N., Provencher, J.F., Raghavan, R., DeRosa, M.C., Lennox, R.J., Robinson Fayek, A., Cristescu, M.E., Murray, S.J., Quinn, J., Cobey, K.D., Browman, H.I.: A harm reduction approach to improving peer review by acknowledging its imper- fections. FACETS9(1), 1â14 (2024)https://doi.org/10.1139/facets-2024-0102 [57] Liang, W., Zhang, Y., Wu, Z., Lepp, H., Ji, W., Zhao, X., Cao, H., Liu, S., He, S., Huang, Z., Yang, D., Potts, C., Manning, C.D., Zou, J.Y.: Mappingthe increasing use of LLMs in scientific papers. In: First Conference onLanguage Modeling (2024).https://openreview.net/forum?id=YX7QnhxESU [58] Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, L., Polosukhin, I.: Attention is all you need. In: Proceedingsof the 31st International Conference on Neural Information Processing Systems. NIPSâ17, p. 6000â6010. Curran Associates Inc., Red Hook, NY, USA (2017) 46