Paper deep dive
Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing
Hyeonchu Park, Gahye Jeong, Bugeun Kim
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/28/2026, 4:00:36 AM
Summary
This study investigates whether AI text detectors flag non-native academic writing as AI-generated due to linguistic style rather than actual authorship. Using 135,389 pairs of original and professionally edited manuscripts, the authors found that editing significantly altered detector scores across 13 different AI detectors. The direction of change varied by detector type: token-statistics and some zero-shot detectors increased AI scores after editing, while classifier-based detectors often decreased them. Score changes correlated with editing intensity, identifying professional editing style as a key confounding variable that raises fairness and reliability concerns in academic settings.
Entities (17)
Relation Signals (16)
Professional Editing → affects → AI Detector Scores
confidence 95% · Professional editing leads to significant score shifts across nearly all detectors
LastDE+ → belongstocategory → Zero-shot
confidence 95% · Zero-shot-based ... LastDE+ Xu et al. (2025)
MAGE → belongstocategory → Classifier-based
confidence 95% · Classifier-based ... MAGE Li et al. ()
RoBERTa → belongstocategory → Classifier-based
confidence 95% · Classifier-based RoBERTa Solaiman et al. (2019a)
RADAR → belongstocategory → Classifier-based
confidence 95% · Classifier-based ... RADAR Hu et al. (2023)
BiScope → belongstocategory → Zero-shot
confidence 95% · Zero-shot-based ... BiScope Guo et al. (2024)
Binoculars → belongstocategory → Zero-shot
confidence 95% · Zero-shot-based ... Binoculars Hans et al. (2024)
DetectLLM-LRR → belongstocategory → Zero-shot
confidence 95% · Zero-shot-based ... DetectLLM-LRR Su et al. (2023)
DiVeye → belongstocategory →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:AI text detectors are increasingly employed in academic settings, but it remains unclear whether their outputs reflect AI authorship itself or broader linguistic features associated with polished academic English. Previous studies have reported high false-positive rates (FPRs) for non-native English writing, but population-level comparisons confound authorship with differences in topic, domain, and writing style. Professional editing provides a useful setting for examining this issue because it changes the linguistic form of manuscripts while preserving authorship and content. We examined 135,389 document pairs from a professional English editing service (2018-2025), comprising non-native manuscripts and their native-edited versions, to assess how editing affects detector responses controlling for content and authorship. For the 13 AI text detectors, FPRs for human-written texts varied widely, from 0.0% to 100.0%. Responses varied across detectors: the same edits increased AI scores in some detectors but decreased them in others. Notably, score changes correlated with the extent of editing. The findings identify professional editing style as a key confounding variable in AI detector outputs, rather than establishing a full separation of text origin from linguistic style, raising concerns about fairness and reliability in academic settings.
Tags
Links
- Source: https://arxiv.org/abs/2608.26710v1
- Canonical: https://arxiv.org/abs/2608.26710v1
Trouble viewing inline? Open PDF directly →
Full Text
62,436 characters extracted from source content.
Expand or collapse full text
Style as a Confound: False Positives in AI Detection of Non-Native Academic Writing Hyeonchu Park Gahye Jeong Bugeun Kim Department of Artificial Intelligence Chung-Ang University Republic of Korea Wordvice Republic of Korea phchu0429 bgnkim@cau.ac.kr gjeong@wordvice.com Abstract AI text detectors are increasingly employed in academic settings, but it remains unclear whether their outputs reflect AI authorship itself or broader linguistic features associated with polished academic English. Previous studies have reported high false-positive rates (FPRs) for non-native English writing, but population-level comparisons confound authorship with differences in topic, domain, and writing style. Professional editing provides a useful setting for examining this issue because it changes the linguistic form of manuscripts while preserving authorship and content. We examined 135,389 document pairs from a professional English editing service (2018–2025), comprising non-native manuscripts and their native-edited versions, to assess how editing affects detector responses controlling for content and authorship. For the 13 AI text detectors, FPRs for human-written texts varied widely, from 0.0% to 100.0%. Responses varied across detectors: the same edits increased AI scores in some detectors but decreased them in others. Notably, score changes correlated with the extent of editing. The findings identify professional editing style as a key confounding variable in AI detector outputs, rather than establishing a full separation of text origin from linguistic style, raising concerns about fairness and reliability in academic settings. 1 Introduction The rise of large language models(LLMs) has increased the use of AI text detectors in academia. While intended to uphold academic integrity, these systems often misclassify human-written text as AI-generated, raising concerns about false positives. Recent studies have found that English text written by non-native speakers is more likely to be labeled as AI-generated text(AIGT) than text written by native speakers Liang et al. (2023). This finding suggests that linguistic properties (e.g., fluency and lexical diversity) may affect detector results, rather than the origin alone. However, prior work has focused on comparing writer populations, which are challenging to interpret because they are confounded by differences in topics, domains, writing purposes, and individual author characteristics. This work addresses this limitation by treating professional academic English editing as a natural experiment. We compare original and professionally edited versions of manuscripts, holding content, topic, and authorship constant, and observe only the effects of linguistic revision. We evaluate 135,389 document pairs from a professional academic editing service and 13 representative AI text detectors. We examine (1) baseline FPRs on human-written academic texts, (2) effects of professional editing on detector results, (3) relationships between editing intensity and detector responses, and (4) whether these patterns differ between the pre-ChatGPT (2018–2022) and post-ChatGPT (2023–2025) periods. The results demonstrate variation among detectors: some assign lower AI scores after professional editing, whereas others assign higher scores to the same text, and score shifts correlate with editing intensity. However, these shifts reach conventional thresholds for practical significance in only 7 of the 13 detectors, indicating that statistical significance and practical relevance diverge across the detector landscape. Using a large-scale paired-document design, this study provides evidence that detector outputs are sensitive to linguistic style even when authorship and content are fixed, raising fairness and reliability concerns for affected detectors. 2 Related Work 2.1 Non-Native Writing Bias Recent studies indicate that AI text detectors exhibit systematic variation in performance across writer populations, with non-native English speakers often facing higher false-positive rates than native speakers Liang et al. (2023). This issue arises across various datasets, such as TOEFL essays and academic writing, where a significant amount of non-native speaker text is misclassified as AI-generated Liang et al. (2023); Weber-Wulff et al. (2023); Giray (2024). These disparities are linked to common features in non-native writing, including lower lexical diversity, simpler syntax, and increased repetition Ai and Lu (2013); Crossley and McNamara (2012); Engber (1995). Many of these traits resemble those of AIGT Liang et al. (2023); Fraser et al. (2025), raising concerns that detectors may focus on surface patterns rather than the text’s true origins. However, existing research mostly compares different writer groups, which limits understanding of the linguistic factors driving detector behavior and complicates analyses due to confounding variables like topic and writing style. As a result, higher detector scores for non-native writing may reflect differences in subject matter, disciplinary conventions, or individual writing habits rather than the effect of linguistic proficiency itself. In this work, we adopt a different perspective by examining how detector outputs vary when the same document is professionally edited. By comparing the original and edited versions, we aim to isolate the effects of linguistic refinement while keeping content and authorship constant, allowing for a clearer assessment of detector responses to text characteristics rather than origin. 2.2 AI-Generated Text Detection Methods for detecting AIGT can be grouped into three broad categories: token-statistic, zero-shot, and classifier approaches. Token-statistics-based detectors. Early approaches exploit statistical properties of token distributions produced by language models. Representative examples include Log-Likelihood, Log-Rank, Entropy, and GLTR Gehrmann et al. (2019), which rely on measures such as token probability, token rank, and prediction entropy. These methods are motivated by the observation that machine-generated text often exhibits more predictable token distributions than human writing. Zero-shot detectors. More recent methods infer text origin directly from language model behavior, without training dedicated classifiers. Some notable examples include DetectGPT Mitchell et al. (2023), Fast-DetectGPT Bao et al. (2024), DetectLLM-LRR Su et al. (2023), LastDE+ Xu et al. (2025), Binoculars Hans et al. (2024), and DiVeye Basani and Chen (2025). These methods use various signals, including probability curvature, likelihood ratios, entropy differences, and comparisons of probabilities across different models. Because they rely on language-model probabilities rather than task-specific training data, zero-shot detectors are often considered more adaptable to unseen generators and domains. However, their predictions still depend on linguistic properties reflected in model likelihoods, raising the possibility that writing proficiency and stylistic refinement may influence detector outputs. Classifier-based detectors. A third family consists of supervised classifiers trained on large collections of human-written and AIGT. Representative examples include RoBERTa-based detectors Solaiman et al. (2019a), MAGE Li et al. (), and RADAR Hu et al. (2023). These models learn decision boundaries from labeled training data and often achieve strong benchmark performance. Because their predictions are derived from patterns learned during training, classifier-based detectors may capture a broad mixture of signals related to both text origin and writing style. As a result, linguistic characteristics commonly associated with non-native writing may inadvertently become predictive cues during training, potentially contributing to the disparities reported in prior work. Although these detector families rely on different underlying principles, most evaluations have been conducted in benchmark settings that compare human-written and AI-generated corpora. Consequently, it remains unclear whether detector outputs primarily reflect the origin of generation or broader linguistic properties that happen to correlate with AIGT. This uncertainty is particularly important in light of the non-native writing bias discussed above, since linguistic features associated with writing proficiency may be mistaken for evidence of AI generation. Understanding which detector paradigms are most sensitive to such stylistic factors is therefore critical for evaluating the fairness and reliability of AI text detection systems. This work addresses this question by comparing detector results across three detection paradigms using professionally edited document pairs. This setting permits the assessment of how detector results change in response to controlled linguistic refinement while holding document content and authorship constant. 3 Method 3.1 Data Source We employed editing-record data from a professional academic English editing service11 1 https://wordvice.ai/. Each record comprises two versions of a document: the original text written by a non-native English speaker and the version professionally edited by a native-speaker editor. Editorial protocol. All editors follow a standardized internal editing-guidelines document (25 pages) that governs the editing process. The protocol is organized around three principles: (1) meaning preservation — editors correct grammar, punctuation, mechanics, clarity, naturalness, and word choice while preserving the author’s meaning; (2) human editorial judgment — stylistic revision, vocabulary selection, and sentence restructuring are performed exclusively by the human editor based on expertise, not by automated rewriting; and (3) restricted AI usage — company policy permits AI-based grammar-checking tools only as a supplementary aid for catching objective mechanical errors (e.g., subject-verb agreement, punctuation), while AI-generated rewriting, paraphrasing, or stylistic generation is explicitly prohibited. Dataset statistics. The final dataset consists of 135,389 original–edited document pairs collected between 2018 and 2025, spanning more than 40 academic disciplines. Table 1 summarizes the key statistics; the complete breakdown by year, service type, subject domain, and English variant is provided in Appendix B. Statistic Value Total document pairs 135,389 Mean word count (original) 2,355 Median word count (original) 950 Std. dev. (word count) 3,038 Subject areas 40+ Docs. w/ 500–5,000 words ∼ 57% Table 1: Summary dataset statistics. Full year-by-year and service-type breakdown in Appendix B Characteristics of professional edits. To characterize the edits themselves (independent of any detector), we compared surface linguistic properties of original and edited texts on a randomly sampled subset of 6,000 document pairs (Table 2). Professional editing introduces modest lexical and readability changes while largely preserving document length and structure: median document length changed by only 6.5 words. Editing intensity, measured in edit units, has a median of 196 (mean 463) per document, indicating that the edits are substantial enough to meaningfully affect wording while leaving overall content intact. Metric Orig. Edited Δ Mean word length 5.196 5.252 +0.055+0.055 Type–token ratio 0.407 0.412 +0.005+0.005 Mean sentence length 24.35 23.51 −0.84-0.84 Flesch–Kincaid grade 14.51 14.42 −0.09-0.09 Table 2: Linguistic characteristics of professional editing (6,000-pair subsample). Edit intensity: median 196 / mean 463 edit units per document. 3.2 AI Detectors We evaluate 13 publicly available AI text detectors representing diverse detection paradigms. These detectors are categorized into three groups: Token-statistics-based Log-rank Gehrmann et al. (2019), Log-likelihood Solaiman et al. (2019b), Entropy Gehrmann et al. (2019), and GLTR Gehrmann et al. (2019): These methods rely on statistical properties of token distributions, such as predictability, token rank, and entropy. Zero-shot-based Fast-DetectGPT Bao et al. (2024), LastDE+ Xu et al. (2025), DetectLLM-LRR Su et al. (2023), DiVeye Basani and Chen (2025), BiScope Guo et al. (2024), Binoculars Hans et al. (2024): These approaches estimate the likelihood of AI generation using language model probabilities or likelihood ratios across different language models. Classifier-based RoBERTa Solaiman et al. (2019a), MAGE Li et al. (), RADAR Hu et al. (2023): These detectors learn decision boundaries through supervised training on human-written and AIGT corpora. We evaluated all detectors in their publicly available, off-the-shelf configurations, without detector-specific fine-tuning or additional training on our dataset. This wide coverage enables a comparative analysis of how different detection paradigms respond to non-native English writing and whether certain approaches are more vulnerable to false positives. Appendix C describes all assessed detectors. 3.3 Experimental Design 3.3.1 Baseline False Positive Analysis We first evaluate the baseline false-positive behavior of each detector on the Pre-ChatGPT dataset. Since these documents were written before the widespread adoption of modern generative AI systems, they can be reasonably treated as human-written texts with negligible AI involvement. Consequently, any AI-generated classification is considered a false positive For each detector, we computed the FPR, accuracy, and F1-score to characterize baseline detector performance. For paired analyses of original and edited documents, we additionally report the change in false-positive rate (Δ ), defined as the difference in FPR between the edited and original versions, to quantify how professional editing alters binary detector decisions. 3.3.2 Effectiveness of Professional Editing Next, we investigate how professional editing affects detector outputs. Because each original-edited pair corresponds to the same document, differences in detection results can be attributed primarily to linguistic revisions rather than differences in topic, content, or authorship. Δscore=scoredit−scoreorig =score_edit-score_orig (1) For analyses involving temporal trends, score shifts were computed separately for the pre-ChatGPT (2018–2022) and post-ChatGPT (2023–2025) subsets. This approach examines whether the effects of professional editing changed the detector results after the widespread adoption of LLM-assisted writing. Negative values indicate that the detector assigns a lower AI score after editing, whereas positive values indicate an increase in AI likelihood. If detectors primarily capture text origin, score shifts should remain relatively small because both versions are human-written. In contrast, if detectors are sensitive to linguistic fluency or stylistic sophistication, systematic changes in detector outputs may emerge following professional editing. 3.3.3 Editing Intensity Analysis To investigate the potential mechanisms underlying detector responses, we measure editing intensity using the edit ratio: redit=|Cdel|+|Cadd||Corig|r_edit= |C_del|+|C_add||C_orig| (2) where CdelC_del, CaddC_add, and CorigC_orig are the number of deleted, added, and initial characters, respectively, in the document. To leverage information about the extent of human revision, we use the edit ratio rather than the normalized edit distance. Because normalized edit distance is based on the minimum number of editing operations required to transform one text into another, it may not accurately reflect the actual editing actions performed by professional editors. As a result, it can distort the magnitude of true editorial intervention. The edit ratio instead measures the volume of added and removed content, making it more suitable for quantifying editing intensity. 3.4 Statistical Analysis Differences between original and edited texts are evaluated using paired-sample t-tests, with effect sizes reported as Cohen’s d and 95% confidence intervals estimated via 1,000 bootstrap resamples. The paired design accounts for document-level variation by comparing each edited document directly with its original version. We assess the relationship between editing intensity and shifts in detector scores using Spearman’s rank correlations and linear regression models. Documents are additionally grouped into editing-intensity quintiles to examine dose–response patterns. Consistent monotonic trends across quintiles indicate systematic detector sensitivity to increasing levels of editorial intervention. For temporal analyses, outcomes are evaluated separately for the pre-ChatGPT (2018-2022) and post-ChatGPT (2023-2025) periods. We compare detector score shifts and false-positive rates across periods to examine whether detector behavior changed following the widespread adoption of LLM-assisted writing. Bonferroni and Benjamini-Hochberg corrections are applied where appropriate to account for multiple hypothesis testing. 4 Results 4.1 False-Positive Results Vary by Detector Detector FP% Δ p orig edit Token-statistics-based Log-rank 100.0 100.0 ++0.0 ns Log-likelihood 99.8 99.9 ++0.0 <<.05 Entropy 99.8 99.9 ++0.1 <<.001 GLTR 99.9 99.9 ++0.0 <<.01 Zero-shot-based Fast-DetectGPT 25.2 31.6 ++6.4 <<.001 LastDE+ 19.9 27.3 ++7.4 <<.001 DetectLLM-LRR 0.2 0.3 ++0.1 <<.01 DiVeye 0.4 0.6 ++0.2 <<.001 BiScope 0.6 0.3 −-0.4 <<.001 Binoculars 0.0 0.1 ++0.0 ns Classifier-based RoBERTa 93.1 93.9 ++0.8 <<.001 MAGE 16.7 7.6 −-9.1 <<.001 RADAR 88.3 86.4 −-1.9 <<.001 Table 3: Baseline false positive behavior on the Pre-ChatGPT dataset (2018–2022). FPR color: ≥ 50%, 10–49%. Δ color: bias reduction, bias amplification. Full results in Appendix D.3 We assess the baseline false-positive rates of AI detectors using a dataset from before the introduction of ChatGPT (2018–2022), assuming minimal influence from modern generative AI. Since all documents were created before the advent of LLMs, any classification indicating that a document is AI-generated is considered a false positive. Table 3 illustrates variability among the detectors, showing FPR that range from less than 1% for some detectors to nearly 100% for others when analyzing the same set of human-written academic texts. Detectors based on token statistics exhibited the highest false positive rates, often misclassifying the majority of documents as AI-generated. In contrast, zero-shot likelihood-based detectors maintained an FPR close to zero. Classifier-based methods presented mixed results. These inconsistencies suggest that detector performance is influenced by the methodology used: token predictability often mislabels non-native academic writing as AI-generated, while more conservative likelihood-ratio-based approaches tend to be more accurate. Overall, the discrepancies among detectors raise questions about what they are actually measuring. The same human-written documents can be classified by different systems as either predominantly AI-generated or human-written. Detectors may rely on various cues that prompt further analysis of how outputs change with controlled linguistic revisions through professional editing. 4.2 Editing Alters Detector Outputs ALL Pre-ChatGPT Post-ChatGPT Detector Type Δscore ΔFP Δscore ΔFP Δscore ΔFP Entropy Stat. ++0.014∗ ++0.1 ++0.017∗ ++0.1 ++0.005∗ ++0.0 GLTR Stat. ++0.010∗ ++0.0 ++0.012∗ ++0.0 ++0.004∗ ++0.0 Log-rank Stat. ++0.001ns ++0.0 ++0.001∗ ++0.0 ++0.001∗ ++0.0 Log-like Stat. ++0.005ns ++0.0 ++0.006∗ ++0.0 ++0.002∗ ++0.0 BiScope Zero −-0.019∗ −-0.5 −-0.021∗ −-0.6 −-0.008∗ −-0.1 Binoculars Zero ++0.002ns ++0.0 ++0.003∗ ++0.0 ++0.000ns ++0.0 DetectLLM-LRR Zero ++0.003∗ ++0.0 ++0.003∗ ++0.0 ++0.002∗ ++0.0 DiVeye Zero ++0.011∗ −-1.0 ++0.013∗ −-1.2 ++0.005ns −-0.3 Fast-DetectGPT Zero ++0.032∗ ++4.6 ++0.038∗ ++5.4 ++0.013∗ ++1.6 LastDE+ Zero ++0.027∗ ++5.8 ++0.031∗ ++6.9 ++0.012∗ ++2.0 MAGE Cls. −-0.108∗ −-10.7 −-0.130∗ −-13.0 −-0.031∗ −-3.1 RADAR Cls. −-0.037∗ −-4.7 −-0.044∗ −-5.6 −-0.013∗ −-1.5 RoBERTa Cls. ++0.017∗ ++1.7 ++0.018∗ ++1.8 ++0.014∗ ++1.3 Table 4: Effect of professional editing on AI detector scores and false positive rates. Δscore=scoreedit−scoreorig =score_edit-score_orig (mean score shift); ΔFP=FPedit−FPorig =FP_edit-FP_orig. Green = decrease; Red = increase. ∗p<.001^***p<.001, p∗∗<.01^**p<.01, ∗p<.05^*p<.05, pns≥.05^nsp≥.05. Full results in Appendix D.3 We next examine how professional editing affects detector outputs. Because each original–edited pair corresponds to the same document, observed score differences can be attributed primarily to linguistic revisions rather than differences in topic, content, or authorship. The results are presented in Table 4, with full results available in Appendix D.3. Professional editing led to significant score shifts across nearly all detectors, with varying directions and magnitudes. Some detectors, such as MAGE (Δ = −0.108-0.108, Δ = −10.7-10.7 p), RADAR (−0.037-0.037, −4.7-4.7), and BiScope (−0.019-0.019, −0.5-0.5), were less likely to classify texts as AI-generated post-editing. Conversely, detectors such as Fast-DetectGPT and LastDE+ showed notable increases in false positive rates, with Δ of +4.6+4.6 and +5.8+5.8 p, respectively. A notable pattern is that the same professional editing produced contradictory responses across detector categories. Classifier-based detectors (e.g., MAGE and RADAR) tended toward the human-written results, whereas most token-statistics-based detectors tended toward the AI-generated results. Zero-shot detectors displayed mixed results, with BiScope producing negative shifts and Fast-DetectGPT and LastDE+ generating positive shifts. Notably, these patterns were largely consistent across pre- and post-ChatGPT subsets. Although the effect magnitude was typically smaller in the post-ChatGPT period, detectors that became more permissive after editing in the full dataset tended to remain permissive, whereas detectors that became more restrictive continued to move in that direction. The findings indicate that detector disagreement extends beyond baseline FPRs. Even with the same human revision, detectors yield different interpretations, suggesting reliance on distinct linguistic signals rather than a unified definition of AIGT. 4.3 Responses Scale with Editing Intensity Detector Type ρ R2R^2 Q1 Q2 Q3 Q4 Q5 Entropy Stat. ++0.343∗ 0.037 −-.009 ++.015 ++.020 ++.027 ++.032 Log-rank Stat. ++0.293∗ 0.012 ++.001 ++.001 ++.002 ++.002 ++.002 Log-likelihood Stat. ++0.287∗ 0.005 ++.003 ++.005 ++.006 ++.007 ++.008 GLTR Stat. ++0.257∗ 0.023 ++.006 ++.011 ++.013 ++.016 ++.019 BiScope Zero −-0.418∗ 0.041 ++.009 −-.019 −-.026 −-.033 −-.040 Binoculars Zero −-0.134∗ 0.019 ++.028 −-.002 −-.003 −-.004 −-.004 DetectLLM-LRR Zero ++0.178∗ 0.013 ++.002 ++.003 ++.003 ++.004 ++.005 DiVeye Zero ++0.147∗ 0.012 ++.008 ++.012 ++.013 ++.016 ++.020 Fast-DetectGPT Zero ++0.026∗ 0.001 ++.025 ++.036 ++.032 ++.041 ++.055 LastDE+ Zero ++0.026∗ 0.001 ++.027 ++.030 ++.030 ++.032 ++.037 MAGE Cls. −-0.193∗ 0.028 −-.032 −-.093 −-.133 −-.184 −-.214 RADAR Cls. −-0.143∗ 0.033 −-.014 −-.031 −-.038 −-.048 −-.058 RoBERTa Cls. ++0.029∗ 0.003 ++.012 ++.018 ++.020 ++.024 ++.028 Table 5: Editing intensity and detector score shifts. Spearman ρ denotes the correlation between edit-token ratio and Δscore . Q1–Q5 represent quintiles of editing intensity (Q1 = weakest, Q5 = strongest). Green = score decreases; Red = score increases. ∗p<.001^***p<.001, p∗∗<.01^**p<.01, ∗p<.05^*p<.05. To understand the observed editing effects, we analyzed how detector responses correlate with the extent of editorial revision. If professional editing impacts detector outputs through linguistic refinement, detectors sensitive to editing should show greater score shifts with more extensive revisions. Table 5 reveals that detectors with decreased AI scores after editing exhibited negative correlations between edit ratio and score shift, while those with increased scores showed positive correlations. Notable examples include MAGE and BiScope, which had the largest score reductions and strong negative associations with editing intensity (ρ=−0.193ρ=-0.193 and ρ=−0.418ρ=-0.418). For MAGE, mean score shifts ranged from −0.032-0.032 in the lowest editing quintile (Q1) to −0.214-0.214 in the highest (Q5), indicating that more extensive revisions pushed documents closer to the human-written region. In contrast, detectors such as Entropy (ρ=+0.343ρ=+0.343) and Log-rank (ρ=+0.293ρ=+0.293) showed positive score shifts as editing intensity increased, suggesting that editing shifted texts toward the AI-generated region. Editing intensity did not change the direction of detector responses but amplified existing patterns. Detectors that viewed professional editing as evidence of human authorship became more permissive with increased revisions, while those that interpreted editing as a sign of AI generation became more restrictive. Although the explanatory power of editing intensity was modest (R2≤0.041R^2≤ 0.041), consistent trends across detector families indicate that outputs are systematically influenced by linguistic modifications. These findings suggest that AI detectors respond not only to the origin of generation but also to the extent of linguistic refinement through professional editing. The same qualitative trends are observed when the analysis is restricted to documents with substantial editorial revisions (Appendix D.2). 4.4 Linguistic Drivers of Detector Score Shifts The preceding results show that professional editing shifts detector scores in opposite directions across detector families (Section 4.2) and that these shifts scale with editing intensity (Section 4.3). We next examine which linguistic properties of the edits are associated with these shifts. Correlations with detector score shifts. We correlate editing characteristics with Δ for statistical, LLM-based zero-shot, and supervised detectors on the pre-ChatGPT subset (Table 6). Edit-token ratio shows the strongest association, with a positive correlation for statistical detectors (ρ=+0.286ρ=+0.286) but a negative correlation for supervised detectors (ρ=−0.089ρ=-0.089). Grammar and syntactic ratios exhibit the same qualitative pattern. Feature Statistical LLM-based Supervised Grammar ratio +0.111+0.111 +0.019+0.019 −0.005-0.005 Lexical ratio +0.042+0.042 −0.018-0.018 −0.020-0.020 Syntactic ratio +0.101+0.101 +0.013+0.013 −0.040-0.040 Edit-token ratio +0.286+0.286 +0.022+0.022 −0.089-0.089 Orig. word count −0.099-0.099 −0.002-0.002 +0.067+0.067 # edit units +0.116+0.116 +0.021+0.021 +0.005+0.005 Table 6: Spearman correlation (ρ) between editing characteristics and detector score shifts, by detector family (pre-ChatGPT subset). We additionally examined changes in linguistic complexity on a 6,000-document subsample (Table 7). Changes in function-word ratio show positive correlations across all detector families, whereas changes in type–token ratio show negative correlations. Associations with average word and sentence length are weaker and less consistent. These results suggest that detector score shifts are more closely associated with fluency- and syntax-related changes than with semantic content. Metric Statistical LLM-based Supervised TTR −0.063-0.063 −0.063-0.063 −0.073-0.073 Word length −0.022-0.022 −0.046-0.046 −0.082-0.082 Function-word ratio +0.071+0.071 +0.068+0.068 +0.064+0.064 Sentence length −0.043-0.043 +0.005+0.005 +0.017+0.017 Table 7: Spearman correlation between changes in complexity and detector score shifts, by detector family. Interpretation. Together with the practical-significance analysis in Section 4.2, these results suggest systematic differences across detector families. Statistical detectors show increasing AI-likelihood with greater editing volume, potentially because professional editing reduces token-level perplexity and increases linguistic predictability. Supervised classifiers show the opposite trend, consistent with their reliance on higher-level patterns learned from AI-generated text. LLM-based zero-shot detectors exhibit comparatively weak and mixed relationships. These interpretations are correlational and therefore do not establish causal effects of any individual linguistic feature. 5 Conclusion This study examined the impact of professional academic editing on AI text-detection behavior using 135,389 pairs of original and edited documents from an academic editing platform. By comparing pre- and post-editing versions, we isolated the effects of linguistic changes while controlling for topic, content, and authorship. Results showed significant variability among detectors, with false-positive rates differing widely for the same human-written texts. Editing affected AI-generation scores in varying ways, with some detectors showing lower scores post-edit and others showing higher scores for the same edits. There was also a correlation between detector responses and editing intensity, suggesting a link to linguistic revision rather than random variation. These findings suggest that AI detectors do not represent a unified measure of AI-generated content; rather, they rely on a mix of generation-origin and linguistic-quality signals. Token-statistics-based detectors had high false-positive rates and were often more likely to classify edited texts as AI-generated, whereas likelihood-ratio-based zero-shot detectors showed lower false-positive rates and were more robust to editing. Temporal analysis indicated that the influence of professional editing weakened in the post-ChatGPT period, suggesting contemporary writing increasingly resembles the linguistic characteristics used for detector training. This highlights the need to evaluate detectors under real-world writing conditions. In practice, AI detector outputs should not be treated as definitive evidence of AI use, especially in academic settings where editing is common. Detector scores may indicate linguistic refinement rather than AI generation, so they should be considered alongside contextual information. The main challenge in AI text detection is improving accuracy while differentiating AI-generated content from enhanced human text. Developing these detectors is essential for ensuring fairness and reliability. 6 Limitations Although this study provides large-scale evidence that professional human editing can influence AI detector outputs, several limitations warrant acknowledgment. First, although our linguistic feature analysis identifies several types of revisions and changes in linguistic complexity associated with shifts in detector scores, these analyses do not fully explain the mechanisms underlying detector behavior. In particular, score changes may reflect the combined effects of lexical diversity, syntactic complexity, discourse organization, predictability, and other stylistic properties. Moreover, the observed associations should not be interpreted as evidence that any single linguistic feature causally determines detector outputs. Thus, our analysis provides a more interpretable characterization of the linguistic factors associated with detector responses, while a complete mechanistic explanation remains an important direction for future work. Second, the dataset predominantly comprises academic manuscripts authored by non-native English-speaking researchers. Although this represents a practically important setting for studying the interaction between linguistic polishing and AI detection, the findings may not generalize to other writing domains. Detector behavior may differ for student essays, journalistic articles, creative writing, social media content, and other genres with distinct linguistic and stylistic characteristics. Third, the paired-document design focuses on human-authored manuscripts and their professionally edited versions. This design lets us examine how professional linguistic refinement relates to detector outputs, but it does not fully separate text origin from writing style. In particular, the study does not include AI-generated control texts subjected to comparable editing procedures. Therefore, we cannot determine whether the observed effects generalize to AI-generated text or whether professional editing affects human- and AI-generated texts in systematically different ways. Future work should incorporate matched human- and AI-generated texts, together with controlled levels and types of human editing, to examine these interactions more directly. Fourth, interpreting absolute false-positive rates depends on the thresholding and per-detector normalization procedures used in the evaluation. Although we apply a consistent protocol to enable comparisons across detectors, different threshold choices may produce different absolute FPR estimates. Accordingly, interpret the absolute FPR values in light of the adopted thresholding protocol, while the relative patterns across editing conditions provide more robust evidence of detector sensitivity to linguistic refinement. Finally, the AI detection field continues to evolve rapidly. We evaluated 13 representative publicly available detectors spanning multiple detection paradigms, but the analysis does not cover future systems or proprietary commercial detectors. In addition, we evaluated all detectors off-the-shelf without detector-specific fine-tuning. Therefore, the findings should be interpreted as evidence of a broader phenomenon—that professional linguistic refinement can act as an important confounding factor in AI detector evaluations—rather than as definitive assessments of any individual detector or as evidence that editing alone determines detector decisions. Despite these limitations, this study provides large-scale evidence from more than 135,000 paired documents in a real-world academic editing environment. The results highlight the importance of accounting for professional editing and stylistic refinement when evaluating the fairness and reliability of AI text detectors. Future research incorporating AI-generated controls, broader writing domains, and controlled linguistic interventions can further clarify how text origin, linguistic style, and editing interact to shape detector behavior. The Use of Large Language Models We used AI-assistance tools during the writing of this manuscript. Specifically, we employed Claude-Sonnet 4.6 and Grammarly to polish language and improve clarity of expression. Acknowledgments This research was supported by Basic Science Research Program through the National Research Foundation of Korea(NRF) funded by the Ministry of Education (RS-2025-25434151) and the Institute of Information & Communications Technology Planning & Evaluation (IITP) grant funded by the Korea government (MSIT) [RS-2021-I211341, Artificial Intelligence Graduate School Program (Chung-Ang University)]. References Ai and Lu (2013) H. Ai and X. Lu A corpus-based comparison of syntactic complexity in nns and ns university students’ writing. Automatic treatment and analysis of learner corpus data 59. Cited by: §2.1. Bao et al. (2024) G. Bao, Y. Zhao, Z. Teng, L. Yang, and Y. Zhang Fast-detectgpt: efficient zero-shot detection of machine-generated text via conditional probability curvature. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024, External Links: Link Cited by: §2.2, §3.2. Basani and Chen (2025) A. R. Basani and P. Chen Diversity boosts ai-generated text detection. arXiv preprint arXiv:2509.18880. Cited by: §2.2, §3.2. Crossley and McNamara (2012) S. A. Crossley and D. S. McNamara Predicting second language writing proficiency: the roles of cohesion and linguistic sophistication. Journal of Research in Reading 35 (2), p. 115–135. Cited by: §2.1. Engber (1995) C. A. Engber The relationship of lexical proficiency to the quality of esl compositions. Journal of second language writing 4 (2), p. 139–155. Cited by: §2.1. Fraser et al. (2025) K. C. Fraser, H. Dawkins, and S. Kiritchenko Detecting ai-generated text: factors influencing detectability with current methods. Journal of Artificial Intelligence Research 82, p. 2233–2278. Cited by: §2.1. Gehrmann et al. (2019) S. Gehrmann, H. Strobelt, and A. Rush GLTR: statistical detection and visualization of generated text. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: System Demonstrations, M. R. Costa-jussà and E. Alfonseca (Eds.), Florence, Italy, p. 111–116. External Links: Link, Document Cited by: §2.2, §3.2. Giray (2024) L. Giray The problem with false positives: ai detection unfairly accuses scholars of ai plagiarism. The Serials Librarian 85 (5-6), p. 181–189. Cited by: §2.1. Guo et al. (2024) H. Guo, S. Cheng, X. Jin, Z. Zhang, K. Zhang, G. Tao, G. Shen, and X. Zhang Biscope: ai-generated text detection by checking memorization of preceding tokens. Advances in Neural Information Processing Systems 37, p. 104065–104090. Cited by: §3.2. Hans et al. (2024) A. Hans, A. Schwarzschild, V. Cherepanova, H. Kazemi, A. Saha, M. Goldblum, J. Geiping, and T. Goldstein Spotting llms with binoculars: zero-shot detection of machine-generated text. arXiv preprint arXiv:2401.12070. Cited by: §2.2, §3.2. Hu et al. (2023) X. Hu, P. Chen, and T. Ho RADAR: robust ai-text detection via adversarial learning. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, A. Oh, T. Naumann, A. Globerson, K. Saenko, M. Hardt, and S. Levine (Eds.), External Links: Link Cited by: §2.2, §3.2. [12] Y. Li, Q. Li, L. Cui, W. Bi, Z. Wang, L. Wang, L. Yang, S. Shi, and Y. Zhang Mage: machine-generated text detection in the wild, 2024. URL: https://arxiv. org/abs/2305.13242. Cited by: §2.2, §3.2. Liang et al. (2023) W. Liang, M. Yuksekgonul, Y. Mao, E. Wu, and J. Zou GPT detectors are biased against non-native english writers. Patterns 4 (7). Cited by: §1, §2.1, §2.1. Mitchell et al. (2023) E. Mitchell, Y. Lee, A. Khazatsky, C. D. Manning, and C. Finn DetectGPT: zero-shot machine-generated text detection using probability curvature. In Proceedings of the 40th International Conference on Machine Learning, A. Krause, E. Brunskill, K. Cho, B. Engelhardt, S. Sabato, and J. Scarlett (Eds.), Proceedings of Machine Learning Research, Vol. 202, p. 24950–24962. External Links: Link Cited by: §2.2. Solaiman et al. (2019a) I. Solaiman, M. Brundage, J. Clark, A. Askell, A. Herbert-Voss, J. Wu, A. Radford, G. Krueger, J. W. Kim, S. Kreps, M. McCain, A. Newhouse, J. Blazakis, K. McGuffie, and J. Wang Release strategies and the social impacts of language models. External Links: 1908.09203, Link Cited by: §2.2, §3.2. Solaiman et al. (2019b) I. Solaiman, M. Brundage, J. Clark, A. Askell, A. Herbert-Voss, J. Wu, A. Radford, and J. Wang Release strategies and the social impacts of language models. arXiv preprint arXiv:1908.09203. Cited by: §3.2. Su et al. (2023) J. Su, T. Zhuo, D. Wang, and P. Nakov Detectllm: leveraging log rank information for zero-shot detection of machine-generated text. In Findings of the Association for Computational Linguistics: EMNLP 2023, p. 12395–12412. Cited by: §2.2, §3.2. Weber-Wulff et al. (2023) D. Weber-Wulff, A. Anohina-Naumeca, S. Bjelobaba, T. Foltỳnek, J. Guerrero-Dib, O. Popoola, P. Šigut, and L. Waddington Testing of detection tools for ai-generated text. International Journal for Educational Integrity 19 (1), p. 1–39. Cited by: §2.1. Xu et al. (2025) Y. Xu, Y. Wang, Y. Bi, H. Cao, Z. Lin, Y. Zhao, and F. Wu Training-free llm-generated text detection by mining token probability sequences. In International Conference on Learning Representations, Vol. 2025, p. 19072–19098. Cited by: §2.2, §3.2. Appendix A Environments All experiments were conducted using Python 3.10 with statsmodels 0.14 and scipy 1.11. The experimental platform consisted of an AMD Ryzen Threadripper 3960X 24-Core Processor and four NVIDIA RTX A6000 GPUs. The GPUs were used for running detector models that required GPU acceleration. Appendix B Dataset Item Value Overall Scale Total order pairs 135,389 Collection period 2018–2025 (8 years) Era Split Pre-ChatGPT (2018–2022) 104,752 (77.4%) Post-ChatGPT (2023–2025) 30,637 (22.6%) Service Type Academic editing 71,595 (52.9%) Admission editing 50,971 (37.6%) Business editing 6,252 (4.6%) TOEFL writing 3,531 (2.6%) Translation 2,895 (2.1%) Other 145 (0.1%) Document Length Mean characters 16,573 (median: 6,430) Mean words 2,355 (median: 950) Min–Max 1–610,964 characters Editing Characteristics Median edit ratio (reditr_edit) 0.175 (17.5%) Edit ratio IQR 0.097–0.269 Outliers removed (redit>5r_edit>5) 417 (0.3%) Zero-edit orders (nchanges=0n_changes=0) 5,916 (4.4%) Language & Domain English style en-US 88.8%, en-GB 11.2% Top domains (top 3) Academic Papers (Application Essays, 22,639) Academic Papers (Medicine, 14,125) Academic Papers (Engineering & Technology, 13,201) Table 8: Dataset Statistics. The pre-ChatGPT / post-ChatGPT split is defined by the public release of ChatGPT (November 2022); orders from 2023 onward are classified as post-ChatGPT. Several preprocessing steps were applied to the raw editing records. We removed cases with missing original or edited texts, documents containing severe encoding errors, and extremely short documents for which detector outputs were unstable. After preprocessing, the final dataset consisted of 135,389 original–edited document pairs. Table 8 presents the statistics of the dataset. Each pair includes the original text, the edited text, the service type, the submission year, the English variant (en-US or en-GB), the subject domain, and the editing ratio (reditr_edit). The editing ratio is defined as the proportion of modified tokens to the total number of tokens in the original document. We excluded 417 cases (0.3%) with redit>5r_edit>5, treating them as outliers. In addition, 5,916 orders (4.4%) involved no textual modifications (nchanges=0n_changes=0). These cases were handled separately in analyses of editing effects. The dataset is dominated by the Pre-ChatGPT period (2018–2022), which contains 104,752 document pairs (77.4%), while the Post-ChatGPT period (2023–2025) contains 30,637 pairs (22.6%). With respect to service type, Academic Editing accounts for the largest share of the dataset with 71,595 documents (52.9%), followed by Admission Editing with 50,971 documents (37.6%) and Business Editing with 6,252 documents (4.6%). The three largest subject domains are Application Essays (22,639 documents), Medicine (14,125 documents), and Engineering & Technology (13,201 documents), which together account for approximately 36.9% of the dataset. Edit-type composition. We categorized edits on a representative subsample into grammar, lexical, syntactic, and mixed operations (Table 9). Mixed edits dominate (66.1%), followed by grammar (15.8%) and lexical (14.0%) edits, with purely syntactic edits rare (0.1%). Across all 13 detectors, score shifts differ significantly by edit category (Kruskal–Wallis, p<10−15p<10^-15). Table 10 illustrates this for RoBERTa: syntactic edits produce the largest positive score shift (dz=+0.225d_z=+0.225), while grammar edits produce a small negative shift (dz=−0.023d_z=-0.023). Edit type Share Mixed 66.1% Grammar 15.8% Lexical 14.0% Syntactic 0.1% no-edit 0.4% Table 9: Distribution of edit-operation categories. Edit type Δ Cohen’s dzd_z Syntactic +0.046+0.046 +0.225+0.225 Mixed +0.015+0.015 +0.073+0.073 Lexical +0.011+0.011 +0.056+0.056 Grammar −0.004-0.004 −0.023-0.023 Table 10: RoBERTa score shift by edit-operation category. Appendix C Baseline Models Detector Category Core Signal Detection Principle Token-statistics-based Detectors Log-likelihood Statistical Average likelihood Uses mean token log-likelihood Log-rank Statistical Token rank Uses average token rank under the language model Entropy Statistical Token entropy Measures uncertainty in token predictions GLTR Statistical Rank histogram Analyzes token-rank distributions across the document Zero-shot Detectors Fast-DetectGPT Zero-shot Probability curvature Measures local curvature of model likelihood DetectLLM-LRR Zero-shot Log-rank ratio Compares token-rank statistics under the source model LastDE+ Zero-shot Local entropy difference Measures entropy changes across neighboring contexts DiVeye Zero-shot Diversity–predictability imbalance Detects inconsistencies between lexical diversity and predictability BiScope Hybrid Bidirectional consistency Combines forward and backward likelihood signals Binoculars Zero-shot Cross-model likelihood ratio Compares token probabilities assigned by two language models Classifier-based Detectors RoBERTa Classifier Contextual representations Binary classifier fine-tuned on human vs. AI text MAGE Classifier Multi-model linguistic features Multi-task detector trained on outputs from multiple LLMs RADAR Classifier Adversarially robust features Detector trained against paraphrasing and rewriting attacks Table 11: AI detectors evaluated in this study. The detectors span three major detection paradigms and rely on different linguistic signals, enabling analysis of whether detector responses to professional editing vary across methodological families. We evaluate 13 publicly available AI text detectors across three major detection paradigms: token-statistics-based, zero-shot, and classifier-based. These detectors were selected because they represent widely used methodologies in the AIGT detection literature and rely on substantially different decision mechanisms. Token-statistics-based detectors operate directly on low-level properties of language-model outputs, such as token likelihood, token rank, and entropy. Zero-shot detectors estimate generation origin using language-model probabilities without training dedicated classifiers, often leveraging likelihood ratios, probability curvature, or consistency-based measures. Classifier-based detectors, in contrast, learn decision boundaries from labeled human-written and AI-generated texts through supervised training. We intentionally include detectors that have been reported to achieve strong performance on benchmark datasets, covering both classical statistical methods and recent LLM-era approaches. This methodological diversity allows us to investigate whether detector sensitivity to professional editing is associated with the underlying detection paradigm rather than the performance of any individual detector. Because the three detector families rely on different linguistic signals, they provide a useful test bed for examining whether responses to professional editing are detector-specific or consistent across methodological paradigms. Table 11 summarizes the detectors evaluated in this study together with their core detection signals and underlying detection principles. Detector Public checkpoint Fine-tuned RoBERTa openai-community/ No roberta-base-openai-detector MAGE yaful/MAGE No RADAR TrustSafeAI/RADAR-Vicuna-7B No Table 12: Classifier-based detector checkpoints. All models were loaded via AutoModelForSequenceClassification.from_pretrained() and evaluated in inference mode (eval() with torch.no_grad()), with no additional fine-tuning or adaptation on our dataset. C.0.1 Thresholding and Normalization Protocol Because the 13 detectors produce heterogeneous raw outputs (e.g., probabilities, log-likelihoods, entropy, ranking statistics, or perplexity-based scores), we converted each detector’s raw score to a common [0,1][0,1] AI-likelihood scale following its original formulation or a commonly used normalization, and applied a single fixed decision threshold of 0.5 across all detectors and both document versions (Table 13). We intentionally did not perform per-detector threshold optimization on our evaluation data: the goal was to compare detector behavior under one consistent operating point rather than to maximize the reported performance of any individual detector. In Table 2/Table 9, Δscore=scoreedit−scoreorig =score_edit-score_orig refers to this normalized [0,1][0,1] AI-likelihood score, not the detector’s raw statistic; Δ refers to the corresponding change in the binary decision under the 0.5 threshold. Because both the original and the edited version of a document are scored under the identical decision rule, the paired Δ and Δ estimates are invariant to the specific choice of threshold, even though the absolute FPR values in Table 1 are not; we discuss this dependency further as a limitation in Section 6. Detector Raw score Normalization RoBERTa Prob. [0,1][0,1] As-is RADAR Prob. [0,1][0,1] As-is MAGE Prob. [0,1][0,1] As-is Binoculars CE ratio σ((r−1.0)×5.0)σ((r-1.0)× 5.0) Fast-DetectGPT Log-prob. diff. σ(δ×10.0)σ(δ× 10.0) Log-Likelihood Mean log-lik. σ((L+5.0)×2.0)σ((L+5.0)× 2.0) Log-Rank Mean log-rank σ((LR−5.5)×−1.0)σ((LR-5.5)×-1.0) Entropy Mean entropy σ((Ent−4.0)×−1.0)σ((Ent-4.0)×-1.0) GLTR Raw statistic As-is DetectLLM-LRR Log-rank ratio σ((LRR−1.5)×2.0)σ((LRR-1.5)× 2.0) LastDE+ Sampling discrep. σ(d×0.5)σ(d× 0.5) DiVeye Var. of surprisal σ(−(V−5.0)×0.4)σ(-(V-5.0)× 0.4) BiScope Backward CE σ(−(BCE−9.0)×0.4)σ(-(BCE-9.0)× 0.4) Table 13: Per-detector raw score and normalization to a [0,1][0,1] AI-likelihood scale; σ(⋅)σ(·) denotes the logistic sigmoid. Anchor values follow each detector’s original formulation or representative values reported in prior work. Appendix D Ablation study D.1 Detector Sensitivity Weakens in the post-ChatGPT Era We compare detector behavior between the pre-ChatGPT (2018–2022) and post-ChatGPT (2023–2025) periods. The results are presented in Table 14. Detector Type Orig Score Δ orig FP% (orig) Δ pre post Δscore d pre post Entropy Stat. 0.76 0.79 ++0.03 ++0.42∗ 99.6 99.9 ++0.3 GLTR Stat. 0.74 0.76 ++0.02 ++0.28∗ 99.7 99.8 ++0.1 Log-likelihood Stat. 0.97 0.98 ++0.01 ++0.14∗ 99.8 99.9 ++0.1 Log-Rank Stat. 0.98 0.98 ++0.00 ++0.17∗ 100.0 100.0 ++0.0 BiScope Zero 0.35 0.31 −-0.04 −-0.50∗ 1.9 1.1 −-0.9 Binoculars Zero 0.39 0.38 −-0.01 −-0.15∗ 0.1 0.0 −-0.0 DetectLLM-LRR Zero 0.35 0.38 ++0.02 ++0.62∗ 0.1 0.4 ++0.3 DiVeye Zero 0.21 0.21 ++0.01 ++0.06∗ 2.1 1.0 −-1.1 Fast-DetectGPT Zero 0.32 0.33 ++0.01 ++0.02∗ 24.7 25.8 ++1.1 LastDE+ Zero 0.35 0.36 ++0.01 ++0.03∗ 19.5 22.1 ++2.6 MAGE Cls. 0.22 0.10 −-0.12 −-0.31∗ 21.9 10.1 −-11.8 RADAR Cls. 0.69 0.69 ++0.00 ns 75.8 76.6 ++0.8 RoBERTa Cls. 0.92 0.96 ++0.04 ++0.18∗ 92.7 96.4 ++3.7 Table 14: Comparison of original-text detector scores between the Pre-ChatGPT and Post-ChatGPT periods. Pre-ChatGPT corresponds to 2018–2022 (npre=104,752n_pre=104,752), and Post-ChatGPT corresponds to 2023–2025 (npost=30,637n_post=30,637). FPR color: ≥ 50%, 10–49%. Δ color: decrease, increase. ∗p<.001^***p<.001, p∗∗<.01^**p<.01, ∗p<.05^*p<.05, pns≥.05^nsp≥.05. Across many detectors, the impact of professional editing becomes substantially weaker in the post-ChatGPT era. For example, the average shift in MAGE decreases from -0.130 in the pre-ChatGPT period to -0.031 in the post-ChatGPT period, representing a reduction of approximately 76%. RADAR exhibits a similar pattern. Likewise, detectors that previously showed positive responses to editing also exhibit attenuated effects after 2023. The magnitude of editing-induced score increases observed in Fast-DetectGPT and LastDE+ becomes considerably smaller during the post-ChatGPT period. At the same time, several detectors assign higher AI scores to original human-written texts collected after the widespread adoption of LLMs. For example, the FPR of RoBERTa increases from 92.7% to 96.4%, while DetectLLM-LRR also shows a significant increase in baseline scores. These results suggest that detector behavior is not static but evolves alongside broader changes in writing practices. As highly polished language becomes increasingly common in the post-ChatGPT era, the distinction between professionally edited human writing and detector-internal representations of AIGT may become less pronounced. Consequently, the incremental effect of professional editing appears to diminish over time. D.2 Results on the Strong Editing Subset Detector Type FP% (heavy) Shift Cohen’s d Δd d orig edit full heavy Entropy Stat. 99.6 99.8 ++0.025 ++0.444 ++0.707 ++0.263 GLTR Stat. 99.7 99.8 ++0.017 ++0.335 ++0.576 ++0.242 Log-Likelihood Stat. 99.6 99.6 ++0.007 ++0.190 ++0.197 ++0.006 Log-Rank Stat. 100.0 100.0 ++0.002 ++0.292 ++0.299 ++0.007 BiScope Zero 2.2 1.3 −-0.031 −-0.488 −-0.833 −-0.345 Binoculars Zero 0.1 0.1 −-0.004 ++0.063 −-0.227 −-0.289 DetectLLM-LRR Zero 0.1 0.2 ++0.005 ++0.172 ++0.329 ++0.158 DiVeye Zero 0.7 1.0 ++0.021 ++0.157 ++0.333 ++0.176 Fast-DetectGPT Zero 24.8 29.6 ++0.035 ++0.213 ++0.174 −-0.039 LastDE+ Zero 20.0 26.2 ++0.028 ++0.275 ++0.225 −-0.049 MAGE Cls. 25.3 9.4 −-0.160 −-0.356 −-0.401 −-0.045 RADAR Cls. 73.1 66.0 −-0.056 −-0.268 −-0.311 −-0.043 RoBERTa Cls. 93.0 95.0 ++0.023 ++0.085 ++0.107 ++0.021 Table 15: Heavy-editing subset analysis (redit>0.20r_edit>0.20, n=57,207n=57,207). Δd=dheavy−dfull d=d_heavy-d_full. All comparisons significant at p<.001p<.001. Type: Stat. = token-statistics-based; Zero = zero-shot-based; Cls. = classifier-based. FPR color: ≥ 50%, 10–49%. Color: decrease, increase. In Section 4, we observed a clear dose–response relationship between editing intensity and detector responses. To assess whether this pattern is driven by a small number of heavily edited documents or diluted by the broader dataset, we conduct an additional analysis focusing exclusively on documents that underwent substantial revisions. Specifically, documents with an edit ratio (reditr_edit) greater than 0.20 were classified as heavy-editing cases. This subset contains 57,207 document pairs. For each detector, we computed the pre- and post-editing false positive rates (FPR), average score shifts, and Cohen’s d, and compared these results with those obtained from the full dataset. Table 15 summarizes the results. Overall, detector responses under heavy editing tended to amplify the patterns observed in the full dataset. BiScope exhibited the largest reduction effect, with Cohen’s d increasing from −0.488-0.488 in the full dataset to −0.833-0.833 in the heavy-editing subset. Similar increases were observed for MAGE and RADAR. In particular, the false positive rate of MAGE decreased from 25.3% to 9.4%, indicating that the detector increasingly perceived heavily edited texts as human-written. Conversely, detectors such as Entropy and GLTR exhibited stronger positive editing effects under heavy-editing conditions. For Entropy, the effect size increased from d=0.444d=0.444 in the full dataset to d=0.707d=0.707, while GLTR increased from d=0.335d=0.335 to d=0.576d=0.576. DiVeye and DetectLLM-LRR also showed larger positive effect sizes compared with the full dataset. These results suggest that extensive linguistic revisions can make human-written texts appear more AI-generated according to certain detectors. In contrast, RoBERTa, Fast-DetectGPT, LastDE++, Log-Likelihood, and Log-Rank exhibited relatively small changes in effect size compared with the full dataset. For Log-Likelihood and Log-Rank, this pattern may be attributable to saturation effects, as both detectors had already produced false-positive rates exceeding 99% before editing. Consequently, additional linguistic modifications had limited room to further influence detector outputs. Taken together, the heavy-editing subset analysis provides additional evidence that the relationship between editing intensity and detector response is not a statistical artifact. For most detectors exhibiting editing effects, the magnitude of those effects became substantially larger under heavy-editing conditions, while the direction of change remained consistent with the patterns observed in the full dataset. These findings further support the interpretation that AI text detectors respond not only to the origin of the text but also to the degree of linguistic refinement introduced by human editing. D.3 Full results Original Edited Δ Detector Type Mean FP% F1 Mean FP% F1 Δ Δ 1 p Log-rank Stat. 0.982 100.0 0.000 0.984 100.0 0.000 ++0.0 — ns Log-likelihood Stat. 0.981 99.8 0.003 0.985 99.9 0.003 ++0.0 — <<.05 Entropy Stat. 0.781 99.8 0.003 0.798 99.9 0.002 ++0.1 — <<.001 GLTR Stat. 0.754 99.9 0.002 0.769 99.9 0.002 ++0.0 — <<.01 Fast-DetectGPT Zero 0.329 25.2 0.856 0.375 31.6 0.812 ++6.4 −-.044 <<.001 LastDE+ Zero 0.358 19.9 0.890 0.392 27.3 0.842 ++7.4 −-.048 <<.001 DetectLLM-LRR Zero 0.366 0.2 0.999 0.369 0.3 0.999 ++0.1 — <<.01 DiVeye Zero 0.196 0.4 0.998 0.218 0.6 0.997 ++0.2 — <<.001 BiScope Zero 0.325 0.6 0.997 0.302 0.3 0.999 −-0.4 ++.002 <<.001 Binoculars Zero 0.384 0.0 1.000 0.382 0.1 1.000 ++0.0 — ns RoBERTa Cls. 0.923 93.1 0.129 0.931 93.9 0.115 ++0.8 −-.014 <<.001 MAGE Cls. 0.168 16.7 0.909 0.077 7.6 0.961 −-9.1 ++.052 <<.001 RADAR Cls. 0.800 88.3 0.209 0.783 86.4 0.239 −-1.9 ++.030 <<.001 Table 16: Baseline false positive behavior on the Pre-ChatGPT dataset (2018–2022). FPR color: ≥ 50%, 10–49%. ALL (n=135,389n=135,389) Pre-ChatGPT (n=104,752n=104,752) Post-ChatGPT (n=30,637n=30,637) Detector Type Δscore d Δscore d Δscore d Entropy Stat. ++0.014∗ ++0.40 ++0.017∗ ++0.44 ++0.005∗ ++0.21 GLTR Stat. ++0.010∗ ++0.30 ++0.012∗ ++0.34 ++0.004∗ ++0.16 Log-rank Stat. ++0.001ns ++0.24 ++0.001∗ ++0.29 ++0.001∗ ++0.09 Log-like Stat. ++0.005ns ++0.16 ++0.006∗ ++0.19 ++0.002∗ ++0.06 BiScope Zero −-0.019∗ −-0.45 −-0.021∗ −-0.49 −-0.008∗ −-0.22 Binoculars Zero ++0.002ns ++0.05 ++0.003∗ ++0.06 ++0.000ns — DetectLLM-LRR Zero ++0.003∗ ++0.16 ++0.003∗ ++0.17 ++0.002∗ ++0.12 DiVeye Zero ++0.011∗ ++0.15 ++0.013∗ ++0.16 ++0.005ns — Fast-DetectGPT Zero ++0.032∗ ++0.18 ++0.038∗ ++0.21 ++0.013∗ ++0.08 LastDE+ Zero ++0.027∗ ++0.24 ++0.031∗ ++0.28 ++0.012∗ ++0.11 MAGE Cls. −-0.108∗ −-0.31 −-0.130∗ −-0.36 −-0.031∗ −-0.13 RADAR Cls. −-0.037∗ −-0.24 −-0.044∗ −-0.27 −-0.013∗ −-0.10 RoBERTa Cls. ++0.017∗ ++0.09 ++0.018∗ ++0.09 ++0.014∗ ++0.10 Table 17: Full score shift results following professional editing. Δscore=scoreedit−scoreorig =score_edit-score_orig. Cohen’s d denotes effect size. Green = decrease; Red = increase. ∗p<.001^***p<.001, p∗∗<.01^**p<.01, ∗p<.05^*p<.05, pns≥.05^nsp≥.05. ALL (n=135,389n=135,389) Pre-ChatGPT (n=104,752n=104,752) Post-ChatGPT (n=30,637n=30,637) Detector Type Orig Edit ΔFP Orig Edit ΔFP Orig Edit ΔFP Log-rank Stat. 100.0 100.0 ++0.0 100.0 100.0 ++0.0 100.0 100.0 ++0.0 Log-like Stat. 99.7 99.7 ++0.0 99.6 99.7 ++0.0 99.7 99.7 ++0.0 Entropy Stat. 99.8 99.9 ++0.1 99.8 99.9 ++0.1 99.9 99.9 ++0.0 GLTR Stat. 99.8 99.9 ++0.0 99.8 99.9 ++0.0 99.9 99.9 ++0.0 BiScope Zero 1.7 1.2 −-0.5 1.9 1.3 −-0.6 1.1 1.0 −-0.1 Binoculars Zero 0.1 0.1 ++0.0 0.1 0.1 ++0.0 0.0 0.0 ++0.0 DetectLLM-LRR Zero 0.2 0.2 ++0.0 0.1 0.2 ++0.0 0.4 0.4 ++0.0 DiVeye Zero 1.9 0.8 −-1.0 2.1 0.9 −-1.2 1.0 0.7 −-0.3 Fast-DetectGPT Zero 24.9 29.5 ++4.6 24.7 30.1 ++5.4 25.8 27.4 ++1.6 LastDE+ Zero 20.1 25.9 ++5.8 19.5 26.4 ++6.9 22.1 24.1 ++2.0 MAGE Cls. 19.2 8.5 −-10.7 21.9 8.9 −-13.0 10.1 7.0 −-3.1 RADAR Cls. 76.0 71.3 −-4.7 75.8 70.1 −-5.6 76.6 75.0 −-1.5 RoBERTa Cls. 93.5 95.2 ++1.7 92.7 94.4 ++1.8 96.4 97.7 ++1.3 Table 18: Full false positive rate results following professional editing. ΔFP=FPedit−FPorig =FP_edit-FP_orig (percentage-point change). FPR color: ≥ 50%, 10–49%. ΔFP color: bias reduction, bias amplification. Table 16, Table 17, and Table 18 provide the complete quantitative results for all detectors evaluated in this study. The table includes the metrics underlying the analyses reported in Sections 4.1–4.2, including baseline false-positive behavior, editing-induced score shifts, changes in false-positive rates, and effect-size estimates. The main text highlights only the most important findings for clarity. The complete results are provided here to facilitate detailed inspection of detector-specific behavior and to support reproducibility of the reported analyses. D.4 Inter-Editor Variability To assess whether editor-specific writing style systematically influences detector outputs, we restricted analysis to 280 professional editors with at least 50 edited documents. A Kruskal–Wallis test indicates significant differences in score shift across editors (H=1157.4H=1157.4, p≈3×10−107p≈ 3× 10^-107). However, an OLS regression of score shift on editing characteristics indicates that the practical contribution of editor identity is limited (Table 19). Predictor β Grammar ratio −0.191∗-0.191^*** Lexical ratio −0.131∗-0.131^*** Edit-token ratio +0.001∗+0.001^*** R2R^2 0.005 Table 19: OLS regression of detector score shift on editing characteristics, across 280 editors. ∗p<.001^***p<.001. While editor-specific differences are statistically detectable, editor identity explains only a very small proportion of variance in detector score shifts (R2=0.005R^2=0.005), indicating that the family-level patterns reported in Section 4.4 are not an artifact of a small number of editors’ individual styles.