Paper deep dive
Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks
Guangsheng Bao, Lihua Rong, Yanbin Zhao, Xiao Yu, Qiji Zhou, Yue Zhang
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 7/5/2026, 5:35:48 AM
Summary
Triospect is a novel three-dimensional statistical detection framework designed to robustly identify AI-generated text against humanizing and adversarial attacks. The framework operates by decoupling a text into its 'content' (semantic meaning) and 'expression' (stylistic elements) using LLM-based transformations. By measuring the original text, the content-preserving variant, and the expression-preserving variant, Triospect creates a three-dimensional feature vector. This vector is then processed by a Bayesian classifier using Gaussian Mixture Models (GMM) to distinguish between human-written and AI-generated text. Experimental results on the Humanize-16K and RAID benchmarks show significant improvements in AUROC and TPR01 over existing baseline detectors.
Entities (6)
Relation Signals (5)
Triospect → employs → Gaussian Mixture Model
confidence 100% · We model the distributions of samples in this space using two multivariate Gaussian Mixture Models
Triospect → evaluatedon → Humanize-16K
confidence 100% · It improves the strong baseline by a significant margin of 22.3% (AUROC) and 13% (TPR01) on the Humanize-16K
Triospect → evaluatedon → RAID
confidence 100% · and by 9.1% (AUROC) and 22% (TPR01) on the adversarial RAID.
Fast-DetectGPT → isusedby → Triospect
confidence 100% · We measure the texts using an existing detection metric m(·) (e.g., Fast-DetectGPT)
Triospect → uses → LLM
confidence 100% · we leverage the strong understanding ability of LLMs to regularize its expression
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Existing AI-generated text detectors are vulnerable to attacks that manipulate textual characteristics. In this study, we propose a novel Triospect Detection Framework by using additional perspectives of content (core ideas) and expression (stylistic elements) within a given text. Experiments on two benchmarks involving 17 attacks, 12 domains, and 17 source models demonstrate that Triospect is robust against these attacks. It improves the strong baseline by a significant margin of 22.3% (AUROC) and 13% (TPR01) on the Humanize-16K after-attack subset, and by 9.1% (AUROC) and 22% (TPR01) on the adversarial RAID. This framework marks a pioneering effort in statistical methods to enhance detection reliability against attacks. We release our data and code at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2606.31074v1
- Canonical: https://arxiv.org/abs/2606.31074v1
Trouble viewing inline? Open PDF directly →
Full Text
74,572 characters extracted from source content.
Expand or collapse full text
Triospect: A Three-Dimensional Framework for Robust Statistical AI-Generated Text Detection Against Diverse Attacks Guangsheng Bao 1,3, ∗ , Lihua Rong 2, ∗ , Yanbin Zhao 4 , Xiao Yu 5 , Qiji Zhou 3 , and Yue Zhang 3, † 1 Zhejiang University 2 Zhejiang University of Technology 3 Westlake University 4 Shanghai Polytechnic University 5 University of Science and Technology of China baoguangsheng@westlake.edu.cn ∗ ronglihua1981@zjut.edu.cn ∗ zhangyue@westlake.edu.cn † Abstract Existing AI-generated text detectors are vulnerable to attacks that manipulate tex- tual characteristics. In this study, we pro- pose a novel Triospect Detection Frame- work by using additional perspectives of content (core ideas) and expression (stylis- tic elements) within a given text. Experi- ments on two benchmarks involving 17 at- tacks, 12 domains, and 17 source models demonstrate that Triospect is robust against these attacks. It improves the strong base- line by a significant margin of 22.3% (AU- ROC) and 13% (TPR01) on the Humanize- 16K after-attack subset, and by 9.1% (AU- ROC) and 22% (TPR01) on the adversarial RAID. This framework marks a pioneering effort in statistical methods to enhance de- tection reliability against attacks. We release our data and code athttps://github. com/baoguangsheng/triospect. 1 Introduction Large language models (LLMs) have been widely used in news, academic, story and advertising writ- ing (Christian, 2023; M Alshater, 2022; Yuan et al., 2022; Chen and Chan, 2023), leading to unprece- dented societal risks including spreading misin- formation (Ahmed et al., 2021; Bagdasaryan and Shmatikov, 2022; Chen and Shu, 2023), eroding academic integrity (Perkins, 2023; Lee et al., 2023; Kumar et al., 2024), and blurring accountability in digital communication (Kaur et al., 2022; Sun et al., 2024). These challenges call for reliable AI-generated text detection tools, requiring the re- search community to develop effective detectors. However, at the same time, attacking techniques are also developed (Krishna et al., 2024; Zhou et al., 2024; Ayub et al., 2024), and commercial AI hu- manizing tools (16 listed in Appendix A) can by- * Equal contribution. † Corresponding author. pass state-of-the-art detectors, posing urgent threats to safe usage of AI technology. Existing state-of-the-art detectors are vulnera- ble when facing attacks. Supervised detectors (So- laiman et al., 2019; Fagni et al., 2021; Yan et al., 2023; Li et al., 2024; Verma et al., 2024) may fail in unfamiliar synthesized styles (Bakhtin et al., 2019; Uchendu et al., 2020; Pu et al., 2023). Zero-shot detectors (Gehrmann et al., 2019; Su et al., 2023; Mitchell et al., 2023; Bao et al., 2024; Xu et al., 2024; Hans et al., 2024) may be affected by surface- level textual changes (Krishna et al., 2024; Dugan et al., 2024; Wu et al., 2024; Chen et al., 2025). The watermarks (Kirchenbauer et al., 2023; Zhao et al., 2023; Christ et al., 2024; Zhao et al., 2024b,a) can be removed using new words and syntax (Krishna et al., 2024; Sadasivan et al., 2023). These vul- nerabilities underscore the necessity of advancing robust detection mechanisms that can withstand a wide range of attacks. Such attacks generally modify textual expres- sions (surface forms) while largely preserving un- derlying contents (semantics) (Jia and Liang, 2017; Alzantot et al., 2018; Ribeiro et al., 2018; Jin et al., 2020). By content, we refer to the information conveyed in the text. It is about ‘what’ is being communicated. By expression, we refer to the man- ner in which the content is conveyed. It is about ‘how’ the content is communicated. That is to say, under these attacks, the content of a text remains relatively stable. Thus, if a detector can measure the content directly, it will be resilient against the at- tacks, as adversarial vulnerability often stems from reliance on non-robust surface features rather than semantically meaningful representations (Madry et al., 2017; Ilyas et al., 2019). To separate the content from the expression of a text, we leverage the strong understanding ability of LLMs to regu- larize its expression, leading to a resultant token sequence primarily determined by its content. This content-representing token sequence has unique arXiv:2606.31074v1 [cs.CL] 30 Jun 2026 50510 Measure: m(T) 0.0 0.1 0.2 0.3 0.4 Probability Density Single-Dimensional View Human Humanized AI AI-Generated 50510 Measure: m(T c ) 0.0 0.1 0.2 0.3 0.4 Probability Density Content Dimension 50510 Measure: m(T e ) 0.0 2.5 5.0 Measure: m ( T c ) linear decision boundary Multi-Dimensional View Figure 1: Left: Humanizing attacks cause failure of existing detectors in single-dimensional view. Middle: The failure can be mitigated by more stable content dimension. Right: In multi-dimensional view, we combine the dimensions to achieve robust detection. (Humanize-16K using Fast-Detect as the metric) features for identifying its source, which we find can be captured by existing detection metrics. Based on this idea, we propose Triospect, a three-dimensional statistical detection framework that jointly analyzes the original text, its content, and its expression, enabling complementary signals to be examined across these three distinct dimen- sions. Specifically, given a textT, we perform a textual transformation to convert it into an approx- imate content-preserving text ˆ T c and expression- preserving text ˆ T e , approximately decoupling the two aspects. We measure the texts using an existing detection metricm(·)(e.g., Fast-Detect), obtaining the original text measurem(T), the content mea- surem( ˆ T c ), and the expression measurem( ˆ T e ). As Figure 1 illustrates, humanizing tools shift the distribution of AI-generated texts towards human- written texts in the measurem(T), causing over- lap between ‘Human’ and ‘Humanized AI’. How- ever, this confusion can be mitigated by the content measure, where the distribution remains relatively stable under attacks. Additionally, the expression measure can help when an attack occurs at the word level and does not change the grammatical patterns. Consequently, by combining these measures in the multi-dimensional view, we achieve a more robust detection framework. We construct a Humanize-16K benchmark for the systematic evaluation of detectors dealing with humanizing attacks. We collect human-written texts from 4 data sources and generate correspond- ing AI texts based on the same contexts (titles or prompts) using 6 LLMs. Then, we apply 6 hu- manizing attacks to the AI-generated texts, where the attacks include human editing, commercial AI tools, and LLM simulated attacks. Consequently, we obtain 16K samples, half for development and half for testing. We evaluate detectors on our Humanize-16K and the existing adversarial RAID benchmarks, where they cover a total of 12 domains, 17 LLMs, and 17 attacks. The experimental results show that Triospect is resistant to humanizing and adversarial attacks, leading to an improvement of 22.3% in AUROC and 13% in TPR01 on the Humanize-16K after-attack subset, as well as an improvement of 9.1% in AUROC and 22% in TPR01 on RAID, compared to the strong baseline detector. Analy- sis shows that Triospect is also robust to source models, decoding strategies, text lengths, and lan- guages. In summary, perspectives from content and ex- pression provide beneficial complements to origi- nal texts, leading to more reliable detection under humanizing and adversarial attacks. To our knowl- edge, this is the first statistical method to tackle detection problems against a wide range of attacks, and it achieves new state-of-the-art results among zero-shot detectors. 2 Related Work AI-Generated Text DetectionExisting detectors consist of three types of technology. The first is supervised classifiers (Solaiman et al., 2019; Ip- polito et al., 2020; Fagni et al., 2021; Hu et al., 2023; Yan et al., 2023; Li et al., 2024; Verma et al., 2024; Yu et al., 2024; Chen et al., 2025), which train a binary classifier based on a large collection of AI-generated and human-written text. The sec- ond is zero-shot classifiers, including white-box methods (Gehrmann et al., 2019; Su et al., 2023; Bao et al., 2024; Xu et al., 2024; Hans et al., 2024) and black-box methods (Mitchell et al., 2023; Yang et al., 2023; Bhattacharjee and Liu, 2024; Bao et al., 2025; Chen et al., 2025). These technologies usu- ally use pre-trained language models to extract de- tection metrics. The third is text watermarking technology (Kirchenbauer et al., 2023; Zhao et al., 2023; Christ et al., 2024; Zhao et al., 2024b,a), which identifies AI-generated text by embedding easy-to-detect markers or patterns. While these methods are successful at identifying entirely AI- generated texts, they lack resilience against diverse attack strategies (Gao et al., 2018; Dyrmishi et al., 2023; Krishna et al., 2024; He et al., 2024; Dugan et al., 2024; Wu et al., 2024; Wang et al., 2024; Zhou et al., 2024). As a result, various commer- cial AI tools offer features to “humanize” content, enabling them to evade current detection systems. To tackle this issue, we introduce the Triospect detection framework as a promising solution to counteract such vulnerabilities. There is limited research focused on the human- izing attack, with most efforts relying on super- vised learning. For instance, ImBD (Chen et al., 2025) fine-tunes a LLM on AI-rewritten and refined texts to learn machine-generated styles, using a style-CPC (conditional probability curvature) met- ric for detection. Similarly, DAMAGE (Masrour et al., 2025) develops a binary classifier trained on humanized texts produced by various humanizers. In contrast, Triospect does not require training and can be integrated with existing detectors to enhance their performance, including the supervised ImBD. Theexistingrewrite-detectframeworks (Mitchell et al., 2023; Liu et al., 2024; Mao et al., 2024) and recent repair-based DNA-DetectLLM (Zhu et al., 2025) may appear superficially similar to our Triospect framework since both employ LLMs to transform original texts. However, they differ fundamentally in their hypotheses and principles. The Triospect framework is driven by the idea of separating the content and expression of a text. We propose that content and expression are two crucial aspects of a text, offering stable and clear signals to identify AI sources. By decoupling these elements, we can evaluate the content and expression dimensions independently. In contrast, the existing rewrite-detect frameworks generally assume that AI-generated texts, when rewritten, exhibit greater similarity.They measure the similarity between rewritten versions, for example, using 7 versions by Raidar (Mao et al., 2024), as a means to indicate AI generation. Decoupling Content and Expression The idea of separating content and expression is related to existing studies on the disentanglement of seman- tics and syntax. These studies mainly focus on the disentanglement at the sentence level and discuss it in different contexts, such as recognition science (Caucheteux et al., 2021; Moro et al., 2001), sen- tence representation (Chen et al., 2019), sentence comprehension (Dapretto and Bookheimer, 1999), and sentence generation (Bao et al., 2019). They generally represent semantics and syntax in sepa- rate neural vectors and train a neural network with a specific structure or training objective to obtain disentangled vectors. In contrast, we leverage textual transformation to preserve the content or expression of a text while reducing another. Our approach is different in three ways. First, we focus on the discourse level instead of the sentence level, where the texts are longer and more complex. Second, we represent content and expression still in texts instead of neural vec- tors, which provides us with convenience for under- standing and explaining. Finally, we use LLM with prompting techniques instead of training a model, which simplifies the usage and generalizes better. 3 Triospect The detection framework relies on measuring the content and expression of a text separately. We achieve the separation through textual transforma- tions and measure them using existing detection metrics. Thus, the detection process involves sep- arating, measuring, and classifying steps, as illus- trated in Figure 2. 3.1 A Conceptual Model of Texts The notions of content and expression are re- lated to the classical distinction between seman- tics and grammar (Bickhard, 1993; Dapretto and Bookheimer, 1999; Chen et al., 2019; Caucheteux et al., 2021; Moro et al., 2001), but we define them at the discourse level rather than the sentence level. •Content refers to the abstract, pre-linguistic meaning structure that can be expressed: it consists of semantic units (e.g., events, en- tities, propositions, imagery) and the rela- tions among them (e.g., causal, temporal, con- trastive). It also encompasses the underlying conceptual network in cognition – intentions, emotions, viewpoints, and thematic structures – which may exist independently of any partic- ular linguistic form. •Expression refers to the linguistic realization of this meaning structure: the concrete en- coding of content through grammar, lexical Original Text 푇=퐶∘퐸 Content Preserve 푇 푐 =퐶∘퐸 0 simplify Expression Preserve 푇 푒 =퐶 0 ∘퐸 replace 푚(푇) 푚(푇 푐 ) 푚(푇 푒 ) Fast-DetectGPT Fast-DetectGPT Fast-DetectGPT [푚푇,푚(푇 푐 ),푚(푇 푒 )] GMM Textual transfor- mation based on LLM Figure 2: Overview of the Triospect detection process. The process consists of three steps: (i) Separating, which transforms the original text into content-preserving and expression-preserving variants using an LLM; (i) Measuring, which computes detection scores using an existing detector (e.g., Fast-DetectGPT); and (i) Classifying, which performs binary classification with a Bayesian classifier based on a GMM. choice, discourse organization, stylistic pref- erence, pragmatic strategy, and other surface- level devices. Psychological studies suggest that these two lev- els are partially separable in human text produc- tion (Kintsch and Van Dijk, 1978; DeRose et al., 1997; Clark, 2007; Flower and Hayes, 2016). Writ- ers typically engage in a content-planning stage, where conceptual and relational structures are or- ganized, followed by an expression-design stage, where these structures are rendered into specific linguistic forms to suit communicative goals and audience expectations. Following this conceptual- ization, we assume that a text can be analytically decoupled into its content and expression. Con- versely, given a well-specified content structure and a chosen mode of expression, a concrete text can be constructed as their realization. Formally, we express a textTas a composition of its content C and expression E as T = C◦ E,(1) where◦ denotes the composition operator. AI-Generated Text DetectionTheoretically, an AI-generated textT M can be expressed as a com- position of an AI-planned contentC M and an AI- designed expressionE M , while a human-written textT H can be expressed as a composition of a human-planned contentC H and a human-designed expression E H , as follows: T M = C M ◦ E M , T H = C H ◦ E H .(2) When we consider a humanizing attack by either an automatic tool or a human, the AI-generated texts are transformed to T ′ M = C M ◦ E ′ M or T ′ M = C M ◦ E H , (3) where the expression is converted into either an AI-altered version E ′ M or a human style E H . Thus, AI-generated text detection tasks require the distinction ofT M ,T ′ M fromT H , which can be identified from either the content or the expression of the text if we can measure them sep- arately. The ChallengeSeparating the latent content and expression from observed texts is challenging be- cause it requires a full understanding of what is said and how it is said. As a prototype, we utilize LLMs to approximate the function. 3.2 Separating Content and Expression Textual Transformation We investigate textual transformations using LLMs to preserve either the content or the expression of a text. LLMs have a strong ability to understand the content and ex- pression of a text and manipulate them under the guidance of specific prompts. Thus, we can ap- proximately achieve the goal by exploring suitable prompts. Formally, given an original textT, an idea content-preserving functionf c transforms the text into a new one T c as T c = f c (T), that T c = C◦ E 0 ,(4) and an idea expression-preserving functionf e pro- duces T e as T e = f e (T), that T e = C 0 ◦ E,(5) Original Text (T): Jane was walking through the park when she spotted a small puppy wandering aimlessly. She picked it up, read the collar tag, and took it to the address. The owner was grateful to have the puppy back. Content Preserving: f 1 Prompt: Summarize the main ideas of the text in a clear and concise manner: text Output: Jane found a lost puppy in the park, read its collar tag, and returned it to its grateful owner. f 2 Prompt: Outline the main points of the text to get a clear and concise picture of the content: text Output: 1. Jane was walking through the park. 2. She noticed a small puppy wandering aimlessly. 3. Jane found a collar tag on the puppy. 4. She took the puppy to the address on the tag. 5. The owner was grateful to Jane for returning the puppy. f 3 Prompt: Simplify the text to make it clear and concise while preserving its meaning: text Output: Jane found a puppy in the park and returned it to its grateful owner after reading the collar tag. Expression Preserving: f 4 Prompt: Identify and list the repre- sentative language expressions used in the text: text Output: 1. "walking through the park" - indicates Jane’s action and setting. 2. "spotted a small puppy" - describes Jane noticing the puppy. 3. "wandering aimlessly" - characterizes the puppy’s behavior. 4. "picked it up" - describes Jane’s action of picking up the puppy. f 5 Prompt: Substitute the key ideas in the text with placeholders, ensur- ing the original tone, style, and lan- guage remain intact: text Output: [Person] was walking through [location] when [they] spotted a [sub- ject] wandering [manner]. [They] picked it up, read the [identifier], and took it to the [destination]. The [recipient] was [reaction] to have the [subject] back. f 6 Prompt: Replace the main points of the text with a generic topic while preserving the language expression: text Output: Alex was strolling through the garden when they noticed a tiny kitten meandering without direction. They scooped it up, checked the collar tag, and brought it to the listed location. The caretaker was thankful to have the kitten returned. Table 1: Textual transformations that we explore for approximately preserving content and expression. The example demonstrates the function of each transformation. where the expressionEare contentCare dimin- ished to an uniform E 0 and C 0 . Content PreservingUsing a specific prompt, we guide an LLM to convert a textTinto an uniform expression while preserving its overall meaning, producing an approximation ˆ T c of the idea content- preserving textT c . Specifically, we investigate three intuitive techniques: summarizing, outlining, and simplifying, as illustrated in Table 1 (prompts f 1 tof 3 ). Summarizing produces a summary of the main ideas of a text, outlining generates a list of a text’s main points, while simplifying produces a simplified expression of the text. Basically, all these techniques are able to approximately preserve the content and reduce the stylistic expression. We explore these techniques and their alternatives em- pirically. Expression Preserving Similarly, we guide an LLM to convert a textTto ˆ T e , which approximates the idea expression-preserving textT e while reduc- ing the influence of its content. We explore three approaches, including listing representative syn- taxes of a text, substituting key ideas of a text with placeholders, and replacing the main points of a text with AI-generated ideas, as the promptsf 4 to f 6 in Table 1 illustrate. Listing preserves typical ex- pressions while reducing most ideas. Substituting preserves the expression template while reducing the specific characters. Replacing preserves the pattern of expression while substituting the content with AI-generated ideasC M , regardless of whether the original content was AI-generated or human- planned. DiscussionAll these transformations are not per- fect. For example, even with a placeholder rep- resenting the characters, the expression template produced by substituting still contains concrete, meaningful verbs. Preserving the expression of a text while reducing its content is more challenging since expression naturally relies on specific con- tent. Additionally, LLM transformations may also introduce some errors. However, we demonstrate that these imperfect representations of content and expression are still surprisingly beneficial for de- tection tasks. 3.3 Measuring Content and Expression Based on the transformations introduced in Section 3.2, we obtain multiple rewritten variants of each text that approximately preserve either its content or its expression. We then apply AI-text detectors to these variants and use the resulting scores to measure content and expression signals. The key intuition is that content-preserving trans- formations retain the main ideas of a text while reducing stylistic variation, whereas expression- preserving transformations retain stylistic patterns while weakening the original content. Therefore, detector scores on these transformed texts reflect different aspects of the original text. We find that these latent aspects display identifi- able token distributions for different sources, where existing detection metrics, typically zero-shot de- tection metrics, can be utilized to differentiate be- tween the sources. Due to the statistical nature of these metrics, they are not sensitive to partial vari- ances in the transformed texts, producing relatively stable measures. 3.4 Binary Classifier After converting a text into a measure vectorx = [m(T),m( ˆ T c ),m( ˆ T e )], we project it into three- dimensional space. We model the distributions of samples in this space using two multivariate Gaus- sian Mixture Models:ρ H for human-written texts andρ M for AI-generated texts. Each distribution is characterized by a group of mean vectorsμ k , repre- senting the average measure values, and a group of covariance matricesΣ k , capturing the relationships and variability between measures. The k ∈ [1..K] denotes the index of mixture components, where K is the number of components. Based on a number of development samples, we estimate the distribution parametersμ H k ,Σ H k ,μ M k , andΣ M k , obtaining two probability density func- tions ρ H (x) = K X k=1 φ H k N(μ H k ,Σ H k )(6) and ρ M (x) = K X k=1 φ M k N(μ M k ,Σ M k ),(7) whereφ H k andφ M k are the weights of mixture components, each totaling 1, andN(·)is a sin- gle component of a multivariate Gaussian distri- bution. The component parametersμ H k andμ M k are3-dimensional mean vectors, andΣ H k andΣ M k are3 × 3covariance matrices. The probability densityρ H (x)andρ M (x)describe the likelihood of a text with a measure vectorxunder the distri- butions of human-written and AI-generated texts, respectively. DomainAvg. WordsDev / Test Student Essay (Crossley et al., 2024)2412K / 2K ArXiv Intro (arXiv, 2024)4102K / 2K Creative Writing (Fan et al., 2018)3452K / 2K C News (Hamborg et al., 2017)1482K / 2K C News (French)2582K / 2K C News (Spanish)2852K / 2K C News (Arabic)1522K / 2K Table 2: Humanize-16K and a multilingual patch. Our goal is to determine whether a given text is more likely to be AI-generated or human-written. To do this, we calculate the Bayesian posterior probability as p(M|x) = p(x|M)p(M) p(x|H)p(H) + p(x|M)p(M) (8) = ρ M (x) λ· ρ H (x) + ρ M (x) ,(9) whereλ = p(H)/p(M)denotes the ratio of the prior probabilities. For balanced classes,λis set to 1, whereas for imbalanced classes, it represents the ratio of human samples to machine samples. To arrive at a final decision, a probability threshold is required, which should be determined based on the specific application. DiscussionEmpirically, we observe that a single Gaussian component is typically sufficient across different scenarios, with extra components offering only marginal improvements. This may be due to two factors. First, the detection metric is based on token averages. By the Central Limit Theo- rem, the mean of many random variables tends to approximate a normal distribution, regardless of their original forms. Second, the detection decision is only sensitive to the density functions near the boundary whereρ H andρ M have close values. The decision boundary usually falls within low-density regions, influencing only a small subset of samples. Moreover, in these regions, the density estimated by a single component is often not substantially different from that estimated with multiple compo- nents. 4 Benchmark of Humanizing Attack We create a dataset Humanize-16K to mimic a real- world situation influenced by human editing and humanizing tools, focusing on a difficult scenario not addressed by current benchmarks. Our dataset covers 6 humanizing approaches, 4 domains, 6 text generation models, and various decoding strategies. We describe the detailed construction of the dataset in Appendix B and summarize it as follows. 4.1 Construction of Humanize-16K As the domains listed in Table 2, we collect 2,000 human samplesT H from each domain. For a random halfT H0 of the human samples, we pro- duce corresponding AI-generated textsT M using 6 LLMs, including gpt-3.5-turbo, gpt-4o, claude- 3.5-sonnet, gemini-1.5-pro, llama-3.3-70b-instruct, and qwen-2.5-72b-instruct. We further humanize these generated texts using human editing, commer- cial AI tools, and simulated humanizing process, producing 1,000 humanized samplesT ′ M per do- main. Consequently, we obtain an equal number of human-written texts and AI-generated/humanized texts, resulting in a total of 16,000 samples. We configure the task into three subsets: the before-attack subset for classifying betweenT H0 andT M , the after-attack subset for classifying betweenT H0 andT ′ M , and the full set mix for classifying betweenT H andT M ,T ′ M , where the complete human setT H is used to balance the two classes. We randomly split the samples into a development set and a test set of equal size for each. Additionally, we create a multilingual patch by collecting French, Spanish, and Arabic news from Common Crawl (Hamborg et al., 2017), and follow the same process to produce 4,000 samples per language. In the construction process, we hire human anno- tators for editing and quality checks of the gen- erated texts. Additionally, we call the API for the commercial AI tools. Totally, they cost about $2500. 4.2 Humanizing AI-Generated Texts We humanize AI-generated text via human edit- ing, commercial tools, and automated LLM-based simulations. 4.2.1 Human Editing We employ five annotators from a specialized an- notation company, consisting of three individuals with professional expertise in English and two with a background in computer science. The team in- cludes three women and two men, aged 22 to 41 years, all of whom have Chinese as their native lan- guage. Each annotator is responsible for revising 50 AI-generated texts, resulting in a total of 250 human-edited samples. The editing process is carried out at three levels: word, sentence, and paragraph. At the word level, synonyms are used to replace existing words; at the sentence level, syntax alterations are made; and at the paragraph level, the logical flow of sentences is reorganized. Annotators are asked to apply these three types of edits in equal proportion, ensuring that more than 50% of the original content is modi- fied. In addition, a separate annotator reviews 10% of the texts to verify that the edits preserve the orig- inal meaning while ensuring that the revised texts remain fluent and comprehensible. It costs about $2,000 for human editing. 4.2.2 Commercial AI Tools We use three commercial AI tools, including hum- bot.ai, bypassgpt.ai, and undetectable.ai, from Ta- ble 9 in Appendix. We call these tools through their API, producing 765 humanized documents. The API costs about $500 in total. 4.2.3 Simulation by LLMs We use the following prompts to humanize AI- generated text by LLMs. Prompt for Diversifying: “Revise the text to enrich its linguistic diversity, employing var- ied sentence structures, synonyms, and stylis- tic nuances, while preserving the original meaning: generation” Prompt for Mimicking: “Rewrite the text using the same language style, tone, and expression as the reference text.Focus on capturing the unique vocabulary,sentence structure, and stylistic elements evident in the reference: generation # Reference Text: reference” Diversifying enriches expression while preserv- ing core ideas. Mimicking creates more human-like text by copying style and structure but may intro- duce fabricated content and alter the original text. 5 Experimental Settings 5.1 Datasets We use our Humanize-16K benchmark and the existing benchmark RAID (Dugan et al., 2024) as representatives of wild scenarios under attack. Humanize-16K covers 6 humanizing attacks, 4 do- mains, and 6 models. RAID covers 11 adversarial diversify mimic humbot.ai bypassgpt.ai undetectable.ai human edit Humanizing Attack 0.2 0.4 0.6 0.8 1.0 Similarity Score SimCSE similarityPOS-BLEU similarity Figure 3: Semantic and grammar similarities per humanizing attack. (T, T c )(T, T e ) Transformation 0.7 0.8 0.9 1.0 Similarity Score (T, T c )(T, T e ) Transformation 0.00 0.25 0.50 0.75 1.00 Similarity Score SimCSE similarityPOS-BLEU similarityAttack Mean Figure 4: Semantic and grammar similarities for content-preserving (f 3 ) and expression-preserving (f 6 ) transformations. attacks, 8 domains, 11 models, and various decod- ing strategies, where we sample 4K for testing and another 4K for development. 5.2 Detectors BaselinesWe consider classical and state-of-the- art zero-shot detectors, which generally leverage pre-trained LLMs to compute a detection metric as an indicator of AI-generated text. These metrics can also be used in the Triospect detection frame- work. Specifically, we take log-perplexity, log- rank, LRR (Su et al., 2023), Fast-Detect (Bao et al., 2024), Binoculars (Hans et al., 2024), Glimpse (Bao et al., 2025), and Raidar (Mao et al., 2024) as representatives. We use falcon-7B as the scoring model for the first three detectors, falcon-7B/falcon- 7B-instruct as the scoring models for Fast-Detect and Binoculars. Raidar extracts 56-dimensional features from multiple rewritten texts and trains a binary classifier. We also consider supervised detectors, using RADAR (Hu et al., 2023), RoBERTa (ChatGPT) (Guo et al., 2023), and ImBD (imitate before de- tect) (Chen et al., 2025) as representatives. Typi- cally, RADAR is trained under the concept of ad- versarial learning, which can resist paraphrasing attacks. ImBD uses a fine-tuned gpt-neo-2.7B on AI-rewritten texts, which is optimized for detecting AI texts refined by a humanizing tool. Our Detector We present Triospect combined with a specific existing detection metric. In our main experiments, we utilizef 3 (T)to generate ˆ T c andf 6 (T)to generate ˆ T e , employing Qwen3-4B (with max_tokens = 200) under non-thinking mode as the transformation model. We use a single mix- ture component (K = 1) in the main experiments and compare this setting with alternatives in the ablation study. 5.3 Metrics We use AUROC, the area under the receiver operat- ing characteristic curve, as the major metric to mea- sure the quality of the classifiers. We report ACC (accuracy) with the best threshold found by max- imizing Youden’s J statistics (Youden, 1950). We also report TPR01, a true positive rate at a false pos- itive rate of 1%, for reference. We do significance test with McNemar’s Test (McNemar, 1947) for ACC, and Bootstrap Resampling (Dwivedi et al., 2017) for TPR01, with ap-value less than 0.01. We run all the experiments on a machine with one Tesla A100 GPU, which takes about six hours. 6 Experiments 6.1 Necessity of Content and Expression Measures Triospect assumes that AI-generated and human- written texts differ in content and expression, while content remains relatively stable under attacks. To test this hypothesis, we use two complementary similarity metrics. Content similarity is measured via SimCSE (Gao et al., 2021) with RoBERTa-large (Liu et al., 2019), which computes cosine similarity between sentence embeddings to capture semantic consistency. Ex- pression similarity uses POS-BLEU, which com- bines POS n-grams (Koppel et al., 2009) that re- flect syntactic patterns with BLEU (Papineni et al., 2002) to quantify surface-form overlap. These met- rics allow separate analysis of meaning and form under attacks or transformations. Effect of Attacks We first analyze how human- izing attacks modify texts. As shown in Figure 3, attacks mainly alter surface expression rather than underlying content. Specifically, the coefficient of variation (std/mean) of POS-BLEU scores reaches 23.1%, whereas SimCSE scores vary by only 6.1%. Detector AUROCACCTPR01 BeforeAfterMixBeforeAfterMixBeforeAfterMix RoBERTa(ChatGPT) (Guo et al., 2023)0.6640.5600.56361%60%61%8%0%5% RADAR (Hu et al., 2023)0.8380.7290.78776%67%70%5%3%6% Log-Perplexity0.7640.6470.54972%59%56%0%4%0% Log-Rank0.7790.6580.55273%60%58%0%6%0% LRR (Su et al., 2023)0.8200.6560.57776%62%61%25%2%11% Raidar (Mao et al., 2024)0.9290.911*0.88084%76%80%77%41%19% Binoculars (Hans et al., 2024)0.9160.6290.77291%67%79%83%30%56% Triospect (Binoculars)0.9670.8480.90893%*78%84%85%*35%59% ImBD (Chen et al., 2025)0.9620.8210.89193%74%83%85%34%59% Triospect (ImBD)0.968*0.8790.925*92%80%*85%*85%41%59% Fast-Detect (Bao et al., 2024)0.9130.6270.77090%68%78%81%31%55% Triospect (Fast-Detect) 0.9600.8500.90192%78%84%84%44%*63%* (+0.047)(+0.223)(+0.131)(+2%)(+10%)(+6%)(+3%)(+13%)(+8%) Table 3: Main results before and after humanizing attacks evaluated on Humanize-16K. The significantly improved scores among each group are marked in bold and the best scores are marked with ‘*’. Detector AUROC per Domain of RAIDMixture of Domains NewsBooksWikiAbstracts Reddit Recipes Poetry ReviewsAUROCACCTPR01 Binoculars0.7680.8500.8040.8260.8110.7590.8260.8120.80778%46% Triospect (Binoculars)0.887*0.9300.8810.913*0.898*0.897*0.920*0.8790.901*84%*59% ImBD0.7830.8750.8040.8060.8250.6990.8230.8400.80476%41% Triospect (ImBD)0.8570.954*0.8760.8650.8660.8430.8740.895*0.87382%56% Fast-Detect0.7610.8450.8030.8210.7940.7490.8180.8100.80077%39% Triospect (Fast-Detect) 0.8780.9250.891*0.8970.8710.8950.8970.8800.89183%61%* (+0.117)(+0.080)(+0.088)(+0.076)(+0.077)(+0.146)(+0.079)(+0.070)(+0.091)(+6%)(+22%) Table 4: Main results under adversarial attacks, evaluated on RAID benchmark. The much smaller variance of SimCSE indicates that semantic content remains largely stable under attack, while grammatical realization changes sub- stantially. It suggests that content-based signals are inherently more robust under adversarial rewriting. Content- and Expression-Preserving Transfor- mationsWe next evaluate whether the proposed transformations successfully isolate the two as- pects of text. As shown in Figure 4, the content- preserving transformation (T → ˆ T c ) maintains high SimCSE similarity but exhibit low POS- BLEU similarity, indicating that semantic content is preserved while surface expression changes. This confirms that the transformation effectively per- turbs expression without altering meaning. The expression-preserving transformation (T → ˆ T e ) exhibits higher POS-BLEU similarity but lower SimCSE similarity, demonstrating that expression is preserved while semantics change. However, the similarity scores show larger variance than in the content-preserving case, suggesting that preserving expression while altering meaning is intrinsically more difficult to control. Complementarity of Content and Expression Measures Finally, we examine whether AI and human texts differ under the proposed measures. As shown in Figure 1 (middle), AI-generated and human-written texts exhibit distinguishable distri- butions in the content measure, which complements the differences observed in expression as illustrated by the skewed linear decision boundary in Figure 1 (right). These results indicate that jointly model- ing content and expression is necessary for reliable AI-generated text identification. 6.2 Main Results We first evaluate the detectors on their ability to mitigate humanizing attacks. As Table 3 shows, all detectors experience a significant performance drop after the attacks. Take Fast-Detect as an exam- ple. The attacks reduce the AUROC from 0.913 to 0.627. Compared to the baseline, Triospect signifi- cantly mitigates the impact of the attacks, increas- ing the AUROC from 0.627 to 0.850. Surprisingly, Triospect also enhances the AUROC on texts be- fore the attacks. The effects on multiple baselines are consistent. Typically, upon the strong baseline ImBD, which has been optimized for AI-rewritten texts, the Triospect detector still achieves signifi- cant improvements in AUROC and ACC. We further compare the detectors on their abil- Method Humanize-16KRAIDTime AUROC TPR01 AUROC TPR01 /Sample Raidar0.87719%0.77317%19.5s Fast-Detect0.77055%0.80039%0.2s Triospect (Fast-Detect) using Qwen3-4B max_tokens=250.87062%0.86842%0.5s max_tokens=500.87562%0.87850%0.7s max_tokens=1000.89063%0.88656%1.3s max_tokens=2000.90163%0.89361%2.1s Triospect (Fast-Detect) using GPT-4o 0.90965%0.89159%6.5s Table 5: Detection efficiency, where the detection accuracy and speed can be balanced by the number of tokens generated. ity to mitigate adversarial attacks using the RAID benchmark. As shown in Table 4, the Triospect detectors outperform the baselines by even larger margins. Triospect enhances the AUROC by 9.1% and TPR01 by 22% for Fast-Detect, and the AU- ROC by 9.4% and TPR01 by 13% for Binoculars. The results on the two datasets demonstrate the effectiveness of Triospect in addressing attacks. Although they are training-free, they achieve bet- ter detection performance than existing supervised and zero-shot detectors. Additionally, the improve- ments are consistent across domains and evaluation metrics. Efficiency As shown in Table 5, Triospect requires about 2.1 seconds per sample with max_tokens=200and0.5secondswith max_tokens=25.For a max_tokens of 25, its runtime is comparable to Fast-Detect while still achieving a notable improvement in AUROC and TPR01. In practice, we could strike a balance between detection accuracy and computational efficiency by constraining the maximum number of generated tokens in the transformations. 6.3 Ablation Study Feature Contribution Triospect incorporates two extra features: content and expression mea- sures. Our findings indicate that the content feature drives most of the performance gains. Specifically, the feature set[m(T),m(T c )]yields an AUROC of 0.898 for Fast-Detect on Humanize-16 mix, nearly matching the 0.901 AUROC obtained with the full feature set. However, it may vary against different attacks. In the context of RAID, the expression measure has a marginally greater impact, which increases the AUROC from 0.880 to 0.891. qwen4B gpt4o gpt3.5 dsk-v3 0.7 0.8 0.9 1.0 AUROC 0.901 0.910 0.878 0.891 Humanize-16K Figure 5: Transfor- mation LLMs. f 1 f 2 f 3 0.7 0.8 0.9 1.0 AUROC 0.875 0.888 0.905 RAID f 4 f 5 f 6 0.898 0.896 0.905 RAID Figure 6: Transformation prompts. Textual TransformationThe choice of LLM for transformation affects detection performance. As shown in Figure 5, Triospect (Fast-Detect) achieves higher AUROC with the stronger gpt-4o but lower with gpt-3.5 and deepseek-v3. We also exam- ine sampling strategies and find that setting top-p with0.6improves AUROC by about 1%, whereas greedy decoding reduces it by roughly 3%. The design of transformation prompts also im- pacts performance. Figure 6 shows that Triospect (Binoculars) withf 3 outperformsf 1 andf 2 , while usingf 6 outperformsf 4 andf 5 . We further test alternative versions off 3 andf 6 by asking gpt-4o to rewrite each prompt five times. Results show only marginal variation across alternatives, with AUROC fluctuations within±0.009. Finally, LLM transformations can introduce er- rors, potentially affecting detection accuracy. To as- sess this, we analyze prediction consistency across multiple generations of the same input text. Since sampling introduces varying errors, we estimate their effects on prediction probabilities. Results indicate that these errors lead to prediction scores with a standard deviation of 0.012, which is negli- gible for detection. Detection Metric We also try other detection metrics, such as Log-Perplexity and Log-Rank, ob- taining consistent improvements. We further try trained detectors such as RoBERTa (ChatGPT) and RADAR, using their predictive probabilities as de- tection metrics. However, empirical results show that trained detectors are ineffective in measuring transformed textsT c andT e . This is probably be- cause the trained detectors are not familiar with these text styles. Density Estimation Triospect estimates its pa- rameters using a dev set. In practice, just 20 sample pairs are sufficient for full performance with a sin- gle Gaussian component, achieving an AUROC of 0.910 on Humanize-16K for Triospect (Fast- diversify mimic humbot.ai bypassgpt.ai undetectable.ai human edit Humanizing 0.2 0.4 0.6 0.8 1.0 AUROC Fast-DetectTriospect Figure 7: Against humanizing at- tacks (Humanize-16K). zero space homoglyph whitespace insert paras misspelling alter spelling article del number upper lower paraphrase synonym Adversarial Attack 0.5 0.6 0.7 0.8 0.9 1.0 AUROC sampling greedy Decoding yes no Rep-Penalty Binoculars2D (Triospect)3D (Triospect) Figure 8: Against adversarial attacks and deal with various decoding strategies (RAID). 0100200300400500600 Document Size (# of Words) 0.6 0.7 0.8 0.9 1.0 AUROC Humanize-16K Triospect Fast-Detect Figure 9: AUROC on Humanize-16K mix, grouped by text lengths. Detect), which is even slightly better than the 0.901 AUROC obtained with the full dev set. We evaluate both single and multiple Gaussian components. Results show that using 2 to 5 compo- nents offers no clear advantage over a single com- ponent, with AUROC varying within±0.005. No- tably, the single-component model requires fewer samples for stable estimation. 6.4 Analysis of Robustness Against Humanizing Attacks The analysis in Figure 7 indicates that AI tools are the major source of threat, unlike human editing, which the baseline detector identifies with high accuracy. The base- line detector struggles significantly, often failing against commercial AI tools (e.g., humbot.ai, by- passgpt.ai, undetectable.ai). In contrast, Triospect shows marked improvement in AUROCs, highlight- ing its resilience against current commercial hu- manizing AI. Against Adversarial Attacks and Decoding Strategies As Figure 8 demonstrates, Triospect detectors are robust against adversarial attacks and decoding strategies. Triospect achieves significant improvements in all categories except ‘synonym’ and ‘non-repetition-penalty’. Among them, the ex- pression measure mainly contributes to zero-space, homoglyph, and synonym attacks, as shown by the Detectorqwen2.5 gemini1.5 llama3.3 gpt3.5 claude3.5 gpt4o Fast-Detect0.7500.7560.7750.7160.7520.490 Triospect0.8790.8880.8880.8780.8500.810 Table 6: AUROC per source LLM evaluated on Humanize-16K mix. Detector FrenchSpanishArabic in / outin / outin / out Fast-Detect0.7730.6960.465 Triospect0.841 / 0.8410.791 / 0.7900.616 / 0.614 Table 7: AUROC per language evaluated on mul- tilingual patch, where ‘(in)’ denotes ‘in-domain’ setting and ‘(out)’ denotes ‘out-of-domain’ setting. comparison between the 2D[m(T),m(T c )]and 3D [m(T),m(T c ),m(T e )]. Over Different LengthsWe categorize the sam- ples into buckets such as[0, 100),[100, 200), [200, 300), etc., and compare Triospect with the baseline in Figure 9. Generally, the longer the texts are, the greater the improvement will be, where Triospect is especially beneficial for long texts. Across Source Models We compared detectors on texts generated by different source models. We conduct both ‘in-domain’ and ‘out-of-domain’ set- tings, where we use dev samples including the model for the ‘in-domain’ setting and those ex- cluding the model for the ‘out-of-domain’ setting. Experiments show that the scores are almost identi- cal in the two settings. As Table 6 shows, Triospect outperforms the baseline in all categories. The im- provements are especially significant for stronger models such as gpt-4o. Across Languages We assess detector perfor- mance on French, Spanish, and Arabic news data under two settings: ‘in-domain’ with dev samples from both English and the language, and ‘out-of- Detector WritingNewsEssayArXiv in / outin / outin / outin / out Fast-Detect0.7300.7290.8620.755 Triospect0.915 / 0.913 0.883 / 0.879 0.874 / 0.866 0.934 / 0.933 Table 8:AUROC per domain evaluated on Humanize-16K. domain’ with dev samples from English only. As Table 7 shows, Triospect outperforms the base- line in every language for both settings, exhibit- ing consistent results across the two settings. The relatively lower AUROC for Arabic stems from the poor performance of the underlying Falcon- 7B scoring models in that language. Employing the stronger multilingual model davinci-002 via Glimpse (Bao et al., 2025) achieves significantly higher AUROC scores of 0.678 for Glimpse and 0.792 for Triospect. Across Text DomainsWe evaluated detector per- formance across general writing, news, essays, and scientific articles (ArXiv) under both ‘in-domain’ and ‘out-of-domain’ settings, where GMM param- eters were estimated with or without samples from the target test domain. Table 8 shows that all detec- tors suffered only minor drops in out-of-domain set- ting (average drop of 0.004 in AUROC), indicating strong robustness. Triospect consistently outper- formed the baseline Fast-Detect across all domains, with notable gains except for essays, demonstrating reliable detection even on unseen text domains. Across Transformation Models We evaluated robustness across different transformation models to assess potential self-bias. When the transfor- mation model matches the source generator, per- formance remains stable: using GPT-4o to trans- form GPT-4o–generated texts achieves nearly iden- tical AUROC (0.814) compared to using a different transformation model, Qwen3-4B (0.810). We also swap the transformation model between calibration and testing. Calibrating the GMM with GPT-4o and testing with Qwen3-4B yields nearly identi- cal AUROC (0.900 vs. 0.901) compared to using Qwen3-4B for both stages, with the same pattern observed in reverse. These results suggest that the transformation model introduces little systematic bias, and the learned densities reflect intrinsic con- tent and expression signals rather than rewriting- model artifacts. 6.5 Analysis of Failure Cases As shown in Figure 8, Triospect does not improve performance over the base detector under the syn- onym attack. We further analyze the failure cases to understand this behavior. Specifically, Binoculars makes 108 wrong pre- dictions, all of which are AI-generated texts after the attack, while Triospect makes 118 wrong pre- dictions, including 99 AI-generated texts and 19 human-written texts after the attack. This result suggests that although Triospect reduces some false negatives on AI-generated texts, it also introduces additional false positives on human-written texts. A possible reason is that synonym attacks in- volve only lexical substitutions while preserving the semantic content as well as stylistic and gram- matical patterns. As a result, the distributions of the original text, the content-preserving text, and the expression-preserving text remain highly simi- lar. This is evidenced by a Pearson correlation of 0.80 betweenm(T)andm( ˆ T c ), and 0.81 between m(T)andm( ˆ T e ), compared to 0.69 and 0.75 un- der other attacks. Such strong correlations indi- cate limited complementary information across the three dimensions, reducing the benefit of the multi- dimensional framework. Moreover, the rewriting transformations in Triospect may introduce addi- tional variance, which can slightly distort human- written texts and lead to misclassification. The key intuition is that synonym attacks modify the expression too weakly to produce meaningful differences across the three dimensions, so the ad- ditional transformations provide little benefit and may even introduce noise. 7 Discussion and Limitations Readers might wonder how human-written texts differ from AI-generated ones once both have un- dergone textual transformations that make them resemble outputs from an LLM. While the way they are expressed changes, their core semantics remain the same, and this consistency can reveal their original source. Typically, human texts carry rich, specific meanings, whereas AI texts tend to be simpler and lack striking, vivid details (Holtzman et al., 2019; Gehrmann et al., 2019). Differences at the semantic level will also lead to distinct and recognizable token distributions (Sahlgren, 2008). When we preserve the semantics while surpassing other aspects, this distinct feature is displayed and can be measured using existing detection metrics. Despite the promising results of the proposed Triospect detection framework, two limitations re- main. First, our approach relies on the approxi- mate decoupling of content and expression through textual transformations. While effective in prac- tice, these transformations are imperfect, leaving an open research topic regarding the disentangle- ment of the two aspects. Second, the current im- plementation incurs additional computational costs compared to single-dimensional detectors, as multi- ple transformations and measurements are required. While we explored efficiency trade-offs, practical deployment in large-scale or real-time scenarios may require further optimization. 8 Conclusion Triospect Detection Framework provides a signif- icant advancement in AI-generated text detection, particularly against humanizing and adversarial at- tacks. By introducing distinct content and expres- sion dimensions, the framework overcomes the lim- itations of existing detectors. Experimental results on diverse datasets demonstrate the resilience and improved performance of Triospect, achieving no- table gains in AUROC and TPR01. This novel approach marks a pioneering step in enhancing the robustness and reliability of AI-generated text de- tection tools under attack. Acknowledgement We would like to thank the editors and anonymous reviewers for their valuable feedback. This work is funded by the National Natural Science Foundation of China Key Program (Grant No. 62336006). References Alim Al Ayub Ahmed,Ayman Aljabouh, Praveen Kumar Donepudi, and Myung Suh Choi. 2021. Detecting fake news using machine learn- ing: A systematic literature review.arXiv preprint arXiv:2102.04458. Moustafa Alzantot, Yash Sharma, Ahmed Elgo- hary, Bo-Jhang Ho, Mani Srivastava, and Kai- Wei Chang. 2018.Generating natural lan- guage adversarial examples.arXiv preprint arXiv:1804.07998. arXiv. 2024. arxiv. Cornell University. Taseef Ayub, Rayees Ahmad Malla, Mas- hood Yousuf Khan, and Shabir Ahmad Ganaie. 2024. The art of deception: humanizing ai to outsmart detection. Global Knowledge, Memory and Communication. Eugene Bagdasaryan and Vitaly Shmatikov. 2022. Spinning language models: Risks of propaganda- as-a-service and countermeasures. In 2022 IEEE Symposium on Security and Privacy (SP), pages 769–786. IEEE. Anton Bakhtin, Sam Gross, Myle Ott, Yuntian Deng, Marc’Aurelio Ranzato, and Arthur Szlam. 2019. Real or fake? learning to discriminate ma- chine from human generated text. arXiv preprint arXiv:1906.03351. Guangsheng Bao, Yanbin Zhao, Juncai He, and Yue Zhang. 2025. Glimpse: Enabling white-box methods to use proprietary models for zero-shot llm-generated text detection. The Thirteenth In- ternational Conference on Learning Representa- tions. Guangsheng Bao, Yanbin Zhao, Zhiyang Teng, Linyi Yang, and Yue Zhang. 2024.Fast- detectgpt:Efficient zero-shot detection of machine-generated text via conditional proba- bility curvature. In The Twelfth International Conference on Learning Representations. Yu Bao, Hao Zhou, Shujian Huang, Lei Li, Lili Mou, Olga Vechtomova, Xinyu Dai, and Jiajun Chen. 2019. Generating sentences from disen- tangled syntactic and semantic spaces. In Pro- ceedings of the 57th Annual Meeting of the As- sociation for Computational Linguistics, pages 6008–6019. Amrita Bhattacharjee and Huan Liu. 2024. Fight- ing fire with fire: can chatgpt detect ai-generated text? ACM SIGKDD Explorations Newsletter, 25(2):14–21. Mark H Bickhard. 1993. Representational con- tent in humans and machines. Journal of Ex- perimental & Theoretical Artificial Intelligence, 5(4):285–333. Charlotte Caucheteux, Alexandre Gramfort, and Jean-Remi King. 2021. Disentangling syntax and semantics in the brain with deep networks. In International conference on machine learning, pages 1336–1348. PMLR. Canyu Chen and Kai Shu. 2023. Combating misin- formation in the age of llms: Opportunities and challenges. AI Magazine. Jiaqi Chen, Xiaoye Zhu, Tianyang Liu, Ying Chen, Chen Xinhui, Yiwen Yuan, Chak Tou Leong, Zuchao Li, Long Tang, Lei Zhang, et al. 2025. Imitate before detect: Aligning machine stylistic preference for machine-revised text detection. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 23559–23567. Mingda Chen, Qingming Tang, Sam Wiseman, and Kevin Gimpel. 2019. A multi-task approach for disentangling syntax and semantics in sentence representations. In Proceedings of the 2019 Con- ference of the North American Chapter of the Association for Computational Linguistics: Hu- man Language Technologies, Volume 1 (Long and Short Papers), pages 2453–2464. Zenan Chen and Jason Chan. 2023. Large language model in creative work: The role of collaboration modality and user expertise. Available at SSRN 4575598. Miranda Christ, Sam Gunn, and Or Zamir. 2024. Undetectable watermarks for language models. In The Thirty Seventh Annual Conference on Learning Theory, pages 1125–1139. PMLR. Jon Christian. 2023. Cnet secretly used ai on arti- cles that didn’t disclose that fact, staff say. Futu- rusm, January. Dave Clark. 2007. Content management and the separation of presentation and content. Techni- cal communication quarterly, 17(1):35–60. Scott Crossley, Perpetual Baffour, Jules King, Lauryn Burleigh, Walter Reade, and Maggie Demkin. 2024. Asap 2.0: Automated student assessment prize. Kaggle. Mirella Dapretto and Susan Y Bookheimer. 1999. Form and content: dissociating syntax and se- mantics in sentence comprehension. Neuron, 24(2):427–432. Steven J DeRose, David G Durand, Elli Mylonas, and Allen H Renear. 1997. What is text, really? ACM SIGDOC Asterisk Journal of Computer Documentation, 21(3):1–24. Liam Dugan, Alyssa Hwang, Filip Trhlik, Josh Magnus Ludan, Andrew Zhu, Hainiu Xu, Daphne Ippolito, and Chris Callison-Burch. 2024. Raid: A shared benchmark for robust evaluation of machine-generated text detectors. arXiv preprint arXiv:2405.07940. Alok Kumar Dwivedi, Indika Mallawaarachchi, and Luis A Alvarado. 2017. Analysis of small sample size studies using nonparametric boot- strap test with pooled resampling method. Statis- tics in medicine, 36(14):2187–2205. Salijona Dyrmishi, Salah GHAMIZI, and Maxime Cordy. 2023. How do humans perceive adver- sarial text? a reality check on the validity and naturalness of word-based adversarial attacks. In The 61st Annual Meeting Of The Association For Computational Linguistics. Tiziano Fagni, Fabrizio Falchi, Margherita Gam- bini, Antonio Martella, and Maurizio Tesconi. 2021. Tweepfake: About detecting deepfake tweets. Plos one, 16(5):e0251415. Angela Fan, Mike Lewis, and Yann Dauphin. 2018. Hierarchical neural story generation. In Pro- ceedings of the 56th Annual Meeting of the Asso- ciation for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics. Linda S Flower and John R Hayes. 2016. The dynamics of composing: Making plans and jug- gling constraints. In Cognitive processes in writ- ing, pages 31–50. Routledge. Ji Gao, Jack Lanchantin, Mary Lou Soffa, and Yanjun Qi. 2018. Black-box generation of ad- versarial text sequences to evade deep learning classifiers. In 2018 IEEE Security and Privacy Workshops (SPW), pages 50–56. IEEE. Tianyu Gao, Xingcheng Yao, and Danqi Chen. 2021. Simcse: Simple contrastive learning of sentence embeddings. In Proceedings of the 2021 Conference on Empirical Methods in Natu- ral Language Processing, pages 6894–6910. Sebastian Gehrmann, Hendrik Strobelt, and Alexander M Rush. 2019. Gltr: Statistical de- tection and visualization of generated text. In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Sys- tem Demonstrations, pages 111–116. Biyang Guo, Xin Zhang, Ziyuan Wang, Minqi Jiang, Jinran Nie, Yuxuan Ding, Jianwei Yue, and Yupeng Wu. 2023. How close is chatgpt to human experts? comparison corpus, evaluation, and detection. arXiv preprint arxiv:2301.07597. Felix Hamborg, Norman Meuschke, Corinna Bre- itinger, and Bela Gipp. 2017. news-please: A generic news crawler and extractor. In Proceed- ings of the 15th International Symposium of In- formation Science, pages 218–223. Abhimanyu Hans, Avi Schwarzschild, Valeriia Cherepanova, Hamid Kazemi, Aniruddha Saha, Micah Goldblum, Jonas Geiping, and Tom Gold- stein. 2024. Spotting llms with binoculars: Zero- shot detection of machine-generated text. In Forty-first International Conference on Machine Learning. Xinlei He, Xinyue Shen, Zeyuan Chen, Michael Backes, and Yang Zhang. 2024. Mgtbench: Benchmarking machine-generated text detection. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, pages 2251–2265. Ari Holtzman, Jan Buys, Li Du, Maxwell Forbes, and Yejin Choi. 2019.The curious case of neural text degeneration.arXiv preprint arXiv:1904.09751. Xiaomeng Hu, Pin-Yu Chen, and Tsung-Yi Ho. 2023. Radar: Robust ai-text detection via adver- sarial learning. Advances in Neural Information Processing Systems, 36:15077–15095. Andrew Ilyas, Shibani Santurkar, Dimitris Tsipras, Logan Engstrom, Brandon Tran, and Aleksander Madry. 2019. Adversarial examples are not bugs, they are features. Advances in neural informa- tion processing systems, 32. Daphne Ippolito, Daniel Duckworth, Chris Callison-Burch, and Douglas Eck. 2020. Auto- matic detection of generated text is easiest when humans are fooled. In Proceedings of the 58th Annual Meeting of the Association for Computa- tional Linguistics, pages 1808–1822. Robin Jia and Percy Liang. 2017. Adversarial ex- amples for evaluating reading comprehension systems. arXiv preprint arXiv:1707.07328. Di Jin, Zhijing Jin, Joey Tianyi Zhou, and Peter Szolovits. 2020. Is bert really robust? a strong baseline for natural language attack on text clas- sification and entailment. In Proceedings of the AAAI conference on artificial intelligence, vol- ume 34, pages 8018–8025. Davinder Kaur, Suleyman Uslu, Kaley J Rittichier, and Arjan Durresi. 2022. Trustworthy artificial intelligence: a review. ACM Computing Surveys (CSUR), 55(2):1–38. Walter Kintsch and Teun A Van Dijk. 1978. To- ward a model of text comprehension and produc- tion. Psychological review, 85(5):363. John Kirchenbauer, Jonas Geiping, Yuxin Wen, Jonathan Katz, Ian Miers, and Tom Goldstein. 2023. A watermark for large language models. In International Conference on Machine Learn- ing, pages 17061–17084. PMLR. Moshe Koppel, Jonathan Schler, and Shlomo Arg- amon. 2009. Computational methods in author- ship attribution. Journal of the American Society for information Science and Technology, 60(1):9– 26. Kalpesh Krishna, Yixiao Song, Marzena Karpin- ska, John Wieting, and Mohit Iyyer. 2024. Para- phrasing evades detectors of ai-generated text, but retrieval is an effective defense. Advances in Neural Information Processing Systems, 36. Rahul Kumar, Sarah Elaine Eaton, Michael Mindzak, and Ryan Morrison. 2024. Academic integrity and artificial intelligence: An overview. Second handbook of academic integrity, pages 1583–1596. Jooyoung Lee, Thai Le, Jinghui Chen, and Dong- won Lee. 2023. Do language models plagia- rize? In Proceedings of the ACM Web Confer- ence 2023, pages 3637–3647. Yafu Li, Qintong Li, Leyang Cui, Wei Bi, Zhilin Wang, Longyue Wang, Linyi Yang, Shuming Shi, and Yue Zhang. 2024. Mage: Machine- generated text detection in the wild. In Proceed- ings of the 62nd Annual Meeting of the Associ- ation for Computational Linguistics (Volume 1: Long Papers), pages 36–53. Shengchao Liu, Xiaoming Liu, Yichen Wang, Ze- hua Cheng, Chengzhengxu Li, Zhaohan Zhang, Yu Lan, and Chao Shen. 2024. Does detectgpt fully utilize perturbation? bridging selective per- turbation to fine-tuned contrastive learning detec- tor would be better. In Proceedings of the 62nd Annual Meeting of the Association for Compu- tational Linguistics (Volume 1: Long Papers), pages 1874–1889. Yinhan Liu, Myle Ott, Naman Goyal, Jingfei Du, Mandar Joshi, Danqi Chen, Omer Levy, Mike Lewis, Luke Zettlemoyer, and Veselin Stoyanov. 2019.Roberta: A robustly opti- mized bert pretraining approach. arXiv preprint arXiv:1907.11692. Kyle Lo, Lucy Lu Wang, Mark Neumann, Rodney Kinney, and Daniel Weld. 2020. S2ORC: The semantic scholar open research corpus. In Pro- ceedings of the 58th Annual Meeting of the As- sociation for Computational Linguistics, pages 4969–4983, Online. Association for Computa- tional Linguistics. Muneer M Alshater. 2022. Exploring the role of artificial intelligence in enhancing academic per- formance: A case study of chatgpt. Available at SSRN. Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2017.Towards deep learning models resis- tant to adversarial attacks.arXiv preprint arXiv:1706.06083. Chengzhi Mao, Carl Vondrick, Hao Wang, and Jun- feng Yang. 2024. Raidar: generative ai detection via rewriting. In The Twelfth International Con- ference on Learning Representations. Elyas Masrour, Bradley Emi, and Max Spero. 2025.Damage:Detecting adversarially modified ai generated text.arXiv preprint arXiv:2501.03437. Quinn McNemar. 1947. Note on the sampling error of the difference between correlated proportions or percentages. Psychometrika, 12(2):153–157. Eric Mitchell, Yoonho Lee, Alexander Khazatsky, Christopher D Manning, and Chelsea Finn. 2023. Detectgpt: Zero-shot machine-generated text de- tection using probability curvature. In Interna- tional Conference on Machine Learning, pages 24950–24962. PMLR. Andrea Moro, Marco Tettamanti, Daniela Perani, Caterina Donati, Stefano F Cappa, and Ferruccio Fazio. 2001. Syntax and the brain: disentangling grammar by selective anomalies. Neuroimage, 13(1):110–118. Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. 2002. Bleu: a method for au- tomatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the As- sociation for Computational Linguistics, pages 311–318. Mike Perkins. 2023. Academic integrity consider- ations of ai large language models in the post- pandemic era: Chatgpt and beyond. Journal of University Teaching and Learning Practice, 20(2). Jiameng Pu, Zain Sarwar, Sifat Muhammad Ab- dullah, Abdullah Rehman, Yoonjin Kim, Paran- tapa Bhattacharya, Mobin Javed, and Bimal Viswanath. 2023. Deepfake text detection: Lim- itations and opportunities. In 2023 IEEE Sympo- sium on Security and Privacy (SP), pages 1613– 1630. IEEE. Marco Tulio Ribeiro, Sameer Singh, and Carlos Guestrin. 2018. Semantically equivalent adver- sarial rules for debugging nlp models. In Pro- ceedings of the 56th Annual Meeting of the Asso- ciation for Computational Linguistics (volume 1: long papers), pages 856–865. Vinu Sankar Sadasivan, Aounon Kumar, Sriram Balasubramanian, Wenxiao Wang, and Soheil Feizi. 2023. Can ai-generated text be reliably detected? arXiv preprint arXiv:2303.11156. Magnus Sahlgren. 2008. The distributional hypoth- esis. Italian Journal of linguistics, 20:33–53. Irene Solaiman, Miles Brundage, Jack Clark, Amanda Askell, Ariel Herbert-Voss, Jeff Wu, Alec Radford, Gretchen Krueger, Jong Wook Kim, Sarah Kreps, et al. 2019. Release strate- gies and the social impacts of language models. arXiv preprint arXiv:1908.09203. Jinyan Su, Terry Yue Zhuo, Di Wang, and Preslav Nakov. 2023.Detectllm: Leverag- ing log rank information for zero-shot detec- tion of machine-generated text. arXiv preprint arXiv:2306.05540. Lichao Sun, Yue Huang, Haoran Wang, Siyuan Wu, Qihui Zhang, Chujie Gao, Yixin Huang, Wen- han Lyu, Yixuan Zhang, Xiner Li, et al. 2024. Trustllm: Trustworthiness in large language models. arXiv preprint arXiv:2401.05561. Adaku Uchendu, Thai Le, Kai Shu, and Dongwon Lee. 2020. Authorship attribution for neural text generation. In Proceedings of the 2020 confer- ence on empirical methods in natural language processing (EMNLP), pages 8384–8395. Vivek Verma, Eve Fleisig, Nicholas Tomlin, and Dan Klein. 2024. Ghostbuster: Detecting text ghostwritten by large language models. In Pro- ceedings of the 2024 Conference of the North American Chapter of the Association for Com- putational Linguistics: Human Language Tech- nologies (Volume 1: Long Papers), pages 1702– 1717. Yichen Wang, Shangbin Feng, Abe Bohan Hou, Xiao Pu, Chao Shen, Xiaoming Liu, Yulia Tsvetkov, and Tianxing He. 2024. Stumbling blocks: Stress testing the robustness of machine- generated text detectors under attacks. arXiv preprint arXiv:2402.11638. Junchao Wu, Runzhe Zhan, Derek F Wong, Shu Yang, Xinyi Yang, Yulin Yuan, and Lidia S Chao. 2024. Detectrl: Benchmarking llm-generated text detection in real-world scenarios. In The Thirty-eight Conference on Neural Information Processing Systems Datasets and Benchmarks Track. Yang Xu, Yu Wang, Hao An, Zhichen Liu, and Yongyuan Li. 2024. Detecting subtle differences between human and model languages using spec- trum of relative likelihood. In Proceedings of the 2024 Conference on Empirical Methods in Natu- ral Language Processing, pages 10108–10121. Duanli Yan, Michael Fauss, Jiangang Hao, and Wenju Cui. 2023. Detection of ai-generated es- says in writing assessment. Psychological Test- ing and Assessment Modeling, 65(2):125–144. Xianjun Yang, Wei Cheng, Yue Wu, Linda Ruth Petzold, William Yang Wang, and Haifeng Chen. 2023. Dna-gpt: Divergent n-gram analysis for training-free detection of gpt-generated text. In The Twelfth International Conference on Learn- ing Representations. William J Youden. 1950. Index for rating diagnos- tic tests. Cancer, 3(1):32–35. Xiao Yu, Yuang Qi, Kejiang Chen, Guoqiang Chen, Xi Yang, Pengyuan Zhu, Xiuwei Shang, Weim- ing Zhang, and Nenghai Yu. 2024. Dpic: Decou- pling prompt and intrinsic characteristics for llm generated text detection. Advances in Neural In- formation Processing Systems, 37:16194–16212. Ann Yuan, Andy Coenen, Emily Reif, and Daphne Ippolito. 2022. Wordcraft: story writing with large language models. In Proceedings of the 27th International Conference on Intelligent User Interfaces, pages 841–852. Xuandong Zhao, Sam Gunn, Miranda Christ, Jaiden Fairoze, Andres Fabrega, Nicholas Car- lini, Sanjam Garg, Sanghyun Hong, Milad Nasr, Florian Tramer, et al. 2024a. Sok: Watermark- ing for ai-generated content. arXiv preprint arXiv:2411.18479. Xuandong Zhao, Lei Li, and Yu-Xiang Wang. 2024b. Permute-and-flip: An optimally robust and watermarkable decoder for llms. arXiv preprint arXiv:2402.05864. Xuandong Zhao, Yu-Xiang Wang, and Lei Li. 2023. Protecting language generation models via in- visible watermarking. In International Confer- ence on Machine Learning, pages 42187–42199. PMLR. Ying Zhou, Ben He, and Le Sun. 2024. Humaniz- ing machine-generated content: Evading ai-text detection through adversarial attack. In Proceed- ings of the 2024 Joint International Conference on Computational Linguistics, Language Re- sources and Evaluation (LREC-COLING 2024), pages 8427–8437. Xiaowei Zhu, Yubing Ren, Fang Fang, Qingfeng Tan, Shi Wang, and Yanan Cao. 2025. Dna- detectllm: Unveiling ai-generated text via a dna-inspired mutation-repair paradigm. In The Thirty-ninth Annual Conference on Neural Infor- mation Processing Systems. A Commercial Humanizing Tools There are various AI humanizing tools that are developed to bypass detectors. We list a few in Table 9, where the first three are used to produce our humanized texts. AI ToolURLUsed BypassGPThttps://bypassgpt.ai/Y Humbothttps://humbot.ai/Y Undetectable AIhttps://undetectable.ai/Y Semihuman AIhttps://semihuman.ai/ HIX Bypasshttps://bypass.hix.ai/ AI Humanizerhttps://aihumanizer.ai/ StealthGPThttps://stealthgpt.ai/ GPTinfhttps://stealthgpt.ai/ WriteHumanhttps://writehuman.ai/ StealthWriterhttps://rewritify.ai/ Phrasly LLChttps://phrasly.ai/ HIX.AIhttps://bypass.hix.ai AISEO Humanizerhttps://aiseo.ai/ Humanize AI Prohttps://w.humanizeai.pro/ Smodinhttps://smodin.io/ Rewritifyhttps://w.rewritify.ai Table 9: Commercial humanizing tools. B Construction of Humanize-16K We create the benchmark dataset following a strict construction process and thorough quality assur- ance. B.1 Collection of Human-Written Texts Student EssayWe randomly select 1,000 essays from the Automated Student Assessment Prize (ASAP) 2.0 (Crossley et al., 2024), each accompa- nied by a title and a prompt. These prompts are utilized to prompt LLMs to generate correspond- ing essays. Additionally, metadata such as ‘race ethnicity’, ‘gender’, and ‘grade level’ are recorded for potential future analyses. ArXiv Intro To build this dataset, we collect 1,000 computer science papers from arXiv (arXiv, 2024) by crawling PDFs published between 2020 and 2024, randomly selecting 200 papers per year. Using S2ORC (Lo et al., 2020), the PDFs are parsed to extract titles and introductions. These titles are then used to prompt LLMs to generate new paper introductions. Creative Writing We randomly pull 1,000 sam- ples from WritingPrompts (Fan et al., 2018), with each sample paired with a corresponding prompt. These prompts serve as triggers for LLMs to create new fictional stories. C NewsFor this dataset, we gather 1,000 news articles in English sourced from Common Crawl (Hamborg et al., 2017). The news headlines are used to prompt LLMs to generate full news articles. B.2 Produce AI-Generated Texts We create AI-generated texts using titles or prompts derived from human-written content. For instance, in the case of student essays, we instruct LLMs with a prompt such as: “Write a student essay (no title) in nwords words (split into nparagraphs paragraphs) based on the given title: title”. To ensure that the generated texts closely match the average length of human-written texts, we specify the same number of words (or characters for Chi- nese) and paragraphs in the prompt. The detailed prompts for all domains can be found in Table 10. Table 10: Prompts for data generation, where the field could be either ‘title’ or ‘prompt’ depending on their availability for each data source. Student Essay: Write a student essay (no title) in n_words words (split into n_paragraphs para- graphs) based on the given field: field_value ArXiv Intro: Write an introductory section (no section name) for an academic paper in n_words words (split into n_paragraphs paragraphs) based on the given field: field_value Creative Writing: Write a creative story (no title) in n_words words (split into n_paragraphs para- graphs) based on the given field: field_value C News: Write a news article (no title) in n_words words (split into n_paragraphs para- graphs) based on the given field: field_value Multi-lingual C News: Write a news article (no title) in lang language in n_words words (split into n_paragraphs paragraphs) based on the given field: field_value We use six language models – gpt-3.5-turbo, gpt-4o, claude-3.5-sonnet, gemini-1.5-pro, llama- 3.3-70b-instruct, and qwen-2.5-72b-instruct – to generate data, with a random model selected for each sample. In terms of decoding parameters, a temperature is randomly chosen from the range [0.8, 1.0, 1.2], a top-pfrom[0.96, 1.0], and both frequency and presence penalties from the range [0.0, 1.0] for each sample. B.3 Quality Assurance We evaluate the length of each generation from the LLM output, and if it is significantly longer (more than twice the original length) or shorter (less than half the original length), we prompt the LLM to generate the text again. Additionally, we monitor for issues like repetition or nonsensical responses and address them by regenerating the text. After processing the data, we truncate the texts to ensure that the length distributions are consistent across different types. As a final quality check, we randomly select 100 samples per domain for manual review, achieving an average pass rate of 99.5%. C Ethical Considerations The dataset we create contains AI-generated texts, which may occasionally exhibit bias, offensive lan- guage, or irresponsibility. While our manual review of 100 samples per domain achieved a 100% pass rate, there remains a small possibility that some content in the broader dataset could be less than ideal. However, this risk is mitigated by the rigor- ous quality control measures we have implemented, and any concerns can be addressed with appropri- ate disclaimers or warnings. Additionally, while AI tools are increasingly used to identify whether an article is AI-generated, we believe that these tools should be viewed as helpful aids rather than definitive authorities. Over- reliance on such technology could lead to inaccura- cies or misuse, and we advocate for incorporating human judgment as an essential step in verifying the origin of any piece of work. This balanced approach ensures responsible use of AI detection tools while minimizing potential ethical concerns.