Paper deep dive
The Misclassification of Autistic Writing as AI-Generated
Summer Chambers, Matthew C. Kelley
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 88%
Last extracted: 7/18/2026, 9:43:23 AM
Summary
This study empirically investigates the hypothesis that AI-detection models exhibit bias against autistic writers by flagging their text as AI-generated. Using a corpus of approximately 60,000 Reddit posts divided into 'likely-autistic' and 'general-Reddit' subcorpora, the authors analyzed outputs from the OpenAI GPT-2 detection model. Results indicate that while the overall false positive rate was low (<2%), texts from the likely-autistic subcorpus were significantly more likely to be flagged as AI-generated compared to the general subcorpus, even after controlling for textual features like word count and perplexity. The study highlights ethical concerns regarding the use of such biased detection tools in academic and professional contexts.
Entities (8)
Relation Signals (5)
likely-autistic subcorpus â hashigherfalsepositiverate â OpenAI GPT-2 detection model
confidence 92% · Based on the odds-ratio, posts from the likely-autistic subcorpus had a 25% greater chance of being classified as AI-generated.
OpenAI GPT-2 detection model â exhibitsbiasagainst â autistic writers
confidence 90% · Results showed that while less than two-percent of either subcorpus was flagged as AI-generated by the model, significantly more texts from the likely-autistic subcorpus were flagged.
perplexity â negativelycorrelateswith â AI flagging probability
confidence 85% · The model also showed a significant negative effect of perplexity... Posts with higher perplexity were less likely to be flagged as AI.
word count â negativelycorrelateswith â AI flagging probability
confidence 85% · Shorter posts tended to be flagged as AI more often than longer posts.
autistic writers â producetextwith â higher perplexity
confidence 75% · Perplexity was close to equal with the likely-autistic subcorpus trending only slightly higher in both.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Recent findings suggest that detection models for artificial intelligence (AI) cannot accurately identify AI-generated text and may exhibit bias against certain minority groups. In the present study, anecdotal claims that autistic writers more often have their work flagged as AI-generated are examined empirically. A corpus of approximately 60,000 Reddit posts split into "likely-autistic" and "general-Reddit" subcorpora is used to compare the distribution of probabilities output by the OpenAI GPT-2 detection model. Differences in textual features between subcorpora are observed and compared to reported features of AI-generated text. Results showed that while less than two-percent of either subcorpus was flagged as AI-generated by the model, significantly more texts from the likely-autistic subcorpus were flagged. Connections between features of text with likely-autistic authors and AI-generated text were not straightforward. The widespread use of AI-detection models with a potential bias against autistic writers in their output prompts ethical scrutiny, and the authors recommend further critical examination of the models themselves as well as their use in academic contexts.
Tags
Links
- Source: https://arxiv.org/abs/2607.14729v1
- Canonical: https://arxiv.org/abs/2607.14729v1
Trouble viewing inline? Open PDF directly â
Full Text
41,278 characters extracted from source content.
Expand or collapse full text
The misclassification of autistic writing as AI-generated Summer Chambers and Matthew C. Kelley George Mason University, Fairfax VA 22030, USA schamb3@gmu.edu Abstract. Recent findings suggest that detection models for artificial intelligence (AI) cannot accurately identify AI-generated text and may exhibit bias against certain minority groups. In the present study, anec- dotal claims that autistic writers more often have their work flagged as AI-generated are examined empirically. A corpus of approximately 60,000 Reddit posts split into âlikely-autisticâ and âgeneral-Redditâ sub- corpora is used to compare the distribution of probabilities output by the OpenAI GPT-2 detection model. Differences in textual features be- tween subcorpora are observed and compared to reported features of AI-generated text. Results showed that while less than two-percent of either subcorpus was flagged as AI-generated by the model, significantly more texts from the likely-autistic subcorpus were flagged. Connections between features of text with likely-autistic authors and AI-generated text were not straightforward. The widespread use of AI-detection mod- els with a potential bias against autistic writers in their output prompts ethical scrutiny, and the authors recommend further critical examination of the models themselves as well as their use in academic contexts. Keywords: AI Detection· Autism· Large Language Models· GPT 1 Introduction In the age of ChatGPT and other large language models (LLMs), schools and publishers have begun using tools to detect text written by some kind of Artificial Intelligence (AI). As ethics questions surround the creation and use of models like ChatGPT, fairness issues related to inaccurate or biased AI-detection tools have also become prominent in discourse [15]. Concerns around AI-detection models have been bolstered by studies revealing their very low accuracy rates [6,39,40] as well as evidence of bias in AI detection against certain groups. In a 2023 survey, 10% of teenagers in the US reported having been falsely accused of using AI to write their assignments. Disturbingly, twice as many Black teenagers as white and Latino teenagers reported being falsely accused [27]. Additionally, Liang et al. [25] showed empirically that AI-detection models are more likely to flag texts written by non-native English speakers. While there is evidence of this tendency, anecdotal claims that such models disproportionately flag autistic peopleâs writing as AI-generated have not been formally investigated. In this This version of the contribution has been accepted for publication after peer review but is not the Version of Record and does not reflect post-acceptance improvements, or any corrections. The Version of Record is available online at: https://doi.org/10.1007/978-3-031-98420-4_7. Use of this Accepted Version is subject to the publisherâs Accepted Manuscript terms of use: https://w.springernature.com/gp/open-research/policies/accepted-manuscript-terms. arXiv:2607.14729v1 [cs.CL] 16 Jul 2026 2 paper, we describe an experiment probing the OpenAI GPT-2 detection model for a false positive bias against likely-autistic writersâ posts on Reddit. 1.1 AI Detection Models The GPT-2 detector used in this experiment was released in 2019 by OpenAI, the creators of the GPT series [35]. The classifier purported to detect text generated with OpenAIâs GPT-2 language model with 95% accuracy, though later evalu- ations indicated much lower rates [32]. Newer iterations of AI-detection models have since been released and subsequently deleted by OpenAI as accuracy issues persist and evidence of biases are revealed [22]. While many major universities in the US and UK are now cautioning their professors against using automated AI-detection tools, free and paid tools such as CopyLeaks, Turnitin, GoWinston, and GPTZero boast largely unverified ac- curacy rates up to 99.6% and are still in high demand [15]. Despite empirically supported claims that AI-detection models are âneither accurate nor reliableâ [40, p. 28], their use is not explicitly prohibited in educational or publishing contexts. In fact, a 2024 survey reported that two thirds of American teachers used them regularly [11]. While newer AI-detection technology called watermarking could have the potential to be much more accurate, OpenAI has not publicly released any new tools, claiming to be wary of stigmatizing the use of AI for groups who rely on it to improve their writing, such as non-native English speakers [31]. 1.2 Understanding AI Detector Predictions Unlike plagiarism detectors, current AI-detection models cannot cite evidence of the phenomenon they attempt to detect, so their outputs cannot be cross- checked empirically [7]. It is quite hard to know what features of AI-generated text detection models rely on since they only return a prediction (âRealâ or âFakeâ) and a probability score (often taken as a modelâs degree of certainty in its prediction, though this interpretation can be misleadingâsee [17]). One can learn a lot about a model from the data it was trained on. However, with large, proprietary models, data sources and statistics about those sources are rarely available. After Liang et al. [25] brought to light evidence of a bias against non-native English speakers in several popular AI-detectorsâincluding the OpenAI GPT-2 detectorâAI and plagiarism-detection company Turnitin reported that a sig- nificant difference in false positives for non-native vs. native English-speaking authors only existed for short-form texts (under the 300-word minimum sug- gested by Turnitin) [2,6]. Most AI-detectors recommend a minimum number of characters or words for their input, citing poor reliability under that length. It has also been reported that low values of perplexity and burstinessâfrequently associated with AI-generated textâare common in non-native English speakersâ writing [25]. Similar claims regarding perplexity and burstiness have not been made for autistic peopleâs writing, but there is anecdotal evidence that humans and AI-detectors alike may be prone to mistake autistic writing styles for AI. 3 1.3 Autistic Experiences with AI-Detection Kling [23] reported on a university professorâs experience of being falsely accused of using AI to write their emails. The professor believed the accusation could be attributed to being autistic and mentioned that receiving such allegations is a common experience for autistic people [23]. An autistic student falsely accused by her professor had similar observations, citing the "formulaic" nature of her own writing and its potential similarity to AI [9]. Gegg-Harrison and Quarterman [14] used a small corpus of their own writing to test the false positive rates of several different popular AI-detection tools and saw much higher false positive rates than those reported by the makers of the tools themselves. They go on to discuss their own neurodivergence as a potential factor in these results, noting that a number of neurodivergent students and writers with similar suspicions reached out to them to share stories and fears related to false accusations of AI usage in their writing. 1.4 Exploring Autistic Writing Styles A surprisingly small amount of descriptive research has been conducted on so- ciolinguistic and stylistic differences in autistic language, despite considerable literature on diagnostic linguistic traits of Autism Spectrum Disorder (ASD) largely limited to children and spokenâas opposed to writtenâlanguage. Some of these linguistic traits are mentioned in diagnostic coding for the ADOS-2 [43] and in the DSM-5 [3], advising diagnosticians to look out for language that is âstereotypedâ, âformalâ, ârigidâ, ârepetitiveâ, and âpedanticâ. While many re- searchers group these linguistic traits under the umbrella of pragmatic âdefi- ciencyâ [24], others approach the topic more neutrally, highlighting only the divergence from the norm or majority [8,42]. Recent research spurred by autistic and other disability self-advocacy groups explicitly frames these features as âdif- ferencesâ rather than âdeficienciesâ [30]. The finding that autistic people without language impairment communicate as effectively and efficiently as neurotypical people do when in a peer group of other autistic people [8] bolsters the claim that autism is characterized by stylistic language differences rather than pragmatic deficiencies. At least two recent studies have attempted to make use of differences in writ- ten language to train machine learning models to predict whether or not tweets were authored by autistic users on social media platform X (formerly Twitter) using only the texts of tweets [21,34]. While they initially claimed that the use case for such classifier models would be assisting with early autism diagnosis, Jaiswal & Washington released an addendum to their original paper and a letter to the editor discussing serious ethical concerns with digital phenotyping and the potential for harm involved in creating such models [19,20]. They also ac- knowledge a lack of explicit consent from users who authored the data that was collected, which ultimately motivated them to delete their models and data. 4 1.5 The Present Paper In this paper, we test the hypothesis that false positives in the outputs of the OpenAI GPT-2 detection model are more common for text written by autistic people. We run a variety of Reddit posts from autism-focused and general dis- cussion subreddits through the detector and compare its predicted probabilities of AI-generated content between the two groups of texts. 2 Data While research involving public Reddit posts does not meet the criteria for sub- mission to the authorsâ Institutional Review Board, ethical concerns arise when collecting data without explicit consent from its creators, especially in marginal- ized communities. Though this dataset contains no user-identifying information, it will not be made publicly available given the potential for malicious use of such data, as is discussed in Jaiswal et al. [20]. See additional research [1,12] on ethical considerations for projects using âscrapedâ Reddit posts. 2.1 Data Collection The goal of the data collection step was to gather text written in a similar domain by autistic and non-autistic people. It is difficult to have high certainty about the accuracy of such a division, since not all autistic people identify themselves publicly as such anywhere their writing appears. The dataset used in the current study suffers from this low degree of certainty but was chosen for the quantity of publicly available and easily categorized texts. We identified 13 subreddits dedicated to discussing topics related to autism on the social media platform Reddit. The descriptions and rules of each of these subreddits make it clear that their purpose is to offer autistic people a space to discuss various facets of their lives. While some subreddits ask that only autistic people or only officially-diagnosed autistic people post and comment, others welcome posts from non-autistic people when made in good faith. Still, looking through a random sample of recent posts from each subreddit, the overwhelming majority are written by people who identify themselves as autistic. As in all subreddits, posters who do not follow the rules of each designated subreddit may have their posts deleted by the subreddit moderators. This means we can be slightly more certain that posts in these subreddits are going to âbelongâ or fit the theme of the subreddit. To curate a comparison group representing the general Reddit population, we selected 12 popular subreddits judged to be similar to the autism-related subreddits in terms of post style and format. These posts tended to be relatively long first-person narratives discussing topics related to mental health, social situations, embarrassing or surprising stories, solicitations for advice or comfort, etc. Importantly, a binary distinction of âautisticâ and ânon-autisticâ cannot be drawn between these two subcorpora, since anyone of any neurotype can join 5 and post in any subreddit. There will be some non-autistic people posting in the subreddits targeted towards autistic people and many autistic people posting in the general subreddits. For this reason, we will refer to the two subcorpora as âlikely-autisticâ and âgeneral-Redditâ hereafter. Figure 1 shows the distribution of posts per subreddit present in the corpus, largely dictated by the availability of data. 0 1000 2000 3000 4000 5000 autism aspergers AutismInWomen aspergirls AutisticAdults AutismTranslated SpicyAutism AutisticPride evilautism AskAutism AutisticPeeps AutismCertified Autistic Number of Posts Likely Autistic Corpus 0 1000 2000 3000 4000 5000 AITAH tifu TwoXChromosomes mentalhealth self rant GetMotivated women relationship_advice TrueOffMyChest therapy AskReddit Subreddit General Reddit Corpus Fig. 1. Distribution of subreddits from which Reddit posts were collected for âlikely- autisticâ and âgeneral-Redditâ subcorpora We collected data with the Python PRAW library [5], a wrapper for the Reddit API, which only returned posts from approximately 2021-2024. With the goal of adding a larger quantity of data and less recent data, we downloaded a large set of Reddit posts from a Pushshift archive [4]. This archive provided much more data than what we could get through PRAW alone, including posts from 2010-2020. Unfortunately, far fewer of the 13 subreddits previously identified for the autism-related subcorpus existed in the older archive. For the sake of quantity, we chose to combine the more recent and older corpora, ending up with around 60,000 Reddit posts to work with. This dataset was later halved in size, approximately, after eliminating posts with a word count of 300 or less. Little is known about the demographics of the authors of these Reddit posts. Reddit users in these communities may be of any age, including early adolescents and adults. The overwhelming majority of texts are written in English, though varied generational, socioeconomic, cultural, and regional dialects as well as dif- fering degrees of English proficiency are expected. All of these factors could complicate this experiment, though these issues are not easily mitigated given the anonymity of Reddit posts. 6 2.2 Preprocessing and Filtering We used the RoBERTa AutoTokenizer which pairs with the OpenAI GPT-2 detection model for the word tokenization step. This involved truncating posts to 480 tokens, which is just below the token maximum for inputs to the OpenAI model. We also used the NLTK punkt sentence tokenizer [26] to break posts up into sentences, though this step was purely for collecting descriptive statistics about the posts. In an effort to exclude as many autistic writers from the general-Reddit group as possible, we excluded posts from the general-Reddit group whose authors appeared in the likely-autistic group or which included keywords about autism in the text. In our manual review of randomly sampled posts from each subcorpus, we noticed that several of the older posts from the archived Pushshift source made in the r/autism and r/aspergers subreddits were authored by people who did not actually identify as autistic but were discussing autistic relatives. Because of this, we chose to filter out posts containing keywords about autism paired with phrases like âmy daughter. . . â, âmy nephew. . . â, etc. in an attempt to limit the number of posts from authors who are not necessarily autistic themselves. This filtering process will not catch all cases and will unnecessarily exclude some posts made by autistic people discussing relatives. We limited our dataset to one randomly chosen post from each user to avoid over-representing any one author. We also excluded posts with fewer than 1000 characters (200 words, roughly), since OpenAI claims this is the minimum length required for an accurate prediction. 2.3 Corpus Statistics Motivated by a desire to explore textual differences between the subcorpora and their effect on AI-detection probability, we chose to compute several descriptive statistical measures for each subcorpus. For AI-generated text, both perplexity and burstiness tend to be low [6]. Perplexity can be loosely understood as the modelâs degree of âsurpriseâ or difficulty predicting the next word in the text. Burstiness is a term often used to represent the non-random distribution of a particular word in text, but in the context of AI-detection, it has been described as the degree of variation in sentence lengths and structures. Perplexity was computed using the evaluate Python package created by HuggingFace [41]. Burstiness was calculated as the coefficient of variation of sentence lengths for each post, modeled after the calculation used in the zippy Python package [36]. In addition, we computed the average word length in characters and average sentence length in words for each post in the corpus. Most notably, the general-Reddit subcorpus had a much larger mean word count than the likely-autistic subcorpus. Mean word length and sentence length were slightly higher in the likely-autistic subcorpus, and both perplexity and burstiness were close to equal with the likely-autistic subcorpus trending only slightly higher in both. Figure 2 depicts the differences in textual features be- tween the two subcorpora. 7 350 400 450 General RedditLikely Autistic Word Count 4.0 4.5 5.0 General RedditLikely Autistic Mean Word Length 10 20 30 General RedditLikely Autistic Mean Sentence Length 10 20 30 40 50 60 General RedditLikely Autistic Perplexity 0.2 0.4 0.6 0.8 1.0 General RedditLikely Autistic Subcorpus Number of Words Mean Word Length (Characters) Mean Sentence Length (Words) Perplexity Score Burstiness Score Burstiness Fig. 2. Descriptive statistics regarding textual features for each subcorpus 3 Experiment 1 3.1 Methods While there are several AI-detection tools on the market, we chose to use Ope- nAIâs RoBERTa GPT-2 detector [35] as it is freely available for download on the platform HuggingFace and thus may be more commonly used. After down- loading the model locally, we ran all posts from each Reddit subcorpus through the AI-detection model to generate both binary (Fake/Real) predictions and decimal/percentage probabilities of each post being AI-generated. Predictions labeled âFakeâ indicate an AI-probability value greater than 0.5. There is a real risk that some of these posts actually were generated by AI, and an unequal distribution of truly AI-generated posts between the two subcorpora would compromise this experiment. One method of mitigating this risk would be to limit the data to only posts made before 2020, as GPT-2âone of the first widely available LLMsâwas publicly released at the end of 2019. Still, due to the limitations of the older subcorpus, we chose to combine it with the newer subcorpus and included a binary âdateâ variable in the dataset which tracks whether or not a post was written after January 1, 2020. In this experiment, 1.7% of all 59,947 posts were flagged by the model as AI- generated. When split by subcorpus, the likely-autistic group had 1.9% of posts flagged as AI-generated, and the general-Reddit group had 1.5% flagged as such. To investigate the significance of this difference, as well as the impact of various textual features on the modelâs probability outputs, we fit a logistic regression model with the detection modelâs AI probability score as the outcome variable. The subcorpus (likely-autistic or general-Reddit) and several other variables (mean sentence length, mean word length, perplexity score, burstiness score, 8 and a binary indication of whether the post was made after 2020) were all in- cluded as predictors. To fit this model, we used the glm function in R [33] with a binomial distribution and no random effects or interaction terms. 3.2 Results and Discussion Table 1 shows the regression coefficients and Figure 3 shows the effect plots for significant variables of this model. We see that the subcorpus a post came from was a significant predictor in determining whether or not it was flagged as AI-generated. Based on the odds-ratio, posts from the likely-autistic subcorpus had a 25% greater chance of being classified as AI-generated. The model also showed a significant negative effect of perplexity, which matches expectations. Posts with higher perplexity were less likely to be flagged as AI. Lastly, we see that the length of each post was significant in determining whether or not a post would be flagged as AI-generated. Shorter posts tended to be flagged as AI more often than longer posts. This tracks with the conventional wisdom that AI-detectors are less accurate with shorter texts. If the model tends to more often flag shorter texts as AI-generated, even above the 1000-character minimum threshold suggested, this is something worth exploring in greater depth. Table 1. Table of coefficients for multiple logistic regression model predicting proba- bility of text being AI-generated. VariableEstimateStd. Errorz valueOdds RatioPr(>|z|) Interceptâ4.163 0.061â68.680 0.016 < 0.001 Likely Autistic 0.220 0.067 3.305 1.246 0.001 Word Countâ0.439 0.031â14.165 0.644 < 0.001 Date 0.094 0.067 1.402 1.098 0.161 Mean Word Lengthâ0.014 0.027â0.530 0.986 0.596 Mean Sentence Lengthâ0.016 0.035â0.457 0.984 0.648 Perplexity Scoreâ0.232 0.036â6.435 0.793 < 0.001 Burstiness Scoreâ0.012 0.031â0.406 0.988 0.685 4 Experiment 2 4.1 Methods Given that the likely-autistic corpus had a lower mean word count than the general-Reddit corpus and that the effect of word count was significant, we de- cided to pursue a secondary experiment neutralizing this variable. Based on the lengths of the available data and the recommended minimum lengths from other AI-detectors, we chose to limit the corpus only to posts longer than 300 words and subsequently truncate the text of each post to exactly 300 words. In this 9 0% 2% 4% 6% General RedditLikely Autistic Subcorpus Probability of AI 0% 2% 4% 6% 100200300400 Wordcount 0% 2% 4% 6% 0100200300 Perplexity Score Fig. 3. Effect plots for significant variables in Experiment 1: Subcorpus, Wordcount, Perplexity Score step, the dataset size was effectively halved, resulting in a subset ofâ33,000 of the originalâ60,000 posts. The modelâs 500-token limit constrained our choice of maximum word count, and a larger minimum word count would have de- creased the amount of data available too dramatically. Before running this set of truncated posts through the same AI-detection model as in Experiment 1, we re-computed the relevant descriptive statistics for each subcorpus for the text with normalized word counts. All other descriptive statistics showed the same comparative trends as in the previous dataset. The following results were observed after running the 300-word posts through the same AI-detection model as in Experiment 1: 1.4% of the 33,216 total posts were flagged as AI-generated (a slightly lower percentage than the initial exper- iment). 1.7% of posts from the likely-autistic subcorpus were flagged as AI, and 1.2% of posts from the general-Reddit subcorpus were flagged as such. We fit a new logistic regression model with this data, using the same design, method, and variables as before (with the exception of word count, which has been normalized across all posts). 4.2 Results and Discussion Table 2 is the table of coefficients for this regression, and effect plots for sig- nificant variables are found in Figure 4. This model showed that the subcorpus variable (likely-autistic or general-Reddit) again had a significant effect on the probability of AI-generation returned by the detection model and a slightly larger effect size. Based on the odds-ratio, posts from the likely-autistic subcorpus had a 50% greater chance of being classified as AI-generated. The âdateâ variable, a 10 binary indication of whether or not a Reddit post was submitted after Jan 1, 2020, also showed significance in this experiment with a small effect size. The effect of this variable on AI-probability was negative, indicating that posts writ- ten in or after 2020 were less likely to be flagged as AI compared to older posts. One possible explanation for the decrease in AI-generated predictions after 2020 is that the GPT-2 detector was not trained to detect texts written by newer LLMs such as ChatGPT, so any Reddit posts written with such tools might fly under its radar as false negatives. Table 2. Table of coefficients for multiple logistic regression model, Experiment 2 VariableEstimateStd. Errorz valueOdds RatioPr(>|z|) (Intercept)â4.146 0.075â55.562 0.016 < 0.001 Likely Autistic 0.406 0.096 4.226 1.501 < 0.001 Dateâ0.206 0.095â2.172 0.814 0.030 Mean Word Lengthâ0.023 0.046â0.495 0.977 0.621 Mean Sentence Lengthâ0.083 0.065â1.287 0.920 0.198 Perplexity Score 0.042 0.043 0.965 1.043 0.335 Burstiness Score 0.010 0.044 0.233 1.010 0.816 0.5% 1.0% 1.5% 2.0% 2.5% General RedditLikely Autistic Subcorpus Probability of AI 0.5% 1.0% 1.5% 2.0% 2.5% Before 2020After 2020 Date Fig. 4. Effect plots for significant variables in Experiment 2: Subcorpus, Date 11 5 General Discussion The fact that posts from the likely-autistic corpus were significantly more likely to be flagged as AI by this detection model prompts concern for all domains in which the model is used. False accusations of AI can cause students to suffer in terms of their academic/career standing and psychological well-being. Chaka argued that âany AI content probability percentage or percentage point, however negligible it may be [. . . ] inflicts immeasurable reputational damage to that essay and to the student who produced itâ [6, p. 10]. Globally, autistic people suffer from extremely high rates of unemployment [18], and reputational damage or limited educational opportunities caused by false accusations of AI will only have more devastating effects on employment rates and livelihoods. Gegg-Harrison and Quarterman [14] discuss the severe psychological impact on students caused by false accusations of cheating via AI, noting that autistic peopleâand other neurodivergent people such as those with ADHDâoften suffer from rejection sensitive dysphoria and difficulty regulating emotions, so the impacts of false accusations could be even more damaging. From another viewpoint, schools and companies could open themselves up to ableism and other discrimination lawsuits by using biased technology. If AI-detection tools finally fall out of fashion, there is still concern that individuals will take it upon themselves to decide whether or not a text has been written by AI. It has been shown in multiple contexts and domains that humans are no betterâand often worseâthan automated detection models at identifying AI-generated content [13]. Still, Verma & Tenjarla [37] reported that Ivy League admissions officers use automated AI-detection models as well as their own judgment to decide whether or not an essay was written with AI. One admissions officer detailed a valid set of criteria he used to spot AI-generated papers which eerily echoed descriptions of autistic narrative styles (see [24]). There are several experimental limitations intrinsic to the data collected for this experiment. The casual style of social media text in comparison to the more formal target material of AI-detectors may constrain any generalizations made. A later iteration of this project could also put a more objective focus on matching up the subcorpora by topic. The use of other publicly available corpora will still contend with the underlying uncertainty in identifying autistic and non-autistic authors. It is reasonable to have a high degree of certainty in autistic authorsâ self-identifications but very difficult to know who in the âgeneral populationâ may be autistic without identifying as such publicly. 3% of the US population have an autism diagnosis, according to figures from 2020 [28]. However, diagnosis rates are rising rapidly, particularly in young adults, women, and certain racial and ethnic minority groups [16]. These rising rates seem to be an artifact of the historically widespread under-diagnosis of autism. One study reported that 80% of autistic women were undiagnosed as of age 18 [29]. Given these complexities, a more controlled experiment would verify all participantsâ results of autism evaluations when designating the two comparison groups. The fact that we do see a significant trend of bias in the detection modelâs outputs even with a very noisy dataset inclines us to suspect that the difference 12 in false positive rates between autistic and non-autistic writers may actually be larger than what was observed here. Of course, It would be decidedly more informative to get predictions from more than one AI-detection tool, as mod- eled in several of the previously referenced experiments. Though certain costs and inconveniences prevented the inclusion of other AI-detection tools in this study, one other freely available AI-detection tool was briefly used [10] as an additional source of prediction data. Unfortunately, this tool produced the exact same predictions and probability scores as the OpenAI model for the same texts, rounded to the hundredth decimal place. While this is just one example, it is possible other tools on the market use the OpenAI model on the back end of their services, despite its bias and accuracy issues. Various seemingly incongruous descriptions of autistic language indicate that while autistic language is often more repetitive, stereotypical, and clichĂ©d, it is also observed to be idiosyncratic and contain more neologisms or oddly-worded phrases [38]. The former point prompts the assumption that autistic writing has lower perplexity and burstiness, but in our likely-autistic subcorpus, we found perplexity and burstiness to be roughly the same asâif not slightly higher thanâ the general-Reddit subcorpus. Perhaps the latter observation regarding idiosyn- crasy is the kernel of an explanation for why we did not see lower perplexity and burstiness. A look at the distribution of frequent lexical items and constructions in both subcorpora could give more insight into these questions. 6 Conclusion Considering the prevalence of AI-generated text detection tools, it is important to understand their rates of accuracy as well as how tendencies in their outputs might affect certain groups disproportionately. This study attempted to test the hypothesis that autistic writers are more likely to have their content flagged as AI-generated by a publicly available AI-detection model. Given a large cu- rated corpus of Reddit posts, it was shown that those posted in autism-centric subredditsâpresumably written by autistic peopleâwere more likely than posts from other subreddits to be flagged as AI-generated by OpenAIâs GPT-2 de- tector. These findings add further motivation to examine biases in other AI- detection models and limit or discontinue their use given the potential for harm. Acknowledgments. The authors thank James P. Blevins for his comments and thoughts on earlier versions of this project. Disclosure of Interests. The authors have no competing interests to report regarding the content of this paper. References 1. Adams, N.N.: âScrapingâ Reddit posts for academic research? Addressing some blurred lines of consent in growing internet-based research trend during the time 13 of Covid-19. International Journal of Social Research Methodology 27(1), 47â62 (Jan 2024). https://doi.org/10.1080/13645579.2022.2111816 2. Adamson, D.: New research: Turnitinâs AI detector shows nostatisticallysignificantbiasagainstEnglishLan- guageLearners(Oct2023),https://w.turnitin.com/blog/ new-research-turnitin-s-ai-detector-shows-no-statistically-significant-bias-against-english-language-learners 3. American Psychiatric Association: Diagnostic and Statistical Manual of Mental Disorders. American Psychiatric Association, fifth edition edn. (May 2013). https: //doi.org/10.1176/appi.books.9780890425596 4. Baumgartner, J., Zannettou, S., Keegan, B., Squire, M., Blackburn, J.: The Pushshift Reddit Dataset. Proceedings of the International AAAI Conference on Web and Social Media 14, 830â839 (May 2020). https://doi.org/10.1609/icwsm. v14i1.7347 5. Boe, B.: praw-dev/praw: PRAW, an acronym for "Python Reddit API Wrapper", is a python package that allows for simple access to Redditâs API. (2012), https: //github.com/praw-dev/praw 6. Chaka, C.: Accuracy pecking order â How 30 AI detectors stack up in detecting generative artificial intelligence content in university English L1 and English L2 student essays. Journal of Applied Learning and Teaching 7(1) (Apr 2024). https: //doi.org/10.37074/jalt.2024.7.1.33 7. Chandere, V., Satish, S., Lakshminarayanan, R.: Online Plagiarism Detection Tools in the Digital Age: A Review. Annals of the Romanian Society for Cell Biol- ogy p. 7110â7119 (Mar 2021), https://annalsofrscb.ro/index.php/journal/article/ view/881 8. Crompton, C.J., Ropar, D., Evans-Williams, C.V., Flynn, E.G., Fletcher-Watson, S.: Autistic peer-to-peer information transfer is highly effective. Autism 24(7), 1704â1712 (Oct 2020). https://doi.org/10.1177/1362361320919286 9. Davalos, J., Yin, L.: AI Detectors Falsely Accuse Students ofCheatingâWithBigConsequences.Bloomberg.com(Oct 2024),https://w.bloomberg.com/news/features/2024-10-18/ do-ai-detectors-work-students-face-false-cheating-accusations 10. Detector, F.A.: Free AI Content Detector : Tool To Detect Content Written By Ai or Humans (2024), https://w.freedetector.ai/ 11. Dwyer, M., Laird, E.: Report â Up in the Air: Educa- tors Juggling the Potential of Generative AI with Detection, Discipline,andDistrust(Mar2024),https://cdt.org/insights/ report-up-in-the-air-educators-juggling-the-potential-of-generative-ai-with-detection-discipline-and-distrust/ 12. Fiesler, C., Zimmer, M., Proferes, N., Gilbert, S., Jones, N.: Remember the Human: A Systematic Review of Ethical Considerations in Reddit Research. Proc. ACM Hum.-Comput. Interact. 8(GROUP), 1â33 (Feb 2024). https://doi.org/10.1145/ 3633070 13. Gao, C.A., Howard, F.M., Markov, N.S., Dyer, E.C., Ramesh, S., Luo, Y., Pearson, A.T.: Comparing scientific abstracts generated by ChatGPT to real abstracts with detectors and blinded human reviewers. NPJ Digit Med 6(1), 75 (Apr 2023). https://doi.org/10.1038/s41746-023-00819-6 14. Gegg-Harrison, W., Quarterman, C.: AI Detectionâs High False Positive Rates and the Psychological and Material Impacts on Students. In: Academic Integrity in the Age of Artificial Intelligence, p. 199â219. IGI Global Scientific Publishing (2024). https://doi.org/10.4018/979-8-3693-0240-8.ch011 14 15. Ghaffary, S.: Universities Rethink Using AI Writing Detec- tors to Vet Studentsâ Work. Bloomberg.com (Sep 2023), https://w.bloomberg.com/news/newsletters/2023-09-21/ universities-rethink-using-ai-writing-detectors-to-vet-students-work 16. Grosvenor, L.P., Croen, L.A., Lynch, F.L., Marafino, B.J., Maye, M., Penfold, R.B., Simon, G.E., Ames, J.L.: Autism Diagnosis Among US Children and Adults, 2011-2022. JAMA Netw Open 7(10), e2442218 (Oct 2024). https://doi.org/10. 1001/jamanetworkopen.2024.42218 17. Guo, C., Pleiss, G., Sun, Y., Weinberger, K.Q.: On calibration of modern neural networks. In: Proceedings of the 34th International Conference on Machine Learn- ing - Volume 70. p. 1321â1330. ICMLâ17, JMLR.org, Sydney, NSW, Australia (Aug 2017). https://doi.org/10.48550/arXiv.1706.04599 18. Hong, S.R., Zampieri, M., Hand, B.N., Motti, V., Chung, D., Uzuner, O.: Collab- orative Design for Job-Seekers with Autism: A Conceptual Framework for Future Research (Jul 2024). https://doi.org/10.48550/arXiv.2405.06078 19. Jaiswal, A., Shah, A., Harjadi, C., Windgassen, E., Washington, P.: Addendum: Using #ActuallyAutistic on Twitter for Precision Diagnosis of Autism Spectrum Disorder: Machine Learning Study. JMIR Formative Research 8(1), e59349 (Jul 2024). https://doi.org/10.2196/59349 20. Jaiswal, A., Shah, A., Harjadi, C., Windgassen, E., Washington, P.: Ethics of the Use of Social Media as Training Data for AI Models Used for Digital Phenotyping. JMIR Formative Research 8(1), e59794 (Jul 2024). https://doi.org/10.2196/59794 21. Jaiswal, A., Washington, P.: Using #ActuallyAutistic on Twitter for Precision Diagnosis of Autism Spectrum Disorder: Machine Learning Study. JMIR Form Res 8, e52660 (Feb 2024). https://doi.org/10.2196/52660 22. Kirchner, J.H., Ahmad, L., Aaronson, S., Leike, J.: New AI classi- fier for indicating AI-written text (Jan 2023), https://openai.com/index/ new-ai-classifier-for-indicating-ai-written-text 23. Kling, J.: Prof accused of being AI bot (Jul 2023), https://w.purdueexponent. org/campus/article_2d1826e2-2bfa-11e-84c9-6f34496edb29.html 24. Lam, Y.G.: Pragmatic Language in Autism: An Overview. In: Patel, V.B., Preedy, V.R., Martin, C.R. (eds.) Comprehensive Guide to Autism, p. 533â550. Springer, New York, NY (2014). https://doi.org/10.1007/978-1-4614-4788-7_25 25. Liang, W., Yuksekgonul, M., Mao, Y., Wu, E., Zou, J.: GPT detectors are biased against non-native English writers. Patterns 4(7), 100779 (Jul 2023). https://doi. org/10.1016/j.patter.2023.100779 26. Loper, E., Bird, S.: NLTK: The Natural Language Toolkit. In: Proceedings of the ACL-02 Workshop on Effective Tools and Methodologies for Teaching Natural Language Processing and Computational Linguistics. p. 63â70. Association for Computational Linguistics, Philadelphia, Pennsylvania, USA (Jul 2002). https: //doi.org/10.3115/1118108.1118117 27. Madden, M., Calvin, A., Hasse, A., Lenhart, A.: The dawn of the AI era: Teens, parents, and the adoption of generative AI at home and school (2024) 28. Maenner, M.J.: Prevalence and Characteristics of Autism Spectrum Disorder Among Children Aged 8 Years â Autism and Developmental Disabilities Mon- itoring Network, 11 Sites, United States, 2020. MMWR Surveill Summ 72 (2023). https://doi.org/10.15585/mmwr.s7202a1 29. McCrossin, R.: Finding the True Number of Females with Autistic Spectrum Disor- der by Estimating the Biases in Initial Recognition and Clinical Diagnosis. Children (Basel) 9(2), 272 (Feb 2022). https://doi.org/10.3390/children9020272 15 30. Monk, R., Whitehouse, A.J.O., Waddington, H.: The use of language in autism research. Trends in Neurosciences 45(11), 791â793 (Nov 2022). https://doi.org/ 10.1016/j.tins.2022.08.009 31. OpenAI:Understandingthesourceofwhatwesee andhearonline(Aug2024),https://openai.com/index/ understanding-the-source-of-what-we-see-and-hear-online/ 32. Perkins, M., Roe, J., Postma, D., McGaughran, J., Hickerson, D.: Detection of GPT-4 Generated Text in Higher Education: Combining Academic Judgement and Software to Identify Generative AI Tool Misuse. J Acad Ethics 22(1), 89â113 (Mar 2024). https://doi.org/10.1007/s10805-023-09492-6 33. R Core Team: R: A Language and Environment for Statistical Computing. R Foun- dation for Statistical Computing, Vienna, Austria (2024), https://w.R-project. org/ 34. Rubio-MartĂn, S., GarcĂa-OrdĂĄs, M.T., BayĂłn-GutiĂ©rrez, M., Prieto-FernĂĄndez, N., BenĂtez-Andrades, J.A.: Enhancing ASD detection accuracy: a combined ap- proach of machine learning and deep learning models with natural language processing. Health Inf Sci Syst 12(1), 20 (Mar 2024). https://doi.org/10.1007/ s13755-024-00281-y 35. Solaiman, I., Brundage, M., Clark, J., Askell, A., Herbert-Voss, A., Wu, J., Rad- ford, A., Krueger, G., Kim, J.W., Kreps, S., McCain, M., Newhouse, A., Blazakis, J., McGuffie, K., Wang, J.: Release Strategies and the Social Impacts of Language Models (Nov 2019). https://doi.org/10.48550/arXiv.1908.09203 36. Torrey, J.: thinkst/zippy (Oct 2024), https://github.com/thinkst/zippy 37. Verma, P.: A professor accused his class of using ChatGPT, putting diplomas in jeopardy. Washington Post (May 2023), https://w.washingtonpost.com/ technology/2023/05/18/texas-professor-threatened-fail-class-chatgpt-cheating/ 38. Volden, J., Lord, C.: Neologisms and idiosyncratic language in autistic speak- ers. J Autism Dev Disord 21(2), 109â130 (Jun 1991). https://doi.org/10.1007/ BF02284755 39. Walters, W.H.: The Effectiveness of Software Designed to Detect AI-Generated Writing: A Comparison of 16 AI Text Detectors. Open Information Science 7(1) (Jan 2023). https://doi.org/10.1515/opis-2022-0158 40. Weber-Wulff, D., Anohina-Naumeca, A., Bjelobaba, S., FoltĂœnek, T., Guerrero- Dib, J., Popoola, O., Ć igut, P., Waddington, L.: Testing of detection tools for AI-generated text. Int J Educ Integr 19(1), 1â39 (Dec 2023). https://doi.org/10. 1007/s40979-023-00146-z 41. Werra, L.v., Belkada, Y., Tunstall, L., Beeching, E., Thrush, T., Lambert, N., Huang, S., Rasul, K., GallouĂ©dec, Q.: TRL: Transformer Reinforcement Learning (2020), https://github.com/huggingface/trl 42. Williams, G.L., Wharton, T., Jagoe, C.: Mutual (Mis)understanding: Reframing Autistic Pragmatic âImpairmentsâ Using Relevance Theory. Front. Psychol. 12 (Apr 2021). https://doi.org/10.3389/fpsyg.2021.616664 43. Woodhouse, Emma: ADOS-2 Reliability (Aug 2021), https://compasspsy.co.uk/ reliability-training/ados-2/