Paper deep dive
Next Reply Prediction X Dataset: Linguistic Discrepancies in Naively Generated Content
Simon Münker, Nils Schwager, Kai Kugler, Michael Heseltine, Achim Rettinger
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 92%
Last extracted: 7/20/2026, 8:27:30 PM
Summary
This paper introduces the Next Reply Prediction X Dataset to evaluate linguistic discrepancies between human-generated and LLM-generated content on X (formerly Twitter). The authors propose a history-conditioned reply prediction task using authentic X data to assess the validity of using LLMs as human proxies in social science research. They compare prompt-based and fine-tuned LLM outputs against human posts using quantitative, morphosyntactic, and semantic metrics, finding that while fine-tuned models show better alignment, systematic linguistic signatures allow for reliable detection of synthetic content.
Entities (9)
Relation Signals (8)
Next Reply Prediction X Dataset → usesplatform → X (formerly Twitter)
confidence 98% · history-conditioned reply prediction task on authentic X (formerly Twitter) data
Simon Münker → affiliatedwith → Trier University
confidence 95% · Simon Münker 1 ... 1 Trier University
Michael Heseltine → affiliatedwith → University of Oxford
confidence 95% · Michael Heseltine 2 ... 2 University of Oxford
Next Reply Prediction X Dataset → createdby → Simon Münker
confidence 95% · This paper addresses these limitations by introducing a novel, history-conditioned reply prediction task... Our paper addresses a fundamental question... Simon Münker 1
Next Reply Prediction X Dataset → evaluates → Large Language Models
confidence 92% · dataset designed to evaluate the linguistic output of LLMs against human-generated content
Qwen3-8b → usedin → Next Reply Prediction X Dataset
confidence 90% · We fine-tune Qwen3 8B... for each language variant... The published dataset... consists of... generated columns... produced by applying the generation procedure... to the base and fine-tuned Qwen3 8B models
Next Reply Prediction X Dataset → analyzedwith → XGBoost
confidence 88% · We utilize XGBoost... as our classification algorithm... to investigate if these features improve the identification of synthetically generated examples.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The increasing use of Large Language Models (LLMs) as proxies for human participants in social science research presents a promising, yet methodologically risky, paradigm shift. While LLMs offer scalability and cost-efficiency, their "naive" application, where they are prompted to generate content without explicit behavioral constraints, introduces significant linguistic discrepancies that challenge the validity of research findings. This paper addresses these limitations by introducing a novel, history-conditioned reply prediction task on authentic X (formerly Twitter) data, to create a dataset designed to evaluate the linguistic output of LLMs against human-generated content. We analyze these discrepancies using stylistic and content-based metrics, providing a quantitative framework for researchers to assess the quality and authenticity of synthetic data. Our findings highlight the need for more sophisticated prompting techniques and specialized datasets to ensure that LLM-generated content accurately reflects the complex linguistic patterns of human communication, thereby improving the validity of computational social science studies.
Tags
Links
- Source: https://arxiv.org/abs/2602.19177v1
- Canonical: https://arxiv.org/abs/2602.19177v1
Trouble viewing inline? Open PDF directly →
Full Text
50,641 characters extracted from source content.
Expand or collapse full text
Next Reply Prediction X Dataset: Linguistic Discrepancies in Naively Generated Content Simon Münker 1 , Nils Schwager 1 , Kai Kugler 1 , Michael Heseltine 2 , Achim Rettinger 1 1 Trier University, Computational Linguistics 2 University of Oxford, Sociology Universitätsring 15, 54296 Trier, Germany42-43 Park End Street, Oxford OX1 1JD, England muenker, schwager, kuglerk, rettinger@uni-trier.demichael.heseltine@sociology.ox.ac.uk Abstract The increasing use of Large Language Models (LLMs) as proxies for human participants in social science research presents a promising, yet methodologically risky, paradigm shift. While LLMs offer scalability and cost-efficiency, their "naive" application, where they are prompted to generate content without explicit behavioral constraints, introduces significant linguistic discrepancies that challenge the validity of research findings. This paper addresses these limitations by introducing a novel, history-conditioned reply prediction task on authenticX(formerly Twitter) data, to create a dataset designed to evaluate the linguistic output of LLMs against human-generated content. We analyze these discrepancies using stylistic and content-based metrics, providing a quantitative framework for researchers to assess the quality and authenticity of synthetic data. Our findings highlight the need for more sophisticated prompting techniques and specialized datasets to ensure that LLM-generated content accurately reflects the com- plex linguistic patterns of human communication, thereby improving the validity of computational social science studies. Keywords: human simulacra, synthetic content, linguistic authenticity 1. Introduction The widespread adoption of Large Language Mod- els (LLMs) began with the release of ChatGPT and similar conversational AI systems, fundamentally transforming how humans interact with artificial in- telligence (Aïmeur et al., 2023). This technological advancement created a paradigm shift in computa- tional social science research, with LLMs increas- ingly positioned as viable proxies for human partici- pants in behavioral studies (Park et al., 2023; Pérez et al., 2023). The promise is compelling: LLMs offer unprecedented scalability, cost-efficiency, and the ability to conduct large-scale behavioral research without the traditional constraints of human partic- ipant recruitment, retention, and ethical complexi- ties. However, this anthropomorphic perspective intro- duces significant methodological risks, particularly when researchers employ naive applications that rely exclusively on prompt engineering without ade- quate consideration of underlying model limitations, training biases, and domain-specific validation re- quirements (Larooij and Törnberg, 2025). The quality and representativeness of training datasets become critically important as grounding mecha- nisms, especially for socially sensitive tasks where cultural nuance, contextual understanding, and au- thentic human judgment remain central to meaning- ful analysis. While earlier concerns focused on de- tecting artificial or malicious content, contemporary LLMs produce increasingly sophisticated outputs that superficially mimic human communication pat- terns (Crothers et al., 2023). This evolution makes validation more critical yet paradoxically more chal- lenging: the better LLMs become at generating plausible content, the more crucial it becomes to understand where and how they diverge from au- thentic human behavior. Our paper addresses a fundamental question at the intersection of natural language processing and computational social science: Can current LLMs reliably replicate authentic human social media be- havior patterns when tasked with user modeling applications? This question becomes particularly pressing given the growing reliance on synthetic data in computational social science (Burgard et al., 2017), where the assumption of authentic human- like generation underpins the validity of research findings. We approach this question through a sys- tematic comparison between genuine X content and synthetic posts generated through both prompt- based and fine-tuned approaches, examining lin- guistic discrepancies across multiple analytical di- mensions. 1.1. Research Questions and Hypotheses We investigate three primary research questions us- ing a self-collected German and English X dataset: RQ1To what extent do LLM-generated social me- dia posts exhibit detectable linguistic patterns that distinguish them from authentic human content across quantitative, morphological, and semantic dimensions? arXiv:2602.19177v1 [cs.CL] 22 Feb 2026 RQ2How does fine-tuning on domain-specific so- cial media data improve the linguistic authentic- ity of generated content compared to prompt- based generation approaches? RQ3Can machine learning classifiers reliably dis- tinguish between human and synthetic social media content, and what features prove most discriminative? Building on empirical evidence from related work on LLM limitations in social simulation (Liu et al., 2022; Hershcovich et al., 2022; Münker et al., 2025), we hypothesize that while LLMs can produce indi- vidually plausible social media posts, systematic analysis reveals consistent linguistic signatures that enable reliable detection of synthetic content. Fur- thermore, we anticipate that fine-tuned models will show reduced but still detectable deviation patterns compared to prompt-based approaches. Related research confirms that fine-tuned models outper- form prompt-based approaches in social simula- tions (Lin, 2024) and text annotation tasks (Alizadeh et al., 2025) in human-LLM alignment. 1.2. Our Contributions Our work makes three primary contributions to the language resources and evaluation community: 1. We publish a history-conditioned reply predic- tion dataset for X content, comprising authen- tic human posts alongside corresponding syn- thetic generations using both prompt-based and fine-tuned approaches across English and German languages. (Sec. 3.1) 2.We present a multi-dimensional evaluation framework combining multiple layers of quan- titative linguistics analysis to assess human- machine linguistic alignment. (Sec. 3.2) 3.We conduct a comparison of encoder (tf-idf, static dense, transformer) and feature (see above) combinations for detecting the syn- thetic content. (Sec. 3.3) 2. Background 2.1. LLMs as Human Simulacra The emergence of Large Language Models has fundamentally transformed computational social science research, with contemporary studies in- creasingly positioning LLMs as human simulacra (Park et al., 2023) capable of simulating complex user behaviors through sophisticated text-to-text engagement (Larooij and Törnberg, 2025; Münker et al., 2025). This paradigm shift offers compelling advantages including cost reduction, ethical com- pliance, and enhanced scalability for large-scale behavioral studies (Pérez et al., 2023; Thapa et al., 2025). However, empirical validation reveals significant limitations in the authenticity of LLM-generated social behavior. Studies demonstrate system- atic biases in the diversity of political (Liu et al., 2022; Münker, 2025) and cultural (Hershcovich et al., 2022) positions represented in current LLMs. These limitations challenge the prevalent assump- tion that LLMs can serve as reliable human proxies, particularly when researchers employ naive appli- cations that rely exclusively on prompt engineering without adequate consideration of underlying model limitations, training biases, and domain-specific val- idation requirements. The anthropomorphic perspective introduces methodological risks that become especially prob- lematic in applications requiring nuanced social understanding. While individual LLM-generated texts may appear plausible, systematic analysis often reveals consistent linguistic signatures that distinguish synthetic from authentic content. This detectability gap has important implications for the ecological validity of LLM-based simulations in so- cial research contexts, where the assumption of authentic human-like generation underpins the va- lidity of research findings. 2.2. Synthetic Content Detection in Social Media The field of synthetic content detection has evolved significantly alongside advances in generation ca- pabilities. Traditional approaches to misinformation detection on social media platforms target artifi- cial or malicious content from regular users (Yang et al., 2019) and develop comprehensive bot detec- tion systems (Hayawi et al., 2023). However, the current generation of LLMs produces increasingly sophisticated outputs that closely mimic human communication patterns, making detection more challenging and validation more critical. Recent advances in AI-generated content detec- tion (Chong et al., 2023; Abburi et al., 2024) reveal that even sophisticated generation techniques ex- hibit systematic linguistic patterns across multiple dimensions. These patterns manifest in quantita- tive features (complexity, readability, lexical diver- sity), morphosyntactic structures (part-of-speech distributions, syntactic complexity), and semantic distributions (topic diversity, emotion patterns, sen- timent biases). The persistence of these linguistic signatures across different generation approaches suggests fundamental limitations in current lan- guage modeling techniques for authentic social media simulation. Our work motivates a shift toward multi- dimensional evaluation frameworks that capture the full spectrum of linguistic differences between human and synthetic content. Surface-level plausi- bility assessments prove insufficient for validating LLM-generated social media content, necessitat- ing comprehensive protocols that examine linguis- tic authenticity across quantitative, morphological, and semantic dimensions simultaneously. We build upon these methodological foundations while in- troducing a novel history-conditioned dataset and systematic comparison of detection approaches, addressing the critical gap between generation ca- pability and authentic behavioral replication. 3. Methods 3.1. Data: Authentic vs. Synthetic Collection/Preprocessing Our final dataset is based on two raw data dumps – English and German – collected from X. The sets are col- lected around keywords concerning the political discourses in the US and Germany during the first half of 2023. The samples contain two types of content: a) Tweets (posts) and (b) replies from X users towards these tweets (DE: 3.381.111, EN: 7.790.741). First, we group all first-order replies with the tweets to which they are responding, creating tweet- reply pairs that preserve conversational context. We then reorganize these samples by user, result- ing in subsets containing each user’s complete re- ply history along with the original tweets they re- sponded to. Next, we apply two preprocessing steps to en- sure data quality. First, we remove tweet-reply pairs containing URLs (images, GIFs, and links), as these cannot be properly processed by the LLM and the classifiers. Second, we remove users with the highest reply frequencies (DE: 5%, EN: 1%; Quo- tas result in max DE: 24, EN: 21 samples per user) and split the remaining users into train and test sets. This ensures our analysis captures the model’s abil- ity to learn generalizable linguistic styles across the user population, rather than memorizing patterns of individual users. Transformation We construct a History- Conditioned Reply Prediction Task (Münker et al., 2025), using the native instruction-completion format of instruction-LLMs: three tweet-reply pairs as "history", along with a fourth tweet for the model to respond to. We add a system prompt: You are a social media user responding to conversations. Keep your replies consistent with your previous writing style and the perspectives you have expressed earlier. This conditions the LLM by presenting the tweet-reply history as if it had already generated those replies during prior turns. This approach offers three advantages: (1) the model learns from authentic behavioral patterns without hand-crafted features encoding response characteristics; (2) It allows synthetic sample gener- ation without further training only by prompting (3) during fine-tuning, the withheld fourth reply serves as the supervised target. Fine-Tuning We fine-tune Qwen3 8B (Yang et al., 2025) for each language variant using supervised learning with loss computed exclusively on the last assistant responses. Both training datasets are sub-sampled to 5000 examples. Training uses a warm-up ratio of 0.1 and single-epoch optimiza- tion with otherwise default hyperparameter (e.g.; learning rate of 2e−5). Each model trains for ap- proximately 100 minutes on an NVIDIA L40S GPU with 48GB of VRAM. Generation We generate a single synthetic reply per test prompt using both the base and fine-tuned Qwen3 8B models. Generation uses Qwen3’s default sampling parameters (temperature: 0.6, top_k: 20,top_p: 0.9) with a maximum output length of 200 tokens. No post-generation filtering is applied. All model outputs are retained regardless of length, coherence, or formatting. Published Dataset The published dataset (GitHub repository, see Sec. 3.4) serves as the foundation for all subsequent analyses. It consists of 1000 samples per language, each containing:prompt(three historical tweet-reply pairs plus fourth tweet in chat completion format), authentic reply(ground truth from test users), and two generated columnsbase model reply andft model replyproduced by applying the generation procedure described above to the base and fine-tuned Qwen3 8B models respectively. As a scientific artifact, the dataset serves three potential usages: 1) improving the next reply prediction task given our proposed metrics, 2) developing additional metrics to analyze the LLM-human alignment further, and 3) improving synthetic content detection classifiers. 3.2. Evaluation: Levels of Alignment Quantitative Features We implement the com- plete NeLa feature suite (Horne et al., 2018) through a modular extraction pipeline spanning five linguistic dimensions: complexity, style, bias, af- fect, and moral reasoning patterns. The system extracts linguistic profiles including type-token ra- tios, average sentence length, lexical diversity mea- sures, readability scores (Flesch-Kincaid (Kincaid et al., 1975), Gunning Fog (Gunning, 1968)), and character-level complexity metrics. Morphosyntactic Extraction Using the spaCy processing pipeline (Montani et al., 2023), we ex- tract comprehensive linguistic annotations includ- ing part-of-speech tag distributions following Uni- versal Dependencies standards (De Marneffe et al., 2021), named entity recognition patterns across 18 standard categories (PERSON, ORG, GPE, DATE, etc.), dependency relation frequencies, and syn- tactic complexity measures. Our implementation computes frequency-normalized distributions for both POS categories and NER labels, incorporat- ing lexical diversity metrics and average sentence length measurements. Semantic Classification We employ the Tweet- Eval benchmark (Barbieri et al., 2020) through pre-trained transformer-based classifiers to eval- uate content across multiple semantic dimensions. Our pipeline integrates three specialized mod- els:tweet-topic-21-multi(Antypas et al., 2022) for topic classification,twitter-RoBERTa- base-emotion(Camacho-Collados et al., 2022) for emotion detection, andtwitter-RoBERTa- base-sentiment(Barbieri et al., 2020) for senti- ment analysis. Cluster-based Similarity Utilizing the state-of- the-art instruction-following embedding model Qwen3 (Zhang et al., 2025), we compute semantic similarity distributions within and across content categories. Through cluster analysis using PCA di- mensionality reduction (Pearson, 1901) and Affinity Propagation (Frey and Dueck, 2007), we analyze the proportion of clusters per content category. Feature-Vector Distance Computation To quantify linguistic alignment between human and synthetic content, we implement a distance-based similarity metric. For each corpusCand linguistic feature setF, we compute normalized feature vectors through the following procedure: 1.Extract mean feature scores for each corpus- feature combination: ̄ f i C = 1 |C| X d∈C F i (d) wheredrepresents sample in corpusCand F i denotes the i-th feature in F. 2.Construct corpus vectors:v C = [ ̄ f 1 C , . . . , ̄ f |F | C ] 3.Compute pairwise Cosine similarity between two corpus vectors defined as s(v C 1 ,v C 2 ). 3.3. Validation: Detecting Synthetics As a downstream validation task, we implement a comparison of detection approaches spanning the spectrum from traditional sparse representations to modern dense embeddings. We concatenate the above-described features with the following text embeddings to investigate if these features improve the identification of synthetically generated exam- ples. Encoding Approaches TF-IDF Sparse term frequency-inverse document frequency vectorization (Ramos et al., 2003) with uni-gram features and lowercase normal- ization for, baseline, traditional bag-of-words representation. FastTextDense 300-dimensional word vec- tors (Joulin et al., 2017) using spaCy’s en_core_web_lgandde_core_news_lg model, aggregated through mean pooling for efficient semantic representation. Qwen3 Embedding: State-of-the-art instruction- following embeddings using the Qwen/Qwen3- Embedding-8B model (Zhang et al., 2025) with a specialized authorship detection prompt: "In- struct: Find tweets with similar authorship pat- terns (human vs. AI-generated) based on writing style, vocabulary choice, and content structure". We choose Qwen3 as it shows benchmark-leading performance compared to the number of parameters in text classification tasks (Pan et al., 2025; Heseltine, 2025). Feature Combination Strategy To investigate the complementary nature of different represen- tation types, we systematically evaluate all pos- sible combinations of encoding approaches and extracted features, creating hybrid representations that capture multiple linguistic perspectives simul- taneously. Classification Model We utilize XGBoost (eX- treme Gradient Boosting) (Chen and Guestrin, 2016) as our classification algorithm. We select XGBoost for its promising performance on hetero- geneous feature combinations, robust handling of different feature scales, and interpretability through feature importance analysis. 3.4. Reproducibility and Code Availability All experimental procedures, statistical analyses, and model training protocols are implemented us- ing open-source tools including scikit-learn (Buit- inck et al., 2013), spaCy (Montani et al., 2023), Transformers Reinforcement Learning (TRL) (von Werra et al., 2020) and Sentence Transformers (Reimers and Gurevych, 2019). 4. Results Our results reveal systematic linguistic differences between human and synthetic content across all analytical dimensions, with fine-tuned models con- sistently showing superior alignment to human content compared to prompt-based approaches. These findings address our three research ques- tions through complementary lenses: similarity analysis (RQ1 and RQ2) and classification perfor- mance (RQ3). 4.1. Quantitative Linguistics Analysis Table 1 presents the calculated similarity scores between corpus subsets across all feature extrac- tion approaches. The results show a consistent hierarchy of alignment, with fine-tuned models (F) showing highest similarity to human original content (O), followed by moderate alignment between orig- inal and prompt-based content (P), while prompt- based and fine-tuned models exhibit the lowest mutual similarity. Feat./Lang.s(O, P )s(O, F )s(P, F ) Quantitative Features (NeLa) German0.79080.80480.6995 English 0.84080.89570.8410 Morphosyntactic Extraction (SpaCy) German0.94980.97480.9437 English0.94230.98160.9357 Semantic Classification (TweetEval) German0.96950.98740.9832 English0.97450.98190.9786 Cluster-based Similarity German0.80160.97130.7435 English0.89770.96200.8323 Table 1: Comparison of the calculated similar- itysbetween the corpora subsets human original (O), synthetic only prompted (P) and synthetic fine- tuned (F) across German and English on all fea- tures described in section 3.2. A higher value indi- cates a more aligned model behavior. Quantitative Features The NeLa features reveal substantial differences in linguistic complexity and style patterns. For German, similarity between orig- inal and fine-tuned content reaches 0.8048, signif- icantly higher than the 0.6995 similarity between prompt-based & fine-tuned approaches. English demonstrates even stronger alignment patterns, with original & fine-tuned similarity achieving 0.8957, while original-prompt similarity reaches 0.8408. Morphosyntactic Extraction The Morphosyn- tactic analysis reveals the highest overall similarity scores across all approaches. German shows a high alignment between original & fine-tuned con- tent (0.9748), with prompt-based models achieving 0.9498 similarity to original content. However, de- tailed examination reveals that prompt-based mod- els exhibit distinctive usage patterns, particularly in coordinating (CCONJ) and subordinating conjunc- tions (SCONJ) (Figure 1). 0.00.20.40.60.81.0 score ADJ ADV INTJ NOUN PROPN VERB ADP AUX CCONJ DET NUM PART PRON SCONJ PUNCT SYM X label corpus original synthetic (prompt-only) synthetic (fine-tuned) Figure 1: Locality, spread and skewness (x-axes) of each POS category (y-axes) for the English cor- pus split into subsets. Semantic Classification Semantic classification shows the most consistent alignment across all model types, with similarity scores exceeding 0.97 in all comparisons. German achieves the highest alignment between prompt and fine-tuned models (0.9832), while English shows marginally lower but still substantial similarity (0.9786). Despite these high similarity scores, qualitative analysis reveals that prompt-based models generate more topically diverse content and exhibit significantly higher pro- portions of positive emotion classifications com- pared to human content (Figure 2). Cluster-based Similarity Embedding-based cluster analysis reveals the most pronounced differences between generation approaches. Fine- tuned models achieve high alignment with original content (German: 0.9713, English: 0.9620), while prompt-based models show notably lower similarity to both original content and fine-tuned variants. The substantial gap between original & prompt similarity (German: 0.8016, English: 0.8977) and original & fine-tuned similarity demonstrates that news_&_social_concern diaries_&_daily_life celebrity_&_pop_culture other_hobbies film_tv_&_video sports business_&_entrepreneurs science_&_technology fitness_&_health relationships learning_&_educational youth_&_student_life family music arts_&_culture travel_&_adventure food_&_dining fashion_&_style gaming label 0.0 0.2 0.4 0.6 0.8 1.0 score classifier = topic disgust anger fear sadness joy optimism love surprise pessimism anticipation trust label classifier = emotions original synthetic (prompt-only) synthetic (fine-tuned) Figure 2: Locality, spread and skewness (y-axes) of the TweetEval topic and emotion classifier (x-axes) for the English corpus split into subsets. semantic distributional properties are particularly sensitive. 4.2. Validation Task Table 2 presents the classification results for dis- tinguishing between human original (O), synthetic fine-tuned (F), and synthetic prompted (P) content across various feature combinations and encoding approaches. The results consistently demonstrate that prompt-based synthetic content (P) achieves the highest detection accuracy, while fine-tuned content (F) proves most challenging to distinguish from human original content (O). Feature Combination Performance The most effective approach combines tf-idf, fastText embed- dings, and one or more extracted features (Tweet- Eval, SpaCy, NeLa), achieving macroF1 scores of 0.7301 for German and 0.6972 for English. No- tably, prompt-based content consistently achieves the highest individualF1 scores across all feature combinations (German: 0.8297−0.8510, English: 0.6725− 0.8163). Encoding Approach Analysis Modern embed- ding approaches show competitive but not superior performance compared to traditional methods. The Qwen embedding model alone achieves moder- ate performance (GermanF1: 0.6365, EnglishF1: 0.6549), while fastText embeddings demonstrate strong baseline performance (GermanF1: 0.7007, EnglishF1: 0.6240). Surprisingly, simple tf-idf rep- resentations prove remarkably effective, particu- larly for prompt-based content detection (German: 0.8333, English: 0.7800). Feat./Lang.F 1(O)F 1(F )F 1(P )avg tf–idf + fastText + TweetEval, SpaCy, NeLa German0.66660.69380.82970.7301 English0.65340.63820.80000.6972 tf–idf + fastText + Qwen German0.63360.64760.85100.7107 English0.55310.61360.67250.6240 Qwen German0.58000.58490.74460.6365 English0.55440.59400.81630.6549 fastText German0.67340.64070.78780.7007 English0.58580.61360.67250.6240 tf–idf German0.64000.59610.83330.6898 English0.52000.50000.78000.6000 Table 2: Results of the detection task for the indi- vidualF1 scores per class, human original (O), syn- thetic only prompted (P) and synthetic fine-tuned (F), and the macro average across German and English on a selected range of feature combina- tions with XGBoost as classifier. 5. Discussion 5.1. Implications for Computational Social Science Our findings reveal fundamental challenges for the ecological validity of LLM-based simulations in social research contexts. While current genera- tion techniques can produce individually plausible social media posts, systematic analysis reveals consistent patterns that distinguish synthetic from authentic content across multiple linguistic dimen- sions. The detection accuracies achieved in our validation task, particularly for prompt-based con- tent, indicate persistent linguistic signatures that compromise the authenticity of LLM-generated so- cial media discourse. This detectability gap has important implications for applications in computational social science, where researchers increasingly rely on LLMs as human proxies for behavioral studies. The system- atic differences we observe in quantitative features, morphosyntactic patterns, and semantic distribu- tions suggest that naive deployment of LLMs for social simulation may introduce systematic biases that compromise research validity. The observa- tion that even fine-tuned models, while substantially improved, still exhibit detectable patterns in classi- fication tasks suggests that the challenge extends beyond simple technical optimization to fundamen- tal questions about the nature of human-like gener- ation. These findings align with broader concerns about the anthropomorphism of AI systems (Salles et al., 2020) and highlight the necessity for validation pro- tocols when deploying LLMs in social research con- texts (Møller and Aiello, 2024). The consistent per- formance hierarchy observed across all feature ex- traction approaches, with fine-tuned models show- ing highest alignment to human content, followed by prompt-based models, while the two synthetic ap- proaches exhibit lowest mutual similarity, provides empirical evidence for the complexity of achieving authentic human simulation. 5.2. Linguistic Authenticity and Model Limitations The systematic differences we observe across quantitative linguistics, morphosyntactic patterns, and semantic distributions point to inherent limi- tations in current language modeling approaches. Our analysis reveals that prompt-based models exhibit distinctive linguistic signatures, including more complex sentence structures (evidenced by coordinating and subordinating conjunction usage patterns), more topically diverse content, and sig- nificantly higher proportions of positive emotion classifications compared to human content. Particularly concerning is the cluster-based simi- larity analysis, which shows the most pronounced differences between generation approaches. The substantial gaps between original & prompt similar- ity and original & fine-tuned similarity demonstrate that semantic distributional properties are particu- larly sensitive to generation method. These find- ings suggest that LLMs may be systematically bi- ased toward producing "ideal" rather than authentic communication, potentially missing the natural vari- ation, errors, and stylistic inconsistencies that char- acterize genuine human social media discourse (Thapa et al., 2025). The cross-linguistic consistency of these patterns across English and German corpora strengthens the generalizability of our findings, indicating that the observed limitations are not language-specific artifacts but reflect fundamental characteristics of current language modeling approaches. 5.3. Methodological Considerations for LLM Deployment The superior performance of fine-tuned models compared to prompt-based approaches across all similarity metrics provides strong evidence for the importance of domain adaptation in social media generation tasks. Fine-tuned models consistently achieve higher similarity scores with human content compared to prompt-based models. However, the persistence of detectable patterns even after fine- tuning, suggests that current adaptation techniques may be insufficient for achieving true linguistic au- thenticity (Münker et al., 2025). The effectiveness of different encoding ap- proaches in our validation task reveals important insights about the nature of synthetic content detec- tion. The surprising performance of traditional tf-idf representations, particularly for prompt-based con- tent detection, suggests that surface-level lexical patterns remain highly discriminative despite the sophistication of modern language models. The superior performance of hybrid approaches com- bining tf-idf, fastText embeddings, and extracted linguistic features demonstrates that multiple repre- sentational perspectives are necessary to capture the full spectrum of linguistic differences between human and synthetic content. 6. Conclusion Our paper has examined a fundamental question about the viability of LLMs as human simulacra in computational social science: can current gen- eration techniques produce social media content that reliably replicates authentic human linguistic behavior? Through systematic analysis of a novel history-conditioned dataset spanning English and German X content, we provide evidence-based an- swers to three interconnected research questions. 6.1. Research Questions RQ1: Linguistic Pattern Detection Our results demonstrate that LLM-generated social media posts exhibit systematic and detectable linguistic patterns across quantitative, morphological, and se- mantic dimensions. The similarity analysis reveals that, while individual synthetic posts may appear plausible, aggregate patterns consistently deviate from human norms. Most notably, prompt-based models show distinctive signatures in morphosyn- tactic complexity, with systematic differences in conjunction usage patterns indicating artificially complex sentence structures compared to human originals. Semantic analysis reveals systematic bi- ases toward positive emotion classifications and increased topical diversity compared to authentic human content. RQ2: Fine-tuning versus Prompt-based Ap- proaches Fine-tuned models consistently out- perform prompt-based approaches across all simi- larity metrics, achieving substantially higher align- ment with human content. However, even fine- tuned models remain distinguishable from human content in classification tasks, particularly through cluster-based similarity analysis where the most pronounced differences emerge. This finding con- firms that training models with human data for con- crete, well-defined tasks consistently outperforms general prompt-based usage approaches, align- ing with findings from concurrent work demonstrat- ing the limitations of generic prompting strategies (Münker et al., 2025). RQ3: Machine Learning Detection Capability Our validation task demonstrates reliable classifi- cation performance across multiple encoding ap- proaches and feature combinations. The high- est performing hybrid approach (tf-idf + fastText + extracted features) achieves macroF1 scores of 0.7301 (German) and 0.6972 (English), with par- ticularly strong detection rates for prompt-based content (F1 > 0.8 across multiple configurations). Surprisingly, traditional tf-idf representations prove remarkably effective, suggesting that surface-level lexical patterns remain highly discriminative despite advances in generation sophistication. 6.2. Recommendations for Responsible LLM Deployment Based on our findings, we propose specific guide- lines for the responsible deployment of LLMs in social applications: Mandatory Validation Protocols Researchers employing LLMs for social simulation must imple- ment comprehensive validation protocols that as- sess linguistic authenticity across multiple dimen- sions rather than relying on surface-level plausi- bility assessments. Our multi-dimensional evalua- tion framework provides a template for such vali- dation, combining quantitative linguistics analysis, morphosyntactic profiling, semantic classification, and distributional similarity measures. Domain-Specific Fine-tuning Requirements Our results confirm that fine-tuned models con- sistently outperform prompt-based approaches for social media generation tasks across all similarity metrics. However, fine-tuning alone proves insuf- ficient to achieve complete linguistic authenticity, as evidenced by persistent detectability in classi- fication tasks. This suggests that domain adapta- tion should be considered a minimum requirement rather than a sufficient solution. Multi-dimensional Evaluation Standards The complementary nature of different linguistic anal- ysis approaches in our study demonstrates that single-metric evaluation is insufficient for assess- ing generation quality. Researchers should adopt multi-layered evaluation frameworks that capture quantitative features, morphosyntactic patterns, se- mantic distributions, and embedding-based similar- ity measures simultaneously. 6.3. Future Directions Our findings open several relevant directions for future research. First, investigating the temporal stability of linguistic signatures as generation tech- niques continue to evolve will be essential to under- stand the longevity of current detection methods and to develop robust evaluation frameworks. Sec- ond, examining domain transfer across different social media platforms beyond X will help establish the generalizability of these linguistic signature pat- terns across diverse communication contexts with varying discourse norms and constraints. Third, exploring adversarial training approaches specifically designed to reduce detectability while maintaining content quality and authenticity repre- sents a promising direction for improving genera- tion fidelity. Such approaches could inform the de- velopment of more sophisticated LLMs that better capture the natural variation, errors, and stylistic in- consistencies characteristic of genuine human dis- course on social media. Finally, developing more nuanced evaluation metrics that capture subtle as- pects of human communication patterns beyond current similarity measures could provide deeper in- sights into the fundamental challenges of achieving truly human-like text generation. The cross-linguistic consistency of our findings across English and German corpora suggests that these challenges transcend language-specific arti- facts and reflect fundamental limitations in current language modeling approaches. Limitations Our analysis focuses on X data collected during the first half of 2023, which may not generalize to other social media platforms or communication contexts with different discourse norms and con- straints. The temporal dimension of our dataset may not capture evolving generation capabilities as LLM technology continues to advance rapidly. Addi- tionally, our current framework focuses on English and German languages, and expanding the anal- ysis to include morphologically richer languages, tonal languages, and non-European linguistic fam- ilies would strengthen the cross-linguistic validity of these findings. Beyond these core limitation, we acknowledge several methodological. Analysis Framework Our linguistic analysis framework, while comprehensive across quantita- tive, morphosyntactic, and semantic dimensions, does not capture complex discourse quality metrics such as argumentation coherence, irony detection, or cultural nuance recognition. The focus on indi- vidual post generation rather than multi-turn con- versational dynamics limits our understanding of how synthetic content would perform in sustained social interactions and community discussions. Validation Experiments Our detection validation experiments, while demonstrating reliable classifi- cation performance, are limited to the specific LLM architectures and fine-tuning approaches employed in this study. The rapid evolution of language mod- els means that newer generation techniques may exhibit different linguistic signatures than those cap- tured in our analysis. Additionally, our evaluation framework relies primarily on automated feature ex- traction and classification metrics, which may not capture subtle qualitative differences that human evaluators would detect. Single Model Architecture Our results are based exclusively on Qwen3 8B, which represents only a single model architecture and size configu- ration. The observed linguistic patterns and detec- tion accuracies may vary considerably across dif- ferent model families, model sizes, quantization ap- proaches within the same base model, and model versions. This architectural specificity limits the generalizability of our findings to the broader land- scape of available LLMs. Reply Prediction Task The history-conditioned reply prediction task relies on only three prior tweet- reply pairs as context, which may provide sparse predictive signal for capturing individual user behav- ior patterns and writing styles. This limited historical context may not fully represent the complexity and variation present in users’ broader communication patterns, potentially affecting both the fine-tuning quality and the authenticity of generated content. German vs. English The German and English datasets differ substantially in their collection con- texts, temporal distribution, and underlying dis- course characteristics. These systematic differ- ences make direct cross-linguistic performance comparisons not recommended, as observed varia- tions may reflect dataset-specific properties rather than fundamental linguistic or modeling differences. Each language corpus should be interpreted within its own context rather than as directly comparable benchmarks. Ethical Considerations As is typical for AI methods, the modeling approach presented in this paper is a dual-use technology. While behavior-based user modeling and synthetic content generation are primarily intended for com- putational social science research and platform safety applications, the findings can also be used to develop more sophisticated manipulation tech- niques or improve the convincingness of synthetic social media content for malicious purposes. Privacy and Consent Considerations A signifi- cant ethical concern in our study involves the use of real user data from X to train models that replicate individual behavior patterns. While our dataset consists of publicly available posts from political discourse and replies from regular users, the indi- viduals whose data we used did not provide explicit informed consent for their communication patterns to be learned and replicated by generative models. This raises important questions about digital pri- vacy rights, even when dealing with publicly posted content. Potential for Misuse The detection methodolo- gies developed in this work, while intended to im- prove synthetic content identification, could poten- tially be used adversarially to develop more sophis- ticated generation techniques that evade detection. The detailed analysis of linguistic signatures across quantitative, morphosyntactic, and semantic dimen- sions provides a road-map for improving synthetic content quality, which could enhance both legiti- mate applications and malicious use cases. Broader Implications The development of in- creasingly sophisticated user modeling and syn- thetic content generation capabilities raises broader questions about the boundaries of acceptable re- search practices in computational social science. As these technologies advance, the research com- munity must carefully balance the scientific value of realistic behavioral simulation against the privacy rights and dignity of individuals whose data enables such research, while considering the potential so- cietal impacts of increasingly convincing synthetic social media content. Acknowledgments We thank Simon Werner and Christoph Hau for our constructive discussions. This work is sup- ported by TWON (project number 101095095), a research project funded by the European Union un- der the Horizon framework (HORIZON-CL2-2022- DEMOCRACY-01-07). Bibliographical References Harika Abburi, Nirmala Pudota, Balaji Veeramani, Edward Bowen, and Sanmitra Bhattacharya. 2024. Toward robust generative ai text detec- tion: Generalizable neural model. In 2024 Inter- national Conference on Machine Learning and Applications (ICMLA), pages 1651–1656. IEEE. Esma Aïmeur, Sabrine Amri, and Gilles Brassard. 2023. Fake news, disinformation and misinfor- mation in social media: a review. Social Network Analysis and Mining, 13(1):30. Meysam Alizadeh, Maël Kubli, Zeynab Samei, Shirin Dehghani, Mohammadmasiha Zahedivafa, Juan D Bermeo, Maria Korobeynikova, and Fab- rizio Gilardi. 2025. Open-source llms for text annotation: a practical guide for model setting and fine-tuning. Journal of Computational Social Science, 8(1):17. Dimosthenis Antypas, Asahi Ushio, Jose Camacho- Collados, Vitor Silva, Leonardo Neves, and Francesco Barbieri. 2022. Twitter topic classi- fication. In Proceedings of the 29th International Conference on Computational Linguistics, pages 3386–3400, Gyeongju, Republic of Korea. Inter- national Committee on Computational Linguis- tics. Francesco Barbieri, Jose Camacho-Collados, Luis Espinosa Anke, and Leonardo Neves. 2020. Tweeteval: Unified benchmark and comparative evaluation for tweet classification. In Findings of the Association for Computational Linguistics: EMNLP 2020, pages 1644–1650. Lars Buitinck, Gilles Louppe, Mathieu Blondel, Fabian Pedregosa, Andreas Mueller, Olivier Grisel, Vlad Niculae, Peter Prettenhofer, Alexan- dre Gramfort, Jaques Grobler, Robert Layton, Jake VanderPlas, Arnaud Joly, Brian Holt, and Gaël Varoquaux. 2013. API design for machine learning software: experiences from the scikit- learn project. In ECML PKDD Workshop: Lan- guages for Data Mining and Machine Learning, pages 108–122. Jan Pablo Burgard, Jan-Philipp Kolb, Hariolf Merkle, and Ralf Münnich. 2017. Synthetic data for open and reproducible methodological re- search in social sciences and official statistics. AStA Wirtschafts-und Sozialstatistisches Archiv, 11(3):233–244. Jose Camacho-Collados, Kiamehr Rezaee, Ta- layeh Riahi, Asahi Ushio, Daniel Loureiro, Dimos- thenis Antypas, Joanne Boisson, Luis Espinosa- Anke, Fangyu Liu, Eugenio Martínez-Cámara, et al. 2022. TweetNLP: Cutting-Edge Natural Language Processing for Social Media. In Pro- ceedings of the 2022 Conference on Empirical Methods in Natural Language Processing: Sys- tem Demonstrations, Abu Dhabi, U.A.E. Associ- ation for Computational Linguistics. Tianqi Chen and Carlos Guestrin. 2016. Xgboost: A scalable tree boosting system. In Proceedings of the 22nd acm sigkdd international conference on knowledge discovery and data mining, pages 785–794. Alicia Tsui Ying Chong, Hui Na Chua, Muhammed Basheer Jasser, and Richard TK Wong. 2023. Bot or human? detection of deepfake text with semantic, emoji, sentiment and linguistic features. In 2023 IEEE 13th International Conference on System Engineering and Technology (ICSET), pages 205–210. IEEE. Evan N Crothers, Nathalie Japkowicz, and Herna L Viktor. 2023. Machine-generated text: A compre- hensive survey of threat models and detection methods. IEEE Access, 11:70977–71002. Marie-Catherine De Marneffe, Christopher D Man- ning, Joakim Nivre, and Daniel Zeman. 2021. Universal dependencies. Computational linguis- tics, 47(2):255–308. Brendan J Frey and Delbert Dueck. 2007. Cluster- ing by passing messages between data points. science, 315(5814):972–976. Robert Gunning. 1968. The technique of clear writing. McGraw-Hill Book Company, New York. Kadhim Hayawi, Susmita Saha, Moham- mad Mehedy Masud, Sujith Samuel Mathew, and Mohammed Kaosar. 2023. Social media bot detection with deep learning methods: a systematic review. Neural Computing and Applications, 35(12):8903–8918. Daniel Hershcovich, Stella Frank, Heather Lent, Miryam de Lhoneux, Mostafa Abdou, Stephanie Brandl, Emanuele Bugliarello, Laura Cabello Pi- queras, Ilias Chalkidis, Ruixiang Cui, et al. 2022. Challenges and strategies in cross-cultural nlp. In Proceedings of the 60th Annual Meeting of the Association for Computational Linguistics (Vol- ume 1: Long Papers), pages 6997–7013. Asso- ciation for Computational Linguistics. Michael Heseltine. 2025. Comparing large lan- guage models for text classification: Model se- lection across tasks, texts, and languages. Benjamin D Horne, William Dron, Sara Khedr, and Sibel Adali. 2018. Assessing the news land- scape: A multi-module toolkit for evaluating the credibility of news. In Companion Proceedings of the The Web Conference 2018, pages 235–238. Armand Joulin, Edouard Grave, Piotr Bojanowski, and Tomas Mikolov. 2017. Bag of tricks for ef- ficient text classification. In Proceedings of the 15th Conference of the European Chapter of the Association for Computational Linguistics: Vol- ume 2, Short Papers, pages 427–431. Associa- tion for Computational Linguistics. J Peter Kincaid, Robert P Fishburne Jr, Richard L Rogers, and Brad S Chissom. 1975. Derivation of new readability formulas (automated readabil- ity index, fog count and flesch reading ease for- mula) for navy enlisted personnel. Research Branch Report, Naval Technical Training Com- mand, Millington. Maik Larooij and Petter Törnberg. 2025. Do large language models solve the problems of agent-based modeling? a critical review of generative social simulations. arXiv preprint arXiv:2504.03274. Haocheng Lin. 2024. Designing domain-specific large language models: The critical role of fine- tuning in public opinion simulation. arXiv preprint arXiv:2409.19308. Ruibo Liu, Chenyan Jia, Jason Wei, Guangxuan Xu, and Soroush Vosoughi. 2022. Quantifying and alleviating political bias in language models. Artificial Intelligence, 304:103654. Anders Giovanni Møller and Luca Maria Aiello. 2024. Prompt refinement or fine-tuning? best practices for using llms in computational social science tasks. arXiv preprint arXiv:2408.01346. Ines Montani, Matthew Honnibal, Adriane Boyd, Sofie Van Landeghem, and Henning Peters. 2023. spacy: Industrial-strength nlp. Zenodo. Simon Münker. 2025. Political bias in llms: Un- aligned moral values in agent-centric simulations. Journal for Language Technology and Computa- tional Linguistics, 38(2):125–138. Simon Münker, Nils Schwager, and Achim Rettinger. 2025. Don’t trust generative agents to mimic com- munication on social networks unless you bench- marked their empirical realism. arXiv preprint arXiv:2506.21974. Yuan Pan, Guocong Feng, Kaitian Huang, and Chunmei Zhang. 2025. Qwen3-powered log clas- sification for improved soc decision-making. In 2025 8th International Conference on Computer Information Science and Application Technology (CISAT), pages 651–655. IEEE. Joon Sung Park, Joseph O’Brien, Carrie Jun Cai, Meredith Ringel Morris, Percy Liang, and Michael S Bernstein. 2023. Generative agents: Interactive simulacra of human behavior. In Pro- ceedings of the 36th annual acm symposium on user interface software and technology, pages 1–22. Karl Pearson. 1901. Liii. on lines and planes of clos- est fit to systems of points in space. The London, Edinburgh, and Dublin philosophical magazine and journal of science, 2(11):559–572. Jaime Pérez, Mario Castro, and Gregorio López. 2023. Serious games and ai: Challenges and opportunities for computational social science. IEEE Access, 11:62051–62061. Juan Ramos et al. 2003. Using tf-idf to determine word relevance in document queries. In Proceed- ings of the first instructional conference on ma- chine learning, volume 242, pages 29–48. New Jersey, USA. Nils Reimers and Iryna Gurevych. 2019. Sentence- bert: Sentence embeddings using siamese bert- networks. In Proceedings of the 2019 Confer- ence on Empirical Methods in Natural Language Processing and the 9th International Joint Confer- ence on Natural Language Processing (EMNLP- IJCNLP), pages 3982–3992. Arleen Salles, Kathinka Evers, and Michele Farisco. 2020. Anthropomorphism in ai. AJOB neuro- science, 11(2):88–95. Surendrabikram Thapa, Shuvam Shiwakoti, Sid- dhant Bikram Shah, Surabhi Adhikari, Hari- ram Veeramani, Mehwish Nasim, and Usman Naseem. 2025. Large language models (llm) in computational social science: prospects, current state, and challenges. Social Network Analysis and Mining, 15(1):1–30. Leandro von Werra, Younes Belkada, Lewis Tun- stall, Edward Beeching, Tristan Thrush, Nathan Lambert, Shengyi Huang, Kashif Rasul, and Quentin Gallouédec. 2020. Trl: Transformer re- inforcement learning.https://github.com/ huggingface/trl. An Yang, Anfeng Li, Baosong Yang, Beichen Zhang, Binyuan Hui, Bo Zheng, Bowen Yu, Chang Gao, Chengen Huang, Chenxu Lv, et al. 2025. Qwen3 technical report. arXiv preprint arXiv:2505.09388. Shuo Yang, Kai Shu, Suhang Wang, Renjie Gu, Fan Wu, and Huan Liu. 2019. Unsupervised fake news detection on social media: A generative approach. In Proceedings of the AAAI confer- ence on artificial intelligence, volume 33, pages 5644–5651. Yanzhao Zhang, Mingxin Li, Dingkun Long, Xin Zhang, Huan Lin, Baosong Yang, Pengjun Xie, An Yang, Dayiheng Liu, Junyang Lin, et al. 2025. Qwen3 embedding: Advancing text embedding and reranking through foundation models. arXiv preprint arXiv:2506.05176.