Paper deep dive
Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest
Ramtin Davoudi, Kartik Thakkar, Nazanin Donyapour, Tyler Derr, Hamid Karimi
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 98%
Last extracted: 4/26/2026, 10:31:31 PM
Summary
This study presents a comprehensive evaluation of modern Large Language Models (LLMs) across three core social media analytics tasks: Authorship Verification, Post Generation, and User Attribute Inference, using a Twitter (X) dataset. The researchers benchmarked models like GPT-4, Gemini, DeepSeek, and Llama against traditional machine learning baselines. The study addresses challenges such as 'seen-data' bias in authorship verification and uses standardized taxonomies (IAB Tech Lab and U.S. SOC) for attribute inference, providing a unified framework for assessing how LLMs understand and simulate user identity.
Entities (11)
Relation Signals (5)
GPT-4 â evaluatedfor â Authorship Verification
confidence 100% ¡ We benchmark five LLMsâGPT-4, GPT-3.5-Turbo, Gemini, Llama, and DeepSeek... for social media authorship verification
Gemini â evaluatedfor â Post Generation
confidence 100% ¡ We assess each LLMâs capability to create plausible social media posts... Gemini 1.5 Pro...
Llama 3.2 â evaluatedfor â User Attribute Inference
confidence 100% ¡ For attribute inference, we annotate occupations and interests using two standardized taxonomies... and benchmark LLMs
Twitter (X) â providesdatafor â Authorship Verification
confidence 100% ¡ across three core social media analytics tasks on a Twitter (X) dataset
IAB Tech Lab 2023 â usedfor â User Attribute Inference
confidence 100% ¡ we annotate occupations and interests using two standardized taxonomies (IAB Tech Lab 2023...)
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:In this study, we present the first comprehensive evaluation of modern LLMs - including GPT-4, GPT-4o, GPT-3.5-Turbo, Gemini 1.5 Pro, DeepSeek-V3, Llama 3.2, and BERT - across three core social media analytics tasks on a Twitter (X) dataset: (I) Social Media Authorship Verification, (II) Social Media Post Generation, and (III) User Attribute Inference. For the authorship verification, we introduce a systematic sampling framework over diverse user and post selection strategies and evaluate generalization on newly collected tweets from January 2024 onward to mitigate "seen-data" bias. For post generation, we assess the ability of LLMs to produce authentic, user-like content using comprehensive evaluation metrics. Bridging Tasks I and II, we conduct a user study to measure real users' perceptions of LLM-generated posts conditioned on their own writing. For attribute inference, we annotate occupations and interests using two standardized taxonomies (IAB Tech Lab 2023 and 2018 U.S. SOC) and benchmark LLMs against existing baselines. Overall, our unified evaluation provides new insights and establishes reproducible benchmarks for LLM-driven social media analytics. The code and data are provided in the supplementary material and will also be made publicly available upon publication.
Tags
Links
- Source: https://arxiv.org/abs/2604.18955v1
- Canonical: https://arxiv.org/abs/2604.18955v1
Trouble viewing inline? Open PDF directly â
Full Text
88,675 characters extracted from source content.
Expand or collapse full text
Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest Ramtin Davoudi Utah State University ramtin.davoudi@usu.edu &Kartik Thakkar Utah State University kartik.thakkar@usu.edu &Nazanin Donyapour Independent Researcher nazanin.donyapour@gmail.com Tyler Derr Vanderbilt University tyler.derr@vanderbilt.edu &Hamid Karimi Utah State University hamid.karimi@usu.edu Abstract In this study, we present the first comprehensive evaluation of modern LLMsâincluding GPT-4, GPT-4o, GPT-3.5-Turbo, Gemini 1.5 Pro, DeepSeek-V3, Llama 3.2, and BERTâacross three core social media analytics tasks on a Twitter (X) dataset: (I) Social Media Authorship Verification, (I) Social Media Post Generation, and (I) User Attribute Inference. For the authorship verification, we introduce a systematic sampling framework over diverse user and post selection strategies and evaluate generalization on newly collected tweets from January 2024 onward to mitigate âseen-dataâ bias. For post generation, we assess the ability of LLMs to produce authentic, user-like content using comprehensive evaluation metrics. Bridging Tasks I and I, we conduct a user study to measure real usersâ perceptions of LLM-generated posts conditioned on their own writing. For attribute inference, we annotate occupations and interests using two standardized taxonomies (IAB Tech Lab 2023 and 2018 U.S. SOC) and benchmark LLMs against existing baselines. Overall, our unified evaluation provides new insights and establishes reproducible benchmarks for LLM-driven social media analytics. The code and data are provided in the supplementary material and will also be made publicly available upon publication. Assessing Capabilities of Large Language Models in Social Media Analytics: A Multi-task Quest Ramtin Davoudi Utah State University ramtin.davoudi@usu.edu Kartik Thakkar Utah State University kartik.thakkar@usu.edu Nazanin Donyapour Independent Researcher nazanin.donyapour@gmail.com Tyler Derr Vanderbilt University tyler.derr@vanderbilt.edu Hamid Karimi Utah State University hamid.karimi@usu.edu 1 Introduction Online social media platforms have become integral to modern society, generating vast amounts of user-generated content that offers unique insights into various domains such as marketing Singh and Singh (2018), public health Schillinger et al. (2020), crisis management Saroj and Pal (2020), and so on. This has led to a unique and vibrant research field known as social media analytics (or social media mining). While there has been significant progress in social media analytics since the emergence of popular social networking platforms such as Facebook and Twitter (now X), it still faces critical challenges in fully leveraging the insights and potential of social media data. One of the main challenges is the complexity of user-generated content, especially the text. The social media content is often informal, ambiguous, or irrelevant (e.g., spam and memes), making content analysis complex. Also, social media language changes rapidly (e.g., new slang, memes), so models trained on past data may quickly become outdated. Recent advances in large language models (LLMs), including GPT, DeepSeek, and Gemini, offer new opportunities to address these challenges. Trained on large and diverse corpora, LLMs demonstrate strong capabilities in understanding and generating language in dynamic and noisy environments. They have shown competitive performance across tasks such as sentiment analysis Zhang et al. (2023a), stance detection Gambini et al. (2024), misinformation classification Hu et al. (2024a), topic extraction Mu et al. (2024), and summarization Zhang et al. (2024a), often with minimal task-specific supervision. In this work, we conduct a comprehensive empirical evaluation of LLMs across three core social media analytics tasks: (I) Social Media Authorship Verification, (I) Social Media Post Generation, and (I) User Attribute Inference. Task I aims to determine whether a post was authored by a specific user. While authorship verification has applications in digital forensics and plagiarism detection Tyo et al. (2022), social media settings pose additional challenges due to short, low-quality content Huang et al. (2025). Task I focuses on generating posts that reflect a userâs style and content preferences. Although automated text generation has been widely studied Celikyilmaz et al. (2020), producing realistic and personalized social media content remains difficult Perez-Castro et al. (2023); Du et al. (2023), and LLMs are increasingly being leveraged for personalized content generation Kumar et al. (2024); Salemi et al. (2024); Zhang et al. (2024b). Task I addresses the prediction of user attributes (e.g., occupation, interests) from textual content. Such inference supports personalization and large-scale behavioral analysis Hu et al. (2007); De Bock and Van den Poel (2010); Goel and Goldstein (2014), yet remains challenging due to label ambiguity and fine-grained classification requirements. For a detailed discussion of work related to these three tasks, we refer readers to Appendix G. Why these three tasks together? Social media characterizes user identity through observable traces in user-generated content GĂźndĂźz (2017); Shulman (2022). Similar to previous studies, we conceptualize this identity at three levels: the userâs unique identity through their social media posts (Task I) Yadav and Li (2017), the userâs unique social media signature, style, and tone (Task I) Abbasi and Chen (2008), and latent demographic attributes (Task I) Tigunova et al. (2020). Thus, together, these tasks enable a coherent assessment of how LLMs infer and recognize user identity. In this paper, we evaluate several state-of-the-art LLMsâincluding GPT, Gemini, DeepSeek, Llama, and BERTâalongside traditional ML baselines on a large Twitter (X) dataset. Our contributions are: ⢠To the best of our knowledge, this is the first study to jointly and systematically assess multiple modern LLMs across three fundamental social media analytics tasks under a unified evaluation framework. ⢠For Social Media Authorship Verification (Task I), we design diverse user and post sampling strategies to enable controlled and realistic benchmarking. We further address âseen dataâ (data leakage) bias by evaluating models on newly collected tweets from January 2024 onward, testing generalization beyond training cut-off periods. ⢠For Social Media Post Generation (Task I), we introduce a structured evaluation protocol for user-conditioned generation based on a curated set of verified, active, and profile-rich users, combining lexical, semantic, and diversity-based metrics to characterize trade-offs between authenticity and fluency. ⢠Bridging Tasks I and I, we conduct the first user study measuring real usersâ perceptions of authenticity in LLM-generated tweets, enabling direct comparison between automatic metrics and human judgments. ⢠For User Attribute Inference (Task I), we predict usersâ occupations and interests using two standardized taxonomies: the IAB Tech Lab Content Taxonomy v3.1 and the 2018 U.S. Standard Occupational Classification (SOC). Grounding evaluation in these multi-level ontologies enables reproducible labeling, hierarchical analysis from coarse to fine granularity, and alignment with established real-world classification standards. Collectively, this work establishes a unified and reproducible benchmarking framework for analyzing the capabilities and limitations of modern LLMs in social media analytics. Remark: In this study, we use Gemini 1.5 Pro, DeepSeek-V3, Llama 3.2, and three versions of GPT (4, 4o, and 3.5-Turbo). For brevity, we drop the versions from Gemini, DeepSeek, and Llama. 2 Methodology Figure 1: An overview of the proposed methodology across three social media analytic tasks Figure 1 demonstrates an overview of the proposed methodology for investigating the capabilities of LLMs across three social media analytic tasks. All three tasks are fed from a social media network dataset consisting of the temporal friendship graph and the user-generated posts acquired from Kheiri et al. (2023) (see the Dataset in Appendix A for more information). We focus on Twitter (X) due to its text-centric nature and well-defined relational signals (e.g., follower networks and various interactions), which make it suitable for controlled benchmarking. The dataset includes temporal network structure and user-generated posts, enabling evaluation of authorship verification, post generation, and attribute inference. Although our experiments use Twitter data, the methodology is platform-agnostic and can be adapted to other platforms by redefining relational and interaction signals. Next, we explain each component of this method. 2.1 Social Media Authorship Verification In this task, we frame a binary classification task in which an LLM must distinguish between social media posts genuinely authored by a target user and those authored by other users. Next, we explain each part of this task. Positive User (and Post) Sampling. Positive (or target) users are users whose posts are assessed by an LLM. The purpose of this sampling is twofold: (1) The sheer high number of social media users in our dataset (around 120K) makes it impossible to assess all of the users (mainly due to the cost); and (2) We can sample different users to investigate LLMsâ power across diverse groups. We sample three sets of target (positive) users: ⢠Random (Rnd): A simple random selection of users. ⢠Recent Active (Rec): Users active within the last three weeks in our dataset. ⢠Top Active Users (Top): Users with the highest overall number of posts (tweets). Once a positive user is selected (shown in green color in Figure 1), we chronologically sort their posts (posts shown in green in Figure 1). Then, we sample two sets of posts: (1) Few-shot Example Posts: k older posts and (2) Positive Evaluation Posts: m1m_1 newer postsâSee Figure 1. The former is used as training examples for LLM to "become familiar" with the positive userâs content (i.e., employing few-shot samples), while the latter is used to form the binary classification described below. Negative Social Media Post Sampling. To form the binary classification task for an LLM, for each positive user, we also sample m2m_2 posts from other users called Negative Evaluation Posts (shown in orange color in Figure 1). To make a fair comparison, we sample these posts at the same time interval as the Positive Evaluation Posts. More specifically, let the target userâs m1m_1 Positive Evaluation Posts correspond to a time interval [Ďstart,Ďend] [ _start,\, _end ], where Ďstart _start and Ďend _end are the earliest and latest timestamps of these posts, respectively (see Figure 1). Then, we employ one of the following strategies for Negative Evaluation Posts within the same time interval of [Ďstart,Ďend] [ _start,\, _end ]: ⢠Random Sampling: A simple random draw of m2m_2 posts from a pool of other usersâ posts. ⢠Similar Topics Sampling: We embed each of the m1m_1 Positive Evaluation Posts into vector representations (e.g., using a text embedding model). Similarly, we embed candidate tweets from other users. Then, we calculate each candidate postâs average similarity to the userâs Positive Evaluation Posts and select the top m2m_2 most similar posts to form Negative Evaluation Posts. We call this sampling strategy Topic-similar. ⢠Social Graph-based Sampling: We leverage the social graph to draw Negative Evaluation Posts. For each positive user uiu_i, we randomly select m2m_2 posts from other users who are followees (âv|vâui\â v|vâ u_i\), followers (âv|vâui\â v|vâ u_i\), or Reciprocal (âv|vâui\â v|v u_i\) of uiu_i, making three sample sets. This sampling results in three sets, referred to as Followers-only, Followees-only, and Reciprocal. LLMâs Prompt. After selecting a target (positive) user and their k Few-shot Example Posts and forming both m1m_1 Positive Evaluation Posts, m2m_2 Negative Evaluation Posts, we present each classification instance to the LLM. The exact prompt is in Appendix B (Prompt 1). Unbiased (Unseen) Data Investigation. To evaluate authorship verification without bias from potentially memorized data, we also examine the issue of âseen dataâ by using social media posts (tweets) posted from Jan. 2024 onwardâwell beyond the LLMsâ training cutoff dates (see Appendix C, Table 11 for exact cutoff dates). Specifically, from an initial pool of systematically filtered active users (VAPOR users, described in Section 2.2), we select 50 users who each authored at least k+m1k+m_1 original tweets (excluding retweets) since 2024111Note that the dataset extends only up to 2020; therefore, we used the X API to collect more recent tweets.. This ensures a realistic, unbiased evaluation of model generalization performance on definitively unseen data. The prompt used is identical to Prompt 1 (Appendix B). Evaluation. For this task, we use weighted F1-score as the evaluation metric. 2.2 Social Media Post Generation In this task, we assess each LLMâs capability to create plausible social media posts on behalf of real users. Specifically, the objective is to generate a set of synthetic posts (tweets) that reflect the userâs writing style, topical preferences, and social context, drawing on the userâs historical posts as guiding examples. User Sampling. We first restrict our pool to users who posted at least 200 original tweets (excluding retweets) during the 2018â2020 window. Additionally, we retain only users who have more than 500 followers, more than 100 followees (friends), a non-empty bio with a description longer than 20 characters, and a verified status. These constraints help ensure each selected user is sufficiently active, has meaningful profile information, and demonstrates personal tweeting behavior. Applying the above filters yields a total of 383 users, each of whom is included in our evaluations for the post generation task. We refer to this set of users as VAPOR users (Verified, Active, Profile-rich, Original). Selecting VAPOR users may bias the data toward public-facing accounts; this was intentional for Tasks IâI to ensure high-quality ground truth. In contrast, Task I and the user study (Section 2.3 and Appendix E) include less-active users and casual tweets, ensuring coverage of broader and informal language use. Prompt Posts. To generate n new posts for each VAPOR user, we sample k user-authored posts and combine them into a single prompt, along with the userâs description and follower/following statistics (shown as Other Information in Figure 1). Each model is instructed to produce a fixed number of posts that resemble the userâs style. We explicitly request short posts not exceeding 280 characters (tweetâs character limit), formatted consecutively with minimal additional text. LLMâs Prompt. The prompt is included in Appendix B (Prompt 2). Evaluation. We evaluate the quality of generated posts against two reference sets: Prompt Posts and Non-prompt Posts. Prompt Posts consist of real user tweets that are included in the input prompt and thus visible to the LLM during generation. In contrast, Non-prompt Posts are a separate set of real tweets authored by the same user but withheld from the prompt, serving as an unseen reference set for evaluation. The purpose is to assess LLMsâ capability on âExample" and âEvaluation" sets, respectively, similar to what is usually done in machine learning modeling. As for evaluation measures, we compute standard natural language generation metrics including BLEU Papineni et al. (2002), ROUGE Lin (2004) (specifically ROUGE-1 and ROUGE-L) to quantify the lexical overlap with reference tweets. Additionally, we calculate Perplexity Chen et al. (1998) using GPT-2 Radford et al. (2019) to assess the fluency and naturalness of the generated content. Figure 2: An example of metrics for post generation evaluation Moreover, we numerically represent each of the n LLM-generated posts and each of the k posts in our evaluation sets using the SBERT sentence transformer model (all-MiniLM-L6-v2) Reimers and Gurevych (2019). Then, we form an nĂknĂ k matrix of âłM where entry âłiâjM_ij stores the cosine similarity between the LLM-generated post i and the real userâs post j. Using this matrix, we propose the following metrics (Figure 2 demonstrates an example): Average Gen vs. Real (AGR): For each generated post (row), we retrieve the maximum similarity to any real post. We then take the mean of these maxima across all generated posts. AâGâR=1nââi=1nmaxjâĄâłiâjAGR= 1n _i=1^n _jM_ij. Average Real vs. Gen (ARG): For each real post (column), we retrieve the maximum similarity to any generated post. We then take the mean of these maxima across all real posts. AâRâG=1kââj=1kmaxiâĄâłiâjARG= 1k _j=1^k _iM_ij. Average Overall Similarity (AOS): The mean of all pairwise cosine similarities between generated and real posts. AâOâS=1nĂkââiâjâłiâjAOS= 1nĂ k _i _jM_ij. Gen Dispersion Ratio (GDR): The fraction of real posts uniquely identified as the most similar match across generated posts (row ratio). GDR=|uânâiâqâuâeâ(argâĄmaxââłiâj|i=1,âŚ,n)|kGDR= |unique ( \ j ~M_ij\, |\,i=1,âŚ,n \ ) |k Real Dispersion Ratio (RDR): The fraction of generated posts uniquely identified as the most similar match across real posts (column ratio). RDR=|uânâiâqâuâeâ(argâĄmaxââłiâj|j=1,âŚ,k)|nRDR= |unique ( \ i ~M_ij\, |\,j=1,âŚ,k \ ) |n The last two metrics (GâDâRGDR and RâDâRRDR) assess the breadth of coverage between generated and real posts: how broadly or narrowly generated posts cover different real posts, and vice versa. 2.3 User Study Bridging Task I and Task I, we conducted a user study to evaluate the perceived authenticity of LLM-generated posts. The study was approved by our universityâs IRB (Institutional Review Board), and participants were recruited via an open call. Eligibility required being at least 18 years old, owning a public Twitter/X account, and having at least 50 original tweets. Nineteen users met these criteria and completed the study. Each participant evaluated a personalized set of 20 LLM-generated tweets (five per model from DeepSeek, Gemini, GPT, and Llama) and two authentic tweets randomly sampled from their own timeline, which served as attention checks. Generated tweets were conditioned on the userâs bio, follower/followee counts, and 50 sampled tweets. The exact prompt is provided in Appendix B (Prompt 2). Participants were instructed as follows: âThe following tweets were generated by AI (LLM) using your publicly available tweets. For each of them, rank how likely it is that you would write it.â Participants rated each tweet using a five-point scale: Definitely not me, Probably not me, Unsure, Probably me, and Definitely me. If a participant failed to select Probably me or Definitely me for either of their two real tweets, their responses were deemed unreliable and excluded. This left us 1212 valid participants. Surveys were administered individually via the Qualtrics platform, with each participant receiving a $15 Amazon gift card. We also collected basic demographic and Twitter (X) usage information of the participants, summarized in Appendix E, Table 15. 2.4 User Attribute Inference To categorize user attributes in our study, we utilize two formal taxonomies. First, we draw on the IAB Tech Lab Content Taxonomy v3.1 IAB Tech Lab (2023), a widely adopted standard in digital advertising that provides a hierarchical classification of online content across diverse topics (e.g., news, sports, business, etc.). Specifically, we use the top-level categories of this taxonomy. Second, we use the 2018 Standard Occupational Classification (SOC) System U.S. Bureau of Labor Statistics (2018), which is maintained by the U.S. Bureau of Labor Statistics to systematically classify occupations in the United States. Using these categories, we annotated and labeled the occupations and interests/hobbies of our VAPOR users (described in Section 2.2). Two authors of this paper collaboratively annotated each userâs occupation and interests by closely examining profile descriptions and historical tweets, with disagreements resolved by a third author (a senior researcher). For interests, we initially considered 39 categories from the IAB Tech Lab Content Taxonomy, of which 25 were represented among the selected users. For occupations, we relied on the SOC 2018 hierarchy and annotated at two levelsâLevel 1 (L1) and Level 2 (L2)âcomprising 18 and 38 occupational groups, respectively, to balance granularity and coverage. The full category lists are provided in Appendix F (see Tables 18â20), and the Task I prompt is included in Appendix B (Prompt 3). Evaluation. For each user, an LLM yields a category (for both occupations and interests/hobbies). Then, we evaluate the performance against the ground truth user attributes using accuracy and weighted F1-score metrics. 3 Experiments In this section, we describe our experimental design and present detailed evaluations of the LLMs on three main tasks. The experiments were conducted on a system with an AMD EPYC 7513 CPU, 4 NVIDIA RTX A4000 GPUs, and 1 TB of RAM. Table 1: Experimental results for social media authorship verification, Task I, (metric: weighted F1-score, multiplied by 100 for clarity) Negative Social Media Post Sampling Model Reciprocal Followees-only Followers-only Random Topicâsimilar Avg Unseen Data (Knowledge Cut-off) Avg Rank Positive User Sampling Rnd Rec Top Rnd Rec Top Rnd Rec Top Rnd Rec Top Rnd Rec Top GPT-4 85 83 84 78 86 83 74 87 86 94 93 95 76 83 80 84.5 85 1.07 Gemini 75 81 80 59 78 79 69 78 80 84 85 87 57 72 69 75.5 79 2.87 DeepSeek 73 77 77 68 77 76 72 76 75 86 87 87 56 70 70 75.1 80 3.47 RF 70 80 82 63 78 75 64 76 77 59 67 72 78 65 65 71.4 - 4.33 TF-IDF 60 60 60 60 60 60 57 60 60 71 81 81 57 75 75 65.1 - 5.80 Compression-NCD 59 63 62 59 62 62 57 61 61 65 78 77 55 71 71 64.2 - 6.20 SIAMESE + SBERT 57 60 59 59 60 60 56 60 60 69 74 74 60 34 34 58.0 - 7.67 SIAMESE + GloVe 19 75 77 40 70 68 65 7 76 60 62 62 33 34 49 53.1 - 7.87 Llama 61 74 76 50 70 71 55 72 70 51 51 51 43 47 44 59.1 56 7.87 GPT-3.5-Turbo 58 61 59 57 56 56 55 57 55 62 60 59 57 57 57 57.7 71 8.47 Bert 55 77 4 27 67 41 51 7 73 41 33 36 43 34 33 41.5 - 10.13 USE 28 43 43 41 44 44 28 43 44 49 60 59 42 53 53 44.9 - 10.47 3.1 Social Media Authorship Verification For this task, we benchmark five LLMsâGPT-4, GPT-3.5-Turbo, Gemini, Llama, and DeepSeekâalongside BERT Devlin et al. (2019) and a Random Forest classifier (training details in Appendix C). We further include established authorship verification baselines: Universal Sentence Encoder (USE) cosine similarity Cer et al. (2018), TF-IDF cosine similarity Stamatatos (2009), a compression-based impostors method using normalized compression distance (NCD) Potha and Stamatatos (2017), and two Siamese models based on GloVe-LSTM Boenninghoff et al. (2019) and SBERT embeddings Reimers and Gurevych (2019) (baseline descriptions in Appendix C). Performance is measured using weighted F1-score, pooling predictions across users. We evaluate 15 settings formed by the Cartesian product of three positive user sampling schemesâRandom, Recent Active, and Top Activeâand five negative post sampling strategiesâRandom, Topic-similar, Followers-only, Followees-only, and Reciprocal (Section 2.1). The number of Few-shot Example Posts and Positive/Negative Evaluation Posts is fixed (k=m1=m2=20k=m_1=m_2=20), and each setting includes 50 users. Table 1 reports results, including evaluation on unseen data collected after model knowledge cut-off dates. As shown in Table 1, GPT-4 consistently achieves the strongest performance, with an average F1-score of 0.845 and the top rank (1.07). Gemini and DeepSeek form a strong second tier, with average F1-scores of approximately 0.755 and 0.751, respectively. Traditional baselines (e.g., Random Forest, TF-IDF, and Compression-NCD) achieve intermediate performance, while Siamese models (SBERT and GloVe) show moderate capability and sensitivity to sampling conditions. In contrast, Llama, GPT-3.5-Turbo, BERT, and USE exhibit lower and more variable performance, highlighting the robustness of GPT-4, Gemini, and DeepSeek for authorship verification. On unseen data, GPT-4 again leads (0.85), followed by DeepSeek (0.80) and Gemini (0.79), demonstrating strong generalization beyond the training period. Table 2: Impact of positive user sampling and negative post sampling strategies on model accuracy. Model User Effect (p) Tweet Effect (p) DeepSeek 4.37 20.17 Gemini 6.62 18.26 GPT-3.5-Turbo 4.53 10.62 GPT-4 5.53 14.65 Llama 9.47 21.13 Average 6.10 16.97 Table 2 further shows that negative post sampling has a substantially larger impact on accuracy (â 17 percentage points) than positive user sampling (â 6 points). Llama and DeepSeek are most sensitive to negative sampling variations, while GPT-3.5-Turbo shows the lowest variability. Overall, these results underscore GPT-4âs robustness and the critical role of negative post sampling in authorship verification. Appendix C provides supplementary materials for Task I, including details on the computation of the User and Tweet Effects, qualitative analysis (Table 9), class-wise results, and implementation details such as model access and hyperparameters. 3.2 Social Media Post Generation The evaluated models for this task include GPT-4o, Gemini, DeepSeek, and Llama, along with traditional baselines such as the Markov Chain Shannon (1948); Freitas et al. (2015), BART-large Lewis et al. (2019), and T5-large Raffel et al. (2020). For BART and T5, we used the âlargeâ pretrained variants. We used k=50k=50, n=10n=10, and 5050 Prompt and 5050 Nonprompt posts (Section 2.2). Table 3: Experimental results for social media post generation (Task I) Set Model AOS AGR ARG GDR RDR BLEU ROUGE-1 ROUGE-L Prompt GPT-4o 0.196 0.466 0.333 0.157 0.919 0.049 0.233 0.187 Gemini 0.199 0.442 0.319 0.134 0.900 0.041 0.209 0.178 DeepSeek 0.214 0.493 0.363 0.163 0.926 0.065 0.264 0.226 Llama 0.188 0.489 0.341 0.156 0.908 0.135 0.292 0.253 Markov Chain 0.260 0.750 0.424 0.117 0.695 0.957 0.670 0.668 BART-large 0.257 0.516 0.386 0.140 0.800 0.099 0.241 0.179 T5-large 0.228 0.542 0.372 0.060 0.800 0.029 0.234 0.189 Nonprompt GPT-4o 0.191 0.417 0.315 0.144 0.891 0.029 0.211 0.168 Gemini 0.195 0.424 0.312 0.133 0.893 0.037 0.201 0.171 DeepSeek 0.209 0.434 0.341 0.151 0.885 0.042 0.236 0.204 Llama 0.183 0.408 0.317 0.146 0.875 0.035 0.220 0.181 Markov Chain 0.245 0.504 0.374 0.112 0.658 0.123 0.345 0.320 BART-large 0.215 0.438 0.330 0.120 1.000 0.006 0.183 0.130 T5-large 0.219 0.456 0.351 0.140 0.900 0.017 0.206 0.170 Table 4: Perplexity Scores (Task I) Model Perplexity GPT-4o 73.43 Gemini 129.31 DeepSeek 141.83 Llama 76.78 Markov Chain 263.26 BART-large 1.123 T5-large 3.00 Tables 4 and 4 summarize model behavior in post generation. The Markov Chain baseline shows strong replication (high AOS and BLEU) but very poor fluency (Perplexity: 263.26). Among LLMs, DeepSeek attains the highest overall semantic similarity in the Prompt setting, indicating closer semantic alignment with real posts, yet its high perplexity suggests reduced fluency. Llama achieves the highest BLEU and ROUGE scores among LLMs, reflecting stronger lexical reuse, but records the lowest overall semantic similarity. GPT-4o offers the most balanced profile, combining strong semantic similarity, broad coverage (high RDR), and the lowest perplexity among LLMs. Gemini emphasizes paraphrasing, yielding moderate semantic similarity with comparatively higher perplexity. Transformer baselines (BART-large, T5-large) achieve extremely low perplexityâindicating highly predictable outputsâbut show narrower coverage under Prompt conditioning. Overall, the results reveal trade-offs among semantic similarity, fluency, lexical reuse, and coverage breadth. Figure 3: User survey responses categorized by generated tweets from different LLMs 3.3 User Study Results Table 5: Experimental results for user attribute inference (Task I) Model Interests/Hobbies Occupations (L1) Occupations (L2) Accuracy Weighted F1 Accuracy Weighted F1 Accuracy Weighted F1 GPT-4o 71.54 75.07 67.62 66.01 56.13 51.50 Gemini 76.24 76.49 78.32 76.61 61.87 60.83 DeepSeek 69.19 69.53 71.27 68.19 57.18 51.53 Llama 30.54 36.28 35.77 43.29 2.61 1.35 PreoĹŁiuc-Pietro et al. (2015) 2.60 2.33 22.08 29.27 13.04 10.32 Lewis et al. (2019) (Few-shot) 46.56 31.27 70.13 63.08 50.65 39.47 Michelson and Macskassy (2010) 15.61 8.11 27.42 3.84 12.99 5.56 Pennacchiotti and Popescu (2011) 41.56 37.75 55.84 49.58 32.47 32.12 To assess perceived authenticity, we conducted the user study described in Section 2.3. Figure 3 shows the distribution of ratings across five categories. Gemini and Llama receive the highest concentration of positive judgments (Definitely/Probably me; 44/60 each), followed by GPT-4o (41/60) and DeepSeek (39/60). GPT-4o also receives the most Definitely not me responses, while DeepSeek shows the highest number of Probably not me ratings. The neutral option (Unsure) remains below 15% across models, suggesting that participants generally formed clear opinions. Mean authenticity scores further confirm this pattern: Gemini and Llama achieve the highest ratings (3.95/5), with DeepSeek and GPT-4o slightly lower (3.67â3.68). Additional statistical and qualitative analyses are provided in Appendix E (Tables 14 and 16). Note: All users consented to having both their own tweets and LLM-generated tweets made publicly available. We further compare human ratings (Appendix E, Table 14) with automatic metrics (Tables 4 and 4) to examine their alignment with perceived authenticity. Llama shows the strongest consistency, combining high author-likeness ratings with strong BLEU and ROUGE scores, suggesting that lexical reuse enhances perceived authenticity. Gemini achieves similarly high human ratings despite lower lexical overlap and higher perplexity, indicating that paraphrastic imitation can also be effective. DeepSeek attains high semantic similarity but lower authenticity ratings, implying that topical alignment alone is insufficient. GPT-4o demonstrates moderate performance across both human and automatic measures, reinforcing that no single metric reliably predicts perceived author-likeness. Overall, the results reveal trade-offs across lexical, stylistic, and semantic dimensions. 3.4 User Attribute Inference We evaluate GPT-4o, Gemini, DeepSeek, and Llama on user attribute inference by predicting the occupations and interests of VAPOR users (Section 2.2). For each user, 50 sampled tweets are provided as input to the prompt described in Section 2.4. We compare LLMsâ performance against several baselines, including a Gaussian Process classifier with Word2Vec embeddings PreoĹŁiuc-Pietro et al. (2015), few-shot BART Lewis et al. (2019), an entity-based DBpedia classifier Michelson and Macskassy (2010), and a TF-IDF gradient boosting model Pennacchiotti and Popescu (2011). Table 5 shows the results. Performance is measured using accuracy and the weighted F1-score. Gemini achieves the strongest results across interests and both occupational levels (L1 and L2). GPT-4o and DeepSeek perform comparably at broader levels (L1) but decline at finer granularity (L2), while Llama consistently underperforms. These findings indicate that increasing classification granularity significantly affects LLM accuracy, with Gemini demonstrating the most robust overall performance. Error patterns are further analyzed via confusion matrices (Table 17, Appendix F). Appendix F also provides implementation details, baseline descriptions, category lists, and qualitative analysis. 4 Conclusion In this paper, we presented a comprehensive evaluation of modern LLMs across three core social media analytics tasks: social media authorship verification, social media post generation, and user attribute inference. GPT-4 achieved the strongest results in authorship verification, particularly under the unseen-data regime. In the post generation, DeepSeek showed strong semantic alignment, while Gemini and Llama received the highest human authenticity ratings. For attribute inference, Gemini performed best, especially at finer-grained hierarchical levels. Overall, our findings highlight the importance of controlled sampling and multifaceted evaluation when assessing LLMs in social media contexts. Our study has some limitations. Although no other datasets meeting our criteria were identified, future work could extend our methodology to additional datasets and platforms (e.g., Reddit). Moreover, our user study is not large-scale. Future studies should use larger, more diverse samples. Future work may extend this framework by incorporating multimodal signals (e.g., images, videos, and social graphs) to enrich stylistic and demographic analyses. Modeling reposting and diffusion dynamics could further clarify content virality. More broadly, advancing LLM-based modeling of social network evolutionâpotentially via link prediction that integrates textual, temporal, and structural signalsâwould strengthen comprehensive benchmarks in social media analytics. Finally, we note ethical risks in modeling user identity, including impersonation and privacy-sensitive profiling. This work is intended for benchmarking, not unsafeguarded deployment. References A. Abbasi and H. Chen (2008) Writeprints: a stylometric approach to identity-level identification and similarity detection in cyberspace. ACM Transactions on Information Systems (TOIS) 26 (2), p. 1â29. Cited by: §1. T. Alsanoosy, B. Shalbi, and A. Noor (2024) Authorship attribution for english short texts. Engineering, Technology & Applied Science Research 14 (5), p. 16419â16426. Cited by: Appendix G. B. Boenninghoff, R. M. Nickel, S. Zeiler, and D. Kolossa (2019) Similarity learning for authorship verification in social media. In ICASSP 2019-2019 IEEE international conference on acoustics, speech and signal processing (ICASSP), p. 2457â2461. Cited by: 4th item, §3.1. A. Celikyilmaz, E. Clark, and J. Gao (2020) Evaluation of text generation: a survey. arXiv preprint arXiv:2006.14799. Cited by: §1. D. Cer, Y. Yang, S. Kong, N. Hua, N. Limtiaco, R. S. John, N. Constant, M. Guajardo-Cespedes, S. Yuan, C. Tar, et al. (2018) Universal sentence encoder for english. In Proceedings of the 2018 conference on empirical methods in natural language processing: system demonstrations, p. 169â174. Cited by: 1st item, §3.1. N. V. Chawla, K. W. Bowyer, L. O. Hall, and W. P. Kegelmeyer (2002) SMOTE: synthetic minority over-sampling technique. Journal of artificial intelligence research 16, p. 321â357. Cited by: Appendix C. S. F. Chen, D. Beeferman, and R. Rosenfeld (1998) Evaluation metrics for language models. Cited by: §2.2. K. De Bock and D. Van den Poel (2010) Predicting website audience demographics forweb advertising targeting using multi-website clickstream data. Fundamenta Informaticae 98 (1), p. 49â70. Cited by: §1. J. Devlin, M. Chang, K. Lee, and K. Toutanova (2019) Bert: pre-training of deep bidirectional transformers for language understanding. In Proceedings of the 2019 conference of the North American chapter of the association for computational linguistics: human language technologies, volume 1 (long and short papers), p. 4171â4186. Cited by: §3.1. H. Du, W. Xing, and B. Pei (2023) Automatic text generation using deep learning: providing large-scale support for online learning communities. Interactive Learning Environments 31 (8), p. 5021â5036. Cited by: §1. C. Freitas, F. Benevenuto, S. Ghosh, and A. Veloso (2015) Reverse engineering socialbot infiltration strategies in twitter. In Proceedings of the 2015 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining 2015, p. 25â32. Cited by: 1st item, §3.2. M. Gambini, C. Senette, T. Fagni, and M. Tesconi (2024) Evaluating large language models for user stance detection on x (twitter). Machine Learning 113 (10), p. 7243â7266. Cited by: §1. S. Goel and D. G. Goldstein (2014) Predicting individual behavior with social networks. Marketing Science 33 (1), p. 82â93. Cited by: §1. U. GĂźndĂźz (2017) The effect of social media on identity construction. Mediterranean journal of social sciences 8 (5). Cited by: §1. T. Hong, J. Choi, K. Lim, and P. Kim (2021) Enhancing personalized ads using interest category classification of sns users based on deep neural networks. Sensors 21 (1), p. 199. Cited by: Appendix G. B. Hu, Q. Sheng, J. Cao, Y. Shi, Y. Li, D. Wang, and P. Qi (2024a) Bad actor, good advisor: exploring the role of large language models in fake news detection. In Proceedings of the AAAI Conference on Artificial Intelligence, Vol. 38, p. 22105â22113. Cited by: §1. J. Hu, H. Zeng, H. Li, C. Niu, and Z. Chen (2007) Demographic prediction based on userâs browsing behavior. In Proceedings of the 16th international conference on World Wide Web, p. 151â160. Cited by: §1. Y. Hu, Z. Hu, C. Seah, and R. K. Lee (2024b) InstructAV: instruction fine-tuning large language models for authorship verification. arXiv preprint arXiv:2407.12882. Cited by: Appendix G. B. Huang, C. Chen, and K. Shu (2024) Can large language models identify authorship?. In Findings of the Association for Computational Linguistics: EMNLP 2024, Cited by: Appendix G. B. Huang, C. Chen, and K. Shu (2025) Authorship attribution in the era of llms: problems, methodologies, and challenges. ACM SIGKDD Explorations Newsletter 26 (2), p. 21â43. Cited by: §1. J. Huertas-Tato, A. MartĂn, and D. Camacho (2024) Understanding writing style in social media with a supervised contrastively pre-trained transformer. Knowledge-Based Systems 296, p. 111867. Cited by: Appendix G. IAB Tech Lab (2023) Content taxonomy: v3.1. Note: https://iabtechlab.com/standards/content-taxonomy/Accessed: 2025-05-07 Cited by: §2.4. M. Injadat, F. Salo, and A. B. Nassif (2016) Data mining techniques in social media: a survey. Neurocomputing 214, p. 654â670. Cited by: Appendix G. K. Kheiri, M. F. A. Khan, T. Derr, and H. Karimi (2023) An analysis of the dynamics of ties on twitter. In 2023 IEEE International Conference on Big Data (BigData), p. 5809â5817. Cited by: Table 6, Appendix A, §2. I. Kumar, S. Viswanathan, S. Yerra, A. Salemi, R. A. Rossi, F. Dernoncourt, H. Deilamsalehy, X. Chen, R. Zhang, S. Agarwal, et al. (2024) Longlamp: a benchmark for personalized long-form text generation. arXiv preprint arXiv:2407.11016. Cited by: §1. M. Lewis, Y. Liu, N. Goyal, M. Ghazvininejad, A. Mohamed, O. Levy, V. Stoyanov, and L. Zettlemoyer (2019) BART: denoising sequence-to-sequence pre-training for natural language generation, translation, and comprehension. arXiv preprint arXiv:1910.13461. Cited by: 2nd item, 2nd item, §3.2, §3.4, Table 5. C. Lin (2004) Rouge: a package for automatic evaluation of summaries. In Text summarization branches out, p. 74â81. Cited by: §2.2. X. Liu, B. Peng, M. Wu, M. Wang, H. Cai, and Q. Huang (2024) Occupation prediction with multimodal learning from tweet messages and google street view images. AGILE: GIScience Series 5, p. 36. Cited by: Appendix G. M. Michelson and S. A. Macskassy (2010) Discovering usersâ topics of interest on twitter: a first look. In Proceedings of the fourth workshop on Analytics for noisy unstructured text data, p. 73â80. Cited by: 3rd item, §3.4, Table 5. Y. Mu, C. Dong, K. Bontcheva, and X. Song (2024) Large language models offer an alternative to the traditional approach of topic modelling. arXiv preprint arXiv:2403.16248. Cited by: §1. K. Papineni, S. Roukos, T. Ward, and W. Zhu (2002) Bleu: a method for automatic evaluation of machine translation. In Proceedings of the 40th annual meeting of the Association for Computational Linguistics, p. 311â318. Cited by: §2.2. M. Pennacchiotti and A. Popescu (2011) A machine learning approach to twitter user classification. In Proceedings of the international AAAI conference on web and social media, Vol. 5, p. 281â288. Cited by: 4th item, §3.4, Table 5. A. Perez-Castro, M.R. MartĂnez-Torres, and S.L. Toral (2023) Efficiency of automatic text generators for online review content generation. Technological Forecasting and Social Change 189, p. 122380. External Links: ISSN 0040-1625, Document, Link Cited by: §1. R. G. Pillai, A. Fokkens, and W. van Atteveldt (2025) Engagement-driven persona prompting for rewriting news tweets. In Proceedings of the 31st International Conference on Computational Linguistics, p. 8612â8622. Cited by: Appendix G. N. Potha and E. Stamatatos (2017) An improved impostors method for authorship verification. In International conference of the cross-language evaluation forum for European languages, p. 138â144. Cited by: 3rd item, §3.1. D. PreoĹŁiuc-Pietro, V. Lampos, and N. Aletras (2015) An analysis of the user occupational class through twitter content. In Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Processing (Volume 1: Long Papers), p. 1754â1764. Cited by: §3.4, Table 5. Z. Qiu, H. Lyu, W. Xiong, and J. Luo (2025) Can llms simulate social media engagement? a study on action-guided response generation. arXiv preprint arXiv:2502.12073. Cited by: Appendix G. A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever, et al. (2019) Language models are unsupervised multitask learners. OpenAI blog 1 (8), p. 9. Cited by: §2.2. C. Raffel, N. Shazeer, A. Roberts, K. Lee, S. Narang, M. Matena, Y. Zhou, W. Li, and P. J. Liu (2020) Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of machine learning research 21 (140), p. 1â67. Cited by: 3rd item, §3.2. N. Reimers and I. Gurevych (2019) Sentence-bert: sentence embeddings using siamese bert-networks. arXiv preprint arXiv:1908.10084. Cited by: 5th item, §2.2, §3.1. A. Salemi, S. Mysore, M. Bendersky, and H. Zamani (2024) Lamp: when large language models meet personalization. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), p. 7370â7392. Cited by: §1. A. Saroj and S. Pal (2020) Use of social media in crisis management: a survey. International Journal of Disaster Risk Reduction 48, p. 101584. Cited by: §1. D. Schillinger, D. Chittamuru, and A. S. RamĂrez (2020) From âinfodemicsâ to health promotion: a novel framework for the role of social media in public health. American journal of public health 110 (9), p. 1393â1396. Cited by: §1. C. E. Shannon (1948) A mathematical theory of communication. The Bell system technical journal 27 (3), p. 379â423. Cited by: 1st item, §3.2. D. Shulman (2022) Self-presentation: impression management in the digital age. In The Routledge international handbook of Goffman studies, p. 26â37. Cited by: §1. M. Singh and G. Singh (2018) Impact of social media on e-commerce. International Journal of Engineering & Technology 7 (2.30), p. 21â26. Cited by: §1. E. Stamatatos (2009) A survey of modern authorship attribution methods. Journal of the American Society for information Science and Technology 60 (3), p. 538â556. Cited by: 2nd item, §3.1. A. Tigunova, P. Mirza, A. Yates, and G. Weikum (2020) RedDust: a large reusable dataset of reddit user traits. In Proceedings of the Twelfth Language Resources and Evaluation Conference, p. 6118â6126. Cited by: §1. J. Tyo, B. Dhingra, and Z. C. Lipton (2022) On the state of the art in authorship attribution and authorship verification. arXiv preprint arXiv:2209.06869. Cited by: §1. U.S. Bureau of Labor Statistics (2018) Standard occupational classification (soc) system. Note: https://w.bls.gov/soc/Accessed: 2025-05-07 Cited by: §2.4. H. Wen, Z. Xiao, E. Hovy, and A. G. Hauptmann (2023) Towards open-domain twitter user profile inference. In Findings of the Association for Computational Linguistics: ACL 2023, p. 3172â3188. Cited by: Appendix G. H. Yadav and J. Li (2017) Social media writing style fingerprint. arXiv preprint arXiv:1712.04762. Cited by: §1. E. Yu, J. Li, and C. Xu (2024) RePALM: popular quote tweet generation via auto-response augmentation. In Findings of the Association for Computational Linguistics ACL 2024, p. 9566â9579. Cited by: Appendix G. T. Zhang, F. Ladhak, E. Durmus, P. Liang, K. McKeown, and T. B. Hashimoto (2024a) Benchmarking large language models for news summarization. Transactions of the Association for Computational Linguistics 12, p. 39â57. Cited by: §1. W. Zhang, Y. Deng, B. Liu, S. J. Pan, and L. Bing (2023a) Sentiment analysis in the era of large language models: a reality check. arXiv preprint arXiv:2305.15005. Cited by: §1. X. Zhang, Y. Malkov, O. Florez, S. Park, B. McWilliams, J. Han, and A. El-Kishky (2023b) Twhin-bert: a socially-enriched pre-trained language model for multilingual tweet representations at twitter. In Proceedings of the 29th ACM SIGKDD conference on knowledge discovery and data mining, p. 5597â5607. Cited by: Appendix G. Z. Zhang, R. A. Rossi, B. Kveton, Y. Shao, D. Yang, H. Zamani, F. Dernoncourt, J. Barrow, T. Yu, S. Kim, et al. (2024b) Personalization of large language models: a survey. arXiv preprint arXiv:2411.00027. Cited by: §1. Y. Zhao, Y. Wang, X. Cheng, A. M. Tumlin, Y. Liu, D. Xia, M. Jiang, and T. Derr (2025) Amplifying your social media presence: personalized influential content generation with llms. arXiv preprint arXiv:2505.01698. Cited by: Appendix G. Appendices Appendix A Dataset We performed our evaluations using the Twitter dataset introduced by Kheiri et al. (2023), comprising data from over 120,000 users tracked across a 15-week period. This dataset is particularly suitable due to its large scale, diversity, and comprehensive metadata and content coverage. It includes weekly snapshots of user networks, capturing the dynamic formation and dissolution of follower connections. Additionally, it provides detailed user-generated content (tweets), interactions (such as mentions), and extensive user profile attributes, including verification status and user activity metrics. The dataset captures an average of 1,175,846 new tweets weekly and approximately 2.9 million total social ties, making it ideal for our comprehensive analysis. We have explicitly obtained permission from the original authors to utilize this dataset in our study. The dataset statistics are summarized in Table 6. Table 6: Twitter dataset statistics acquired from Kheiri et al. (2023) Network Property Value Total users 123,829 Total ties 2,922,732 # Verified accounts 3,829 Avg. weekly new followers 10,855 Avg. weekly new unfollowers 465 Avg. weekly new Tweets 1,175,846 Percentage verified users 1.687 Avg. followees count per user 205 Avg. followers count per user 150 Diameter (longest shortest path) 8 Avg. new tweets (w/mentions) 2,021 Appendix B LLMsâ Prompts and Expenses User Description: â¨description⊠Previous Tweets: â¨previous_tweets⊠Tweet to Classify: â¨tweet_to_classify⊠Task: Based on the description and the previous tweets of the user, output â0â if the above tweet was not generated by the user or â1â if it was. Even if the tweet is a URL, output only â0â or â1â as your response. Prompt 1: Social media authorship verification prompt (Task I) Here, description and previous_tweets correspond to the userâs bio and a few-shot list of the userâs previously authored tweets (excluding retweets). Meanwhile, tweet_to_classify is drawn from either the target userâs own future tweets (Positive Evaluation Posts) or from other usersâ tweets (Negative Evaluation Posts). The LLM outputs either 0 or 1, indicating whether it believes the specified tweet belongs to the same user. We have collected K tweets from one user on Twitter, included below: <tweets go here> Below, you can also find the user-written bio (description) as well as the number of followings and followers of the user: User Bio: â¨user_description⊠Number of Followings: â¨following_count⊠Number of Followers: â¨followers_count⊠Now, generate exactly â¨NUM_GENERATED_TWEETS⊠tweets that the user could have posted on their Twitter account. Each tweet should not exceed 280 characters and must be formatted as: 1. <tweet_text> 2. <tweet_text> ⌠Ensure all â¨NUM_GENERATED_TWEETS⊠tweets are included, without extra spacing or missing tweets. Prompt 2: Social media Post generation prompt (Task I) For inferring the usersâ interests and hobbies, we provide the LLMs with the following structured prompt: Your task is to infer the primary â¨ATTRIBUTE⊠of a Twitter/X user based on their tweets. You will be provided with exactly â¨#SAMPLED_TWEETS⊠randomly selected tweets posted by this user. Assign the user to exactly one of the following â¨ATTRIBUTE⊠categories, based on the content of their tweets: â¨CATEGORY_LIST⊠Here are the userâs sampled tweets: â¨SAMPLED_TWEETS⊠Respond strictly and exclusively in the following JSON format: "â¨JSON_KEYâŠ": "â¨one category from the above listâŠ" Do not include any other text, explanation, or formatting. Prompt 3: User attribute inference prompt (Task I) We instantiate Prompt 3 for (i) interests/hobbies, (i) SOC 2018 Level 1 occupations, and (i) SOC 2018 Level 2 occupations by substituting the placeholders in Table 7. Table 7: Placeholders for interest and occupation prompt. Task ATTRIBUTE CATEGORY_LIST JSON_KEY Interests / hobbies interest category IAB categories present in our data interest_category Occupation L1 occupational category (Level--1) 18 SOC major groups occupation_category Occupation L2 occupational category (Level--2) 38 SOC minor groups occupation_category LLMsâ Costs Table 8 provides estimated expenses incurred from using LLMs in our experiments. The total estimated cost for all model experiments across all three tasks is $220. Table 8: Estimated cost of using LLMs in our experiments. Model Cost (USD) GPT-4 $150 GPT-4o $20 GPT-3.5-Turbo $20 Gemini $20 DeepSeek $10 Llama Free Total $220 Appendix C Social Media Authorship Verification Model Access and Hyperparameters We accessed LLMs through their respective APIsâOpenAIâs ChatCompletion endpoint for GPT-4o, Googleâs GenerativeModel API for Gemini, and the ollama.chat interface for Llama models. For consistency across models, we fix the temperature parameter to 0 in all LLM calls, ensuring deterministic outputs in the binary classification setting. Qualitative Analysis for Task I (Case Studies) As shown in Table 9, the userâs previous tweets reveal a clear and consistent pattern of sharing educational YouTube content explicitly linked to the brand âLetstute." Gemini, GPT-4o, and DeepSeek effectively recognize this pattern, correctly identifying tweets associated with âLetstute" educational videos as authentic. GPT-3.5 struggles because it relies heavily on exact wording matches and misses broader topical connections. Llama mistakenly accepts short conversational tweets unrelated to the userâs known educational content, indicating difficulty in grasping the userâs overall topical theme. Table 9: Authorship verification example with model predictions. (GT: Ground Truth, â: correct, â: incorrect) Previous tweets from the user: ⢠I liked a @YouTube video https://youtube.com/watch?v=IWX2MobGsnE Amazing Trick To Understand Arithmetic Progression Formula. ⢠How to present an Answer | Smart Answer Presentation | Letstute: https://youtube.com/watch?v=UH5fspzdRDc ⢠GST (Goods & Service Tax) | Problem Solving | LetsTute: https://youtube.com/watch?v=g_HIDyeAmDI Tweet Example GT Gem. 4o 3.5 Lla. DS T1 Revise your syllabus CBSE class 10th as per NCERT solution https://youtube.com/watch?a&v=aPGgIvHGuQM 1 â â â â â T2 3 tips to speak English fluently https://youtube.com/watch?a&v=bQgQSF5HP_c 1 â â â â â T3 @tvserieshub She can not! 0 â â â â â T4 @evanwolf No, I was still doing the actual writing on the computer. 0 â â â â â Random Forest Baseline Implementation For the authorship verification task, we trained a Random Forest (RF) classifier separately for each dataset configuration, using the same user and tweet splits as in the LLM experiments. Specifically, for each user, we constructed a balanced, pairwise training set by (i) pairing each of the userâs k=20k=20 Few-shot Example Posts with every other Few-shot Example Post from the same user, yielding kĂ(kâ1)=380kĂ(k-1)=380 positive pairs per user (labeled as 1), and (i) pairing each of these k Few-shot Example Posts with m2=20m_2=20 Negative Evaluation Posts sampled from other users within the same timeframe, resulting in an additional kĂm2=400kĂ m_2=400 negative pairs per user (labeled as 0). For the test set, we paired each userâs k=20k=20 Few-shot Example Post with their m1=20m_1=20 Positive Evaluation Posts (labeled as 1) and another set of m2=20m_2=20 Negative Evaluation Posts from other users (labeled as 0), ensuring no overlap and preventing trainâtest leakage. Each tweet was embedded once using SBERT all-MiniLM-L6-v2, and embeddings of each tweet pair were concatenated into a single 768-dimensional feature vector. For dataset configurations involving graph-based sampling (Followers-only, Followees-only, Reciprocal), potential class imbalance in the training set was addressed using the Synthetic Minority Over-Sampling Technique (SMOTE) Chawla et al. (2002). Each RF model was tuned independently via 3-fold grid search, optimizing the hyperparameters detailed in Table 10. At test time, given a candidate tweet f, we formed k=20k=20 pairs by individually pairing it with each prior tweet: (t1,f),(t2,f),âŚ,(tk,f)(t_1,f),(t_2,f),âŚ,(t_k,f). The RF predicted labels independently for each pair, resulting in a set of pairwise predictions P=p1,p2,âŚ,pkP=\p_1,p_2,âŚ,p_k\, where each piâ0,1p_iâ\0,1\. The final classification for the candidate tweet was determined by majority voting: y^f=1ifâi=1kpi>k20otherwise y_f= cases1&if _i=1^kp_i> k2\\[6.0pt] 0&otherwise cases The predicted label y^f y_f indicates whether the candidate tweet f is considered authored by the target user (y^f=1 y_f=1) or by a different user (y^f=0 y_f=0). Similar to the LLM experiments, we evaluated this modelâs performance using accuracy and weighted F1-scores, ensuring direct comparability with the results obtained from the LLMs. Table 10: Search space for the 3-fold RF hyper-parameter tuning Hyperparameter Grid values nestimatorsn_estimators 100, 200 max_depth None, 10, 30 min_samples_split 2, 5 min_samples_leaf 1, 3 Descriptions of Additional Baseline Models To complement our LLM-based evaluation, we implemented a diverse set of established authorship verification baselines from the literature. Below is a summary of each: ⢠Universal Sentence Encoder (USE) + Cosine Similarity: This semantic similarity approach uses Googleâs pretrained Universal Sentence Encoder to embed tweets. An authorâs representation is formed by averaging the embeddings of their prior tweets. Cosine similarity is then computed between this profile and each candidate tweet, and a similarity threshold determines the classification Cer et al. (2018). ⢠TF-IDF + Cosine Similarity: A traditional lexical matching method that encodes tweets using TF-IDF vectors with unigrams, bigrams, and trigrams. The averaged TF-IDF vector of a userâs previous tweets forms the author profile, and cosine similarity is used to compare it to test tweets Stamatatos (2009). ⢠GZip Compression Distance (NCD): This compression-based baseline constructs a profile by concatenating all of a userâs previous tweets. It computes the Normalized Compression Distance (NCD) between the profile and each test tweet using GZip. A dynamic threshold based on median distances is used for classification Potha and Stamatatos (2017). ⢠Siamese Network with GloVe + LSTM: This neural architecture uses pretrained GloVe embeddings and an LSTM encoder to represent tweets. A Siamese network computes the cosine similarity between a pair of encoded tweets, and the result is passed through a sigmoid output layer to determine authorship likelihood Boenninghoff et al. (2019). ⢠Siamese Network with SBERT Embeddings: Each tweet is encoded using a pretrained SBERT model to capture semantic meaning. The Siamese architecture computes cosine similarity between tweet embeddings and uses a dense layer to classify whether the tweets belong to the same author Reimers and Gurevych (2019). LLMs Knowledge Cut-off Dates The knowledge cut-off dates of the evaluated LLMs for assessing their generalization performance on âunseen data" (Section 2.1 and 3.1) are presented in Table 11. As it can be observed, all dates are before January 2024 onward, the date for which we we collected the new (âunseen data"). Table 11: LLMs and their knowledge cutoff dates Model Cut-off Date GPT-4 Oct. 2023 GPT-4o Sep. 2023 GPT-3.5-Turbo Aug. 2021 Gemini 1.5 Pro Nov. 2023 DeepSeek-V3 Oct. 2023 Llama 3.2 Dec. 2023 Impact of Sampling Strategies The values reported in Table 2 for the User Effect and Tweet Effect were computed as follows: ⢠User Effect: For each negative post sampling strategy, we calculated the maximum and minimum accuracy across the three user post sampling strategies (Rnd, Rec, Top), and averaged these differences across all negative post sampling strategies: Ueffect=1|T|ââtâT(maxuâUâĄ(At,u)âminuâUâĄ(At,u))U_effect= 1|T| _tâ T ( _uâ U(A_t,u)- _uâ U(A_t,u) ) ⢠Tweet Effect: For each positive user sampling strategy, we calculated the maximum and minimum accuracy across the five negative post sampling strategies (Reciprocal, Followees-only, Followers-only, Random, Topic-similar), and averaged these differences across all positive user sampling strategies: Teffect=1|U|ââuâU(maxtâTâĄ(At,u)âmintâTâĄ(At,u))T_effect= 1|U| _uâ U ( _tâ T(A_t,u)- _tâ T(A_t,u) ) Where: ⢠T represents the set of negative post sampling strategies: Reciprocal, Followees-only, Followers-only, Random, Topic-similar. ⢠U represents the set of positive user sampling strategies: Random (Rnd), Recently active (Rec), Top-active (Top). ⢠At,uA_t,u denotes the accuracy obtained for a specific negative post sampling strategy t and positive user sampling strategy u. Average performance analysis (Figure. 4 and Table 12) indicates that GPTâ4 consistently leads, delivering high and balanced performance across both Negative Evaluation Posts and Positive Evaluation Posts (genuine userâs tweets) classes. Gemini and DeepSeek represent a robust second tier, achieving strong accuracy but with a clearer bias toward the Positive class. While GPTâ3.5âTurbo stands out as the most balanced model, it does so at a noticeably lower overall performance level. Traditional models like Random Forest exhibit substantial Positive-class bias, and other baselines such as BERT and Llama highlight significant performance gaps and limitations. Similarly, classical baseline methods (e.g., USE, TF-IDF, Gzip, and SIAMESE variants) show varying degrees of accuracy and imbalance. Figure 4: Mean class-wise weighted F1-scores for each model. Points above the dashed line (y=xy=x) indicate better performance on Positive Evaluation Posts; points below indicate better performance on Negative Evaluation posts. Table 12: Mean for weighted F1âscores per class and absolute balance gap. Model F1 (Neg.) F1 (Pos.) Gap GPTâ4 7171 8989 1818 Gemini 1.5 Pro 5858 8282 2424 DeepSeek 5555 8282 2727 GPTâ3.5âTurbo 5454 5656 22 RF 4848 7979 3131 Llama 3.2 2323 7676 5353 BERT 2929 4444 1515 USE 5454 3636 1818 TF-IDF 5757 6666 99 Gzip 5252 6868 1616 SIAMESE + SBERT 4444 6565 2121 SIAMESE + GloVe 3333 6060 2727 Figure 5 details the class-specific precision and recall for the authorship verification task. GPTâ4 maintains a high precision (0.803) for Class 0, effectively reducing false positives, alongside an excellent recall (0.939) for Class 1, correctly identifying most user-authored tweets. Gemini and DeepSeek show strong balanced recall (approximately 0.87 and 0.86) for Class 1, but lower precision (0.66 and 0.62) for Class 0, reflecting moderate false-positive rates. While GPTâ3.5âTurbo achieves high recall (0.930) for Class 0, its precision is notably low (0.391), indicating aggressive labeling with many false positives; inversely, it shows high precision (0.931) but low recall (0.402) for Class 1. Traditional methods such as Random Forest, TF-IDF, Gzip, SIAMESE+SBERT, SIAMESE+GloVe, and USE exhibit varied levels of performance, with notable biases and moderate precision-recall trade-offs. In contrast, BERT and Llama demonstrate significantly limited performance across both precision and recall metrics. Overall, these precision-recall analyses reinforce GPTâ4âs robustness and superior capability in accurately verifying authorship with balanced performance. Figure 5: Precision and recall comparison across models for the authorship verification task. Figure 6 summarizes the mean evaluation metrics (accuracy, precision, recall, and F1-score) across evaluated models for the authorship verification task. GPTâ4 achieves the highest overall performance, with accuracy (85.08%) and F1-score (0.8898), reflecting a robust balance between precision (85.16%) and recall (93.82%). Gemini and DeepSeek exhibit strong and balanced results, with Gemini slightly outperforming DeepSeek in accuracy (76.25%) and F1-score (0.8240). Conversely, GPT-3.5-Turbo demonstrates an imbalanced performance, with exceptional precision (93.17%) but notably low recall (40.11%), indicative of conservative classification behavior. Traditional methods like Random Forest provide moderate and balanced performance (accuracy: 73.25%), whereas baseline methods such as Llama 3.2 (accuracy: 64.41%, high recall but lower precision) and BERT (accuracy: 51.44%) show clear limitations. Additional classical baselines (USE, TF-IDF, Gzip, SIAMESE+SBERT, and SIAMESE+GloVe) also exhibit varied levels of accuracy and balance, underscoring GPTâ4âs superior performance and robustness in authorship verification. Figure 6: Heatmap illustrating mean evaluation metrics (accuracy, precision, recall, and F1-score) across different models evaluated on the authorship verification task. Higher scores (darker colors) indicate stronger performance. Appendix D Social Media Post Generation Model Access and Hyperparameters For tweet generation, we accessed OpenAIâs GPT-4o using a batch API to reduce cost by submitting prompts in bulk. In contrast, Gemini, DeepSeek, and Llama (via Ollama) were accessed through individual non-batched calls. The generation temperature was uniformly set to 0.7 across all models to promote content diversity while avoiding degenerate or repetitive outputs. Descriptions of Post Generation Baselines To complement our evaluation of LLMs on the social media post generation task, we implemented a set of traditional baseline models that reflect diverse approaches to content generation. Below, we provide a brief description of each method. ⢠Markov Chain: A classical probabilistic model that generates new tweets by modeling transition probabilities between word sequences. We train a second-order Markov chain using a userâs prompt tweets and sample new sentences constrained to tweet length. This method captures frequent local patterns but lacks global coherence or semantic understanding. Shannon (1948); Freitas et al. (2015) ⢠BART-large: A denoising autoencoder for sequence-to-sequence generation that uses a bidirectional encoder and a left-to-right decoder. Pretrained on large-scale corpora with text corruption tasks, BART is well-suited for conditional generation tasks, including style imitation and text rewriting. Lewis et al. (2019) ⢠T5-large: A unified text-to-text transformer that reformulates all NLP tasksâincluding generationâas text-to-text problems. The large variant is pretrained on the C4 corpus with a multi-task objective and fine-tuned for conditional generation. Raffel et al. (2020) Qualitative Analysis for Task I (Case Studies) Table 13 compares tweets originally posted by one user to tweets generated by four LLMs (DeepSeek, GPT-4o, Gemini, and Llama). The table highlights two main topics from the userâs tweets: appreciation for volunteers and recognition of alumni scholarship achievements. Among the LLM models, DeepSeek and GPT-4o deliver the closest matches. Both accurately reflect the volunteer appreciation theme with similar enthusiasm, although GPT-4o slightly lacks specifics regarding the alumni scholarship. Gemini effectively praises volunteers but becomes too general when mentioning scholarships. Lastly, Llama maintains relevance to Science Olympiad but often overlooks specific details, shifting attention to general team accomplishments and different scholarships. Overall, DeepSeek and GPT-4o provide tweets most faithful to the userâs original content, while Gemini and Llama are thematically relevant but less precise. Table 13: Comparison of one userâs tweets and LLM-generated tweets. Theme A â Volunteer appreciation blue!12Real: @MDSciOly Congrats to Russ for being recognized for his incredible Science Olympiad volunteer work! We are so proud! @SOAlumniNetwork Gen (DeepSeek): Shoutout to all the incredible coaches and volunteers who make Science Olympiad possible. Your passion fuels the next generation of innovators! #ThankYou Comment: Matches closely; same focus on volunteers with similar enthusiasm. Gen (GPT-4o): Shoutout to all the incredible volunteers who make Science Olympiad possible! Your dedication inspires future scientists every day. Thank you! #VolunteerAppreciation Comment: Excellent match; closely mirrors topic, appreciation tone, and enthusiasm. Gen (Gemini): Volunteers are the heart of Science Olympiad! Thank you for your dedication. #volunteerappreciation Comment: Good match on volunteers, but a bit general compared to the original. Gen (Llama): Meet our incredible teams of 7,800+ students from across 50 states gearing up for the ultimate STEM competition - Science Olympiad! Comment: Stays related to Science Olympiad but misses the volunteer focus. Theme B â Alumni success & scholarships blue!12Real: This Science Olympiad alum and valedictorian candidate just won the prestigious $40,000 Edison Scholarship! Tanya is MIT-bound and plans to major in chemical engineering. Way to go! #STEM #scienceolympiad Gen (DeepSeek): Meet Tanya, a Science Olympiad alum and valedictorian who just won the $40K Edison Scholarship! Your future starts here. #STEMSuccess Comment: Good match; includes the same name, scholarship amount, and excitement. Gen (GPT-4o): Big congrats to our alumni who are making waves in the STEM world! Your achievements showcase the lasting impact of Science Olympiad. Keep shining! #ScienceOlympiadAlumni Comment: Good topical alignment about alumni success, though it omits specific scholarship details. Gen (Llama): Congratulations to our Division C winners who each won $1,000 college scholarships from DuPont! Your hard work pays off! Comment: On-topic with scholarships but uses different details and lacks personal specifics. Gen (Gemini): Science Olympiad alumni, where are you now? Share your success stories! #sciolyalumni Comment: Mentions alumni generally but misses details about the specific scholarship and winner. Appendix E User Study Table 14 summarises the 60 ratings collected for each model. For every system we report three descriptive statistics: the mean score on the fiveâpoint âdefinitely not meâ (1) â âdefinitely meâ (5) scale, its associated 95% confidence interval (CI), and the Topâ2âbox proportionâi.e., the share of responses falling in the two most positive categories, an intuitive acceptance rate. Gemini and Llama obtain the strongest reception: almost threeâquarters of their tweets are marked probably me or definitely me, and both achieve a mean rating of 3.95/5. DeepSeek and GPTâ4o lag by roughly 0.30 points in mean score and by about 5â8 percentage points in the positive bins metric, implying a modestâbut not dramaticâloss of perceived author likeness. Because the 95 % CIs overlap across all four systems, these differences should be viewed as suggestive rather than conclusive. To test whether the full fiveâcategory rating distributions differ by model, we conducted a Pearson Ď2Ď^2 test on the complete 4Ă54Ă 5 contingency table (four models, five response categories). The result, Ď2=7.29Ď^2=7.29 with 12 degrees of freedom (p=0.84p=0.84), fails to reject the null hypothesis of identical distributions. In practical terms, the available 60 judgments per model do not provide sufficient power to claim a statistically significant winner, even though Gemini and Llama trend higher on both descriptive metrics. Table 14: Participantâlikeness ratings on a 1â5 scale (higher = more authentic). Model Mean Âą 95% CI Positive Bins % DeepSeek 3.68Âą0.343.68Âą 0.34 65.0 Gemini 3.95Âą0.323.95Âą 0.32 73.3 GPT-4o 3.67Âą0.363.67Âą 0.36 68.3 Llama 3.95Âą0.293.95Âą 0.29 73.3 Table 15 presents detailed demographic characteristics and additional information about the user study participants. Table 15: Demographics of User Study Participants Age Range ⢠[22â26): 63.64% ⢠[26â30]: 36.36% Gender ⢠Male: 91.67% ⢠Female: 8.33% Tweet Activity ⢠Daily: 66.67% ⢠A few times/week: 25.00% ⢠Weekly: 8.33% Main Tweet Topics ⢠Entertainment: 91.67% ⢠Technology/AI: 75.00% ⢠Sports: 58.33% ⢠Science: 50.00% ⢠Politics: 41.67% ⢠Personal updates: 33.33% ⢠E-commerce: 33.33% ⢠Religion: 8.33% Qualitative Analysis for User Study (Case Studies) Table 16 compares original tweets from one participant with tweets generated by four different LLMs (GPT-4o, DeepSeek, Gemini, and Llama), focusing on two specific topics: fintech skepticism and car purchasing decisions. Among these models, GPT-4o best captures the userâs style, accurately reflecting the original tweetsâ topics and conversational tone, though it sometimes sounds slightly more formal. DeepSeek does very well in fintech discussions, matching the userâs viewpoint closely, but struggles with other topics like cars, where it drifts into unrelated general tech content. Gemini often misses the userâs specific interests entirely, defaulting instead to generic technology-related statements. Lastly, Llama picks up on contextual details (e.g., Canadian outdoors) but frequently misses the userâs main points about fintech and specific vehicle choices, shifting instead toward broader, less relevant subjects. Table 16: Comparison of one participantâs original tweets and LLM-generated tweets. Theme A â Fintech skepticism An overview of the user: Male, 26 years old, Software Developer (USA). Tweets weekly, mainly about Technology/AI and Entertainment. blue!12Real sample tweet: Are there any cons to @Wealthsimple? So far I donât see why anyone would use a traditional bank. Gen (DeepSeek): @Wealthsimple Your app is a game-changer. Traditional banks should be worried. Comment: Closely matches topic and tone; clearly shares the userâs viewpoint. Userâs survey response: Definitely me Gen (GPT-4o): Letâs talk smart investments. With the rise of fintech platforms, do you think traditional banks will become obsolete? #Fintech #Investing Comment: Keeps the casual question-and-answer style and fintech focus, but slightly more formal. Userâs survey response: Definitely me Gen (Gemini): Is it just me, or is AI getting smarter every day? Kind of exciting, kind of terrifying. #AI Comment: Completely off-topic; shifts to general AI instead of finance. Userâs survey response: Definitely me Gen (Llama): Thinking of switching to a Linux distribution for my next project â any recommendations? #Linux Comment: Misses the financial topic completely; instead emphasizes a general tech persona. Userâs survey response: Definitely not me Theme B â Car decisions blue!12Real: First time Subaru is stuck. Hmm should I get a Lexus GX or a Land Cruiser? Gen (GPT-4o): Thinking about switching from Subaru to Lexus GX. Need something more rugged for those off-road adventures. Any thoughts? #CarTalk Comment: Excellent match; closely follows userâs original question and specific vehicle brands. Userâs survey response: Probably me Gen (Llama): Planning a road trip to explore the Canadian wilderness â any recommendations for off-road gear? #offroadadventures #Canada Comment: Captures outdoor theme but doesnât reflect userâs original car-brand comparison. Userâs survey response: Unsure Gen (Gemini): Electric cars are the future. Just sayinâ. #EV #electriccar Comment: Briefly touches on cars, but too general; misses userâs interest in specific brands and off-road vehicles. Userâs survey response: Probably me Gen (DeepSeek): Just pre-ordered the new Framework laptop. Modular design is the future. #RightToRepair Comment: Completely unrelated topic; shifts to a generic tech subject rather than cars. Userâs survey response: Probably not me Appendix F User Attribute Inference Misclassification Patterns Table 17 summarizes the most frequent trueâpredicted label confusions in Task I, along with their occurrence counts. Across both interest and occupation inference, most errors occur between semantically adjacent categories with overlapping topical or stylistic signals, rather than between unrelated domains. This suggests that misclassifications largely reflect boundary ambiguity in social media language rather than systematic model failure. Table 17: Most Frequent Confusion Pairs in User Attribute Inference (Task I) Category Type True Category Predicted Category Count Interest (IAB) Entertainment Pop Culture 45 Pop Culture Entertainment 18 Occupation (L1L_1) Education, Training, and Library Arts, Design, Entertainment, Sports, and Media 33 Management Arts, Design, Entertainment, Sports, and Media 22 Occupation (L2L_2) Entertainers and Performers Media and Communication Workers 43 Media and Communication Workers Entertainers and Performers 41 Model Access and Hyperparameters For the user attribute inference task, we set the generation temperature to 0 for all models to ensure consistent outputs and reproducibility. GPT-4o was accessed via OpenAIâs ChatCompletion endpoint, Gemini through Googleâs GenerativeModel API, DeepSeek using its REST interface, and Llama via local inference through Ollama. All models were queried individually in non-batch mode, and responses were parsed to extract structured JSON-formatted predictions. Table 18: Interest Categories Based on IAB Content Taxonomy v3.1 ID Interest Category I1 Attractions I2 Automotive I3 Books and Literature I4 Business and Finance I5 Careers I6 Communication I7 Crime I8 Disasters I9 Education I10 Entertainment I11 Fine Art I12 Food & Drink I13 Hobbies & Interests I14 Home & Garden I15 Law I16 Medical Health I17 Pets I18 Politics I19 Pop Culture I20 Science I21 Sports I22 Style & Fashion I23 Technology & Computing I24 Travel I25 Video Gaming Table 19: Occupational Categories (Level 1) Based on SOC 2018 ID Occupation Category (Level 1) L1-1 Accommodation and Food Services L1-2 Arts, Design, Entertainment, Sports, and Media Occupations L1-3 Community and Social Service Occupations L1-4 Computer and Mathematical Occupations L1-5 Education, Training, and Library Occupations L1-6 Healthcare Practitioners and Technical Occupations L1-7 Legal Occupations L1-8 Life, Physical, and Social Science Occupations L1-9 Management Occupations L1-10 Management, Business, Science, and Arts Occupations L1-11 Management, Business, and Financial Occupations L1-12 Office and Administrative Support Occupations L1-13 Production Occupations L1-14 Professional and Related Occupations L1-15 Professional, Scientific, and Technical Services L1-16 Protective Service Occupations L1-17 Sales and Office Occupations L1-18 Sales and Related Occupations Table 20: Occupational Categories (Level 2) Based on SOC 2018 ID Occupation Category (Level 2) L2-1 Accommodation L2-2 Advertising, Marketing, Promotions, Public Relations, and Sales Managers L2-3 Arts and Design Workers L2-4 Assemblers and Fabricators L2-5 Business and Financial Operations Occupations L2-6 Computer Occupations L2-7 Computer and Information Systems Managers L2-8 Computer and Mathematical Occupations L2-9 Counselors, Social Workers, and Other Community and Social Service Specialists L2-10 Educational Instruction and Library Occupations L2-11 Entertainers and Performers, Sports and Related Workers L2-12 Financial Clerks L2-13 Firefighting and Prevention Workers L2-14 Health Technologists and Technicians L2-15 Information and Record Clerks L2-16 Law Enforcement Workers L2-17 Lawyers, Judges, and Related Workers L2-18 Legal Occupations L2-19 Librarians, Curators, and Archivists L2-20 Life Scientists L2-21 Life, Physical, and Social Science Occupations L2-22 Management Occupations L2-23 Mathematical Science Occupations L2-24 Media and Communication Occupations L2-25 Media and Communication Workers L2-26 Office and Administrative Support Occupations L2-27 Other Management Occupations L2-28 Other Protective Service Workers L2-29 Physical Scientists L2-30 Postsecondary Teachers L2-31 Professional, Scientific, and Technical Services L2-32 Retail Sales Workers L2-33 Sales Representatives, Services L2-34 Sales Representatives, Wholesale and Manufacturing L2-35 Scientific Research and Development Services L2-36 Social Science Occupations L2-37 Social Scientists and Related Workers L2-38 Top Executives Descriptions of User Attribute Inference Baselines To complement our evaluation of LLMs on the user attribute inference task, we implemented a set of traditional baseline models. Below, we briefly describe each method along with their references. ⢠(PreoĹŁiuc-Pietro et al. 2015): A probabilistic classifier that employs Word2Vec embeddings to represent tweets and spectral clustering to group semantically related words into clusters. It handles class imbalance through random oversampling, followed by classification using a Gaussian Process with an Automatic Relevance Determination (ARD) kernel. ⢠Lewis et al. (2019): A fine-tuning approach leveraging a pre-trained BART model (bart-large-mnli) on a small subset of labeled user tweets (20%). It employs random oversampling for balancing classes, evaluating performance on an independent test set. ⢠Michelson and Macskassy (2010): This method extracts named entities from tweets and maps them to DBpedia categories to build user interest profiles. It infers user attributes based on the frequency of categories associated with these entities, capturing topical interests rather than textual semantics. ⢠Pennacchiotti and Popescu (2011): A gradient boosting classifier that utilizes textual features extracted via TF-IDF vectorization of user tweets. The classifier addresses class imbalance using sample weighting to optimize predictive performance. Qualitative Analysis for Task I (Case Studies) Table 21 shows occupation predictions by four LLMs (Gemini, GPT-4o, DeepSeek, and Llama) based solely on tweets from a user involved in emergency management. Among the models, Gemini correctly identifies the occupation as âOther Protective Service Workers," accurately capturing specialized professional cues such as âCertified Emergency Manager," mentions of FEMA, and emergency response drills. In contrast, GPT-4o incorrectly emphasizes general community support themes and hashtags, leading to a prediction in social and community services. DeepSeek misinterprets hazard-related language as indicating frontline firefighting roles, failing to distinguish between emergency response coordination and direct hazard management. Finally, Llama provides an overly broad classification (âPublic Service"), overlooking precise professional indicators and domain-specific terminology found clearly within the userâs tweets. Table 21: Occupation (L2) inference from tweets. User Bio (not shown to models): Disaster technologist, inclement weather enthusiast, tenderâhearted public servant. Just trying to make the world a safer place. Views expressed are my own. Sample Tweets: â âI am an internationally Certified Emergency Manager with a Bachelor of Science in Emergency Management.â â âIâm wrapping up the last bit of workâfun for the week by reviewing the new @FEMA ICS/NIMS courses.â â âTaught an Incident Command System (ICS) course with the brilliant @IDICworld today!â â âJust wrappingâup a fun day helping test our statewide response capability to a complex coordinated cyber attack.â â âI can not say enough good things about the Emergency Management Accreditation Program! Urban or rural, private or public âŚâ â âThose are some pretty serious polygons! Stay safe friends! #KSWXâ â â#Preach! Everyone is entitled to #SelfCare in any form, as long as it isnât hurting anyone else.â GroundâTruth Occupation (SOC 2018): Other Protective Service Workers (33â9099) Model Predictions and Diagnostic Analysis: Gemini: Other Protective Service Workers Comment: Correct. The model accurately identified clear evidence, such as âCertified Emergency Manager,â âICS training,â mentions of âFEMA,â and disaster-response drills, matching closely with the protective-service occupation. GPTâ4o: Counselors, Social Workers, and Other Community & Social Service Specialists Comment: Incorrect. The model was influenced heavily by general expressions of well-being and hashtags like #SelfCare, overlooking specific professional terms related to emergency management. DeepSeek: Firefighting and Prevention Workers Comment: Incorrect. The model overly focused on tweets about hazards and severe weather (e.g., âpolygonsâ in weather warnings), mistakenly interpreting coordination and training roles as frontline firefighting. Llama: Public Service Comment: Incorrect. Too broad and vague, this prediction missed the specific professional cues (e.g., certifications and specialized acronyms such as âICS,â âNIMS,â and âCEMâ) clearly indicating emergency management work. Interest and Occupation Category Taxonomies Tables 18, 19, and 20 list the interest categories and the Level 1 and Level 2 occupation categories, respectively, as defined by the IAB Content Taxonomy and the 2018 SOC classification. These tables document the label space used for user attribute inference and support reproducibility of our experiments. Appendix G Related Work The research in social media analytics has been conducted using conventional machine learning and statistical methods Injadat et al. (2016). The introduction of transformer-based LLMs specially versatile ones such as Gemini and GPT, has marked a significant advancement in social media analytics. Next, we briefly review the recent studies that utilize LLMs in the three tasks of our interest. Recent authorship attribution and verification studies increasingly use LLMs to capture distinctive writing styles from text alone Huertas-Tato et al. (2024). Fine-tuned transformer encoders (e.g., BERT, RoBERTa) applied to authorship tasks have achieved state-of-the-art accuracy, surpassing traditional stylometric approaches that rely on handcrafted features Hu et al. (2024b). More recent works explore both fine-tuning and prompting strategies with advanced LLMs. For example, (Huang et al., 2024) demonstrated that GPT-based LLMs can accurately verify authorship (and even attribute texts to the correct author among many candidates) in a zero-shot setting without task-specific training, essentially establishing new performance benchmarks. Other researchers have proposed prompt-based techniques to harness LLMsâ knowledge; for instance, a âPromptAVâ method uses step-by-step stylometric cues to improve GPT-3.5âs verification accuracy and explainability, and a linguistically informed prompting approach similarly guides GPT-3.5/4 models to strong authorship verification results even without fine-tuning Hu et al. (2024b). However, these LLM-driven approaches typically consider only textual content, omitting valuable contextual signals such as user profile bios or social network features. In the social media domain, ignoring such metadata can be limiting â social media posts are short and rife with slang, often making it difficult to identify the author from text alone Alsanoosy et al. (2024). In contrast to prior studies, our work integrates rich contextual metadata (e.g., profile descriptions and network-derived features) into the authorship verification pipeline. We introduce systematic and robust user/post sampling strategies to construct a diverse evaluation set, and we mitigate potential data leakage biases by using tweet content posted LLMsâ knowledge cut-off. A growing body of work uses LLMs to write (generate) social media posts, yet each study tackles a very specific goal. RePALM fineâtuned GPTâ3.5 with a reinforcementâlearning reward that predicts likes and retweets, so it generates quoteâtweets optimized purely for popularity Yu et al. (2024). Pillai et al. prompted GPTâ4 to rewrite news headlines into tweets in three fixed persona styles (formal, casual, factual) to boost engagement Pillai et al. (2025). Qiu et al. first predicted whether a user will retweet, quote or reply to a trending post and then let GPTâ4 craft the corresponding response, but the model still handled one interaction type at a time Qiu et al. (2025). Across these efforts, the LLM sees little more than the source post (plus an optional style tag); richer cues such as the authorâs bio, follower network, or recent tweets are ignoredâeven though adding social signals during preâtraining is known to improve tweet representations Zhang et al. (2023b); Zhao et al. (2025). Our study fills this gap by conditioning generation on a compact userâcontext blockâbio, follower/followee counts, and representative past tweetsâand by rating the outputs on four complementary dimensions: (i) semantic fidelity (how closely each tweetâs meaning matches the userâs real posts), (i) output diversity (coverage of topics and phrasings), (i) stylistic congruence (faithfulness to the userâs voice or brand tone), and (iv) perceived authenticity (how natural and humanâlike the tweets sound). More importantly, we assess LLMsâ post-generation capability by asking real users to rate tweets the models create from their own timelines. Recent LLM-based approaches to social media user profiling typically target a single user attribute at a time â for example, classifying only a personâs occupation or their political leaning Liu et al. (2024); Wen et al. (2023). To improve predictive power, many studies leverage auxiliary user data beyond the social media posts themselves. It is common to incorporate profile descriptions, full timelines, or social network cues along with the post text Hong et al. (2021); Wen et al. (2023). For instance, users often self-report their job roles or hobbies in their bios Hong et al. (2021), and some profiling methods feed such metadata (and even friendship information) into models alongside tweet content Wen et al. (2023). Moreover, prior work seldom applies standard taxonomies for labeling user attributes. Instead, researchers usually define task-specific or coarse-grained categories â e.g. grouping occupations into a few broad classes Liu et al. (2024) or using ad-hoc sets of interest topics â rather than mapping to established schemas. In contrast, our approach infers both occupation and personal-interest profiles simultaneously using only each userâs tweet content, without relying on any self-description or network features. We further constrain and explain the modelâs outputs by grounding them in official taxonomies (the SOC for occupations and the IAB Tech Lab content taxonomy for interests), which enables more standardized, interpretable predictions in comparison to previous methods.