Paper deep dive
Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media
Yijie Xu, Chao Wang, Hui Xiong
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 8/22/2026, 1:54:44 AM
Summary
The paper introduces TSN4PI, a unified framework for tracking the evolution of political ideologies on social media. It comprises two modules: PIDN, which uses large language models with style transfer and unsupervised domain adaptation to detect ideology and filter noise, and PIPN, which employs temporal graph neural networks to predict future ideological shifts. The authors release two large-scale datasets (news sources and X/Twitter user posts) and validate the framework on X and Truth Social, finding that users' ideologies often trend toward the center rather than polarizing.
Entities (14)
Relation Signals (11)
Chao Wang → affiliatedwith → University of Science and Technology of China
confidence 95% · Affiliation: School of Artificial Intelligence and Data Science, University of Science and Technology of China
Yijie Xu → affiliatedwith → The Hong Kong University of Science and Technology (Guangzhou)
confidence 95% · Affiliation: Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou)
TSN4PI → containsmodule → PIDN
confidence 95% · It includes two core modules. The PIDN uses large language models...
TSN4PI → containsmodule → PIPN
confidence 95% · It includes two core modules... The PIPN employs temporal graph neural networks...
TSN4PI → validatedonplatform → X
confidence 95% · Extensive case studies on multiple platforms (X and Truth Social) validate the effectiveness of TSN4PI
TSN4PI → validatedonplatform → Truth Social
confidence 95% · Extensive case studies on multiple platforms (X and Truth Social) validate the effectiveness of TSN4PI
TSN4PI → addressesproblem → political polarization
confidence 90% · Against Political Polarization: A Unified Framework... highlighting the need to understand individual political ideologies
PIDN → usestechnique →
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideologies and their temporal dynamics. This task faces challenges such as data scarcity, abundant non-political content, costly and bias-prone manual annotation, and difficulty in modeling future ideological inclinations. To address these issues, we propose TSN4PI, a unified framework for tracking the evolution of political ideologies on social media. It includes two core modules. The PIDN uses large language models with style transfer and unsupervised domain adaptation to enable robust ideology detection and filter irrelevant content from noisy, cross-domain data. The PIPN employs temporal graph neural networks to predict future ideological shifts, enabling comprehensive analysis of ideology presence, intensity, and evolution. We release two large-scale datasets for noncommercial research use to facilitate further work. Extensive case studies on multiple platforms (X and Truth Social) validate the effectiveness of TSN4PI and provide empirical insights into political polarization and the evolution of online ideologies. Our findings offer a nuanced perspective, advancing both methodological development and empirical understanding in this field.
Tags
Links
- Source: https://arxiv.org/abs/2608.17987v1
- Canonical: https://arxiv.org/abs/2608.17987v1
Trouble viewing inline? Open PDF directly →
Full Text
126,344 characters extracted from source content.
Expand or collapse full text
Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social MediaDOI: X.XXXXXXXJournal: TISTCCS: Computing methodologies Natural language processingCCS: Computing methodologies Graph neural networksCCS: Human-centered computing Social network analysis Yijie Xu email: yxu409@connect.hkust-gz.edu.cn OrcID: 0009-0008-4529-2701 Affiliation: Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou) , Guangzhou , China , Chao Wang Note: Corresponding authors. email: wangchaoai@ustc.edu.cn Affiliation: School of Artificial Intelligence and Data Science, University of Science and Technology of China , Hefei , China and Hui Xiong email: xionghui@ust.hk Affiliation: Thrust of Artificial Intelligence, The Hong Kong University of Science and Technology (Guangzhou) , Guangzhou , China Affiliation: Department of Computer Science and Engineering, The Hong Kong University of Science and Technology , Hong Kong , Hong Kong SAR, China 2026; © , 2026; Received 22 July 2026 Abstract. The rapid growth of social media has greatly influenced political discourse, highlighting the need to understand individual political ideologies and their temporal dynamics. This task faces challenges such as data scarcity, abundant non-political content, costly and bias-prone manual annotation, and difficulty in modeling future ideological inclinations. To address these issues, we propose TSN4PI, a unified framework for tracking the evolution of political ideologies on social media. It includes two core modules. The PIDN uses large language models with style transfer and unsupervised domain adaptation to enable robust ideology detection and filter irrelevant content from noisy, cross-domain data. The PIPN employs temporal graph neural networks to predict future ideological shifts, enabling comprehensive analysis of ideology presence, intensity, and evolution. We release two large-scale datasets for noncommercial research use to facilitate further work. Extensive case studies on multiple platforms (X and Truth Social) validate the effectiveness of TSN4PI and provide empirical insights into political polarization and the evolution of online ideologies. Our findings offer a nuanced perspective, advancing both methodological development and empirical understanding in this field. Keywords: Ideology Detection, Social Media, Natural Language Processing 1. Introduction Social media has reshaped political discourse, enabling individuals to express, reinforce, and disseminate ideological views at scale (29). A concern in this environment is the formation of echo chambers and the increasing polarization of politics. These effects are particularly visible on platforms where algorithmic curation and social clustering amplify existing biases. In the United States, political opinions are often divided along partisan lines—commonly labeled as “left” (Democratic) and “right” (Republican). These distinctions influence not only electoral behavior but also patterns of engagement on digital platforms (12; 11). This scenario underscores the importance of examining these ideological differences, which have a significant impact on voting behavior and social media activism. In this paper, we aim to contribute a comprehensive methodological framework that not only elucidates these phenomena but also aids in informed policy-making for digital political environments, thereby addressing a vital need in contemporary political discourse. Four original vector panels illustrate political ideology dynamics: two asset-free mockups of contrasting social posts, directed-network snapshots at time steps 1 and 2, and custom scales placing ABC News at -2.4 (Lean Left) and Fox News at 4.0 (Right). Scale labels L, L, C, LR, and R denote Left, Lean Left, Center, Lean Right, and Right. No platform logos or profile photographs are used. Figure 1. Illustration of political ideology dynamics. (a)–(b) Original vector mockups of contrasting X-style and Truth Social-style posts; identities, text, avatars, and engagement counts are illustrative. (c) Temporal graph snapshots. (d) Original five-category scales of AllSides-derived outlet scores.Four original vector panels illustrate political ideology dynamics: two asset-free mockups of contrasting social posts, directed-network snapshots at time steps 1 and 2, and custom scales placing ABC News at -2.4 (Lean Left) and Fox News at 4.0 (Right). Scale labels L, L, C, LR, and R denote Left, Lean Left, Center, Lean Right, and Right. No platform logos or profile photographs are used. The study of political polarization and echo chambers has received increasing attention. Early work investigated ideological clustering on blogs and Facebook (30; 64), while more recent studies explored polarization during major political events (25). However, the significance and mechanisms of echo chambers remain debated (33). Yet current research confronts several challenges. First, the issue of data accessibility is prominent; platforms like Twitter impose restrictions on data collection, leading to a scarcity of suitable datasets that align with the size and focus required for comprehensive research. Moreover, the sheer volume of posts on these platforms presents a significant hurdle for manual annotation, a process further complicated by annotator bias. Another challenge is the prevalence of irrelevant content: many social media posts focus on everyday life, lacking clear political signals, and their misclassification can lead to inaccurate predictions of ideological leanings. Effectively filtering this noise to isolate genuine political discourse is paramount for accurate modeling. Additionally, conventional research methods often overlook the dynamic nature of political engagement, missing temporal variations in individual users’ political ideologies. LLMs have shown broad text-processing capabilities (9). Recent work includes retrieval-utility evaluation for retrieval-augmented generation (19), test-time adaptation (86), cross-domain sequential recommendation (84), and multimodal video and spatial reasoning (34; 48). However, direct in-context learning (ICL) (59) over social media data at scale (23; 53) can incur substantial latency and performance limitations. Watermarking has also been studied for detecting the misuse of open-source LLMs (85). These constraints motivate a resource-efficient approach tailored to the textual and graph structures of social media. To address these challenges, we propose a unified framework, TSN4PI (Temporal Social Networks for Political Ideologies), for tracing evolving political ideologies on social media. This framework consists of two core components: the Political Ideology Detection Network (PIDN) and the Political Ideology Prediction Network (PIPN). PIDN identifies and quantifies the intensity of political ideology in media content. It combines LLM-based style transfer with Unsupervised Domain Adaptation (UDA), aligning structured news (the source) with dynamic, noisy social media (the target). This approach enhances data utility, particularly when annotated social media data is scarce. By focusing on stylistic and semantic cues indicative of political leaning, the PIDN improves the discernment of political ideologies amidst a high volume of irrelevant posts, effectively aiding in filtering non-political content. Complementing the detection capabilities, PIPN leverages Temporal Graph Neural Networks (TGNNs) to forecast the evolution of users’ ideological intensity over time. The TGNN-based PIPN captures the dynamic nature of social media interactions and individual ideological trajectories, enabling prediction of future inclinations. This focus on dynamic, user-level ideological intensity distinguishes our work. While previous long-term analyses typically characterize broader patterns of political communication at the macro level, our approach models the fine-grained, temporal evolution of individual user ideologies. Furthermore, the graph-based nature of our model inherently captures social interaction patterns, offering a more holistic perspective than methods that analyze text in isolation or assume fixed ideological spectrums. Overall, TSN4PI represents a pioneering effort in concurrently determining both the existence and intensity of political ideology among social media users and predicting its evolution. Our contributions in this paper are multi-fold: • We propose TSN4PI, a novel unified framework designed to overcome key challenges in tracing political ideologies on social media. The framework features two integrated modules: 1) The PIDN, which robustly detects ideology by strategically using LLMs for offline style transfer and UDA to bridge the gap between news and noisy social media data, while effectively filtering non-political content. 2) The PIPN, which leverages TGNNs to model and predict the evolution of user ideologies, capturing their presence, intensity, and future trajectory. • We contribute two novel, large-scale, publicly released datasets to the research community, significantly addressing the challenge of data accessibility in political discourse analysis on social media platforms. These include: 1) a comprehensive dataset of news sources with curated political ideology scores from AllSides, and 2) a massive, user-centric corpus of nearly 77 million posts from 4,545 X (Twitter) users, spanning 16 years (2007–2022). Both datasets and our code are released for noncommercial research use at https://github.com/yeahjack/TSN4PI. • We conduct extensive experiments on diverse social media platforms (X and Truth Social) to demonstrate the effectiveness and robustness of TSN4PI, validating its capability to analyze complex political discourse and track ideological evolution in real-world settings. Our analysis reveals compelling insights into online political behavior. Contrary to the widely accepted echo chamber hypothesis, the users whose ideological positions move the most over time trend predominantly toward the center rather than away from it. This challenges the often-assumed trajectory of increasing polarization over time, highlighting instead a more dynamic and complex landscape of online discourse that warrants empirical, data-driven investigation. 2. Related Work User-centric Political Ideology Detection has evolved from traditional surveys and manual coding (1) to data-driven methods enabled by social media. Early studies used Twitter to infer user leanings (18), while Xiao et al. introduced TIMME, a graph-based multi-relational embedding method (83). Despite progress, prior work often relies on small datasets (5), heuristic or manual annotations, and single data modalities, typically text (18) or graphs (83; 40). Many models also model ideology as binary (left/right) or coarse labels, such as agreement/disagreement (69; 50), which limits their ability to capture nuanced stances. In contrast, our TSN4PI adopts a granular approach using floating ideological scores, integrating textual and graph-based signals to improve prediction. We acknowledge that political ideology is inherently multidimensional, beyond a single left-right axis. For example, Colacrai et al. (16) analyzed Reddit’s political compass, disentangling ideology into economic (left-right) and social (libertarian-authoritarian) axes. Inspired by this, we adopt the widely-studied one-dimensional continuum as a tractable starting point for modeling polarization, while recognizing multidimensional extensions as a promising direction for future work. Our framework further supports expansion to issue-level or temporal perspectives on behavior. Opinion Evolution and Political Dynamics on Social Media involve phenomena such as echo chambers (where users are exposed predominantly to ideologically aligned content) and political polarization, often reinforced by platform algorithms. These dynamics are central to understanding how public opinion forms and shifts in online environments. For example, Chen et al. (15) proposed a graph-clustering method to model the evolution of public opinion topics on social media during Pelosi’s visit to Taiwan, revealing causal transitions between topics across event stages. Other studies focused on long-term communication patterns of political figures. Oliveira et al. (57) analyzed six years of Brazilian politicians’ communications, addressing concept drift by designing classifiers to distinguish political from non-political posts. Their findings showed that communication strategies and public reactions varied around major events (57). While prior work offers insights into topic-level discourse and macro trends, our approach differs in both focus and granularity. Instead of classifying content types or tracking broad topic shifts, we model continuous, user-level ideological scores over time, a task methodologically related to trajectory prediction (77) and user modeling from implicit feedback (76; 79; 78). By combining semantic signals with dynamic interaction patterns, our framework enables fine-grained tracking of individual ideological evolution. Semantic Text Analysis for Political Ideology has advanced rapidly with NLP developments. Early work employed RNNs for partisan classification (39), laying the groundwork for computational ideology detection. The advent of transformer-based models like BERT (21) enabled more precise text representations and improved classification performance. Later variants—RoBERTa (49), SBERT (65), BERTweet (56), and TWHIN-BERT (87)—further enhanced effectiveness on social media data. More recently, LLMs like GPT-3 (9) have enabled zero-/few-shot political text understanding via in-context learning. However, their direct use for large-scale, fine-grained prediction remains limited by inference cost and stability. Our framework adopts a modular design, using an LLM not as an end-to-end predictor but as a style transfer module. This bridges the domain gap between structured political news and noisy social media posts, as detailed in Section 4. Stance Detection is a related task in NLP that involves identifying an expressed position—such as “favor”, “against”, or “neutral”—towards a specific target (54; 45; 71). While both stance detection and political ideology detection analyze opinions, they differ in scope. Stance is target-specific and often transient, reflecting an individual’s position on a particular issue (4; 45). In contrast, political ideology is broader and more stable, representing a cross-topic belief system that shapes multiple stances (5). Our work focuses on identifying and modeling this fundamental ideological attribute over time, rather than classifying positions on specific targets. 3. Data Two pie charts show news-article and media-outlet shares across five political leanings, with Center the largest category in both. A bar chart below ranks the top 20 outlets by article share, with The Hill highest. Figure 2. Distribution of articles (left) and outlets (right) by political leaning, based on data from AllSides. The bar graph shows the top 20 sources by share of total articles, colored by ideology.Two pie charts show news-article and media-outlet shares across five political leanings, with Center the largest category in both. A bar chart below ranks the top 20 outlets by article share, with The Hill highest. We use three datasets in this study. The first, from allsides.com, provides ideological labels for news sources and supports cross-domain adaptation. The second and third contain social media posts from X (formerly Twitter) and Truth Social, respectively. These platforms are used for model training, evaluation, and cross-platform analysis of political discourse and user ideology across distinct online communities. Privacy considerations are detailed in Appendix A.1.1.11 1 All appendices are provided in the supplementary material submitted alongside this article. 3.1. News Political Ideology Dataset We construct a political ideology dataset using media bias ratings from allsides.com (3), which aggregates partisan perspectives across the ideological spectrum.22 2 AllSides Media Bias Ratings™ by AllSides.com, used with permission under C BY-NC 4.0. This dataset supports modeling and transferring ideological signals across domains, based on the assumption that media affiliation reflects consistent leanings. allsides.com applies a multi-partisan methodology, including editorial review, blind bias surveys, third-party analysis, and community feedback. Each news source is assigned a continuous bias score from -6 (most left) to 6 (most right), grouped into five categories: Left ([−6,−3][-6,-3]), Lean Left ((−3,−1](-3,-1]), Center ((−1,1](-1,1]), Lean Right ((1,3](1,3]), and Right ((3,6](3,6]). Using requests and Newspaper3k (58), we retrieved rated outlets and crawled ∼ 230,000 articles from their main pages, covering August 3, 2017 to December 22, 2022. Each article inherits its source’s label, yielding a large-scale, weakly supervised dataset with balanced ideological coverage. Score preprocessing details are in Appendix A.1.2. Preliminary Data Analysis. We conducted a preliminary analysis of the AllSides dataset to examine its ideological distribution and linguistic characteristics. The upper part of Figure 2 shows the proportions of news articles and media outlets across five political categories. The dataset displays a relatively balanced Left-Right distribution, supporting class balance during model training. The lower part of Figure 2 illustrates the share of total articles held by the top 20 news outlets. Among them, “The Hill,” labeled as Center, contributes the largest share. Despite a Right-leaning tilt among the very largest outlets, these 20 sources span all five ideological categories. Additional distribution details and word cloud visualizations appear in Table 1 and Figure 1 of that material. Beyond word cloud visualizations, we conducted a frequency analysis of full article texts to examine ideological differences in lexical usage. Table 1 lists the top 10 most frequent content words per political leaning after stopword removal. These patterns reveal how lexical emphasis varies across ideological contexts. Trump is the most frequent term in both Left (96,727) and Right (60,993) articles, indicating shared focus but likely divergent framing. Biden also appears frequently in Lean Right (45,403) and Right (50,717), reflecting sustained—often critical—attention from right-leaning outlets. An asymmetry emerges with covid, ranked 6th in Lean Left (86,424) but absent elsewhere, suggesting greater pandemic focus among moderately left-leaning sources. Center-leaning content foregrounds governing bodies—state (122,494), house (95,068), senate (82,457)—indicating a governance-oriented tone. Lean Right and Right texts instead emphasize oversight and platform arenas, with congress (34,186), committee (25,647), and twitter (34,956), reflecting political critique and media discourse. These trends support our hypothesis that bias emerges not only in topic selection, but also in lexical framing—even when events are shared. Table 1. Top 10 most frequent content words per political leaning (full text, stopwords removed). Rank Left Lean Left Center Lean Right Right 1 trump (96,727) people (130,164) state (122,494) president (54,261) trump (60,993) 2 people (62,773) president (96,049) trump (121,089) biden (45,403) biden (50,717) 3 election (58,163) state (94,501) president (115,675) house (42,055) house (50,398) 4 republican (47,839) new (92,173) people (109,190) new (41,471) president (47,401) 5 even (47,647) trump (88,297) new (105,062) state (40,075) news (47,398) 6 president (46,673) covid (86,424) house (95,068) congress (34,186) people (40,715) 7 political (43,828) biden (79,017) year (84,825) year (28,292) new (37,368) 8 house (42,453) years (72,749) states (83,646) former (28,068) state (35,696) 9 democrats (41,652) states (72,083) senate (82,457) last (26,605) twitter (34,956) 10 republicans (38,995) election (65,009) election (82,358) committee (25,647) government (33,186) 3.2. A User-Centric Twitter Dataset on U.S. Presidential Elections Table 2. X Dataset Overview. Characteristic Value Number of Users 4,545 Total Tweets 76,718,924 Avg. Tweets/User 16,880 Avg. Retweets/Tweet 20.50 Avg. Likes/Tweet 88.00 Avg. Quotes/Tweet 2.40 (a) Basic Statistics Field Relations Target ID Content ID Retweets RT ID Text Username Quotes QT ID Hashtags Time Likes Reply ID Conv. ID Reply User Mentions (b) Attributes To support research on political ideology and discourse—particularly in the context of U.S. presidential elections—we construct a large-scale dataset from X (formerly Twitter). Using a keyword seed query (“presidential election”), we collected ∼ 10,000 tweets and identified 4,545 unique users. We then retrieved their full tweet histories, yielding a user-centric corpus of 76,718,924 tweets from January 1, 2007 to December 31, 2022 (Table 2(a)). This dataset enables longitudinal analysis of political discourse and user behavior. Table 2(a) summarizes dataset statistics, and Table 2(b) lists tweet-level fields: tweet IDs, user metadata, timestamps, conversation IDs, engagement metrics (retweets, quotes, likes), reference links (e.g., replies, retweets), and full text with hashtags. The dataset’s feature richness supports detailed analysis of user activity, engagement patterns, and content evolution. Preprocessing details appear in Appendix A.1.3. Comparison with Existing Datasets. Table 3 compares recent large-scale social media datasets. While many are valuable, they vary in platform, scale, granularity, and thematic scope. For example, Dreaddit (74) and the Reddit Corpus (36) focus on Reddit, providing post- or dialogue-level data (190k and 727M entries, respectively). Weibo-COV (37) offers 40M COVID-19-related posts from Sina Weibo, but is also post-centric. Koo (52) (72M posts from 1.4M users) and iDRAMA-Scored (60) (57M posts from 226k users) target Indian and fringe-platform discourse, respectively, and are likewise post-level. Even Twitter-based datasets such as TweetIntent@Crisis (2) and News Feed Ranking (8) involve small user bases (e.g., 79 and 46 users) and are not user-centric. In contrast, our dataset offers full tweet histories for a fixed cohort of 4,545 users over 16 years. This supports fine-grained, longitudinal modeling of U.S. presidential election discourse. Unlike post-level snapshots, our dataset retains full threads and engagement, allowing for the analysis of ideological evolution and stance shifts over time. Its scale, temporal span, and user continuity make it uniquely suited for studying political polarization and ideology formation on social media. Table 3. Summary of Recent Social Media Datasets Dataset Name Platform Count Users Duration Data Level Theme Dreaddit (74) Reddit ∼ 190k N/A 2017–2018 Post Social media stress Reddit Corpus (36) Reddit 727M N/A 2015–2018 Dialogue Conversational AI Weibo-COV (37) Sina Weibo >40M 20M 2019–2020 Post COVID-19 discourse News Feed Rank. (8) Twitter 26k 46 2021–2022 News Personalized ranking MetaHate (62) Mixed 1.2M N/A N/A Post Hate speech detect. IsamasRed (14) Reddit 8M N/A 2023 Dialogue Israel-Hamas discourse Koo Dataset (52) Koo 72M 1.4M 2020–2023 Post Indian microblogging TweetIntent Crisis (2) Twitter ∼ 18k 79 2022–2023 Post Russia-Ukraine narratives SocialDrought (70) Twitter&News 1.5M N/A 2012–2023 Integrated Drought societal impact LGBTQ+ MiSSoM (10) Reddit 27.7k N/A 2012–2021 Post LGBTQ+ stress research AiGen-FoodRev. (27) Yelp 20k N/A N/A Post AIGC Detection iDRAMA-Scored (60) Scored 57M 226k 2020–2023 Post Fringe communities Unintended Offense (73) Twitter 2.4k N/A N/A Dialogue Unintentional offense PulseReddit (35) Reddit ∼ 73k N/A 2024–2025 Post MAS HFT Crypto Trad. Ours Twitter ∼ 77M ∼ 5k 2007–2022 User Presidential Election Preliminary Data Analysis. To better understand the dataset’s temporal growth and topical focus, we carried out exploratory analyses of tweet volume over time and hashtag frequency. (a) Number of tweets per month from January 2007 to September 2022. (b) Top 30 hashtags in the Twitter dataset by frequency. Figure 3. Temporal growth and topical focus of the X (Twitter) dataset.Two panels summarize the X dataset. The left line chart shows monthly tweet counts rising from near zero in 2007 to roughly two million by late 2022, with intermediate spikes. The right horizontal bar chart ranks 30 hashtags, led by p2, trump, news, smartnews, and socialmedia. Tweet Volume Over Time. Figure 3(a) shows monthly tweet counts from Jan. 2007 to Sep. 2022. The overall trajectory is upward, with tweet activity growing steadily through late 2022 and accelerating especially rapidly after 2018. Four peaks punctuate this trend: a modest bump during the 2012 U.S. presidential election, a more pronounced surge in 2016, a third spike during the COVID-19 outbreak in early 2020, and the largest peak surrounding the 2020 presidential election, reflecting heightened online political activity. Afterward, tweet volume continues to rise sharply, suggesting a sustained expansion of political engagement on the platform. Top Hashtags. Figure 3(b) lists the 30 most frequent hashtags. News- and media-oriented tags, such as #news, #smartnews, and #socialmedia, appear most prominently, underscoring Twitter’s role as a live information channel. Excluding these politically neutral news tags, partisan motifs—#gop, #maga, and #resist—still command a substantial share of the remaining distribution, confirming the platform’s ideological polarity. International-politics hashtags (#ukraine, #auspol, #cdnpoli) and global-health hashtags (#covid19, #coronavirus) are likewise present, highlighting the dataset’s geographic breadth and its coverage of worldwide events. 3.3. Truth Social Dataset We also incorporated a dataset from Truth Social (28), launched in February 2022 following the January 2021 suspension of then-U.S. President Donald Trump from multiple platforms, including Twitter (now known as X), Facebook, and others. From Figure 1 (a) and (b), Truth Social was designed to closely mirror X in terms of functionality, with almost identical features. For instance, X’s “Retweet” functionality is analogous to “ReTruth” on Truth Social. Supporters of Trump predominantly frequent this platform and generally exhibit a right-leaning political bias. With this dataset, we aim to explore the dissemination of information and track the evolution of user political ideologies within a social network that inherently carries a political bias from the outset. Table 4. Truth Social Dataset Overview Field # of Records Users 454,458 Truths 823,927 Quotes 10,508 Replies 506,276 Structurally, Truth Social bears similarities to Twitter. In contrast to Twitter’s “Tweets,” posts on this platform are referred to as “Truths,” and reposts are termed “ReTruths.” The dataset comprises over 454,000 users and exceeds 823,000 “Truths.” Additionally, it includes the complete post history of the 65,536 most active users. Subtables within this dataset, such as “Truths” and “Users,” serve as primary resources for our exploration. An overview of this dataset is presented in Table 4. Details of data preprocessing are provided in Appendix A.1.4. 4. Methods This section introduces our proposed framework and consists of both PIDN and PIPN modules. A two-stage pipeline. The upper PIDN transforms news into social-media style, combines it with daily-life and unlabeled social-media posts, encodes them with a BERT-based language model, and optimizes classification, regression, and domain-alignment losses. The lower PIPN processes temporal social graphs with a temporal graph neural network for link prediction and node classification and regression. Figure 4. The overview of our proposed framework: TSN4PI. The upper part of the figure represents the PIDN module, while the lower part represents the PIPN module.A two-stage pipeline. The upper PIDN transforms news into social-media style, combines it with daily-life and unlabeled social-media posts, encodes them with a BERT-based language model, and optimizes classification, regression, and domain-alignment losses. The lower PIPN processes temporal social graphs with a temporal graph neural network for link prediction and node classification and regression. 4.1. Political Ideology Detection Network Political ideology detection involves analyzing social media posts to determine 1) whether they are politically relevant and, if so, 2) to what extent they align along the political spectrum. Rather than assigning scores to all inputs indiscriminately, the task first filters out non-political content and then estimates an ideology score for posts identified as political. Our proposed network addresses these dual challenges through a combination of unsupervised domain adaptation (UDA), text style transfer (TST) using LLMs, and an efficient model architecture. While LLMs show strong ICL capabilities (9; 22), their direct use in large-scale tasks like ideology detection faces high inference latency. Adding in-context examples lengthens the input and increases attention computation (88; 63; 82). LLMs can also underperform fine-tuned models on well-defined, high-resource tasks (61), limiting their suitability for real-time or high-throughput use. We therefore adopt a BERT-based architecture, which offers a stronger trade-off between task performance and efficiency—especially in inference speed and resource usage (68). LLMs are reserved for the offline TST stage, where latency is less critical, rather than the core detection pipeline. Figure 5. Embedding across PLMs.Four UMAP scatter plots compare AllSides news, social-media posts, and style-transferred news under BERT, XLM-R, Longformer, and TWHIN-BERT embeddings. Across models, the style-transferred news shifts toward the social-media cluster. Building on this architectural decision, our detection pipeline integrates offline LLM-based style adaptation with an online BERT-based classifier. Specifically, given the substantial stylistic differences between formal news articles and informal social media posts, we apply a TST module to transform news content into a style resembling that of platforms such as X. This stylistic alignment narrows the domain gap between the labeled source domain (news) and the unlabeled target domain (social media), thereby improving transferability. To further bridge this gap, we incorporate UDA to align the latent representations of political news and unlabeled social media posts. We also augment training with a mixture of political news articles and non-political social media content, allowing the model to jointly learn 1) political relevance classification and 2) ideology scoring for relevant posts. This unified approach enables robust detection in noisy, real-world social media environments, while leveraging structured ideological signals from traditional media. LLMs for Text Style Transfer. To bridge the stylistic gap between formal news articles and the colloquial nature of social media posts, we use LLMs to perform TST. This approach is especially suitable given the absence of parallel corpora and the strong instruction-following capabilities of modern LLMs. Specifically, for each news article n in the source corpus N, we generate a style-aligned version nsn_s by prompting the LLM with: ns=LLM([prompt,n]),n∈N.n_s=LLM([prompt,n]),n∈ N. To intuitively validate our approach, we sampled texts from three groups: the source domain (AllSides news), the target domain (Truth Social posts), and our generated TST outputs. These were encoded using four diverse BERT-based PLMs: bert-base-cased (21), xlm-roberta-base (49), longformer-base-4096 (6), and twhin-bert-base (87)—and projected to 2D via UMAP (67). As shown in Figure 5, original news articles (blue) are widely scattered, while social media posts (orange) form a denser cluster. Crucially, style-transferred texts (green) align closely with the target domain, reducing the distributional gap. This pattern holds across all four PLMs, supporting the effectiveness and robustness of our TST strategy. A quantitative and comprehensive automatic evaluation follows in Section 5.2.2. Unsupervised Domain Adaptation for Distribution Alignment. While LLM-based TST narrows the stylistic gap between news and social media texts, residual distributional shifts remain. To address this, particularly without requiring labeled social media data, we incorporate an optional UDA module designed to enhance domain generalization. Our primary approach employs Maximum Mean Discrepancy (MMD), a non-parametric measure that quantifies the divergence between two distributions in a reproducing kernel Hilbert space (RKHS), without requiring paired or balanced samples. Given a text sample T, we extract its embedding E=L(T)E=L(T) using a PLM L. Let EXE_X and EYE_Y denote embeddings from the style-transferred source and target domains, respectively. The MMD loss is computed as: (1) ℒMMD(EX,EY)=[k(x,x′)]+[k(y,y′)]−2[k(x,y)],L_ MMD(E_X,E_Y)= E[k(x,x )]+E[k(y,y )]-2E[k(x,y)], where k(x,y)=exp(−∥x−y∥2/(2σ2))k(x,y)= (-\|x-y\|^2/ (2σ^2 ) ) is the Gaussian RBF kernel, and x,x′∈EXx,x ∈ E_X, y,y′∈EYy,y ∈ E_Y. This kernel maps samples into a space where their similarity can be effectively measured and optimized. In addition to MMD, we also consider Kullback-Leibler (KL) divergence as an alternative distribution alignment method. KL divergence is particularly effective when model outputs can be interpreted as probability distributions—such as softmax scores or normalized embeddings—allowing explicit alignment of posterior distributions between domains. Given source domain prediction P and target domain prediction Q, the KL loss is computed as: (2) ℒKL(P∥Q)=∑iP(i)logP(i)Q(i).L_ KL(P\,\|\,Q)= _iP(i) P(i)Q(i). Compared to MMD, KL divergence assumes stronger alignment in probability space and is more suitable when output distributions are well-calibrated. We evaluate both metrics to assess robustness under varying domain conditions. To incorporate domain adaptation during training, we define a composite loss that jointly optimizes task performance and representation alignment: (3) ℒPIDN=αℒc+βℒr+(1−α−β)ℒUDA,L_PIDN= _c+ _r+(1-α-β)L_UDA, where ℒcL_c and ℒrL_r are the political-relevance classification and ideology-regression losses, respectively; ℒUDA∈ℒMMD,ℒKLL_UDA∈\L_ MMD,L_ KL\ is the selected domain-alignment loss; and α,β≥0α,β≥ 0 with α+β≤1α+β≤ 1 control their relative contributions. When UDA is disabled in an ablation, the ℒUDAL_UDA term is omitted. By aligning embedding distributions across domains, the UDA module enhances generalization to informal, unlabeled social media content while preserving task-specific learning. A comprehensive complexity analysis of the UDA module is provided in Appendix A.4. 4.2. Political Ideology Prediction Network Political ideology prediction models how users’ stances evolve over time—a task fundamentally different from static detection due to the absence of future posts. Traditional post-level embedding methods are thus inapplicable. We adopt Temporal Graph Neural Networks (TGNNs), which are well-suited for ideological forecasting. This choice is motivated by two factors: social media platforms naturally produce dynamic user interaction graphs, and TGNNs are designed to model temporal dependencies in such evolving structures. Temporal Graph Construction. We construct a dynamic, directed, and weighted graph G=(V,E,t)G=(V,E,t), where V denotes users, E represents interactions, and t is the timestamp associated with each edge. Each post is modeled as a temporal edge whose features are derived from PLM-based text embeddings. If user A re-posts content from user B, we introduce a directed edge from A to B, weighted by the frequency of such reposts within a defined time window. Self-loop edges from A to A represent original posts, capturing the user’s individual content generation. This construction encodes both interpersonal influence and independent ideological expression. To capture temporal evolution, we partition the graph chronologically into a sequence of snapshots G1,G2,…,Gt−1G_1,G_2,…,G_t-1, and train models to predict ideological properties in the future snapshot GtG_t. This forward-looking setup mirrors real-world forecasting scenarios, where a user’s historical activity (e.g., during years 1-5) informs predictions about their future ideological stance (e.g., in year 6). Temporal Graph Modeling. To capture spatio-temporal dynamics, we evaluate three TGNN models: JODIE, APAN, and TGN. JODIE (46) uses coupled RNNs to update user and item embeddings from interactions. A projection operator estimates future user embeddings via time-conditioned scaling: (4) ^(t+Δt)=(1+Δt)⊙(t), u(t+ t)=(1+w_ t) (t), where Δtw_ t encodes the elapsed-time context. APAN (81) uses asynchronous updates with an attention-based encoder. Upon interaction, it updates the embedding based on the last state and a “mailbox” storing recent neighbor interactions: (5) (t)=MLP(LayerNorm(MultiHead((t−),(t))+(t−))).z(t)=MLP(LayerNorm(MultiHead(z(t^-),M(t))+z(t^-))). TGN (66) maintains time-evolving memory states. Upon interaction between nodes i and j, it generates a message: (6) i(t)=msg(i(t−),j(t−),Δt,ij(t)),m_i(t)=msg(s_i(t^-),s_j(t^-), t,e_ij(t)), which updates memory via: i(t)=mem(i(t),i(t−))s_i(t)=mem(m_i(t),s_i(t^-)) . Final node embeddings i(t)z_i(t) are computed using temporal attention over neighbors. Self-Supervised Link Prediction. As a self-supervised representation-learning stage, temporal link prediction first trains the TGNN to learn node representations by forecasting future edges E^′ E from historical graph sequences. For a candidate node pair (u,v)(u,v), the link score is computed as: (7) s(u,v)=σ(u⊤v),s(u,v)=σ (h_u Wh_v ), where uh_u and vh_v are time-specific embeddings, W is a learnable projection matrix, and σ(⋅)σ(·) denotes the sigmoid function. This stage produces temporally contextualized node embeddings nth_n^t, which are then used by the downstream ideology-prediction heads. Node Classification and Regression. To predict users’ future political ideologies, we attach classification and regression heads to temporal node embeddings nth_n^t learned via TGNN. PIDN-derived targets are aggregated from each user’s historical posts up to time t. Let Pn,tP_n,t denote this post set, and let r^p r_p and s^p s_p denote the PIDN-derived relevance label and ideology score for post p, respectively. Political relevance R(n)tR(n)_t is defined as the majority post-level relevance label: (8) R(n)t=moder^p:p∈Pn,t,R(n)_t=mode\ r_p:p∈ P_n,t\, and, for relevant users (R(n)t=1R(n)_t=1), let Pn,t+=p∈Pn,t:r^p=1P^+_n,t=\p∈ P_n,t: r_p=1\. Their ideology target is the mean PIDN score over relevant posts: (9) S(n)t=1|Pn,t+|∑p∈Pn,t+s^p.S(n)_t= 1|P^+_n,t| _p∈ P^+_n,t s_p. To jointly optimize both tasks, we define a unified loss: (10) ℒPIPN=λℒc+(1−λ)ℒr,L_PIPN= _c+(1-λ)L_r, where λ∈[0,1]λ∈[0,1] balances classification and regression. This multi-task setup guides the model to learn temporally grounded node representations that capture ideological relevance and intensity. 5. Experiments This section introduces the experiments conducted to validate the framework’s efficacy in detecting and predicting political ideologies. We first outline the experimental setup, including datasets, baselines, and implementation details. We then present a detailed evaluation of the PIDN, followed by an analysis of the PIPN. Implementation details are provided in Appendix A.2. 5.1. Experimental Setup 5.1.1. Datasets and Data Preparation Training Data Construction for PIDN The PIDN module is trained on a hybrid dataset of politically relevant and irrelevant content. Relevant samples are derived from style-transferred news articles with annotated ideology scores and are labeled according to their political relevance. To balance the data, we include an equal number of daily-life tweets from our Twitter corpus, labeled as politically irrelevant (relevance 0), with a placeholder ideology score that is excluded from the regression objective. To enhance domain adaptation, we add non-political tweets from the MajidTweets dataset (51), a large-scale corpus of real-world tweets. These are selected to be disjoint from our Twitter corpus, introducing broader linguistic and stylistic diversity. Importantly, the PIDN regression head is only activated for posts classified as politically relevant, ensuring that ideology scores are not assigned to irrelevant content. We also normalized target scores before training and evaluation. Graph Construction for PIPN The experiments for the PIPN module are conducted on two datasets: Truth Social and our large-scale X (Twitter) corpus. The temporal graphs for TGNNs were constructed using posts identified as “Politically Relevant” and subsequently scored by our PIDN-KL model. For the static GNN baselines, we aggregate all temporal interactions into a single, weighted static graph for each dataset. This setup intentionally discards temporal information, allowing us to directly quantify the performance gains attributable to modeling temporal dynamics. 5.1.2. Models, Baselines, and Evaluation Metrics Political Ideology Detection (PIDN) Our evaluation involves: 1) PLM backbone selection, 2) hyperparameter tuning, and 3) ablation of TST and UDA components. We benchmark PIDN against the Qwen2.5 model family (72), spanning 0.5B to 72B parameters, in both zero-shot and few-shot ICL settings. Our analysis encompasses the full suite of Qwen2.5 models under a zero-shot setting, as well as Qwen2.5-7B with varying numbers of ICL shots, given its widespread use and strong representativeness. All ICL experiments are conducted on 1,000 randomly sampled test instances, repeated five times to ensure robustness. Evaluation metrics include classification performance—reported as F1 scores for both political/non-political classes and macro-average—as well as class-specific precision and recall for the political class. Regression performance is assessed via Mean Absolute Error (MAE) and Root Mean Squared Error (RMSE), with results summarized in Table 5. Inference latency (in milliseconds) is reported in Figure 6. A comprehensive summary of all performance metrics across models appears in Tables 6 and 7 of that material. To evaluate the TST module, we adopt two forms of automated assessment. First, we compute semantic similarity between original and style-transferred texts using cosine similarity over sentence embeddings. Second, we apply an LLM-as-a-Judge framework to score outputs based on style appropriateness and semantic preservation. All scores are normalized to the [0,1][0,1] range for consistency. Political Ideology Prediction (PIPN) For the PIPN module, we employ TGNNs as the backbone, which are specifically designed to capture the temporal evolution of interactions within dynamic networks. Given that our datasets model user interactions over time, TGNNs are inherently more suitable for our research objectives. The temporal pipeline is evaluated in two successive stages: 1) self-supervised link prediction, which forecasts future user interactions to learn temporal user embeddings; and 2) downstream political ideology prediction, which combines node classification (determining whether a user’s posts are political) and node regression (quantifying the user’s ideological score). We conducted a comparative study of three representative TGNNs: TGN (66), JODIE (46), and APAN (81). We also establish two strong static GNN baselines, GCN (44) and GAT (75), to validate the necessity of a temporal approach. Performance is measured using average precision (AP) and area under the curve (AUC) for link prediction, and accuracy and MSE for political ideology prediction. The static baselines (GCN, GAT) are trained end-to-end directly on the ideology prediction task, as they do not inherently perform temporal link prediction. 5.2. Evaluation of the Political Ideology Detection Network (PIDN) This section presents an evaluation of PIDN, focusing on its performance and efficiency compared to LLMs, followed by a detailed component analysis and an ablation study. Table 5. Unified performance analysis of the Qwen2.5 model family. Top: cross-model scaling analysis in a zero-shot setting. Bottom: few-shot analysis for the Qwen2.5-7B-Instruct model. Within each block, bold marks the best and underlining the second best per metric. Classification (F1 ↑ ) Political Class Detail ↑ Regression Error ↓ Model / Setting Non-political Political Macro Avg. Precision Recall MAE RMSE Cross-Model Scaling (0-shot) Qwen2.5-0.5B-Instruct 0.510±\,±\,0.022 0.727±\,±\,0.007 0.618±\,±\,0.014 0.601±\,±\,0.007 0.918±\,±\,0.006 0.413±\,±\,0.007 0.478±\,±\,0.005 Qwen2.5-3B-Instruct 0.946±\,±\,0.003 0.943±\,±\,0.003 0.945±\,±\,0.003 0.945±\,±\,0.008 0.942±\,±\,0.006 0.283±\,±\,0.005 0.333±\,±\,0.005 Qwen2.5-7B-Instruct 0.927±\,±\,0.002 0.905±\,±\,0.003 0.916±\,±\,0.002 0.992±\,±\,0.003 0.833±\,±\,0.004 0.238±\,±\,0.004 0.288±\,±\,0.005 Qwen2.5-14B-Instruct 0.918±\,±\,0.003 0.900±\,±\,0.004 0.909±\,±\,0.003 0.986±\,±\,0.001 0.827±\,±\,0.006 0.227±\,±\,0.004 0.285±\,±\,0.003 Qwen2.5-32B-Instruct 0.946±\,±\,0.001 0.937±\,±\,0.002 0.942±\,±\,0.002 0.991±\,±\,0.001 0.889±\,±\,0.004 0.225±\,±\,0.002 0.283±\,±\,0.001 Qwen2.5-72B-Instruct 0.955±\,±\,0.002 0.950±\,±\,0.003 0.953±\,±\,0.002 0.994±\,±\,0.002 0.910±\,±\,0.004 0.216±\,±\,0.004 0.275±\,±\,0.004 Few-Shot for Qwen2.5-7B-Instruct 0-shot 0.927±\,±\,0.002 0.905±\,±\,0.003 0.916±\,±\,0.002 0.992±\,±\,0.003 0.833±\,±\,0.004 0.238±\,±\,0.004 0.288±\,±\,0.005 2-shot 0.936±\,±\,0.021 0.928±\,±\,0.026 0.932±\,±\,0.024 0.986±\,±\,0.005 0.879±\,±\,0.049 0.242±\,±\,0.005 0.292±\,±\,0.007 4-shot 0.920±\,±\,0.032 0.905±\,±\,0.044 0.912±\,±\,0.038 0.986±\,±\,0.006 0.840±\,±\,0.079 0.230±\,±\,0.007 0.282±\,±\,0.012 6-shot 0.918±\,±\,0.024 0.903±\,±\,0.035 0.911±\,±\,0.030 0.987±\,±\,0.008 0.835±\,±\,0.063 0.243±\,±\,0.007 0.301±\,±\,0.007 8-shot 0.906±\,±\,0.024 0.886±\,±\,0.035 0.896±\,±\,0.029 0.990±\,±\,0.007 0.804±\,±\,0.061 0.246±\,±\,0.007 0.305±\,±\,0.011 10-shot 0.898±\,±\,0.019 0.874±\,±\,0.028 0.886±\,±\,0.024 0.995±\,±\,0.002 0.781±\,±\,0.047 0.236±\,±\,0.012 0.291±\,±\,0.015 20-shot 0.923±\,±\,0.031 0.908±\,±\,0.043 0.916±\,±\,0.037 0.995±\,±\,0.002 0.838±\,±\,0.074 0.244±\,±\,0.016 0.306±\,±\,0.019 50-shot 0.961±\,±\,0.014 0.958±\,±\,0.015 0.960±\,±\,0.014 0.996±\,±\,0.003 0.924±\,±\,0.031 0.241±\,±\,0.018 0.304±\,±\,0.020 100-shot 0.964±\,±\,0.008 0.962±\,±\,0.009 0.963±\,±\,0.009 0.998±\,±\,0.000 0.929±\,±\,0.017 0.237±\,±\,0.010 0.307±\,±\,0.012 200-shot 0.974±\,±\,0.008 0.973±\,±\,0.008 0.973±\,±\,0.008 0.998±\,±\,0.002 0.950±\,±\,0.017 0.243±\,±\,0.012 0.315±\,±\,0.018 Figure 6. Inference latency comparison. Left: 0-shot latency of PIDN (≈ 0.3B) vs. LLMs of varying sizes. Right: Latency of PIDN vs. Qwen2.5-7B with increasing ICL shots (0–10). Latencies are reported in milliseconds (ms).The left bar chart shows PIDN latency at 4.39 milliseconds versus Qwen2.5 models ranging from 92.90 milliseconds for 0.5B to 1145.71 milliseconds for 72B. The right plot shows Qwen2.5-7B latency remaining around 341 to 342 milliseconds across zero to ten shots, while PIDN remains at 4.39 milliseconds. 5.2.1. Performance and Efficiency vs. Large Language Models Performance Analysis. To comprehensively evaluate the effectiveness of our model, we first compared our fine-tuned PIDN against the Qwen2.5 model family (Table 5). For few-shot ICL, each setting was rigorously repeated five times using randomly drawn, class-balanced examples from a held-out pool. Our primary benchmark, PIDN, achieved a near-perfect classification accuracy of 99.25% and an impressively low regression MSE of 0.0048—an RMSE of 0.069 (Table 7), demonstrating high precision and fine-grained scoring capability. Cross-model scaling analysis reveals a generally positive, albeit complex and nuanced, trend in zero-shot LLM performance. The largest model, Qwen2.5-72B-Instruct, achieves the highest Macro F1-score of 0.953, establishing it as the top performer. However, the trend is not strictly monotonic, as mid-sized models (7B to 32B) exhibit a strong bias towards precision over recall. They reach near-perfect precision (up to 0.992) but at the cost of lower recall (down to 0.827), reflecting conservative decision boundaries that minimize false positives. For regression, scaling yields more consistent gains; yet, the best LLM’s RMSE of 0.275 is still nearly four times worse than PIDN ’s 0.069, underscoring the limits of prompting without task-specific tuning. The few-shot performance of Qwen2.5-7B-Instruct further illustrates these stark limitations. We observe a surprising “performance trough”: after an initial gain at 2 shots, adding more demonstrations (specifically, 4–10 shots) unexpectedly degrades classification performance, with the F1-score dropping below the zero-shot baseline of 0.916. A consistent improvement over the baseline requires 50 shots or more, which partially corrects the earlier conservative bias by improving recall. Even at its best—a Macro F1 of 0.973 at 200 shots, and a lowest RMSE of 0.282 at 4 shots—ICL still falls well short of our fine-tuned PIDN. This significant gap stems from core methodological differences, as PIDN is explicitly fine-tuned on a large, domain-specific dataset with ideological supervision. This allows it to learn nuanced signals that generic few-shot prompting cannot capture, highlighting the clear superiority of task-specific tuning for precise sociopolitical inference. Latency Analysis. To complement the performance evaluation, we conducted a comparative latency analysis to validate the efficiency of our architectural design. We benchmarked our PIDN, built on the twhin-bert-base backbone (≈ 0.3B parameters), against the same Qwen2.5 LLM baselines. Latency is a known limitation of LLMs, especially with larger models or in-context examples. Figure 6 presents the latency comparison. Our PIDN achieves a low inference latency of 4.39 ms. In the left panel, even the smallest Qwen2.5 model (0.5B) exhibits a 0-shot latency of 92.90 ms—over 21× slower than PIDN. Latency increases sharply with model size, reaching 1145.71 ms for the 72B variant, which is more than 260 times slower. The right panel shows the impact of increasing ICL shots on the 7B model. While latency rises only slightly from 340.73 ms (0-shot) to 342.29 ms (10-shot), it remains consistently high—about 78× slower than PIDN, even under minimal ICL. Discussion on Model Architecture. Our evaluation challenges the common assumption that there is a trade-off between the efficiency of task-specific models and the performance of LLMs. In high-resource settings, our fine-tuned PIDN significantly outperforms LLMs—delivering higher accuracy with orders of magnitude lower inference latency. Even the best ICL setups incur prohibitive computational costs, making LLMs unsuitable for real-time or large-scale deployment. While PIDN is the clear choice when labeled data is abundant and performance critical, LLMs remain useful in low-resource scenarios where annotation is infeasible. Their zero- and few-shot capabilities offer a practical fallback. Nonetheless, our results strongly support the development of lightweight, specialized models in settings where scalability, precision, and efficiency are most critical. These findings support a broader insight: when conditions allow, architectural specialization often outperforms general-purpose flexibility in real-world analytical tasks. 5.2.2. Component Analysis and Ablation Studies Table 6. Comparison of PLM backbones and hyperparameter tuning results. Best values are in bold. (a) Backbone Comparison Model Acc. (Kind) (%) ↑ MSE (Score) ↓ bert-base-cased (21) 95.56±\,±\,0.21 0.0847±\,±\,0.0015 xlm-roberta-base (17) 96.67±\,±\,0.15 0.0686±\,±\,0.0011 covid-twitter-bert-v2 (55) 95.53±\,±\,0.25 0.0874±\,±\,0.0018 twhin-bert-base (87) 97.79±\,±\,0.12 0.0606±\,±\,0.0008 (b) Hyperparameter Tuning Hyperparameters Performance Metrics α β Acc. (%) ↑ MSE ↓ 0.4 0.4 97.12±\,±\,0.18 0.073±\,±\,0.002 0.3 0.3 96.98±\,±\,0.20 0.071±\,±\,0.002 0.4 0.3 97.79±\,±\,0.15 0.067±\,±\,0.001 0.3 0.4 97.52±\,±\,0.17 0.069±\,±\,0.002 Table 7. Ablation study of TST and UDA methods. All experiments use the twhin-bert-base backbone and the selected hyperparameter setting (α=0.4,β=0.3α=0.4,β=0.3). Best performance is in bold. Configuration Acc. (Kind) (%) ↑ MSE (Score) ↓ Backbone Only 97.79±\,±\,0.15 0.0670±\,±\,0.0012 Backbone + UDA (KL) 98.15±\,±\,0.11 0.0350±\,±\,0.0009 Backbone + UDA (MMD) 98.35±\,±\,0.10 0.0320±\,±\,0.0010 Backbone + TST 99.12±\,±\,0.08 0.0060±\,±\,0.0004 Backbone + TST + UDA (KL) (PIDN-KL) 99.25±\,±\,0.07 0.0048±\,±\,0.0003 Backbone + TST + UDA (MMD) (PIDN-MMD) 99.52±\,±\,0.05 0.0057±\,±\,0.0004 Backbone Selection. We began by evaluating four prominent PLM backbones to identify the most suitable encoder for our task. To isolate the backbone’s inherent capacity, this comparison excluded both TST and UDA. We used a fixed hyperparameter setting (α=0.4,β=0.2α=0.4,β=0.2) for all models, with standard deviations from 5 runs. As shown in Table 6(a), twhin-bert-base achieved the best results for both political relevance classification (97.79% accuracy) and ideology score regression (0.0606 MSE). This performance underscores the benefit of pre-training on large-scale social media data with user- and network-aware objectives. In contrast, although covid-twitter-bert-v2 is also social media-oriented, its domain specificity toward COVID-19 limited its effectiveness in capturing the broader political discourse. Interestingly, all models yield comparably high classification accuracy, suggesting that identifying whether a tweet expresses political relevance is a relatively easy task for PLMs. In other words, classifying the presence of political inclination poses little challenge to modern language models, with minimal variation in performance across different backbones. In conclusion, we adopted twhin-bert-base as the backbone for all subsequent experiments. Hyperparameter Tuning. Next, we fine-tuned the hyperparameters α and β, which weight the three terms of the PIDN objective αℒc+βℒr+(1−α−β)ℒUDA _c+ _r+(1-α-β)L_UDA (classification, regression, and domain alignment; see Figure 4), using our selected twhin-bert-base backbone, again without engaging TST or UDA. Table 6(b) reports the four configurations evaluated in this dedicated tuning stage. Among these settings, α=0.4α=0.4 and β=0.3β=0.3 achieved the highest classification accuracy (97.79%) and the lowest MSE (0.067), and we therefore selected this setting for the subsequent ablation study. The (α=0.4,β=0.2)(α=0.4,β=0.2) setting in Table 6(a) was used only to hold the loss weights fixed during the preliminary backbone comparison and was not part of this tuning grid. Accordingly, we do not claim a global optimum across the two stages. The limited variation within the Table 6(b) grid suggests that the framework is not highly sensitive to these loss weights. Ablation Study. To rigorously evaluate the individual and synergistic contributions of our core components, we conducted a detailed ablation study. Starting with the twhin-bert-base backbone and the selected hyperparameter setting (α=0.4,β=0.3α=0.4,β=0.3), we incrementally added TST and the two UDA methods: Kullback-Leibler (KL) divergence and Maximum Mean Discrepancy (MMD). The comprehensive results are presented in Table 7. The most dramatic improvement stems from TST. By adding TST to the backbone, the regression MSE plummets from 0.0670 to 0.0060—an order-of-magnitude reduction. This starkly demonstrates that bridging the stylistic gap via LLM-based data augmentation is the single most critical factor for achieving high-fidelity ideological scoring. UDA also yields significant gains. When applied directly to the backbone, MMD reduces the MSE by over 50% to 0.032, while KL achieves a comparable reduction to 0.035. When combined with TST, both full models deliver the best results among the evaluated configurations. The PIDN-MMD model achieves the highest classification accuracy at 99.52%, while the PIDN-KL model secures a marginally better regression performance with the lowest MSE of 0.0048. Given that both models demonstrate highly competitive and nearly equivalent performance, the deciding factor shifts to a practical consideration: computational efficiency. As detailed in Appendix A.4, there is a stark contrast in the overhead of the two UDA methods. MMD’s quadratic time complexity (O(N2D)O(N^2D)) makes it compute-bound, limiting its scalability with increasing batch size. In contrast, KL divergence’s linear complexity (O(ND)O(ND)) makes it significantly more efficient and scalable, posing no practical constraints on large-batch training. Combining top-tier performance with superior computational efficiency, we therefore identify PIDN-KL as our selected configuration. Figure 7. Evaluation of Text Style Transfer. Top: Histograms of similarity scores between original news and tweet-style outputs, computed using four sentence embedding models. Bottom: Histograms of Style Appropriateness and Semantic Preservation scores (normalized to 0–1) assigned by LLM judges.Eight histograms evaluate text style transfer. The top row shows news-to-tweet similarity distributions from four embedding models, concentrated mostly between 0.6 and 0.9. The bottom row shows style and semantic scores from DeepSeek and OpenAI judges, with most probability mass at higher scores. Validation of the TST Module. To rigorously assess the contribution and quality of our LLM-based TST module, we conducted a multifaceted evaluation. The goal is to verify that TST effectively adapts formal news into tweet-style social media text while preserving semantic meaning, which is crucial for downstream ideology analysis. The evaluation comprises two complementary components: Semantic Similarity Analysis and LLM-as-a-Judge Evaluation. • For Semantic Similarity Analysis, to ensure robustness and reduce model bias, we employed four diverse embedding models: bge-m3 (13), Linq-Embed-Mistral (43), multilingual-e5-large-instruct (80), and all-MiniLM-L6-v233 3 https://huggingface.co/sentence-transformers/all-MiniLM-L6-v2. These span multiple architectures (e.g., MiniLM, XLM-R, Mistral, BGE) and sizes (22M–7.1B), providing an architecture-agnostic similarity estimate. High scores across models suggest strong semantic preservation. • For the LLM-as-a-Judge Evaluation, we employed DeepSeek-V3 (20) and GPT-4o (38) as automated evaluators (89; 32) to score outputs on Style Appropriateness and Semantic Preservation, complementing embedding-based metrics. Figure 7 summarizes the TST evaluation results. The top row shows cosine similarity distributions across four embedding models. bge-m3 and all-MiniLM-L6-v2 concentrate in the 0.6–0.9 range, indicating solid semantic alignment. Linq-Embed-Mistral yields a similar pattern, slightly skewed lower. Notably, multilingual-e5-instruct peaks near 0.9, suggesting strong semantic consistency. Overall, the results indicate good to excellent semantic preservation. The bottom row shows LLM-based scores. For Style Appropriateness, both DeepSeek-V3 and GPT-4o peak around 0.9, confirming alignment with tweet-style norms. For Semantic Preservation, DeepSeek-V3 yields scores tightly clustered near 0.9–1.0; GPT-4o peaks more broadly around 0.7–0.8, still indicating strong fidelity. In summary, both embedding- and LLM-based evaluations confirm that our TST module produces tweet-style outputs with preserved meaning. This validates the PIDN pipeline and ensures downstream tasks operate on stylistically consistent, semantically faithful inputs. Bias Analysis and Mitigation Although the TST module is effective, LLM-based style transfer may still bias the data if political cues are shifted unevenly across ideology groups or platforms (26). We focus on three risks: ideology drift, where stance-bearing markers are amplified or softened more for one side; lexical homogenization, where frequent social-media tokens replace minority or platform-specific expressions and reduce linguistic variety for TGNNs; and source leakage, where platform–ideology correlations from pretraining are copied into transferred texts. We mitigate these effects in three ways. First, we constrain the TST prompt to retain stance phrases and entity mentions, and to modify only surface style (e.g., emojis, contractions, tweet formatting). Second, we perform post-hoc filtering and balancing: pairs with good style but noticeably lower semantic scores from the LLM judges are discarded or regenerated, and PIPN batches are balanced across labels and platforms to prevent artefacts from concentrating in one class. This approach aligns with methods for mitigating task-specific biases (7). Third, the evaluation setup itself acts as a safeguard: using four sentence embedding models and two independent LLM-as-a-judge evaluators (DeepSeek-V3 and GPT-4o) lowers reliance on any single model, and their agreement indicates that the TST step preserves meaning without systematic distortion. 5.3. Evaluation of the Political Ideology Prediction Network (PIPN) This section evaluates the PIPN module by comparing the evaluated TGNNs with static GNN baselines on link prediction and political ideology prediction. Model Comparison and Analysis Table 8 presents the results of our comparative analysis. Findings reveal distinct strengths across models, establishing the advantage of temporal modeling. On the Truth Social dataset, the results reveal a performance trade-off among TGNNs. JODIE achieves the best link-prediction performance among the evaluated models (99.73% AP and 99.69% AUC). However, TGN demonstrates significantly better performance on the critical task of ideology classification, achieving the highest accuracy (72.42%), and it ties with APAN for the lowest regression MSE (0.036). In stark contrast, the static baselines underperform substantially. GAT, the stronger of the two, achieves only 61.81% accuracy, lagging behind TGN by over 10 percentage points. This considerable performance gap underscores the need for temporal modeling to capture the evolving patterns of ideological expression on this platform. On the X (Twitter) dataset, TGN consistently outperforms all other models. Here, the advantage of temporal modeling becomes even more pronounced. The static GCN and GAT models reach accuracies of 96.85% and 95.89%, respectively—substantially lower than those of the TGNNs (all exceeding 99.9%). Moreover, the regression performance of the static models is inferior, with MSEs of 0.0793 (GCN) and 0.0629 (GAT), both considerably higher than TGN’s top-performing 0.0560. We attribute the performance gap to the dataset’s extensive 16-year span and high interaction density, which allows temporal models to learn evolving ideological trajectories that are inherently lost in the time-aggregated graphs used by static models. This demonstrates that omitting temporal information results in a loss of classification fidelity even with strong topical signals. Table 8. Comparative results of static (✗) and temporal (✓) models. The evaluation covers Link Prediction (AP, AUC) and Ideology Prediction (Accuracy, MSE). Best performance for each metric is in bold; second best is underlined. Static models are not applicable (N/A) for the link prediction task. Dataset Model Temporal Link Prediction Ideology Prediction AP(%) ↑ AUC(%) ↑ Acc.(%) ↑ MSE ↓ Truth Social GCN (44) ✗ N/A N/A 58.35±\,±\,0.85 0.045±\,±\,0.002 GAT (75) ✗ N/A N/A 61.81±\,±\,0.72 0.043±\,±\,0.002 TGN (66) ✓ 99.25±\,±\,0.08 99.14±\,±\,0.09 72.42±\,±\,0.35 0.036±\,±\,0.001 JODIE (46) ✓ 99.73±\,±\,0.04 99.69±\,±\,0.05 66.28±\,±\,0.41 0.039±\,±\,0.001 APAN (81) ✓ 97.56±\,±\,0.15 97.24±\,±\,0.18 65.14±\,±\,0.45 0.036±\,±\,0.001 X (Twitter) GCN ✗ N/A N/A 96.85±\,±\,0.22 0.0793±\,±\,0.0015 GAT ✗ N/A N/A 95.89±\,±\,0.28 0.0629±\,±\,0.0011 TGN ✓ 98.82±\,±\,0.05 98.70±\,±\,0.06 99.99±\,±\,0.01 0.0560±\,±\,0.0002 JODIE ✓ 97.83±\,±\,0.09 97.72±\,±\,0.10 99.98±\,±\,0.01 0.0564±\,±\,0.0003 APAN ✓ 98.78±\,±\,0.06 98.65±\,±\,0.07 99.98±\,±\,0.01 0.0564±\,±\,0.0004 This decision is further solidified by the substantial performance gap exhibited by the static GCN and GAT baselines. Their inability to match the performance of temporal models on our core tasks underscores that a temporal modeling approach is not merely advantageous but essential for accurately predicting political ideology in dynamic networks. 6. Results and Findings This section integrates experimental results to derive novel observations in social media platforms. 6.1. A Validated Framework for Ideology Detection and Prediction Our study validates an integrated two-stage framework for analyzing political ideology on social media. In Stage I, the NLP-based PIDN transforms raw textual content into reliable ideological annotations. Building on this foundation, Stage I employs the TGNN-based PIPN to capture the temporal evolution of user interactions and ideological structure dynamics. Together, these components form a robust bridge between micro-level textual expression and macro-level network dynamics. 6.2. Findings about Echo Chambers and Political Polarization We conducted data visualization of social media post content across different levels and objectives. Our goal was to identify shifts in political ideology across social media platforms and explore insights into echo chambers and political polarization. Different Social Media Platforms exhibit distinct overarching Political Ideologies. We used PIDN to map each platform’s ideological landscape. As illustrated in Figure 8(a), despite differences in dataset size, the ideological distributions align with prevailing public perceptions: Twitter exhibits a modest leftward tilt, while Truth Social skews markedly to the right. Twitter, as a widely used mainstream platform, draws from an internet-active American population that tends to lean slightly left. In contrast, Truth Social—founded by former U.S. President Donald J. Trump, a Republican—naturally attracts a predominantly conservative user base. We further hypothesize that the pronounced rightward shift on Truth Social stems from a combination of factors: the platform’s appeal to influential right-wing figures and its broader user community, which may already possess a mild conservative orientation. (a) (b) Figure 8. Ideological structure and temporal dynamics across platforms. (a) shows how users on the two platforms are distributed along the left–right scale, which reflects cross-platform differences in baseline audience composition. (b) tracks how highly unstable users move on the same scale over time.Panel (a) compares ideology-score histograms: Twitter is concentrated left of zero near -0.25, whereas Truth Social is concentrated on the positive, right-leaning side. Panel (b) plots yearly mean scores for the 20 most volatile Twitter users, colored by initial ideology, with a dashed line at the 0.5 midpoint. Ideological trajectories of highly volatile users on Twitter do not support the echo chamber or polarization hypothesis. To further examine the dynamics of ideological transformation at the individual level, we identified the top 20 users in the X (Twitter) dataset with the largest spread between their highest and lowest yearly mean ideology scores. To ensure robust estimation, we restricted the analysis to users who posted at least 10 tweets per year and a minimum of 50 tweets in total during the observation period. Each user was assigned an initial political leaning (Left, Center, or Right) based on the mean ideology score in their earliest available year. As shown in Figure 8(b), we then visualized each user’s ideological trajectory over time, with colors denoting their initial political stance. We used the Twitter dataset for its broader time span; Truth Social was excluded because its data covers only 2022, too short for the long-term shifts central to this inquiry. Contrary to the classical “echo chamber” hypothesis—which posits that users tend to reinforce and radicalize their pre-existing beliefs—we observed complex and multidirectional shifts. Many initially right-leaning users trended towards the ideological center, with several even crossing the midpoint into left-leaning territory over time. This pattern suggests a form of ideological moderation or de-radicalization, rather than an amplification of partisan viewpoints. Meanwhile, the few users in our sample who were initially left-leaning also failed to show a consistent pattern of leftward polarization. One such user, despite significant fluctuations, ultimately returned to an ideological position near their starting point, while another exhibited a clear trajectory toward the right. Most strikingly, users classified as centrist did not remain stable but instead exhibited significant volatility, with their ideological scores fluctuating widely across the political spectrum. In this cohort, an initial centrist position did not predict future stability. Table 9. Case study of three Twitter users illustrating non-polarizing trajectories within the U.S. political context. For each user, representative tweets are shown from two distinct periods. Key phrases are bolded. User Year Representative tweet (truncated) Score 1 2009 Is there such a thing as a “Lie Factory”? If so, its obvious right-wing Republicans & their talking heads hold a great deal of stock in it. 0.073 3 days left. Vote 4 my rescue group 2 save lives! Retweet! 0.369 2022 As you know, currently the price per barrel of oil is $120. In this chart, note how low the price has to be for corps to profit by drilling. It’s the greed. 0.515 Next time a [Republican] points out that more have died from Covid under [President A] than [President B], remind them that correlation doesn’t equal causation. 0.521 This is where I want my taxes to go, investing in America and future generations. A population without access to higher education or drowning in debt cannot compete. 0.319 2 2017 There’s been some great addresses about #ScriptureStudy. What will you begin to study tomorrow? #[ReligiousConference] 0.321 #[ReligiousConference] #goals: Study EVERYTHING about Christ & be a better #Disciple. Who’s with me? 0.431 We need our own testimony in these difficult times. Testimonies of others will only carry you so far. #[ChurchLeader] 0.240 2022 Remember this started under [the former President] and is just being continued under [the current President]. 0.457 In 2020 I saw Republicans screaming “you have to vote for [Candidate A]!” But I also saw Dems saying “you have to vote for [Candidate B] or we will end up with 2016 all over again!” 0.666 I’ve been saying the “low unemployment” and “employee shortage” have been partly due to the great amount of disabled from Covid no longer in the work force. I was told I was fear-mongering. 0.455 3 2012 [The incumbent’s] Administration Closing 9 Border Patrol Stations […]. WTF??!!! We need MORE, not LESS! 0.579 Voter fraud? Boston reports 129% voter turnout, 79% for [the incumbent]; 74% for [a Senate candidate]. 0.518 In #Iran, one can be jailed, fined, & whipped for defending #HumanRights. Dark reality for [a Christian pastor’s] attorney. 0.363 2021 We also have a proud history of our laws & Constitution being upheld. But today we have an illegal two-tiered system & a Constitution that the Dems & RINOs don’t obey. 0.874 THIS is why people are dying when they go into a hospital for “C-19 treatment”!!! STAY OUT OF HOSPITALS!!! Signed…a nurse! 0.899 We’re in the End Days, people. Accept Jesus […]. If ever we needed Jesus’s protection, it’s NOW! 0.714 Qualitative analysis reveals diverse, non-polarizing ideological trajectories. To complement our quantitative findings, we conducted a case study to examine user-level ideological dynamics—specifically, whether individuals with high variability in predicted ideology scores exhibit patterns consistent with polarization or show more nuanced shifts. We measured each user’s ideological volatility as the difference between their maximum and minimum yearly mean scores and selected the top 20 most volatile users. From these, we analyzed three representative cases in detail. None of these users align with the classical echo chamber hypothesis of unidirectional radicalization; Table 9 presents them. • User 1 began using partisan rhetoric in 2009, referring to Republicans as a “Lie Factory” (score 0.073) alongside unrelated everyday appeals such as promoting an animal rescue group (score 0.369). By 2022, their tone had shifted toward moderation and analytical reasoning—referencing economic data like “price per barrel of oil” (score 0.515) and cautioning that “correlation doesn’t equal causation” (score 0.521). This trajectory reflects a reduction in rhetorical extremity rather than an intensification of ideology. • User 2 transitioned from a non-political participant to a politically engaged yet ideologically balanced commentator. In 2017, their posts focused on religious topics (e.g., #ScriptureStudy, score 0.321). By 2022, they discussed policies spanning multiple administrations (score 0.457) and observed commonalities in partisan messaging strategies (score 0.666). These expressions reflect a critical, cross-cutting awareness rather than alignment with a single political pole. • User 3 maintained consistent conservative views over a decade. In 2012, they raised concerns about border security (score 0.579) and election integrity (score 0.518). By 2021, they continued expressing skepticism toward institutions, citing constitutional violations (score 0.874) and rejecting public health guidance (score 0.899). Though topics shifted, the ideological stance remained stable, reflecting continuity rather than radicalization. Together, these cases challenge the universality of the polarization narrative. Table 10. Representative Topics by Political Leaning and Year (2009–2022). Year Left-leaning Topics Center-leaning Topics Right-leaning Topics 2009 [hcr, health, reform] [copenhagen, cop15] [stimulus, job, economy] [hcr, health, reform] [iran, iranelection] [spending, stimulus, govt] [hcr, gop, tcot] [iran, iranelection] 2010 [health, reform, hcr] [haiti, earthquake] [oil, bp, oilspill] [haiti] [election, mn2010, tea, teaparty] [hcr, obamacare, repeal] [tea, party, teaparty] [flgov, scott, spending] 2011 [egypt, libya] [medicare, cuts, deficit] [wiunion] [debt, debtceiling, boehner] [egypt, mubarak, libya] [bin laden, pakistan] [libya, nato, war] [wiunion, gov, walker] [debt, obama, gop] 2012 [waronwomen, women] [romney, mitt, returns] [health, medicare] [obama, romney, debate] [syria] [gun, guncontrol, newtown, nra] [libya, benghazi, attack] [obama, romney, debate] [gun, shooting, aurora] 2013 [standwithwendy, txlege, abortion] [gun, guns, guncontrol] [blacklivesmatter] [snowden, leak, nsa] [boston, marathon, bombing] [shutdown, government] [irs, lerner, targeting] [benghazi, cia] [syria, assad, military] 2014 [ferguson, police, racism] [acaworks, aca] [climate, change] [ferguson, michael brown] [ukraine, russia, putin] [ebola] [mh370] [isis, syria, iraq] [benghazi] [bridgegate, christie] [irs, lerner, emails] 2015 [blacklivesmatter] [demdebate, berniesanders] [scottwalker, antiunion] [trump, realdonaldtrump, gopdebate] [iran, irandeal] [paris, parisattacks] [iran, deal, irandeal] [plannedparenthood, standwithpp] [gopdebate, trump] 2016 [feelthebern, berniesanders] [flintwatercrisis] [demsinphilly] [trump, hillaryclinton, debate] [brexit, euref] [scalia, scotus] [trump, realdonaldtrump, maga] [wikileaks, emails, clinton] [benghazi] 2017 [muslimban, travelban, nobannowall] [aca, healthcare] [impeachment] [comey, mueller, trumprussia] [charlottesville] [metoo, sexual harassment] [gorsuch, scotus] [fakenews, cnn] [tax, taxreform] [obamacare, repeal] 2018 [guncontrolnow, nra, parkland] [bluewave, midterms] [metoo, kavanaugh] [kavanaugh, ford, kavanaughhearings] [border, children, familiesbelongtogether] [khashoggi] [witch hunt, mueller, fbi] [caravan, border, buildthewall] [kavanaugh, confirmkavanaugh] 2019 [impeachment, impeach, repadamschiff] [greennewdeal] [demdebate, warren] [impeachment, ukraine, whistleblower] [muellerreport, obstruction] [hong kong] [impeachment, sham, hoax] [biden, hunter, ukraine] [border, wall, emergency] 2020 [covid19, pandemic, trumpvirus] [george floyd, blm, defundthepolice] [usps, mailin] [coronavirus, covid19] [biden, trump, election] [george floyd, protests] [obamagate, spygate] [reopenamerica, reopen] [voterfraud, electionfraud, stopthesteal] 2021 [jan 6, insurrection, coup] [votingrights, filibuster] [buildbackbetter, b] [capitol, riot, jan 6] [afghanistan, withdrawal] [vaccine, mandate, delta] [crt, critical race theory] [arizona, audit] [antimask, antivax, bordercrisis] 2022 [abortion, codifyroe, womensrights] [uvalde, gunreform] [ukraine, standwithukraine] [ukraine, war, russia] [maralago, search, documents] [roe, wade, overturn] [inflation, gas] [fbi, doj] [border, migrants] [hunter, laptop] Comparative Thematic Analysis Reveals Divergent Discursive Arenas To complement our quantitative findings, we conducted a thematic analysis of political discourse on Twitter and Truth Social—two ideologically distinct platforms. We grouped users into Left-leaning, Centrist, and Right-leaning categories based on aggregated post-level ideology scores, and applied BERTopic (31) to extract dominant themes. BERTopic embeds documents contextually, reduces their dimensionality, clusters them by density, and extracts representative keywords per topic. Results are shown in Table 10 (Twitter, 2009–2022) and Table 11 (Truth Social, 2022). Twitter reveals a dynamic we term a “shared agenda with competing frames.” Across the political spectrum, users engage with the same high-salience topics—from [hcr] in 2009 to [ukraine, war] in 2022—suggesting a common baseline of topical attention. However, ideological divergence emerges in framing. For example, during the 2019 impeachment, the Left referenced institutional actors ([repadamschiff]), while the Right used delegitimizing terms ([sham, hoax]). Similarly, the events of January 6th surfaced as [insurrection, coup] on the Left and [capitol, riot] at the Center, while the Right’s 2021 agenda instead foregrounded [arizona, audit] and [crt, critical race theory]. This is a contested discursive space, not a set of isolated echo chambers. In contrast, Truth Social exhibits a “high-conflict ideological arena” with a right-shifted discursive center. Themes categorized as far-right on Twitter—e.g., [hunter, laptop]—appear under the “Centrist” label on Truth Social, alongside claims such as [2000mules, ballot] that have no Twitter counterpart, reflecting that partisan narratives constitute the platform’s mainstream. Rhetorical intensity is notably elevated: economic concerns are framed as [bidenflation], the Mar-a-Lago search is portrayed as a partisan [witchhunt] by a [corrupt] FBI, and the term [crimefamily] denounces the Bidens. Such vocabulary suggests a shift from political discourse to moral indictment. While Left-leaning users exist on Truth Social, they act more as ideological foils than dialogue participants. Topics like [jan6, hearings] and [gun control] appear, but often in oppositional framing. Instead of reducing polarization, their presence intensifies it by provoking adversarial responses. This dynamic is compounded by conspiratorial clusters—e.g., [stoptheshots, pureblood]—which anchor its ideological extremity. Table 11. Representative Topics on Truth Social by Political Leaning (2022). Year Left-leaning Topics Center-leaning Topics Right-leaning Topics 2022 [abortion, roe, wade, overturned] [uvalde, shooting, gun, control] [ukraine, standwithukraine, kyiv] [jan6, committee, hearings, insurrection] [maralago, raid, fbi] [ukraine, russia, war, putin] [maralago, raid, fbi, doj, warrant] [roe, wade, overturned, supreme, court] [inflation, recession, economy, prices] [hunter, laptop, bidens, fbi] [2000mules, ballot, harvesting, fraud] [inflation, gas, prices, bidenflation] [fbi, doj, raid, corruption, witchhunt] [border, crisis, migrants, illegal] [hunter, laptop, biden, crimefamily] [stoptheshots, diedsuddenly, pureblood] [abortion, prolife, unborn, baby] Overall, polarization manifests differently across platforms: Twitter hosts a fragmented but interconnected public sphere where competing frames coexist, whereas Truth Social functions as a self-reinforcing enclave whose right-shifted center, moralized rhetoric, and reactive dynamics amplify division rather than bridge it. 7. Conclusion This paper presents TSN4PI, a unified framework that integrates natural language processing with temporal graph neural networks for tracking political ideologies on social media. Methodologically, we provide a robust, end-to-end pipeline that effectively bridges the domain gap between news and social content, modeling the dynamic evolution of user interactions. We also contribute two large-scale, publicly released datasets from X (formerly Twitter) and allsides.com. Empirically, our analysis of these platforms reveals that the most ideologically volatile users tend to converge toward the center, a finding contrary to widely held assumptions about increasing online polarization. Future work includes extending to multi-dimensional ideological models and cross-platform comparative studies. Acknowledgements. This work was supported in part by the National Natural Science Foundation of China (Grant Nos. 92370204 and 62506348), the National Key R&D Program of China (Grant No. 2023YFF0725001), the New Generation Artificial Intelligence-National Science and Technology Major Project (Grant No. 2025ZD0122601), Guangdong Provincial Key Laboratory of Frontier Basic Science for All-domain Intelligence, the Guangdong Basic and Applied Basic Research Foundation (Grant No. 2023B1515120057), the Key-Area Special Project of Guangdong Provincial Ordinary Universities (Grant No. 2024ZDZX1007), the Natural Science Foundation of Anhui Province (Grant No. 2508085QF211), and the Opening Foundation of State Key Laboratory of Cognitive Intelligence, iFLYTEK (Grant No. COGOS-2025HE02). Supplementary Material This document is the online-only supplementary material for the article “Against Political Polarization: A Unified Framework for Tracing Evolving Political Ideologies on Social Media.” It collects dataset details, implementation settings, prompt templates, complexity analysis, and the full per-model results referenced from the main text. Section, table, and figure numbers below are those cited in the main text as belonging to the supplementary material. Appendix A Appendix A.1. Details of the Datasets A.1.1. Data Privacy Given the sensitivity of political ideology, privacy is central to our data handling. All news data were sourced from public websites. Social media data were anonymized to meet ethical standards. On X and Truth Social, personal identifiers (e.g., names, handles) were replaced with numeric IDs. Only public posts were retained; no private user data was used. Data Release Format. To respect the terms of the underlying sources, the released datasets are distributed in a dehydrated form. For the X (Twitter) and Truth Social corpora we release post and user identifiers together with our derived annotations (political relevance labels, ideology scores, and the temporal interaction edges used to build the graphs), rather than the original post text; users can rehydrate the text through the respective platform APIs subject to their terms of service. For the news corpus we release article URLs, outlet names, publication dates, and the AllSides ideology score of each outlet, rather than the full article text, which remains the property of the publishing outlets. The AllSides ratings redistributed in this way are covered by the C BY-NC 4.0 license and are provided with the attribution required by that license; the released collection is therefore available for noncommercial research use only. Scripts for rehydration and for reproducing every table in the paper are provided in the repository at https://github.com/yeahjack/TSN4PI. A.1.2. The allsides.com Dataset The overview of the allsides.com dataset is as follows: Table 1. Overview of the allsides.com news with annotated political ideology dataset. Characteristic # of News Articles # of Media Outlets Total Articles/Outlets 233,570 466 Political Ideology Center 76,023 158 Lean Left 53,532 106 Right 45,506 75 Lean Right 32,555 62 Left 25,954 65 Data Preprocessing. As with the X dataset described below, we retained only English articles to ensure linguistic consistency. We also normalized ideology scores to align with other framework components. For an article N with original score Sraw∈[−6,6]S_raw∈[-6,6], the adjusted score SadjustedS_adjusted is computed as: (1) Sadjusted(N)=Sraw+612.S_adjusted(N)= S_raw+612. This normalization maps the score to the [0,1][0,1] interval, facilitating uniform scale interpretation during model training and evaluation. Figure 1. Word Clouds Highlighting Common Terms in News Reports Titles by Political Leanings. The visualization spans from ‘Left’ to ‘Right’, with ‘Center’ in the middle, reflecting the language usage patterns in political news reporting.Five side-by-side word clouds compare title terms for Left, Lean Left, Center, Lean Right, and Right outlets. Prominent phrases include Capitol Riot and Taylor Greene on the left, COVID vaccine and Supreme Court in the center, and criminal referrals and vaccine mandate on the right. We also created word clouds for news titles from different political ideology factions, as shown in Figure 1. We found that media with a left-leaning bias generally focus on topics such as the Capitol riot, the Supreme Court, and COVID, while right-leaning media often focus on Elon Musk, Paul Pelosi, the Committee Criminal, and vaccine mandates, with center media reporting on both. This reflects not only the different focus points of media across the political spectrum but also the distinct use of language in reporting the same topics, aligning with our research hypothesis: even for the same events, the choice of words can still reveal their political ideologies. A.1.3. The X (Twitter) Dataset Data Preprocessing. To ensure linguistic consistency with our analytical focus of presidential elections in the United States, we retained only English-language tweets. Language filtering was performed using the fasttext-langdetect library (41; 42). A.1.4. The Truth Social Dataset Data Preprocessing. Preliminary cleaning and preprocessing were applied to the raw dataset. Given that some “Truths” had missing or misaligned timestamps, we utilized web scraping tools to supplement and correct the timestamps for these accessible “Truths.” A.2. Implementation Details Models were trained with a 90/10 train-test split. For PIDN, we trained for 3 epochs with a learning rate of 2×10−52× 10^-5 and a weight decay of 0.01. For the TST task, we used the Llama-3.1-8B-Instruct model to generate style-transferred texts with a temperature of 0.7, top_p of 0.95, and a maximum generation length of 512 tokens, terminating on the “</Tweet>” token. The prompts for directly using the LLM-based method to assess political relevance and assigning scores, TST and LLM-as-a-judge for evaluation are provided in Tables A.3, A.3, and A.3, respectively. To facilitate automatic evaluation, we leverage tool calling to parse model responses into a structured JSON format. For PIDN, we used (α=0.4,β=0.2)(α=0.4,β=0.2) during the preliminary backbone comparison. After the dedicated tuning stage reported in the main paper, we selected (α=0.4,β=0.3)(α=0.4,β=0.3) for the ablation study and final models. For PIPN, we adopted the default sampling, memory, and network configurations provided by TGL (90) for the different TGNNs. Mixed-precision training and inference were adopted for efficiency. All experiments were conducted on eight machines, each equipped with 8 Nvidia A800 GPUs (80GB VRAM each), running Ubuntu 22.04 with Python 3.11. We used the vllm (47) framework for text generation tasks, and TGL (90) for TGNNs. For static graph tasks, we used PyTorch Geometric (24). The training and inference were performed using PyTorch 2.4.0. A.3. Prompt Details Prompt for LLM-based method Here we provide the prompt template used for the LLM-based method, which is designed to generate relevance and scores for posts. We enabled tool calling to allow the model to output structured results after reasoning about the political relevance and leaning of a given tweet. Prompt Template for Political Tweet Analysis with Function Calling System Prompt: You are an expert political analyst specializing in U.S. politics. Your task is to analyze a given tweet and provide a structured analysis by using the record_political_analysis tool. Carefully determine the following parameters for the tool: • relevance: Set to ’1’ if the tweet is politically relevant to the U.S. context, ’0’ otherwise. • score: If relevant, assign a float from 0.0 (left-leaning) to 1.0 (right-leaning). This parameter should only be provided for relevant tweets. • reasoning: Provide a clear, step-by-step explanation for your decisions. Escape all double quotes with \. You must directly use the record_political_analysis function to output your findings. Guidelines & Tool Schema (record_political_analysis): Function Name: record_political_analysis Description: Records the political analysis of a tweet after careful reasoning. Parameters: - ‘reasoning‘ (string, required): Detailed step-by-step thinking. - ‘relevance‘ (string, required): ’1’ if political, ’0’ otherwise. - ‘score‘ (number, optional): Score 0.0-1.0. Only required if relevance==’1’. Example Input Tweet: BREAKING: The Senate just passed the new infrastructure bill. A major win for the administration, but critics on both sides are already voicing concerns about the spending. Example Output (Tool Call Arguments): ⬇ "reasoning": "The tweet discusses a major piece of U.S. legislation, the infrastructure bill, and mentions the Senate and the administration. This is clearly a political topic. The tweet presents a balanced view, noting it as a ’win’ but also mentioning ’critics on both sides’, suggesting a centrist take. Therefore, the score is set to 0.5.", "relevance": "1", "score": 0.5 Input Format: Analyze this tweet: tweet_text Start of Response: The model should directly output the tool call. The arguments should be a single, well-formed JSON string. Prompt for Text Style Transfer. Here we provide the prompt template used for the text style transfer (TST) task, which is designed to convert news articles into concise tweets. This prompt is crucial for training our model to understand the nuances of tweet generation from longer news content. Prompt Template for News-to-Tweet Style Transfer System Prompt: You will be given a news report. Your task is to rewrite it as a tweet that captures the core message. Guidelines: Be concise and stay within 280 characters. Use common abbreviations (e.g., ICYMI, IMO). Add hashtags to highlight key topics (e.g., #Breaking). Mention relevant accounts using @ when applicable. Always produce a non-empty tweet in ENGLISH. End your generated tweet with the token: </Tweet> Example News: In an extraordinary face-to-face meeting at the White House on Wednesday, President Joe Biden told Ukrainian President Volodymyr Zelenskyy that… Example Tweet: President Biden meets with Ukrainian President Zelenskyy at the White House, pledging “unequivocal and unbending support” to Ukraine. #UkraineWar #WhiteHouseMeeting Input Format: <Tweet>news_text</Tweet> Start of Response: <Tweet> Prompt for Text Style Transfer Evaluation. This prompt is used to evaluate the quality of the generated tweets from the TST task by LLM-as-a-judge. It is designed to assess the relevance and quality of the generated tweet in relation to the original news article. We enabled tool calling to allow the model to output structured scores after reasoning about the tweet’s style and semantic preservation. Prompt for LLM-as-a-Judge Evaluation of Text Style Transfer System Prompt: You are a strict grader. Use the rubric below to give two integer scores from 0–10: • style_score - tweet naturalness on social media • semantic_score - key-information preservation Rubric examples (do not output these examples in your answer): 10 - flawless style AND all facts present 8 - minor issues in style or small missing details 5 - several awkward phrases OR half the facts missing 2 - hard to read or mostly irrelevant 0 - gibberish or unrelated to the article Your job is to think step-by-step (Chain-of-Thought) about style and semantics, then call the function grade_tweet with two integer scores. Tool Schema (grade_tweet): Function Name: grade_tweet Description: Returns integer style and semantic scores for a tweet given a news article. Parameters: - ‘style_score‘ (integer, required): Tweet naturalness on social media, from 0 to 10. - ‘semantic_score‘ (integer, required): Key-information preservation score, from 0 to 10. Example Output (Tool Call Arguments): ⬇ "style_score": 9, "semantic_score": 10 Input Format:News article: <ARTICLE> [Original News Article Text Here] </ARTICLE> Generated tweet: <TWEET> [Style-Transferred Tweet Text Here] </TWEET> Start of Response: The model should first provide its Chain-of-Thought reasoning, then output the tool call. The final response must include the call to the grade_tweet function. A.4. UDA Complexity Analysis To provide a comprehensive understanding of the computational overhead associated with our UDA module, we present a detailed complexity analysis of the Maximum Mean Discrepancy (MMD) and Kullback-Leibler (KL) divergence methods. This analysis examines theoretical time complexity, floating-point operations (FLOPs), and practical characteristics of GPU implementations. We assume two sets of embeddings, EXE_X and EYE_Y, each containing N samples of dimension D, corresponding to a batch from the source and target domains, respectively. Maximum Mean Discrepancy (MMD) The computational core of the MMD loss is the construction of three kernel matrices, KXXK_X, KYYK_Y, and KXYK_XY, each of size N×N× N. Time Complexity. The calculation is dominated by the pairwise evaluation of the kernel function. For the Gaussian RBF kernel, computing the squared Euclidean distance ‖x−y‖2\|x-y\|^2 between two D-dimensional vectors requires O(D)O(D) operations. Since this must be done for all N×N× N pairs to construct each kernel matrix, the complexity of forming one matrix is O(N2D)O(N^2D). The final summation over the kernel matrices is O(N2)O(N^2). Therefore, the overall time complexity is dominated by kernel matrix computation, yielding O(N2D)O(N^2D). Floating-Point Operations (FLOPs). We can estimate the FLOPs for the RBF kernel. The calculation of ‖x−y‖2\|x-y\|^2 involves approximately 3D3D FLOPs (D subtractions, D multiplications, D−1D-1 additions). Constructing the three N×N× N kernel matrices thus requires roughly 3×N2×3D=9N2D3× N^2× 3D=9N^2D FLOPs. This quadratic dependence on N makes MMD computationally demanding. GPU Latency and Bottlenecks. On GPU architectures, the pairwise distance calculations can be efficiently mapped to highly optimized Generalized Matrix-Matrix Multiplication (GEMM) operations. However, the O(N2D)O(N^2D) complexity makes the algorithm fundamentally compute-bound. Latency grows quadratically with the batch size N, which rapidly becomes the primary performance bottleneck. Furthermore, storing the intermediate kernel matrices requires O(N2)O(N^2) memory, which can exhaust the GPU’s VRAM for large N (e.g., N>4096N>4096), thereby imposing a practical limit on the batch size. Kullback-Leibler (KL) Divergence The KL divergence loss is computed element-wise on the distribution-like outputs P and Q, which are tensors of shape (N,D)(N,D). Time Complexity. The computation involves element-wise division, logarithm, and multiplication across two (N,D)(N,D) tensors, followed by a summation. Each of these element-wise operations has a time complexity proportional to the number of elements. Thus, the overall time complexity is linear with respect to both N and D, resulting in O(ND)O(ND). Floating-Point Operations (FLOPs). The sequence of operations (division, logarithm, multiplication, and summation) results in approximately 4ND4ND FLOPs. This is substantially lower than MMD’s quadratic complexity. GPU Latency and Bottlenecks. The element-wise nature of the KL divergence calculation makes it an “embarrassingly parallel” problem, ideally suited for GPU execution. The operations exhibit high data parallelism and low computational intensity. Consequently, the algorithm is typically memory-bound; its latency is limited not by the speed of computation but by the GPU’s memory bandwidth for reading the input tensors P and Q. The overall latency is minimal and scales linearly with the total number of elements, N×DN× D. Summary and Implications The choice between MMD and KL divergence for domain adaptation introduces a critical trade-off in computational efficiency. Table 5 summarizes the computational characteristics of each method. Table 5. Computational complexity comparison of UDA methods. Metric Time Complexity FLOPs (Approx.) GPU Bottleneck MMD O(N2D)O(N^2D) 9N2D9N^2D Compute-Bound KL Divergence O(ND)O(ND) 4ND4ND Memory-Bound In practice, the quadratic complexity of MMD restricts its application to smaller batch sizes (N), whereas KL divergence remains efficient even for large-batch training. This computational difference necessitates careful management of batch sizes when using MMD to maintain feasible training times, while KL divergence offers a more scalable alternative from a computational standpoint. Table 6. Unified performance analysis of the Qwen2.5 model family (Part 1: 0.5B, 3B, 7B). All models are evaluated from 0-shot to 200-shot settings. Classification (F1 ↑ ) Political Class Detail ↑ Regression Error ↓ Model / Setting Non-political Political Macro Avg. Precision Recall MAE RMSE Few-Shot for Qwen2.5-0.5B-Instruct 0-shot 0.510±\,±\,0.022 0.727±\,±\,0.007 0.618±\,±\,0.014 0.601±\,±\,0.007 0.918±\,±\,0.006 0.413±\,±\,0.007 0.478±\,±\,0.005 2-shot 0.568±\,±\,0.132 0.717±\,±\,0.035 0.643±\,±\,0.082 0.629±\,±\,0.094 0.855±\,±\,0.061 0.297±\,±\,0.055 0.364±\,±\,0.054 4-shot 0.607±\,±\,0.132 0.751±\,±\,0.051 0.679±\,±\,0.089 0.662±\,±\,0.126 0.901±\,±\,0.088 0.330±\,±\,0.031 0.395±\,±\,0.038 6-shot 0.577±\,±\,0.098 0.750±\,±\,0.032 0.664±\,±\,0.062 0.622±\,±\,0.040 0.951±\,±\,0.052 0.305±\,±\,0.016 0.367±\,±\,0.018 8-shot 0.525±\,±\,0.233 0.753±\,±\,0.058 0.639±\,±\,0.143 0.628±\,±\,0.088 0.960±\,±\,0.067 0.257±\,±\,0.024 0.317±\,±\,0.032 10-shot 0.619±\,±\,0.099 0.773±\,±\,0.023 0.696±\,±\,0.061 0.647±\,±\,0.053 0.970±\,±\,0.046 0.287±\,±\,0.056 0.349±\,±\,0.062 20-shot 0.824±\,±\,0.090 0.867±\,±\,0.046 0.845±\,±\,0.068 0.797±\,±\,0.097 0.964±\,±\,0.038 0.285±\,±\,0.019 0.357±\,±\,0.015 50-shot 0.896±\,±\,0.063 0.916±\,±\,0.043 0.906±\,±\,0.053 0.849±\,±\,0.074 0.999±\,±\,0.001 0.282±\,±\,0.027 0.352±\,±\,0.027 100-shot 0.988±\,±\,0.007 0.988±\,±\,0.006 0.988±\,±\,0.006 0.980±\,±\,0.016 0.996±\,±\,0.006 0.274±\,±\,0.018 0.338±\,±\,0.017 200-shot 0.988±\,±\,0.009 0.988±\,±\,0.008 0.988±\,±\,0.009 0.978±\,±\,0.018 0.998±\,±\,0.003 0.245±\,±\,0.023 0.309±\,±\,0.027 Few-Shot for Qwen2.5-3B-Instruct 0-shot 0.946±\,±\,0.003 0.943±\,±\,0.003 0.945±\,±\,0.003 0.945±\,±\,0.008 0.942±\,±\,0.006 0.283±\,±\,0.005 0.333±\,±\,0.005 2-shot 0.977±\,±\,0.004 0.976±\,±\,0.005 0.977±\,±\,0.005 0.987±\,±\,0.005 0.965±\,±\,0.012 0.258±\,±\,0.016 0.299±\,±\,0.021 4-shot 0.970±\,±\,0.015 0.968±\,±\,0.018 0.969±\,±\,0.016 0.984±\,±\,0.025 0.954±\,±\,0.042 0.271±\,±\,0.035 0.321±\,±\,0.046 6-shot 0.981±\,±\,0.012 0.979±\,±\,0.014 0.980±\,±\,0.013 0.994±\,±\,0.006 0.966±\,±\,0.027 0.260±\,±\,0.021 0.312±\,±\,0.029 8-shot 0.984±\,±\,0.008 0.982±\,±\,0.010 0.983±\,±\,0.009 0.993±\,±\,0.004 0.972±\,±\,0.022 0.275±\,±\,0.024 0.335±\,±\,0.032 10-shot 0.980±\,±\,0.014 0.978±\,±\,0.016 0.979±\,±\,0.015 0.997±\,±\,0.003 0.961±\,±\,0.032 0.268±\,±\,0.030 0.331±\,±\,0.038 20-shot 0.987±\,±\,0.002 0.986±\,±\,0.002 0.986±\,±\,0.002 0.998±\,±\,0.003 0.974±\,±\,0.004 0.260±\,±\,0.028 0.325±\,±\,0.036 50-shot 0.985±\,±\,0.007 0.983±\,±\,0.009 0.984±\,±\,0.008 1.000±\,±\,0.000 0.967±\,±\,0.018 0.291±\,±\,0.034 0.368±\,±\,0.032 100-shot 0.978±\,±\,0.005 0.972±\,±\,0.009 0.975±\,±\,0.006 1.000±\,±\,0.000 0.947±\,±\,0.017 0.283±\,±\,0.025 0.367±\,±\,0.029 200-shot 0.992±\,±\,0.002 0.989±\,±\,0.003 0.990±\,±\,0.002 0.998±\,±\,0.002 0.980±\,±\,0.006 0.273±\,±\,0.015 0.353±\,±\,0.014 Few-Shot for Qwen2.5-7B-Instruct 0-shot 0.927±\,±\,0.002 0.905±\,±\,0.003 0.916±\,±\,0.002 0.992±\,±\,0.003 0.833±\,±\,0.004 0.238±\,±\,0.004 0.288±\,±\,0.005 2-shot 0.936±\,±\,0.021 0.928±\,±\,0.026 0.932±\,±\,0.024 0.986±\,±\,0.005 0.879±\,±\,0.049 0.242±\,±\,0.005 0.292±\,±\,0.007 4-shot 0.920±\,±\,0.032 0.905±\,±\,0.044 0.912±\,±\,0.038 0.986±\,±\,0.006 0.840±\,±\,0.079 0.230±\,±\,0.007 0.282±\,±\,0.012 6-shot 0.918±\,±\,0.024 0.903±\,±\,0.035 0.911±\,±\,0.030 0.987±\,±\,0.008 0.835±\,±\,0.063 0.243±\,±\,0.007 0.301±\,±\,0.007 8-shot 0.906±\,±\,0.024 0.886±\,±\,0.035 0.896±\,±\,0.029 0.990±\,±\,0.007 0.804±\,±\,0.061 0.246±\,±\,0.007 0.305±\,±\,0.011 10-shot 0.898±\,±\,0.019 0.874±\,±\,0.028 0.886±\,±\,0.024 0.995±\,±\,0.002 0.781±\,±\,0.047 0.236±\,±\,0.012 0.291±\,±\,0.015 20-shot 0.923±\,±\,0.031 0.908±\,±\,0.043 0.916±\,±\,0.037 0.995±\,±\,0.002 0.838±\,±\,0.074 0.244±\,±\,0.016 0.306±\,±\,0.019 50-shot 0.961±\,±\,0.014 0.958±\,±\,0.015 0.960±\,±\,0.014 0.996±\,±\,0.003 0.924±\,±\,0.031 0.241±\,±\,0.018 0.304±\,±\,0.020 100-shot 0.964±\,±\,0.008 0.962±\,±\,0.009 0.963±\,±\,0.009 0.998±\,±\,0.000 0.929±\,±\,0.017 0.237±\,±\,0.010 0.307±\,±\,0.012 200-shot 0.974±\,±\,0.008 0.973±\,±\,0.008 0.973±\,±\,0.008 0.998±\,±\,0.002 0.950±\,±\,0.017 0.243±\,±\,0.012 0.315±\,±\,0.018 Table 7. Unified performance analysis of the Qwen2.5 model family (Part 2: 14B, 32B, 72B). All models are evaluated from 0-shot to 200-shot settings. Classification (F1 ↑ ) Political Class Detail ↑ Regression Error ↓ Model / Setting Non-political Political Macro Avg. Precision Recall MAE RMSE Few-Shot for Qwen2.5-14B-Instruct 0-shot 0.918±\,±\,0.003 0.900±\,±\,0.004 0.909±\,±\,0.003 0.986±\,±\,0.001 0.827±\,±\,0.006 0.227±\,±\,0.004 0.285±\,±\,0.003 2-shot 0.949±\,±\,0.022 0.941±\,±\,0.029 0.945±\,±\,0.026 0.992±\,±\,0.003 0.896±\,±\,0.054 0.219±\,±\,0.007 0.273±\,±\,0.011 4-shot 0.955±\,±\,0.009 0.949±\,±\,0.011 0.952±\,±\,0.010 0.994±\,±\,0.001 0.908±\,±\,0.021 0.216±\,±\,0.002 0.271±\,±\,0.007 6-shot 0.963±\,±\,0.014 0.959±\,±\,0.017 0.961±\,±\,0.016 0.991±\,±\,0.004 0.929±\,±\,0.034 0.235±\,±\,0.025 0.299±\,±\,0.032 8-shot 0.961±\,±\,0.019 0.956±\,±\,0.024 0.959±\,±\,0.022 0.992±\,±\,0.005 0.924±\,±\,0.047 0.232±\,±\,0.008 0.295±\,±\,0.014 10-shot 0.962±\,±\,0.015 0.958±\,±\,0.018 0.960±\,±\,0.017 0.988±\,±\,0.003 0.930±\,±\,0.036 0.236±\,±\,0.009 0.305±\,±\,0.013 20-shot 0.980±\,±\,0.006 0.978±\,±\,0.006 0.979±\,±\,0.006 0.990±\,±\,0.002 0.967±\,±\,0.013 0.235±\,±\,0.010 0.300±\,±\,0.015 50-shot 0.982±\,±\,0.002 0.981±\,±\,0.002 0.981±\,±\,0.002 0.991±\,±\,0.002 0.971±\,±\,0.005 0.239±\,±\,0.006 0.309±\,±\,0.011 100-shot 0.968±\,±\,0.007 0.965±\,±\,0.008 0.966±\,±\,0.008 0.988±\,±\,0.004 0.943±\,±\,0.015 0.241±\,±\,0.007 0.311±\,±\,0.006 200-shot 0.979±\,±\,0.004 0.979±\,±\,0.005 0.979±\,±\,0.004 0.995±\,±\,0.001 0.963±\,±\,0.009 0.223±\,±\,0.004 0.301±\,±\,0.004 Few-Shot for Qwen2.5-32B-Instruct 0-shot 0.946±\,±\,0.001 0.937±\,±\,0.002 0.942±\,±\,0.002 0.991±\,±\,0.001 0.889±\,±\,0.004 0.225±\,±\,0.002 0.283±\,±\,0.001 2-shot 0.956±\,±\,0.008 0.955±\,±\,0.009 0.956±\,±\,0.008 0.993±\,±\,0.002 0.921±\,±\,0.017 0.206±\,±\,0.009 0.263±\,±\,0.012 4-shot 0.962±\,±\,0.008 0.962±\,±\,0.008 0.962±\,±\,0.008 0.994±\,±\,0.001 0.933±\,±\,0.015 0.214±\,±\,0.005 0.275±\,±\,0.006 6-shot 0.968±\,±\,0.005 0.969±\,±\,0.005 0.968±\,±\,0.005 0.992±\,±\,0.003 0.946±\,±\,0.011 0.212±\,±\,0.005 0.272±\,±\,0.008 8-shot 0.976±\,±\,0.006 0.977±\,±\,0.006 0.976±\,±\,0.006 0.995±\,±\,0.002 0.959±\,±\,0.010 0.210±\,±\,0.007 0.272±\,±\,0.011 10-shot 0.973±\,±\,0.009 0.974±\,±\,0.010 0.973±\,±\,0.009 0.994±\,±\,0.001 0.954±\,±\,0.018 0.207±\,±\,0.006 0.270±\,±\,0.010 20-shot 0.980±\,±\,0.004 0.980±\,±\,0.004 0.980±\,±\,0.004 0.993±\,±\,0.002 0.968±\,±\,0.009 0.219±\,±\,0.011 0.284±\,±\,0.016 50-shot 0.989±\,±\,0.002 0.989±\,±\,0.002 0.989±\,±\,0.002 0.995±\,±\,0.002 0.984±\,±\,0.004 0.218±\,±\,0.007 0.285±\,±\,0.008 100-shot 0.990±\,±\,0.001 0.990±\,±\,0.001 0.990±\,±\,0.001 0.996±\,±\,0.001 0.985±\,±\,0.002 0.225±\,±\,0.007 0.296±\,±\,0.007 200-shot 0.992±\,±\,0.002 0.993±\,±\,0.002 0.992±\,±\,0.002 0.996±\,±\,0.001 0.990±\,±\,0.003 0.212±\,±\,0.011 0.284±\,±\,0.011 Few-Shot for Qwen2.5-72B-Instruct 0-shot 0.955±\,±\,0.002 0.950±\,±\,0.003 0.953±\,±\,0.002 0.994±\,±\,0.002 0.910±\,±\,0.004 0.216±\,±\,0.004 0.275±\,±\,0.004 2-shot 0.946±\,±\,0.015 0.935±\,±\,0.020 0.940±\,±\,0.017 0.990±\,±\,0.003 0.887±\,±\,0.037 0.206±\,±\,0.004 0.265±\,±\,0.003 4-shot 0.972±\,±\,0.007 0.969±\,±\,0.009 0.971±\,±\,0.008 0.989±\,±\,0.003 0.950±\,±\,0.018 0.206±\,±\,0.009 0.268±\,±\,0.011 6-shot 0.967±\,±\,0.012 0.962±\,±\,0.014 0.964±\,±\,0.013 0.989±\,±\,0.002 0.937±\,±\,0.029 0.212±\,±\,0.008 0.277±\,±\,0.011 8-shot 0.972±\,±\,0.010 0.969±\,±\,0.012 0.971±\,±\,0.011 0.990±\,±\,0.003 0.949±\,±\,0.022 0.225±\,±\,0.016 0.292±\,±\,0.022 10-shot 0.965±\,±\,0.014 0.959±\,±\,0.017 0.962±\,±\,0.015 0.993±\,±\,0.002 0.928±\,±\,0.032 0.217±\,±\,0.018 0.285±\,±\,0.022 20-shot 0.974±\,±\,0.006 0.971±\,±\,0.007 0.972±\,±\,0.006 0.992±\,±\,0.002 0.950±\,±\,0.015 0.214±\,±\,0.011 0.279±\,±\,0.013 50-shot 0.988±\,±\,0.005 0.988±\,±\,0.006 0.988±\,±\,0.006 0.991±\,±\,0.002 0.985±\,±\,0.012 0.213±\,±\,0.006 0.284±\,±\,0.009 100-shot 0.987±\,±\,0.003 0.992±\,±\,0.001 0.989±\,±\,0.002 0.992±\,±\,0.001 0.993±\,±\,0.003 0.220±\,±\,0.014 0.297±\,±\,0.016 200-shot 0.985±\,±\,0.003 0.993±\,±\,0.001 0.989±\,±\,0.002 0.994±\,±\,0.002 0.991±\,±\,0.002 0.212±\,±\,0.007 0.290±\,±\,0.008 References Achen (1975) C. H. Achen Mass political attitudes and the survey response. Am. Political Sci. Rev.. Cited by: §2. Ai et al. (2024) L. Ai, S. Gupta, S. Oak, Z. Hui, Z. Liu, and J. Hirschberg TweetIntent@Crisis: A Dataset Revealing Narratives of Both Sides in the Russia-Ukraine Crisis. In ICWSM, Cited by: §3.2, Table 3. AllSides (2024) AllSides AllSides Media Bias Ratings. Note: https://w.allsides.com/media-bias/media-bias-ratingsLicensed under C BY-NC 4.0 Cited by: §3.1. Augenstein et al. (2016) I. Augenstein, T. Rocktäschel, A. Vlachos, and K. Bontcheva Stance detection with bidirectional conditional encoding. In EMNLP, Cited by: §2. Baly et al. (2019) R. Baly, G. Karadzhov, A. Saleh, J. Glass, and P. Nakov Multi-task ordinal regression for jointly predicting the trustworthiness and the leading political ideology of news media. In NAACL, Cited by: §2, §2. Beltagy et al. (2020) I. Beltagy M. E. Peters et al. Longformer: the long-document transformer. arXiv:2004.05150. Cited by: §4.1. Bhardwaj et al. (2024) R. Bhardwaj, S. Kumar, and V. Pudi CAT-Gen: a context-aware teleology-guided generative model for mitigating task-specific bias. In FAccT, Cited by: §5.2.2. Boukhalfa et al. (2022) K. Boukhalfa, S. Belkacem, and O. Boussaid Ranking social media news feeds: a comparative study of personalized and non-personalized prediction models. In AIAP, Cited by: §3.2, Table 3. Brown et al. (2020) T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. D. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei Language models are few-shot learners. In NeurIPS, Cited by: §1, §2, §4.1. Cascalheira et al. (2024) C. J. Cascalheira, S. Chapagain, R. E. Flinn, D. Klooster, D. Laprade, Y. Zhao, E. M. Lund, A. Gonzalez, K. Corro, R. Wheatley, A. Gutierrez, O. G. Villanueva, K. Saha, M. D. Choudhury, J. R. Scheer, and S. M. Hamdi The LGBTQ+ Minority Stress on Social Media (MiSSoM) Dataset: A Labeled Dataset for Natural Language Processing and Machine Learning. In ICWSM, Cited by: Table 3. Center (2014) P. R. CenterPolitical polarization in the american public(Website) External Links: Link Cited by: §1. Center (2021) P. R. CenterBoth republicans and democrats prioritize family, but they differ over other sources of meaning in life(Website) External Links: Link Cited by: §1. Chen et al. (2024a) J. Chen, S. Xiao, P. Zhang, K. Luo, D. Lian, et al. BGE m3-embedding: multi-lingual, multi-functionality, multi-granularity text embeddings through self-knowledge distillation. arXiv:2402.03216. Cited by: 1st item. Chen et al. (2024b) K. Chen, Z. He, K. Burghardt, J. Zhang, and K. Lerman IsamasRed: a public dataset tracking reddit discussions on israel-hamas conflict. In ICWSM, Cited by: Table 3. Chen et al. (2024c) T. Chen, B. Zhang, X. Wang, W. Zhang, C. Hon, W. Di, L. Chen, Q. Li, and F. Wang Public opinion evolution in cyberspace: a case analysis of pelosi’s visit to taiwan. IEEE TCSS. Cited by: §2. Colacrai et al. (2024) E. Colacrai, F. Cinus, G. De Francisci Morales, and M. Starnini Navigating multidimensional ideologies with reddit’s political compass: economic conflict and social affinity. In W, Cited by: §2. Conneau et al. (2019) A. Conneau, K. Khandelwal, N. Goyal, V. Chaudhary, G. Wenzek, F. Guzmán, E. Grave, M. Ott, L. Zettlemoyer, and V. Stoyanov Unsupervised cross-lingual representation learning at scale. arXiv:1911.02116. Cited by: 6(a). Conover et al. (2011) M. D. Conover, B. Goncalves, J. Ratkiewicz, A. Flammini, and F. Menczer Predicting the political alignment of twitter users. In Proc. IEEE PASSAT and SocialCom, Cited by: §2. Dai et al. (2025) L. Dai, Y. Xu, J. Ye, H. Liu, and H. Xiong SePer: measure retrieval utility through the lens of semantic perplexity reduction. In ICLR, Cited by: §1. DeepSeek-AI (2024) DeepSeek-AI DeepSeek-v3 technical report. arXiv:2412.19437. Cited by: 2nd item. Devlin et al. (2018) J. Devlin, M. Chang, K. Lee, and K. Toutanova BERT: pre-training of deep bidirectional transformers for language understanding. arXiv:1810.04805. Cited by: §2, §4.1, 6(a). Dong et al. (2023) Q. Dong, L. Li, D. Dai, C. Zheng, Z. Wu, B. Chang, X. Sun, J. Li, and Z. Sui A survey on in-context learning. AI Open. Cited by: §4.1. Edwards and Camacho-Collados (2024) A. Edwards and J. Camacho-Collados Language models for text classification: is in-context learning enough?. In LREC-COLING 2024, Cited by: §1. Fey et al. (2019) M. Fey et al. Fast graph representation learning with pytorch geometric. In ICLR GRL Workshop, Cited by: §A.2. Flamino et al. (2023) J. Flamino, A. Galeazzi, S. Feldman, M. W. Macy, B. Cross, Z. Zhou, M. Serafino, A. Bovet, H. A. Makse, and B. K. Szymanski Political polarization of news media and influencers on twitter in the 2016 and 2020 us presidential elections. Nature Human Behaviour. Cited by: §1. Gallegos et al. (2023) I. O. Gallegos, A. Urman, S. Annagar, C. Kacperski, I. Lupu, and I. Ktena Bias and fairness in large language models: a survey. In FAccT, Cited by: §5.2.2. Gambetti and Han (2024) A. Gambetti and Q. Han AiGen-FoodReview: A Multimodal Dataset of Machine-Generated Restaurant Reviews and Images on Social Media. In ICWSM, Cited by: Table 3. Gerard et al. (2023) P. Gerard, N. Botzer, and T. Weninger Truth social dataset. Cited by: §3.3. Gil de Zúñiga et al. (2012) H. Gil de Zúñiga, N. Jung, and S. Valenzuela Social media use for news and individuals’ social capital, civic engagement and political participation. Journal of Computer-Mediated Communication. Cited by: §1. Gilbert et al. (2009) E. Gilbert T. Bergstrom et al. Blogs are echo chambers: blogs are echo chambers. In HICSS, Cited by: §1. Grootendorst (2022) M. Grootendorst BERTopic: neural topic modeling with a class-based tf-idf procedure. arXiv preprint arXiv:2203.05794. Cited by: §6.2. Gu et al. (2024) J. Gu, X. Jiang, Z. Shi, H. Tan, X. Zhai, C. Xu, W. Li, Y. Shen, S. Ma, H. Liu, et al. A survey on llm-as-a-judge. arXiv:2411.15594. Cited by: 2nd item. Guess et al. (2018) A. Guess, B. Nyhan, B. Lyons, and J. Reifler Avoiding the echo chamber about echo chambers. Knight Foundation. Cited by: §1. Guo et al. (2025) W. Guo, Z. Chen, S. Wang, J. He, Y. Xu, J. Ye, Y. Sun, and H. Xiong Logic-in-frames: dynamic keyframe search via visual semantic-logical verification for long video understanding. In NeurIPS, Cited by: §1. Han et al. (2025) Q. Han, Q. Wang, A. Yoshikawa, and M. Yamamura PulseReddit: A Novel Reddit Dataset for Benchmarking MAS in High-Frequency Cryptocurrency Trading. arXiv preprint. Cited by: Table 3. Henderson et al. (2019) M. Henderson, I. Casanueva, N. Mrkšić, P. Su, I. Vulić, and T. Wen A repository of conversational datasets. In ConvAI Workshop at ACL 2019, Cited by: §3.2, Table 3. Hu et al. (2020) Y. Hu, H. Huang, A. Chen, and X. Mao Weibo-cov: a large-scale covid-19 social media dataset from weibo. In EMNLP Workshop on NLP for COVID-19, Cited by: §3.2, Table 3. Hurst et al. (2024) A. Hurst, A. Lerer, A. P. Goucher, A. Perelman, A. Ramesh, A. Clark, A. Ostrow, A. Welihinda, A. Hayes, A. Radford, et al. GPT-4o system card. arXiv:2410.21276. Cited by: 2nd item. Iyyer et al. (2014) M. Iyyer, P. Enns, J. Boyd-Graber, and P. Resnik Political ideology detection using recursive neural networks. In ACL, Cited by: §2. Jiang et al. (2023) J. Jiang, X. Ren, and E. Ferrara Retweet-bert: political leaning detection using language features and information diffusion on social networks. In ICWSM, Cited by: §2. Joulin et al. (2016a) A. Joulin, E. Grave, P. Bojanowski, M. Douze, H. Jégou, and T. Mikolov FastText.zip: compressing text classification models. arXiv preprint arXiv:1612.03651. Cited by: §A.1.3. Joulin et al. (2016b) A. Joulin, E. Grave, P. Bojanowski, and T. Mikolov Bag of tricks for efficient text classification. arXiv preprint arXiv:1607.01759. Cited by: §A.1.3. Kim et al. (2024) J. Kim, S. Lee, J. Kwon, S. Gu, Y. Kim, M. Cho, J. Sohn, and C. Choi Linq-embed-mistral: elevating text retrieval with improved gpt data through task-specific control and quality refinement. Note: Linq AI Research Blog External Links: Link Cited by: 1st item. Kipf and Welling (2017) T. N. Kipf and M. Welling Semi-supervised classification with graph convolutional networks. In ICLR, Cited by: §5.1.2, Table 8. Küçuk and Can (2020) D. Küçuk and F. Can Stance detection: a survey. ACM CSUR. Cited by: §2. Kumar et al. (2019) S. Kumar, X. Zhang, and J. Leskovec Predicting dynamic embedding trajectory in temporal interaction networks. In ACM SIGKDD, Cited by: §4.2, §5.1.2, Table 8. Kwon et al. (2023) W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, et al. Efficient memory management for large language model serving with pagedattention. In SOSP, Cited by: §A.2. Li et al. (2025) P. Li, P. Song, W. Li, H. Yao, W. Guo, Y. Xu, D. Liu, and H. Xiong See&Trek: training-free spatial prompting for multimodal large language model. In NeurIPS, Cited by: §1. Liu et al. (2019) Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov RoBERTa: a robustly optimized bert pretraining approach. arXiv:1907.11692. Cited by: §2, §4.1. Lo et al. (2021) K. Lo, S. Dai, A. Xiong, et al. Escape from an echo chamber. In W Companion, Cited by: §2. Majid (2023) I. Majid Tweets dataset. External Links: Link Cited by: §5.1.1. Mekacher et al. (2024) A. Mekacher, M. Falkenberg, and A. Baronchelli The koo dataset: an indian microblogging platform with global ambitions. In ICWSM, Cited by: §3.2, Table 3. Milios et al. (2023) A. Milios, S. Reddy, and D. Bahdanau In-context learning for text classification with many labels. arXiv:2309.10954. Cited by: §1. Mohammad et al. (2016) S. M. Mohammad, S. Kiritchenko, P. Sobhani, X. Zhu, and C. Cherry SemEval-2016 task 6: detecting stance in tweets. In SemEval 2016, Cited by: §2. Müller et al. (2020) M. Müller, M. Salathé, and P. E. Kummervold COVID-twitter-bert: a natural language processing model to analyse covid-19 content on twitter. arXiv:2005.07503. Cited by: 6(a). Nguyen et al. (2020) D. Q. Nguyen et al. BERTweet: a pre-trained language model for english tweets. In EMNLP Demo, Cited by: §2. Oliveira et al. (2021) L. S. d. Oliveira, M. S. Amaral, and P. O.S. Vaz-de-Melo Long-term characterization of political communications on social media. In IEEE/WIC/ACM WI, Cited by: §2. Ou-Yang (2014) C. Ou-Yang Newspaper3k: article scraping & curation. Note: Python library. Available at: https://github.com/codelucas/newspaper Cited by: §3.1. Ouyang et al. (2022) L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray, et al. Training language models to follow instructions with human feedback. In NeurIPS, Cited by: §1. Patel et al. (2024) J. Patel, P. Paudel, E. De Cristofaro, G. Stringhini, and J. Blackburn iDRAMA-Scored-2024: A Dataset of the Scored Social Media Platform from 2020 to 2023. In ICWSM, Cited by: §3.2, Table 3. Pérez et al. (2025) J. M. Pérez, P. Miguel, and V. Cotik Exploring large language models for hate speech detection in Rioplatense Spanish. In NAACL Findings, Cited by: §4.1. Piot et al. (2024) P. Piot, P. Martín-Rodilla, and J. Parapar MetaHate: a dataset for unifying efforts on hate speech detection. In ICWSM, Cited by: Table 3. Pope et al. (2023) R. Pope, S. Li, H. Li, M. Durdan, M. Patwary, G. Cheng, M. Casarsa, D. Kalamkar, J. Laurenzo, J. Adkins, et al. Efficiently scaling transformer inference. In MLSys, Cited by: §4.1. Quattrociocchi et al. (2016) W. Quattrociocchi, A. Scala, and C. R. Sunstein Echo chambers on facebook. SSRN. Cited by: §1. Reimers and Gurevych (2019) N. Reimers and I. Gurevych Sentence-bert: sentence embeddings using siamese bert-networks. arXiv:1908.10084. Cited by: §2. Rossi et al. (2020) E. Rossi, B. Chamberlain, F. Frasca, D. Eynard, F. Monti, and M. Bronstein Temporal graph networks for deep learning on dynamic graphs. In ICML GRL Workshop, Cited by: §4.2, §5.1.2, Table 8. Sainburg et al. (2020) T. Sainburg, L. McInnes, and T. Q. Gentner Parametric umap: learning embeddings with deep neural networks for representation and semi-supervised learning. arXiv:2009.12981. Cited by: §4.1. Sanh et al. (2019) V. Sanh, L. Debut, J. Chaumond, and T. Wolf DistilBERT: a distilled version of bert: smaller, faster, cheaper and lighter. arXiv:1910.01108. Cited by: §4.1. Sapiro-Gheiler (2019) E. Sapiro-Gheiler Examining political trustworthiness through text-based measures of ideology. In AAAI, Cited by: §2. Shang et al. (2024) L. Shang, B. Chen, A. Vora, Y. Zhang, X. Cai, and D. Wang SocialDrought: a social and news media driven dataset and analytical platform towards understanding societal impact of drought. In ICWSM, Cited by: Table 3. Sobhani et al. (2016) P. Sobhani, D. Inkpen, and S. Matwin A dataset for stance detection in tweets. In LREC 2016, Cited by: §2. Team (2025) Q. Team Qwen2: the next generation of qwen language models. Technical report Alibaba Cloud Computing Ltd.. External Links: Link Cited by: §5.1.2. Tsai et al. (2024) C. Tsai, Y. Huang, T. Liao, D. F. Salazar Estrada, R. Latifah, and Y. Chen Leveraging conflicts in social media posts: unintended offense dataset. In EMNLP, Cited by: Table 3. Turcan and McKeown (2019) E. Turcan and K. McKeown Dreaddit: a reddit dataset for stress analysis in social media. In EMNLP Workshop on LOUHI, Cited by: §3.2, Table 3. Veličković et al. (2018) P. Veličković, G. Cucurull, A. Casanova, A. Romero, P. Lio, and Y. Bengio Graph attention networks. In ICLR, Cited by: §5.1.2, Table 8. Wang et al. (2018) C. Wang, Q. Liu, R. Wu, E. Chen, C. Liu, X. Huang, and Z. Huang Confidence-aware matrix factorization for recommender systems. In AAAI, Cited by: §2. Wang et al. (2021a) C. Wang, H. Zhu, Q. Hao, K. Xiao, and H. Xiong Variable interval time sequence modeling for career trajectory prediction: deep collaborative perspective. In W, Cited by: §2. Wang et al. (2021b) C. Wang, H. Zhu, P. Wang, C. Zhu, X. Zhang, E. Chen, and H. Xiong Personalized and explainable employee training course recommendations: a bayesian variational approach. ACM TOIS. Cited by: §2. Wang et al. (2020) C. Wang, H. Zhu, C. Zhu, C. Qin, and H. Xiong Setrank: a setwise bayesian approach for collaborative ranking from implicit feedback. In AAAI, Cited by: §2. Wang et al. (2024) L. Wang, N. Yang, X. Huang, L. Yang, R. Majumder, and F. Wei Multilingual e5 text embeddings: a technical report. arXiv:2402.05672. Cited by: 1st item. Wang et al. (2021c) X. Wang, D. Lyu, M. Li, Y. Xia, Q. Yang, X. Wang, X. Wang, P. Cui, Y. Yang, et al. Apan: asynchronous propagation attention network for real-time temporal graph embedding. In SIGMOD, Cited by: §4.2, §5.1.2, Table 8. Wu et al. (2025) W. Wu, Z. Pan, K. Fu, C. Wang, L. Chen, Y. Bai, T. Wang, Z. Wang, and H. Xiong TokenSelect: efficient long-context inference and length extrapolation for LLMs via dynamic token-level KV cache selection. In EMNLP, Cited by: §4.1. Xiao et al. (2020) Z. Xiao, W. Song, H. Xu, Z. Ren, and Y. Sun TIMME: twitter ideology-detection via multi-task multi-relational embedding. In ACM SIGKDD, Cited by: §2. Xin et al. (2025) H. Xin, Y. Sun, C. Wang, and H. Xiong LLMCDSR: enhancing cross-domain sequential recommendation with large language models. ACM TOIS. Cited by: §1. Xu et al. (2025a) Y. Xu, A. Liu, X. Hu, L. Wen, and H. Xiong Mark your LLM: detecting the misuse of open-source large language models via watermarking. In ICLR 2025 Workshop on GenAI Watermarking, Cited by: §1. Xu et al. (2025b) Y. Xu, H. Yao, Z. Guo, P. Li, A. Liu, X. Hu, W. Guo, and H. Xiong You only need 4 extra tokens: synergistic test-time adaptation for LLMs. arXiv preprint arXiv:2510.10223. Cited by: §1. Zhang et al. (2023) X. Zhang, Y. Malkov, O. Florez, S. Park, B. McWilliams, J. Han, et al. TwHIN-bert: a socially-enriched pre-trained language model for multilingual tweet representations at twitter. In ACM SIGKDD, Cited by: §2, §4.1, 6(a). Zhao et al. (2023) W. X. Zhao, K. Zhou, J. Li, T. Tang, X. Wang, Y. Hou, Y. Min, B. Zhang, J. Zhang, Z. Dong, et al. A survey of large language models. arXiv:2303.18223. Cited by: §4.1. Zheng et al. (2023) L. Zheng, W. Chiang, Y. Sheng, S. Zhuang, Z. Wu, Y. Zhuang, Z. Lin, Z. Li, D. Li, E. Xing, et al. Judging llm-as-a-judge with mt-bench and chatbot arena. In NeurIPS, Cited by: 2nd item. Zhou et al. (2022) H. Zhou, D. Zheng, I. Nisa, V. Ioannidis, X. Song, and G. Karypis TGL: a general framework for temporal gnn training on billion-scale graphs. arXiv:2203.14883. Cited by: §A.2.