Paper deep dive
Multiclass Sentiment Analysis for Identifying Political Viewpoints
Girma Yohannis Bade, Olga Kolesnikova, Jose Luis Oropeza, Grigori Sidorov
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 94%
Last extracted: 8/16/2026, 3:45:12 AM
Summary
This paper presents a multiclass sentiment analysis framework for identifying political viewpoints in Tamil social media posts. The authors evaluate two machine learning approaches: XGBoost (using TF-IDF features) and BERT (using BertTokenizer). The models were trained on a dataset from the DravidianLangTech 2025 shared task, consisting of 4,352 training samples and 544 test samples across seven sentiment classes. The XGBoost model achieved a macro F1-score of 0.2835, slightly outperforming the BERT model's score of 0.2806. The study highlights the challenges of classifying complex political discourse in low-resource languages and provides a baseline for future research.
Entities (10)
Relation Signals (7)
XGBoost → achievedscore → 0.2835
confidence 95% · The XGBoost model reaches an F1-score of 0.2835
BERT → achievedscore → 0.2806
confidence 95% · the BERT-based model reaches an F1-score of 0.2806
Tamil → istargetlanguageof → Sentiment Analysis
confidence 95% · classifying the political opinions expressed in Tamil tweets
XGBoost → usesfeatureextraction → TF-IDF
confidence 95% · we employed TF-IDF... for the Logistic Regression... [and XGBoost]
BERT → usesfeatureextraction → BertTokenizer
confidence 95% · we employed... BertTokenizer for the... BERT models
XGBoost → outperformed → BERT
confidence 90% · XGBboost outperformed bert in this usecase.
DravidianLangTech 2025 → provideddatasetfor → Sentiment Analysis
confidence 90% · we utilize the dataset collected as part of the DravidianLangTech 2025 shared task
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:The rapid growth of social media has created vast amounts of political discourse, which provides valuable opportunities to analyze public opinions and identify different political perspectives. Sentiment Analysis (SA) is a core task in Natural Language Processing (NLP) that allows the computational study of attitudes and opinions in textual data, and has become increasingly important for understanding political discourse. In this work, we investigate multiclass sentiment analysis of political view- points on social media, that is to automatically discriminate multiple sentiment classes over political issues and figures. To solve this task we design and evaluate two machine-learning approaches based on XGBoost and BERT. We train and evaluate the models on a labeled dataset of political social media posts using standard classification metrics. The experimental results show that the XGBoost model reaches an F1-score of 0.2835 and the BERT- based model reaches an F1-score of 0.2806 on the test set. These results demonstrate the challenge of classifying complex and contextualized political discourse sentiment and provide a baseline for future research in multiclass political sentiment analysis.
Tags
Links
- Source: https://arxiv.org/abs/2608.11049v1
- Canonical: https://arxiv.org/abs/2608.11049v1
Trouble viewing inline? Open PDF directly →
Full Text
26,967 characters extracted from source content.
Expand or collapse full text
Multiclass Sentiment Analysis for Identifying Political Viewpoints Girma Yohannis Bade, Olga Kolesnikova, Jose Luis Oropeza, Grigori Sidorov Centro de Investigaciones en Computación (CIC), Instituto Politécnico Nacional (IPN), Miguel Othon de Mendizabal, Ciudad de México, 07320, México Abstract The rapid growth of social media has created vast amounts of political discourse, which pro- vides valuable opportunities to analyze pub- lic opinions and identify different political per- spectives. Sentiment Analysis (SA) is a core task in Natural Language Processing (NLP) that allows the computational study of attitudes and opinions in textual data, and has become increasingly important for understanding po- litical discourse. In this work, we investigate multiclass sentiment analysis of political view- points on social media, that is to automatically discriminate multiple sentiment classes over po- litical issues and figures. To solve this task we design and evaluate two machine-learning ap- proaches based on XGBoost and BERT. We train and evaluate the models on a labeled dataset of political social media posts using standard classification metrics. The experi- mental results show that the XGBoost model reaches an F1-score of 0.2835 and the BERT- based model reaches an F1-score of 0.2806 on the test set. These results demonstrate the challenge of classifying complex and contextu- alized political discourse sentiment and provide a baseline for future research in multiclass po- litical sentiment analysis. K ̨ eywords:Multiclass,Sentiment analysis,Political view,NLP 1 Introduction Social media has gained increasing popularity and is a valuable source of diverse, expressive, and real-time political discourse, especially X (former Twitter). These platforms provide an opportunity for users to freely express their thoughts, feelings, opinions and comments on a variety of topics. Sen- timent analysis (SA) plays an important role in areas such as political affairs, social awareness, in- ternational conflicts, movie reviews, education sys- tems, and feedback on products (Bade et al., 2026). It enables the extraction of valuable insights from tweets to understand public sentiment and opinions, and has consequently been increasingly adopted by companies, government agencies, and other organi- zations to correct their stands (Ramanathan et al., 2021; Kolesnikova et al.). SA is a fundamental task in natural language processing (NLP) that aims to identify and classify opinions and attitudes expressed in textual data into predefined categories. In the political domain, sentiment analysis is particularly valuable for un- derstanding public opinion, capturing diverse polit- ical perspectives, identifying societal concerns, and supporting evidence-based policymaking. Beyond politics, the ability to automatically identify sen- timents toward products, organizations, services, and other entities has become increasingly impor- tant across a wide range of applications (Mullen and Malouf, 2006; Yigezu et al.). For instance, Mullen and Malouf (2006); Yigezu et al. (2024a) demonstrated the potential of NLP- based sentiment analysis for examining political trends, complementing traditional opinion polls, identifying political bias in news and other ostensi- bly objective texts, and characterizing the views as- sociated with particular texts and individuals. Such analyses can provide valuable insights for targeted communication and outreach activities, including political campaigns, public engagement, donation appeals, and petition initiatives. Building on these applications, this paper presents a multiclass sentiment analysis framework for identifying and classifying political viewpoints expressed in social media discourse for Tamil lan- guage. Most speakers of Dravidian languages (Tamil, Kannada, Malayalam, Telugu, Tulu and others) are found in South India and Sri Lanka represent a rich linguistic and cultural heritage of the region. The workshop is organized by Dra- vidianLangTech to promote research in tackling the challenges arising from the lack of progress in language technology for these languages. Every arXiv:2608.11049v1 [cs.CL] 11 Aug 2026 year, it provides a gold standard dataset for the se- lected tasks and organizes a workshop (Duraphe et al., 2022; Coelho et al., 2023; Bade et al., 2024a; Mersha et al., 2024). We leveraged dataset provided for Dravidian- LangTech@NAACL 2025 (Bade et al., 2025c) to classifying the political opinions expressed in Tamil tweets into seven distinct classes, such as: 1) Substantiated, 2) Sarcastic, 3) Opinionated, 4) Positive, 5) Negative, 6) Neutral, and 7) None of the above. the main contribution of this paper is summarized as follows: •We review the existing literature on sentiment analysis in the political domain and examine how sentiment analysis can be applied to mon- itor and analyze political discourse. •We investigate and select appropriate fea- ture extraction methods for the given datasets and prepare the resulting representations for model training. • We select state-of-the-art language models, fine-tune them on the target datasets, and eval- uate their performance on held-out test data using standard evaluation metrics. The paper is structured as follows: Section 2 discusses the related works. Section 3 mentions the methodology including data sets, approaches and experimentation. Finally Section 4 discusses the results obtained from the experiments. 2 Related works SA is an interesting area of research in NLP (Bade et al., 2024a,b; Yigezu et al., 2024b; Bade et al., 2024c). It is about a classification task that catego- rizes or predicts the linguistic input features based on the patterns trained through AI algorithms dur- ing the training phase. The concept can be adopted for numerous languages while employing AI ap- proaches such as transformer-based, deep learning, and machine learning (Yigezu et al., 2023d). According to Chan et al. (2023), many sentiment analysis problems, including as emotion detection, cross-domain sentiment classification, multimodal sentiment analysis, aspect-based sentiment analy- sis, and multilingual sentiment analysis, are useful for knowledge adaptation strategies. Basically, sen- timent analysis is a broad umbrella for many NLP downstream tasks like hope speech, abusive and hate detection, stress identification, emotional anal- ysis, and so on (Yigezu et al., 2023c,a). For instance,Ghosh et al. (2023); Mersha et al. (2025), attempted the first sentiment and emotional state recognition challenge, which triggered other researchers to contribute in under-resourced lan- guages. A transformer-based multitask framework has been utilized by them to identify emotions and detect sentiment in code-mixed datasets. SA is not limited to text datasets; it can also operate on custom datasets that include political and film reviews. Code-mixed speech-sentiment classifi- cation has also been attempted in (Keshav et al., 2023), utilizing the 3-shot, few-shot learning (FSL) framework and a fully connected neural network (FCNN) model. Although sentiment analysis (SA) is widely researched for languages with abun- dant resources(Bade and Seid, 2018; Yigezu et al., 2023b), developing trustworthy systems for low- resource languages is still difficult because there is a lack of training data for this kind of work(Bade and Afaro, 2018). 3 Methodology This section provides comprehensive details about the dataset and the experimental settings adopted in this study. It also describes the overall system ar- chitecture, data encoding and preprocessing proce- dures, as well as the data format and configuration used for the selected models. 3.1 Datasets For research in the NLP domain, a well-articulated collection of datasets are driving fuel to produce insightful language models. For this study, we utilize the dataset collected as part of the Dravid- ianLangTech 2025 shared task (Roy et al., 2025). The dataset is organized into three subsets: training, development, and test sets. The training and de- velopment sets are accompanied by corresponding sentiment labels, whereas the test set is unlabeled. A summary of these dataset splits is presented in Table 1. Table 1: Dataset Statistics NoDatasetsSample Sizes 1Train Dataset4,352 2Development Dataset544 3Test Dataset (Unlabeled)544 As we can see from the Table 1, it consisted three datasets. The training dataset is main dataset we use it for training our chosen algorithm. In case when we choose a supervised machine learn- ing algorithm, it learns the pattern from this train- ing data. The second set is development dataset, which is used to tune the behavior of our model dur- ing experimentation. It’s mostly used to validate the model performance before applying the test data. In works like shared task where the model performance is tested by third party (organizer), validating the model with the development dataset gives the confidence before sending the final test predictions to the organizer. The test dataset is one that determines the final performance of the model. This data is separate and never been seen during training. The Table 2 shows the class label distribution of training and development datasets. Table 2: Class label distribution statistics Labels # Count in Train # Count in Dev Opinionated1,361153 Sarcastic790115 Neutral63784 Positive57569 Substantiated41252 Negative40651 None of the above17120 3.2 Preprocessing The training, development, and test datasets were subjected to a standardized preprocessing proce- dure to improve the consistency and quality of the input data. The primary objectives of this prepro- cessing step were to remove punctuation marks, emojis, and user mentions, which may introduce noise or contribute limited semantic information in wrong way (Bade et al., 2025a). In particular, user mentions were removed to minimize the in- fluence of author-specific or account-specific infor- mation on model predictions. The built-in Python re (regular expression) module was employed to identify and remove usernames and punctuation marks from the text. Emojis were also removed to ensure a consistent textual representation across the datasets. The same preprocessing procedure was applied uniformly to all dataset splits to maintain consistency between training, development, and test data and to prevent discrepancies arising from different preprocessing strategies. 3.3 Feature Extraction Since practically all AI algorithms operate on nu- merical data, it is essential to appropriately encode language inputs into their corresponding numer- ical representations (Bade et al., 2024d, 2025b). The process of transforming textual input into a numerical form is commonly referred to as data encoding or feature extraction (Bade et al., 2024d). Although various feature extraction techniques are available, we employed TF-IDF and BertTokenizer for the Logistic Regression and BERT models, re- spectively. This approach ensures that the input text is represented in a format compatible with the re- quirements of each model. Furthermore, the same feature extraction strategy was consistently applied across the relevant datasets to maintain compara- bility between the experimental settings. 3.4 Model Selection Once the NLP part is ready, the next step is to choose and apply AI algorithms (Yigezu et al., 2023b). Thus, we have started our experiment with one of the traditional machine learning algorithms known as XGBClassifier. XGBoost is a widely used traditional machine learning algorithm that has demonstrated strong performance across a vari- ety of classification tasks. To make it effective, we employed IF-IDF to extract the features from the text data. In our second experiment, we chose the bert- base-uncased transformer model. For this model, there is its own BertTokenizer to tokenize and con- vert text data into numeric form. Table 3 presents its hyperparameters. Table 3: BERT Hyperparameters HyperparametersValues Learning Rate1e-5 Evaluation StrategyEpoch Epochs5 Batch Size32 Activation@output levelsoftmax As we can see from the table 3, the learning rate indicates the number of times the execution took place to improve the model performance. The epoch refers to a complete pass through the entire training data set by the learning algorithm. Thus, we set the epoch to be 5,.i.e the execution did pass 5 complete times. The batch size refers to dividing the total data size into 32 and bringing the divided batch one a time for the execution. This helps the execution to be fast. Figure 1: The work flow of proposed model 3.5 Results and Discussion The experimental results provide an empirical com- parison of the machine learning and transformer- based approaches considered in this study. We first employed the XGBClassifier algorithm (Chang et al., 2022) to train the model on the provided training dataset and evaluated its performance on a separate test set. The XGBoost-based approach achieved a macro F1 score of 0.2835, indicating its performance in identifying the target classes under the given experimental setting. Similarly, we conducted a second experiment us- ing a BERT-based transformer model. The BERT model achieved a macro F1 score of 0.2806 on the same evaluation setting. The results of both approaches are presented in Tables 4 and 6, re- spectively, providing a comparative view of their performance. Table 4: The result statistics of XGBboost on test data Labels PrReF1Sup Opinionated0.17390.08700.115946 Sarcastic0.12000.08570.100070 Neutral0.75000.72000.734725 Positive0.36070.64330.4622171 Substantiated0.32560.18670.237375 Negative0.30950.24530.2737106 None of0.13330.03920.060651 Accuracy0.3309 Macro Avg0.31040.28670.2835544 Weighted Avg0.29570.33090.2934544 Table 5: The BERT model on test data LabelsPrReF1Sup Opinionated0.22220.08700.125046 Sarcastic0.12500.07140.090970 Neutral0.72730.64000.680925 Positive0.36210.63740.4619171 Substantiated0.26320.26670.264975 Negative0.34250.23580.2793106 None0.14290.03920.061551 Accuracy0.3327 Macro Avg0.31220.28250.2806544 Weighted Avg0.29850.33270.2955544 In Tables 4 and 6, the column headings Pr, Re, F1, and Sup represent precision, recall, F1 score, and support, respectively. Support indicates the number of true instances in each labels. These metrics, along with accuracy, macro average, and weighted average, are used to evaluate performance. Among these, the macro average F1 score is often the most significant metric, as it relies on the har- monic mean of recall and precision, providing a balanced measure of performance. Therefore, our work is ranked based primarily on the F1 score values. Table 6: The number of true instances and predicted instances’ statistics in the labels. Labels#Actual#PredictedRemark Opinionated46305over Sarcastic7084over Neutral2550over Positive17143under Substantiated7515under Negative10623under None5124under Total544544— Table 6 makes a figurative comparison between actual number of lables and predicted number of labels. For instance, the label ’Opinionated’ had 46 instances in test set which are manually annotated but our model (XGBboost) made 305 prediction for it. Thus, the predicted values are greater than actual and hence it is marked as ’over’ in the remark column. The reason for this might be the algorithm has learned more number of this class, see Table 2. Similarly, the label ’Positive’ had 171 instances but our model’s prediction is 43. Therefore, the number of predicted instances are less than actual, and hence it marked as ’under’. Figures 2 and 3 visualize the results in confusion matrix. Figure 2: Confusion matrix that shows results of XGB- boost algorithm In Figure 2, the diagonal elements are corret predictions. For example the label ’Positive’ has been correctly classified 110 times, holding the first position. Next, the label ’Negative’ has been clas- sified moderately 26 times. In third position, the label ’Neutral’ is predicted 18 times correctly. On the other hand, the model misclassified the label ’Negative’ as positive 56 times,holding first posi- tion. Next, ’Substantiated’ is classified as ’positive’ wrongly 43 times. In third osition, the label ’sarcas- tic’ is misclassified as positive 38 times. Similarly, ’None of the above’ and ’Opinionated’ are misclas- sified as 29 times and 24 times respectively. Finally, the cell with 0 values indicate that the label in the XY coordinate have never been mixed up. For instance, the label ’Opinionated’ has never been classified to ’None of the above’ and vice versa. Likewise, ’Neutral’ is not misclassified to ’Non of the above’ and ’Non of the above’is not misclassi- fied to ’Neutral’. All others that are not explicitly mentioned can be defined in similar fashion. Figure 3: Confusion matrix that shows results of BERT model In Figure 3, diagonal elements are correct predic- tions. For instance, ’Positive’ is correctly classified the most (109 times), suggesting the model per- forms best in identifying Positive sentiment. ’Neu- tral’ (16 times) and ’Negative’ (25 times) have mod- erate correct classifications. ’None of the Above’ label (2 times) is poorly classified, meaning the model struggles with identifying this category. The most common misclassifications ’Opinionated’ is misclassified as ’Positive’ (23 times). Sarcastic is frequently misclassified as ’Positive’ (37 times) and ’Substantiated’ (12 times). ’Substantiated’ is also mostly misclassified as ’Positive’ (43 times). ’Negative’ has 56 cases misclassified as ’Positive’. ’None of the Above’ is often confused with ’Posi- tive’ (27 times). 4 Conclusion and Future Work In this task, we have developed social media clas- sifying models and evaluated their performance using various metrics. The developed model is able to classify social media posts into into seven mul- ticlasses as expected. The two AI algorithms we employed here are XGBboost and Bert model. As our result show, XGBboost outperformed bert in this usecase. As a direction for future research, similar studies should be conducted for a wider range of languages, particularly those that are underrepresented in exist- ing research, since political opinions and public dis- course are increasingly expressed through multilin- gual social media platforms. Extending the analysis to additional languages would provide a broader understanding of political opinion across different linguistic and cultural contexts. Furthermore, the performance of the proposed models could be im- proved by incorporating and comparing additional machine learning and deep learning algorithms for the languages considered in this study, as well as by increasing the size and diversity of the datasets. Such extensions could contribute to more robust, generalizable, and language-independent models for political opinion analysis. Acknowledgments The work was done with partial support from the Mexican Government through the grant A1- S-47854 of CONACYT, Mexico, and grants 20241816, 20241819, and 20240951 of the Sec- retaría de Investigación y Posgrado of the Insti- tuto Politécnico Nacional, Mexico. The authors thank the CONACYT for the computing resources brought to them through the Plataforma de Apren- dizaje Profundo para Tecnologías del Lenguaje of the Laboratorio de Supercómputo of the INAOE, Mexico and acknowledge the support of Microsoft through the Microsoft Latin America PhD Award. Limitation and Ethics Statement Since Tamil is a language with limited resources and the model was trained using a tiny dataset, the performance observed may not generalize well to all unseen data. Despite the challenges of lim- ited resources and competition, our model demon- strated strong performance in classifying multiclass political sentiments in Tamil social media posts. Furthermore, our work adheres to the ethical prin- ciples outlined for computational research and pro- fessional conduct 1 . References Girma Bade, Olga Kolesnikova, José Oropeza, Grig- ori Sidorov, and Mesay Yigezu. 2025a. Amado at 1 https://w.aclweb.org/portal/content/ acl-code-ethics semeval-2025 task 11: Multi-label emotion detection in amharic and english data. In Proceedings of the 19th International Workshop on Semantic Evaluation (SemEval-2025), pages 1406–1410. Girma Bade, Olga Kolesnikova, Grigori Sidorov, and José Oropeza. 2024a. Social media hate and offen- sive speech detection using machine learning method. In Proceedings of the Fourth Workshop on Speech, Vision, and Language Technologies for Dravidian Languages, pages 240–244. Girma Yohannis Bade and Akalu Assefa Afaro. 2018. Object oriented software development for artificial intelligence. American Journal of Software Engineer- ing and Applications, 7(2):22–24. Girma Yohannis Bade, O Koleniskova, José Luis Oropeza, Grigori Sidorov, and Kidist Feleke Bergene. 2024b. Hope speech in social media texts using trans- former. In Proceedings of the Iberian Languages Evaluation Forum (IberLEF 2024), co-located with the 40th Conference of the Spanish Society for Natu- ral Language Processing (SEPLN 2024), CEURWS. org. Girma Yohannis Bade, Olga Kolesnikova, and Jose Luis Oropeza. 2024c. Evaluating the quality of data: Case of sarcasm dataset. Girma Yohannis Bade, Olga Kolesnikova, José Luis Oropeza, and Grigori Sidorov. 2024d. Lexicon-based language relatedness analysis. Procedia Computer Science, 244:268–277. Girma Yohannis Bade, Olga Kolesnikova, Jose Luis Oropeza, and Moein Shahiki Tash. 2026. Evaluat- ing the capability of base and large-scale language models for multilingual sarcasm detection. PeerJ Computer Science, page e3584. Girma Yohannis Bade, Jose Luis Oropeza, and Olga Kolesnikova. 2025b. Pragmatic generalization in llms: Insights from fine-tuning and evaluating on multilingual sarcasm. In Mexican International Con- ference on Artificial Intelligence, pages 218–230. Springer. Girma Yohannis Bade and Hussien Seid. 2018. Devel- opment of longest-match based stemmer for texts of wolaita language. vol, 4:79–83. Girma Yohannis Bade, Muhammad Tayyab Zamir, Olga Kolesnikova, José Luis Oropeza, Grigori Sidorov, and Alexander Gelbukh. 2025c. Girma@ dravid- ianlangtech 2025: Detecting ai generated product reviews. In Proceedings of the Fifth Workshop on Speech, Vision, and Language Technologies for Dra- vidian Languages, pages 133–138. Jireh Yi-Le Chan, Khean Thye Bea, Steven Mun Hong Leow, Seuk Wai Phoong, and Wai Khuen Cheng. 2023. State of the art: a review of sentiment analy- sis based on sequential transfer learning. Artificial Intelligence Review, 56(1):749–780. Chih-Chi Chang, Yu-Zhen Li, Hui-Ching Wu, and Ming- Hseng Tseng. 2022. Melanoma detection using xgb classifier combined with feature extraction and k- means smote techniques. Diagnostics, 12(7):1747. Sharal Coelho, Asha Hegde, G Kavya, and Hosa- halli Lakshmaiah Shashirekha. 2023. Mucs@ dra- vidianlangtech2023: Malayalam fake news detection using machine learning approach. In Proceedings of the Third Workshop on Speech and Language Tech- nologies for Dravidian Languages, pages 288–292. Ankita Duraphe, Ratnavel Rajalakshmi, and Antonette Shibani. 2022. Dlrg@ dravidianlangtech-acl2022: Abusive comment detection in tamil using multilin- gual transformer models. In Proceedings of the Sec- ond Workshop on Speech and Language Technologies for Dravidian Languages. Association for Computa- tional Linguistics (ACL). Soumitra Ghosh, Amit Priyankar, Asif Ekbal, and Push- pak Bhattacharyya. 2023. Multitasking of senti- ment detection and emotion recognition in code- mixed hinglish data. Knowledge-Based Systems, 260:110182. S Keshav, G Jyothish Lal, and B Premjith. 2023. Multi- modal approach for code-mixed speech sentiment classification. In Advances in Signal Processing, Embedded Systems and IoT: Proceedings of Seventh ICMEET-2022, pages 553–563. Springer. Olga Kolesnikova, Mesay Gemeda Yigezu, Alexander Gelbukh, Selam Abitte, and Grigori Sidorov. De- tecting multilingual hate speech targeting immigrants and women on twitter. Journal of Intelligent & Fuzzy Systems, (Preprint):1–10. Melkamu Abay Mersha, Jugal Kalita, et al. 2024. Semantic-driven topic modeling using transformer- based embeddings and clustering algorithms. Proce- dia Computer Science, 244:121–132. Melkamu Abay Mersha, Mesay Gemeda Yigezu, and Jugal Kalita. 2025. Evaluating the effectiveness of xai techniques for encoder-based language models. Knowledge-Based Systems, page 113042. Tony Mullen and Robert Malouf. 2006. A preliminary investigation into sentiment analysis of informal po- litical discourse. In AAAI spring symposium: com- putational approaches to analyzing weblogs, pages 159–162. Vallikannu Ramanathan, T Meyyappan, and SM Thama- rai. 2021.Sentiment analysis: an approach for analysing tamil movie reviews using tamil tweets. Re- cent Advances in Mathematical Research and Com- puter Science, 3:28–39. Billodal Roy, Souvik Bhattacharyya, Pranav Gupta, and Niranjan Kumar. 2025. Lexilogic@ dravidian- langtech 2025: Political multiclass sentiment analysis of tamil x (twitter) comments and sentiment analysis in tamil and tulu. In Proceedings of the Fifth Work- shop on Speech, Vision, and Language Technologies for Dravidian Languages, pages 557–561. Mesay Yigezu, Olga Kolesnikova, Grigori Sidorov, and Alexander Gelbukh. 2024a. Habesha@ dravidian- langtech 2024: Detecting fake news detection in dra- vidian languages using deep learning. In Proceedings of the Fourth Workshop on Speech, Vision, and Lan- guage Technologies for Dravidian Languages, pages 156–161. Mesay Gemeda Yigezu, Girma Yohannis Bade, Olga Kolesnikova, Grigori Sidorov, and Alexander Gel- bukh. 2023a. Multilingual hope speech detection using machine learning. Mesay Gemeda Yigezu, Girma Yohannis Bade, At- nafu Lambebo Tonja, Olga Kolesnikova, Grigori Sidorov, and Alexander Gelbukh. 2023b. Bilingual word-level language identification for omotic lan- guages. In International Conference on Advances of Science and Technology, pages 63–77. Springer. Mesay Gemeda Yigezu,Selam Kanta,Olga Kolesnikova, Grigori Sidorov, and Alexander Gelbukh. 2023c.Habesha@ dravidianlangtech: Abusive comment detection using deep learning approach. In Proceedings of the Third Workshop on Speech and Language Technologies for Dravidian Languages, pages 244–249. Mesay Gemeda Yigezu, Olga Kolesnikova, Alexander Gelbukh, and Grigori Sidorov. Odio-bert: Evaluating domain task impact in hate speech detection. Journal of Intelligent & Fuzzy Systems, (Preprint):1–12. Mesay Gemeda Yigezu, Olga Kolesnikova, Grig- ori Sidorov, and Alexander Gelbukh. 2023d. Transformer-based hate speech detection for multi- class and multi-label classification. Mesay Gemeda Yigezu,Melkamu Abay Mer- sha, Girma Yohannis Bade, Jugal Kalita, Olga Kolesnikova, and Alexander Gelbukh. 2024b. Ethio-fake: Cutting-edge approaches to combat fake news in under-resourced languages using explainable ai. arXiv preprint arXiv:2410.02609.