Paper deep dive
Echoes of Unrest: A Multimodal NLP Framework for Early Warning of Fake News and Violence-Driven Mob Activity
Md. Maruf Bangabashi, Tahmid Hasan, Golam Mahmud, Md. Mostafijur Rahman, Md. Toufiqur Rahman, Jahanur Biswas
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/7/2026, 4:26:15 AM
Summary
The paper presents a multilingual, multimodal NLP framework named AttentionFusion for early detection of fake news and violence-driven mob activity. It integrates text (XLM-RoBERTa), visual (CLIP), and metadata (sarcasm, geospatial) features via a multi-head attention mechanism. Trained on a fused dataset of 138,256 Bangla and English samples, the model achieves 98% test accuracy, demonstrating strong potential for real-time misinformation monitoring and geospatial hotspot visualization to prevent social unrest.
Entities (12)
Relation Signals (10)
AttentionFusion model → detects → Fake News
confidence 96% · framework for early detection of misinformation and violence-prone dynamics
AttentionFusion model → achieves → 98% Test Accuracy
confidence 95% · Experiments on a stratified 30% subset achieved 98% test accuracy
AttentionFusion model → integrates → XLM-RoBERTa
confidence 95% · Text written in either English or Bangla is processed using XLM-RoBERTa
AttentionFusion model → integrates → CLIP
confidence 95% · images are passed through a CLIP-based feature extractor
AttentionFusion model → leverages → Geospatial Metadata
confidence 94% · enhanced with auxiliary features such as sarcasm and geospatial metadata
AttentionFusion model → utilizes → Multi-head attention mechanism
confidence 93% · These representations are then fused through a multi-head self-attention mechanism
Fused Dataset → combines → Kaggle Fake News
confidence 90% · The proposed dataset combines four major sources: ... Kaggle Fake News
Fused Dataset → combines → Sarcasm Headlines
confidence 90% · The proposed dataset combines four major sources: ... Sarcasm Headlines
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Rapid growth in social media has transformed global communication by enabling fast information exchange, but it has also accelerated the spread of misinformation. Fake news, manipulated content, and provocative narratives are increasingly linked to social unrest, political instability, and mob violence. Incidents in South Asia and elsewhere demonstrate how false information disseminated via platforms such as Facebook and WhatsApp can trigger real-world harm, often spreading faster than fact-checking efforts can respond. To address this challenge, this chapter presents a multilingual, multimodal Natural Language Processing (NLP) framework for early detection of misinformation and violence-prone dynamics. A fused dataset of 138,256 Bangla and English samples was created by combining multiple benchmark datasets. The framework integrates XLM-RoBERTa for multilingual text representation, CLIP for visual embedding, and a multi-head attention mechanism for multimodal fusion, enhanced with auxiliary features such as sarcasm and geospatial metadata. Experiments on a stratified 30% subset achieved 98% test accuracy with strong precision and recall. The outcomes show the efficacy of multimodal approaches in early misinformation detection and highlight the added value of geospatial signals for anticipating real-world escalation.
Tags
Links
- Source: https://arxiv.org/abs/2607.02734v1
- Canonical: https://arxiv.org/abs/2607.02734v1
Trouble viewing inline? Open PDF directly →
Full Text
31,116 characters extracted from source content.
Expand or collapse full text
Taylor and Francis Book Chapter ©2001 by CRC Press LLC arXiv:2607.02734v1 [cs.CL] 2 Jul 2026 Contents Chapter 1 ■ Echoes of Unrest: A Multimodal NLP Framework for Early Warning of Fake News and Violence-Driven Mob Activity1 Md. Maruf Bangabashi, Tahmid Hasan, Golam Mahmud, Md. Mostafijur Rahman, Md. Toufiqur Rahman , and Jahanur Biswas* 1.1 INTRODUCTION2 1.2 RELATED WORK3 1.2.1 Text-Based Approaches3 1.2.2 Multimodal Approaches4 1.2.3 Geospatial Awareness in Recognizing False Information4 1.3 MATERIALS AND METHODS5 1.3.1 Dataset5 1.3.2 User Analysis and Design Considerations6 1.3.3 Proposed Method7 1.3.4 Architecture Development8 1.4 EXPERIMENTAL RESULTS AND DISCUSSION10 1.4.1 Experimental Setup10 1.4.2 Performance Evaluation11 1.5 CONCLUSION AND FUTURE WORK13 i C H A P T E R 1 Echoes of Unrest: A Multimodal NLP Framework for Early Warning of Fake News and Violence-Driven Mob Activity Md. Maruf Bangabashi Department of Computer Science and Engineering, Dhaka International University, Dhaka- 1212, Bangladesh. Tahmid Hasan Department of Computer Science and Engineering, Bangladesh University of Engineering and Technology, Dhaka-1000, Bangladesh. Golam Mahmud Department of Computer Science and Engineering, Dhaka International University, Dhaka- 1212, Bangladesh. Md. Mostafijur Rahman Department of Computer Science and Engineering, Dhaka International University, Dhaka- 1212, Bangladesh. Md. Toufiqur Rahman Department of Computer Science and Engineering, Dhaka International University, Dhaka- 1212, Bangladesh. Jahanur Biswas* Department of Computer Science and Engineering, Dhaka International University, Dhaka- 1212, Bangladesh. CONTENTS 1.1 Introduction 1.2 Related Work 1.2.1 Text-Based Approaches 1 1.2.2 Multimodal Approaches 1.2.3 Geospatial Awareness in Recognizing False Information 1.3 Materials and Methods 1.3.1 Dataset 1.3.2 User Analysis and Design Considerations 1.3.3 Proposed Method 1.3.4 Architecture Development 1.4 Experimental Results and Discussion 1.4.1 Experimental Setup 1.4.2 Performance Evaluation 1.5 Conclusion and Future Work R apid growth in social media has transformed global communication by enabling fast information exchange, but it has also accelerated the spread of misinfor- mation. Fake news, manipulated content, and provocative narratives are increasingly linked to social unrest, political instability, and mob violence. Incidents in South Asia and elsewhere demonstrate how false information disseminated via platforms such as Facebook and WhatsApp can trigger real-world harm, often spreading faster than fact-checking efforts can respond. To address this challenge, this chapter presents a multilingual, multimodal Natural Language Processing (NLP) framework for early detection of misinformation and violence-prone dynamics. A fused dataset of 138,256 Bangla and English samples was created by combining multiple benchmark datasets. The framework integrates XLM-RoBERTa for multilingual text representation, CLIP for visual embedding, and a multi-head attention mechanism for multimodal fusion, enhanced with auxiliary features such as sarcasm and geospatial metadata. Experi- ments on a stratified 30% subset achieved 98% test accuracy with strong precision and recall. The outcomes show the efficacy of multimodal approaches in early misinfor- mation detection and highlight the added value of geospatial signals for anticipating real-world escalation. 1.1 INTRODUCTION Social media usage has grown exponentially and changed the way people communi- cate around the world. They make the exchange of information fast and provide a means of sharing misleading information. In recent years, false news and photographs, sarcastic or provocative comments have been associated with social unrest, political instability, and mob violence in many countries. The dangers of online misinformation are evident in incidents of mob lynchings in South Asia, fueled by false content spread via Facebook and WhatsApp. Similar crises have happened in most regions across the world, where the information that spreads is more likely than the verification procedures. Current fake news detection systems have made considerable progress through text-based natural language processing methods. However, these systems remain lim- ited in three important ways. First, they are largely language-focused and rely pri- ©2001 by CRC Press LLC marily on textual data, thereby overlooking low-resource languages such as Bangla [1]. Second, they frequently fail to capture multimodal misinformation that combines text with images and stylistic cues to strengthen deceptive claims [2]. Third, they gen- erally do not consider geospatial factors, which restrict the identification of localized misinformation hotspots that may provoke mob actions [3]. Recent studies have shown the effectiveness of multimodal misinformation detec- tion frameworks [1, 2, 3, 4]. Nevertheless, no unified framework adequately combines multilinguality, multimodality, sarcasm awareness, and geospatial signals for violence monitoring. This chapter therefore addresses the following question: how can a multi- lingual, multimodal NLP framework be developed to accurately detect fake news and misinformation across diverse languages, modalities, and geospatial contexts while also providing early warning for the escalation of misinformation into violence-driven mob activity? The proposed framework addresses this gap by integrating multilingual text, vi- sual cues, stylistic signals such as sarcasm, and geospatial metadata into a unified mis- information detection system. It is particularly relevant for low-resource language con- texts, including Bangla, where misinformation can have serious social consequences but remains understudied in existing systems. Geospatial visualizations of misinformation hotspots can be valuable for policy- makers by helping them identify regions at risk of violence or unrest and enabling proactive intervention. The system can also help security officials allocate resources more effectively on the basis of the geographical distribution of misinformation. Fact- checkers and journalists may benefit from real-time misinformation warnings that support timely verification before publication. As a result, the proposed framework provides a real-time, automated misinformation detection mechanism that combines multimodal analysis with geospatial awareness and thus supports informed decision- making and public safety. This chapter’s main contributions are as follows: • A heterogeneous fused dataset of more than 138,000 records is constructed by integrating four resources: Bangla and English misinformation-related resources with differing modality coverage. • An attention-based multimodal fusion model is introduced to jointly integrate textual, visual, stylistic, and geospatial features for misinformation detection. • The proposed framework achieves an accuracy of 98%, demonstrating strong predictive performance on the fused corpus. This is how the rest of the chapter is structured. In Section 2, the suggested framework is placed in context, and relevant work is reviewed. The materials and techniques, such as system design and dataset construction, are described in Section 3. The experimental setup and evaluation findings are shown in Section 4. The chapter is finally concluded, and future research possibilities are outlined in Section 5. 1.2 RELATED WORK ©2001 by CRC Press LLC Several research studies have been done on the identification of fake news and misin- formation, taking into account the increasing prevalence of social media sites. How- ever, the initial approaches focused mainly on the use of text data. Specifically, Singh et al. [1] utilized text data to detect fake news written in English. Despite contribut- ing positively to the field, these studies cannot be considered comprehensive, since they neglect other useful data sources, including images and geographic metadata. 1.2.1 Text-Based Approaches The conventional approaches employed in the identification of misinformation con- tent were based largely on text-based characteristics like linguistics, sentiments, and lexical aspects. In their research, Singh et al. [1] applied linguistic and sentiment cues to determine any misleading information. Similarly, Segura-Bedmar and Alonso- Bartolome [5] suggested a system that considered text-based characteristics to detect disinformation. Though the techniques may be quite efficient in some situations, their weaknesses cannot be overlooked. According to Alam et al. [2], these techniques often overlook significant visual characteristics and cannot perform efficiently in instances where sarcasm is present. 1.2.2 Multimodal Approaches A number of studies in recent years have explored multimodal solutions for the detec- tion of fake news. In Qi et al. [6] an entity-centric multimodal approach to integrate text and visual data was offered to improve detection outcomes. Also, Xu et al. [7] offered MDAM3, a deep learning method to integrate text, images, and videos. Even though these methods represent significant achievements in terms of fake news detec- tion, they require high computational power and tend to ignore geospatial attributes. This becomes especially important because of potential misinformation leading to violence in certain regions. Moreover, Bansal et al. [8] presented an approach to caption-aware multimodal analysis of low-resource Indic languages. On the other hand, Sharma and Arya [9] developed a multimodal multiclass approach to detecting Hindi fake news using contrastive learning. However, even though this development represents significant progress, its practical applicability is limited due to language specificity and a lack of geospatial information about the violent tendencies of certain pieces of news. Furthermore, Wang et al. [10] suggested a multimodal fake news detection method using cross-modal contrastive learning. They also noted that multimodal fake news detection models can outperform unimodal baselines. Chen et al. [11] also examined multimodal approaches to fake news detection with respect to robustness. However, there are still limitations regarding multilingual fake news detection and sarcasm- aware models. Unlike those, the multimodal framework developed in this chapter uses multilingual text and geospatial metadata. 1.2.3 Geospatial Awareness in Recognizing False Information ©2001 by CRC Press LLC Many existing studies that focus on identifying fake news primarily consider text and images, neglecting the role played by spatial information. However, some recent studies have emphasized the importance of spatial cues. Jing et al. [12] developed a method that leverages spatial and textual information to identify the presence of spatial metadata that helps recognize the spread of misinformation and predict future flashpoints before they cause societal unrest. Alam et al. [2] also emphasized the potential value of geospatial signals in identify- ing misinformation hotspots and tracking the spread of misleading narratives. Wang et al. [10] further observed that regional context can improve fake news identification by revealing local tendencies in misinformation dissemination. Nevertheless, even re- cent work in low-resource language settings, such as MMCFND [8] and MMHFND [9], remains language-specific and does not incorporate violence-aware geospatial mod- elling. 1.3 MATERIALS AND METHODS 1.3.1 Dataset The proposed dataset combines four major sources: BanFakeNews (50,000 Bangla articles) [13], Kaggle Fake News (45,000 English articles) [14], Sarcasm Headlines (30,000 headlines) [15], and Twitter Multimodal Fake News (13,000 posts containing text, images, and geotags) [16]. A standardized format consisting of headline, content, image URL, label, sarcasm, latitude, longitude, and timestamp was applied across all datasets. The merged collection contains 138,256 Bangla and English records. Statistics on the integrated fusion corpus are presented in Table 1.1. The fused corpus is heterogeneous rather than uniformly multimodal. BanFak- eNews, Kaggle Fake News, and Sarcasm Headlines contribute text-only records, whereas the Twitter subset contributes records with image references in addition to text. To preserve a unified model input format, missing image inputs were treated as unavailable visual modality and replaced during training with a placeholder im- age. Similarly, sarcasm was retained as an auxiliary stylistic signal. Because the four source datasets differ in both modality coverage and label semantics, the fused dataset should be interpreted as a heterogeneous benchmark for multimodal-aware misinfor- mation analysis rather than a fully image-complete multimodal corpus. Because the source datasets differ in annotation purpose and label semantics, a unified binary training space was constructed for model development. The original fake/real labels from the misinformation datasets were retained as the core veracity signal, while the sarcasm dataset was incorporated as a stylistic component and its sarcasm annotation was also preserved separately through the metadata field. This design allows the framework to capture rhetorical cues that may co-occur with misleading or provocative content, while also acknowledging that sarcasm and factual veracity are not identical constructs. Figure 1.1 presents the distribution of samples across data sources and languages in the fused corpus. Figure 1.2 illustrates the distribution of class labels and sarcasm indicators across the benchmark, highlighting the multilingual and heterogeneous nature of the collection. ©2001 by CRC Press LLC Table 1.1 Statistics of datasets integrated into the fusion corpus. DatasetRecords LanguageModalitiesLabels BanFakeNews [13]49,977BanglaTextReal/Fake Fake News Detection [14] 44,898EnglishTextReal/Fake Sarcasm Dataset [15]28,619EnglishTextSarcastic/Non-sarcastic Twitter and Weibo [16]14,762EnglishText+ImageReal/Fake Fusion Dataset138,256 Multilingual Text + Partial Image + Metadata Unified Binary Label + Style Metadata Figure 1.1 Distribution of samples across the four integrated data sources and the two languages represented in the fused corpus. Figure 1.3 displays a word cloud created from the most common terms found in the fused dataset’s headlines. The visualization highlights commonly occurring terms from both Bangla and English content, reflecting the multilingual nature of the corpus. 1.3.2 User Analysis and Design Considerations Journalists, non-governmental organizations, security agencies, and policymakers are among the target users of the proposed system. Policymakers require early warnings to plan timely interventions, while hotspot maps may help security personnel focus resources on areas susceptible to instability. Fact-checkers and journalists need to identify false information rapidly in order to ensure reporting accuracy. Therefore, real-time interpretability, multilingual processing capability, and multimodal integra- tion were central design considerations in the proposed framework. 1.3.3 Proposed Method The suggested approach combines linguistic, visual, stylistic, and geospatial informa- tion to identify misleading claims and monitor their potential escalation into violence. Figure 1.4 illustrates the six main stages of the framework: integrating data sources, preprocessing, feature extraction, projection into a common latent space, attention- based fusion and classification, and geospatial hotspot visualization. ©2001 by CRC Press LLC Figure 1.2 Distribution of class labels and sarcasm annotations across the fused mul- tilingual corpus. A combination of BanFakeNews [13], Kaggle Fake News [14], Sarcasm Headlines [15], and multimodal tweets [16] was merged into one unified collection containing text, images, sarcasm cues, and geospatial details. Before feature extraction, the data were standardized through preprocessing, including text normalization, image resizing, and metadata encoding. XLM-RoBERTa was used to process text, CLIP to extract visual representations, and numerical encoding to represent metadata. These features were projected into a common latent space and combined through an attention-based fusion mechanism that dynamically weighed multiple cues. A dense layer then generated real or fake predictions using a softmax function. These predic- tions were subsequently aggregated by location and visualized as hotspot maps in order to provide early warning signals. The full process for misinformation detection and unrest monitoring is summarized in Table 1.2. 1.3.4 Architecture Development Figure 1.5 shows the structure of the AttentionFusion model. Instead of treating inputs separately, the framework combines text, images, and metadata such as geolo- cation and sarcasm indicators. Text written in either English or Bangla is processed using XLM-RoBERTa, while images are passed through a CLIP-based feature extrac- ©2001 by CRC Press LLC Figure 1.3 Word cloud of the most frequent words in the headlines of the fused mul- tilingual dataset. Table 1.2 Proposed attention-based multimodal fusion framework for fake news de- tection and violence monitoring. Input: Dataset D = T, I, M, where T denotes text, I denotes image, and M = [s, g] includes sarcasm flag s and geospatial coordinates g. Output: Predicted label ˆy ∈ 0, 1 (Real/Fake), with optional geospatial hotspot map. Step 1: Preprocessing. Normalize and tokenize T; resize and normalize I; encode sarcasm flag s; clean geospatial metadata g. Step 2: Feature Extraction. Compute X t = f text (T) using XLM-RoBERTa; compute X i = f image (I) using CLIP; compute X m = f meta ([s, g]). Step 3: Projection. H t = W t X t , H i = W i X i , H m = W m X m . Step 4: Attention-Based Fusion. Z = Attention([H t , H i , H m ]). Step 5: Classification. ˆy = Softmax(W Z + b). Step 6: Optional Geospatial Visualization. When valid coordinates are available, Aggregate g | ˆy = 1, cluster results, and visualize misinformation intensity. tor. Metadata are converted into numerical form and projected into a shared hidden ©2001 by CRC Press LLC Figure 1.4 Workflow of the proposed multimodal misinformation detection framework. space through a dense neural layer. These representations are then fused through a multi-head self-attention mechanism that dynamically adjusts feature importance across modalities. For records without available images, a placeholder visual input was used during preprocessing so that a consistent multimodal interface could be maintained across heterogeneous data sources. After fusion, the integrated representation is processed through a dense neural network with ReLU activations and a final softmax classification layer trained to distinguish fake and genuine news. When coordinate metadata are available and suf- ficiently informative, predicted misinformation cases may also be aggregated spatially to support exploratory hotspot analysis. The architectural configuration of the proposed AttentionFusion model, including output dimensions and trainable parameters of its main components, is shown in Table 1.3. 1.4 EXPERIMENTAL RESULTS AND DISCUSSION 1.4.1 Experimental Setup To evaluate the effectiveness of the suggested attention-based fusion framework, ex- periments were conducted on a stratified random sample comprising 30% of the fused ©2001 by CRC Press LLC Figure 1.5 Proposed architecture of the attention-based fusion model. Table 1.3 Architectural parameter configuration of the proposed AttentionFusion model. LayerOutput Shape Trainable Parameters XLM-RoBERTa(batch, 768)278043648 CLIP Image Encoder (batch, 512)11176512 Metadata FC(batch, 256)1024 Text Projection(batch, 256)196864 Image Projection(batch, 256)131328 Multi-head Attention (3, batch, 256)263168 Classifier(batch, 2)66306 dataset. From the original 138,256 records, 41,477 samples were used for model de- velopment. This sampling strategy was adopted to maintain computational feasibil- ity while preserving class diversity and multilingual coverage across the integrated benchmark. The experiment was conducted with a fixed random seed of 42. There were 29,033 training instances, 6,222 validation instances, and 6,222 test instances after the sam- pled dataset was split into training, validation, and test subsets in a stratified 70:15:15 ratio. XLM-RoBERTa was used to tokenize textual inputs with a maximum sequence length of 128. For 10 epochs, the model was trained with a batch size of 8 and a learn- ©2001 by CRC Press LLC ing rate of 2× 10 −5 , a dropout of 0.2, a hidden dimension of 256, and the Adam optimizer. Visual features were extracted using CLIP and combined with textual and meta- data features through the proposed attention-based fusion mechanism. Metadata were represented through sarcasm, latitude, and longitude fields. When an image was un- available, a placeholder visual input was used to preserve a consistent model interface across heterogeneous source datasets. Performance was assessed using a confusion ma- trix, precision, recall, F1-score, ROC curve, AUC, and training-validation accuracy and loss curves. Because the fused benchmark integrates heterogeneous sources with different an- notation purposes and modality coverage, the reported results should be interpreted as a benchmark evaluation on the constructed corpus rather than as a claim of uni- versal real-world generalization. 1.4.2 Performance Evaluation As shown in Figure 1.6(a), the confusion matrix provides a clear demonstration of the predictive dynamics of the model on a class-by-class basis. It correctly identified 4,051 Class 0 instances and 2,046 Class 1 instances. Meanwhile, 84 instances of Class 0 were incorrectly identified as Class 1 and 41 instances of Class 1 were incorrectly identified as Class 0. The small number of misclassifications in both categories indicates that the proposed framework is highly consistent and reliable in distinguishing between fake news and violence-motivated mob behaviour. A more detailed quantitative evaluation is presented in Table 1.4. For Class 0, the model obtained a precision of 0.99, a recall of 0.98, and an F1-score of 0.98 on 4,135 test samples. For Class 1, the respective values were 0.96, 0.98, and 0.97 on 2,087 samples. The overall test accuracy was 0.98, and both the macro-average F1-score and weighted-average F1-score were approximately 0.98. These results show that the model has robust and balanced predictive accuracy across both classes, rather than merely favoring the dominant category. Table 1.4 Classification report of the proposed model on the test set. ClassPrecision Recall F1-Score Support Class 0 (Real)0.990.980.984135 Class 1 (Fake)0.960.980.972087 Accuracy0.980.980.980.98 Macro Avg.0.970.980.976222 Weighted Avg.0.980.980.986222 Although the overall training process was stable, the fluctuation in validation loss during later epochs suggests the possibility of mild overfitting after mid-training. This behaviour did not substantially reduce held-out test performance, but it indi- cates that future versions of the framework would benefit from stricter regularization, ©2001 by CRC Press LLC Figure 1.6 Performance evaluation of the proposed multimodal fusion framework for detecting fake news and violence-driven mob activity: (a) confusion matrix showing class-wise prediction outcomes, (b) training and validation accuracy and loss curves over 10 epochs, and (c) ROC curve with an AUC of 0.998, demonstrating strong discriminative ability and stable convergence. early stopping, and additional robustness checks. In addition, the present study eval- uates a single fusion architecture; future work should include direct comparison with text-only, text-plus-metadata, and alternative multimodal baselines in order to more clearly quantify the advantage of the suggested model. The discriminative strength of the proposed framework is further supported by the ROC curve shown in Figure 1.6(c), where the model achieved an area under the curve (AUC) of approximately 0.99. This high AUC value indicates excellent separability ©2001 by CRC Press LLC between the two classes across a wide range of threshold settings, confirming the robustness of the classifier beyond a single operating point. Figure 1.6(b) shows the training and validation accuracy and loss curves, provid- ing additional insight into the model’s training behaviour. Validation accuracy in- creased from 94.49% to 98.10%, while training accuracy rose from 91.33% to 99.01% over 10 epochs. Simultaneously, training loss decreased from 0.1871 to 0.0277, and validation loss generally declined from 0.1117 to 0.0641. These trends indicate stable convergence and effective parameter optimization throughout training. Overall, the experimental findings show the effectiveness of the suggested multi- modal framework and validate its suitability for accurately identifying fake news and mob activity motivated by violence. Comparative performance against recent related studies is presented in Table 1.5. Table 1.5 State-of-the-art comparative analysis of the proposed method with recent related studies StudyLanguageModalitiesMethodAccuracyLimitations Ref. [1]EnglishText + VisualMultimodal Analysis 85.25%English-only Ref. [6]Chinese+EnglishText + VisualEntity Fusionup to 97.5% Weak sarcasm cues Ref. [7]EnglishText + Image + VideoMDAM385.3%High compute cost Ref. [10]EnglishText + VisualRobustness Evalup to 92.5% Adversarial fragility Ref. [8]IndicText + VisualCaption-Aware99.6%Low-resource Ref. [9]HindiText + VisualContrastive Learning ∼98.6%Language-specific Proposed Model Bangla+English+Meta Text+Image+Sarcasm+Geo Fusion Model98.0%Violence-aware 1.5 CONCLUSION AND FUTURE WORK In this chapter, we present a model developed through multi-language and multi- modal inputs that is used to detect misinformation likely to instigate violence. The model achieves accurate classification by using textual information, visual informa- tion, sarcasm, and geographical data. One of the major advantages of this model is that it can help map misinformation hot spots, which makes it applicable for law en- forcement officers, media personnel, and policymakers in locating areas where there is a higher probability of violence. Despite the encouraging outcomes, there are still a few restrictions. First, the merged dataset contains varied content, but only a portion of it is accompanied by images, making it difficult to fully explore the multimodal aspects of the dataset. Secondly, each individual source dataset has different purposes for its labels, in- cluding sarcasm. Third, the experiments were conducted on a stratified 30% sample of the corpus for computational reasons rather than on the full dataset. Fourth, the present evaluation focuses on internal benchmark performance and does not yet include source-isolated validation, external real-world testing, or formal statistical significance analysis. These limitations mean that the reported results should be in- terpreted as encouraging but still preliminary. Future work will focus on constructing more fully aligned fused datasets, eval- uating the framework under stricter split strategies, and comparing it against stronger unimodal and multimodal baselines. Extending the framework to low- ©2001 by CRC Press LLC resource and underrepresented languages remains a major priority. Incorporating additional modalities such as audio and video, together with more reliable location- aware metadata and broader social media signals, may further strengthen perfor- mance. Another important direction is the deployment of scalable real-time mon- itoring pipelines and interactive dashboards. With these extensions, the proposed framework may evolve into a more comprehensive misinformation tracking system for early detection and intervention. Bibliography [1] Singh, V.K., Ghosh, I. and Sonagara, D. (2021) Detecting fake news stories via multimodal analysis. Journal of the Association for Information Science and Technology, 72(1), 3–17. [2] Alam, F., Cresci, S., Chakraborty, T., Silvestri, F., Dimitrov, D., Da San Martino, G., Shaar, S., Firooz, H. and Nakov, P. (2021) A survey on multi- modal disinformation detection. arXiv preprint arXiv:2103.12541. Available at: https://doi.org/10.48550/arXiv.2103.12541. [3] Tufchi, S., Yadav, A. and Ahmed, T. (2023) A comprehensive survey of multi- modal fake news detection techniques: advances, challenges, and opportunities. International Journal of Multimedia Information Retrieval, 12(2), 28. [4] Mostafa, M., Almogren, A.S., Al-Qurishi, M. and Alrubaian, M. (2024) Modality deep-learning frameworks for fake news detection on social networks: a system- atic literature review. ACM Computing Surveys, 57(3), 1–50. [5] Segura-Bedmar, I. and Alonso-Bartolome, S. (2022) Multimodal fake news de- tection. Information, 13(6), 284. [6] Qi, P., Cao, J., Li, X., Liu, H., Sheng, Q., Mi, X., He, Q., Lv, Y., Guo, C. and Yu, Y. (2021) Improving fake news detection by using an entity-enhanced framework to fuse diverse multimodal clues. In: Proceedings of the 29th ACM International Conference on Multimedia, p. 1212–1220. [7] Xu, Q., Du, H., Łukasik, S., Zhu, T., Wang, S. and Yu, X. (2025) MDAM3: A misinformation detection and analysis framework for multitype multimodal media. In: Proceedings of the ACM Web Conference 2025, p. 5285–5296. [8] Bansal, S., Singh, N.S., Dar, S.S. and Kumar, N. (2024) MMCFND: Multimodal multilingual caption-aware fake news detection for low-resource Indic languages. arXiv preprint arXiv:2410.10407. [9] Sharma, R. and Arya, A. (2024) MMHFND: Fusing modalities for multimodal multiclass Hindi fake news detection via contrastive learning. ACM Transactions on Asian and Low-Resource Language Information Processing, 23(11), 1–25. ©2001 by CRC Press LLC [10] Wang, L., Zhang, C., Xu, H., Xu, Y., Xu, X. and Wang, S. (2023) Cross-modal contrastive learning for multimodal fake news detection. In: Proceedings of the 31st ACM International Conference on Multimedia, p. 5696–5704. [11] Chen, J., Jia, C., Zheng, H., Chen, R. and Fu, C. (2023) Is multi-modal neces- sarily better? Robustness evaluation of multi-modal fake news detection. IEEE Transactions on Network Science and Engineering, 10(6), 3144–3158. [12] Jing, J., Wu, H., Sun, J., Fang, X. and Zhang, H. (2023) Multimodal fake news detection via progressive fusion networks. Information Processing & Manage- ment, 60(1), 103120. [13] Hossain, M.Z., Rahman, M.A., Islam, M.S. and Kar, S. (2020) BanFakeNews: A dataset for detecting fake news in Bangla. arXiv e-prints. [14] Jain, P. Fake News Detection Dataset. Kaggle dataset. Available at: https: //w.kaggle.com/datasets/jainpooja/fake-news-detection/data (ac- cessed 1 March 2026). [15] Misra, R. and Arora, P. (2023) Sarcasm detection using news headlines dataset. AI Open, 4, 13–18. [16] Ojo, A. (2025) Fake news multimodal datasets (Twitter and Weibo). Figshare dataset. Available at: https://doi.org/10.6084/m9.figshare.28516655.v2. ©2001 by CRC Press LLC