Paper deep dive
BanglaMemeEvidence: A Multimodal Benchmark Dataset for Explanatory Evidence Detection in Bengali Memes
Fatema Tuj Johora Faria, Mukaffi Bin Moin, Md. Mahfuzur Rahman, Pronay Debnath, Asif Iftekher Fahim, Faisal Muhammad Shah
Intelligence
Status: succeeded | Model: Gemma-4-26B-A4B | Prompt: intel-v1 | Confidence: 93%
Last extracted: 7/7/2026, 10:45:56 AM
Summary
This paper introduces BanglaMemeEvidence, a multimodal benchmark dataset of 2,917 Bengali memes annotated with contextual evidence and relevance scores. It proposes MemeEvidenceDetect, a hybrid task for explanatory evidence detection, and BengaliMemeEvidenceNet, a multimodal framework integrating early, late, and intermediate fusion techniques to achieve an F1 score of 0.74.
Entities (10)
Relation Signals (8)
BengaliMemeEvidenceNet â achievesmetric â F1 Score
confidence 97% ¡ Our experiments demonstrate the effectiveness of BengaliMemeEvidenceNet, achieving an F1 score of 0.74.
BanglaMemeEvidence â supportstask â MemeEvidenceDetect
confidence 96% ¡ To support this task, we present BanglaMemeEvidence, a curated dataset of 2,917 Bengali memes
BengaliMemeEvidenceNet â implements â MemeEvidenceDetect
confidence 95% ¡ we propose BengaliMemeEvidenceNet, a hybrid multimodal framework that integrates textual and visual features for comprehensive meme representation.
MemeEvidenceDetect â usesmetric â Relevance Score
confidence 94% ¡ This task involves detecting relevance scores, with 0 indicating âNot relevant,â 1 representing âPartially relevant,â and 2 signifying âRelevant.â
BanglaMemeEvidence â coverscategory â Sports
confidence 93% ¡ We have taken Politics, Sports, Entertainment, Education, Technology, and Others as meme categories.
BanglaMemeEvidence â coverscategory â Politics
confidence 93% ¡ We have taken Politics, Sports, Entertainment, Education, Technology, and Others as meme categories.
BengaliMemeEvidenceNet â employstechnique â Early Fusion
confidence 91% ¡ we explore multiple fusion techniques, including Early Fusion, Late Fusion, and Intermediate Fusion.
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Memes have become influential communication tools on social media, combining viral visuals with concise messaging to convey impactful ideas. While substantial research has examined the affective dimensions of memes, key challenges such as detecting harmful content, identifying cyberbullying, and performing accurate sentiment analysis remain critical, largely due to the need for deeper contextual understanding. In this paper, we introduce MemeEvidenceDetect, a hybrid task aimed at analyzing a meme and its contextual information to identify specific sentences that explain or elucidate its meaning and humor. To support this task, we present BanglaMemeEvidence, a curated dataset of 2,917 Bengali memes, emphasizing its significance as a resource for the Bangla language. Each meme is annotated with natural language explanations, including Meme OCR, Meme Context, and Evidence Sentences, alongside relevance scores that reflect the relationship between a meme and its corresponding annotations. To address the gap in dynamically inferring a meme's context, we propose BengaliMemeEvidenceNet, a hybrid multimodal framework that integrates textual and visual features for comprehensive meme representation. Our experiments demonstrate the effectiveness of BengaliMemeEvidenceNet, achieving an F1 score of 0.74. To the best of our knowledge, this is the first study to focus on evidence detection in Bengali memes, marking a notable step forward in the analysis of memes in low-resource languages.
Tags
Links
- Source: https://arxiv.org/abs/2607.03981v1
- Canonical: https://arxiv.org/abs/2607.03981v1
Trouble viewing inline? Open PDF directly â
Full Text
44,628 characters extracted from source content.
Expand or collapse full text
Accepted at 6th International Conference on Innovations in Computational Intelligence and Computer Vision (ICICV 2026). BanglaMemeEvidence: A Multimodal Benchmark Dataset for Explanatory Evidence Detection in Bengali Memes Fatema Tuj Johora Faria 1 , Mukaffi Bin Moin 1 , Md. Mahfuzur Rahman 1 , Pronay Debnath 1 , Asif Iftekher Fahim 1 , and Faisal Muhammad Shah 1 Ahsanullah University of Science and Technology, Dhaka-1208, Bangladesh fatema.faria142@gmail.com, mukaffi28@gmail.com, pronaydebnath99@gmail.com, mahim1066@gmail.com, fahimthescientist@gmail.com, faisal.cse@aust.edu Abstract. Memes have become influential communication tools on so- cial media, combining viral visuals with concise messaging to convey impactful ideas. While substantial research has examined the affective dimensions of memes, key challenges such as detecting harmful content, identifying cyberbullying, and performing accurate sentiment analysis remain critical, largely due to the need for deeper contextual under- standing. In this paper, we introduce MemeEvidenceDetect, a hybrid task aimed at analyzing a meme and its contextual information to iden- tify specific sentences that explain or elucidate its meaning and humor. To support this task, we present BanglaMemeEvidence, a curated dataset of 2,917 Bengali memes, emphasizing its significance as a re- source for the Bangla language. Each meme is annotated with natural language explanations, including Meme OCR, Meme Context, and Evi- dence Sentences, alongside relevance scores that reflect the relationship between a meme and its corresponding annotations. To address the gap in dynamically inferring a memeâs context, we propose BengaliMemeEv- idenceNet, a hybrid multimodal framework that integrates textual and visual features for comprehensive meme representation. Our experiments demonstrate the effectiveness of BengaliMemeEvidenceNet, achieving an F1 score of 0.74. To the best of our knowledge, this is the first study to focus on evidence detection in Bengali memes, marking a notable step forward in the analysis of memes in low-resource languages. Keywords: Meme analysis¡ Bengali memes¡ Early fusion¡ Late fusion ¡ Multimodal learning¡ Multimodal fusion techniques 1 Intoduction Social media has become a key way for people to communicate, changing how we interact in society. The content shared on social media comes in many forms, including text, audio, and images, or a mix of them. Memes, often shared with humor or sarcasm, are a common example. While memes help spread complex arXiv:2607.03981v1 [cs.CL] 4 Jul 2026 2F.T.J. Faria et al. ideas about society, culture, or politics, they often lack the context needed for full understanding, which is important for both humans and computers to interpret them accurately. [1] [2]. Relevance_Score: 2 Meme_Category: Politics Meme_OCR: āϤāĻŋāĻŽ āĻŋāĻ āĻāĻŽāĻžāϰ āĻŋāĻŦā§āϰāĻžāϧ⧠āĻžāĻĨ Meme_Context: ⧍ā§Ļ⧍ā§Ē āĻ āĻŋāύāĻŦāĻžāĻā§āύ āĻāĻā§āĻžāĻŽā§ āϞā§āĻ āϏāϰāĻāĻžāϰ āĻĄāĻžāĻŋāĻŽ āĻžāĻĨ āĻāĻŦāĻ āĻŋāĻŦā§āϰāĻžāϧ⧠āĻžāĻĨāϰ āύāĻžā§āĻŽ āĻŋāĻŦāĻŋāĻ āϧāϰā§āύāϰ āĻŽā§āύāĻžāύā§āύ āĻŋāĻĻā§ā§ āĻĨāĻžāĻā§āϞāĻ āĻŽā§āϞāϤ āϤāĻžāϰāĻž āϏāĻŦāĻžāĻ āύā§āĻāĻžāϰ āĻŽā§āύāĻžāύā§āύ āĻžāĻĨāĨ¤ āĻāϰ āĻŽāĻžāϧā§āĻŽ āĻŽā§āϞāϤ āĻāĻāĻžāĻŋāϤāĻ āĻŋāĻŦāώā§ā§āĻ āĻāĻāĻ āĻŋāύāĻŦāĻžāĻāύ āĻā§ā§āĻžāĻāύ āĻā§āϰ āĻĒā§āύāϰāĻžā§ āĻŽāϤāĻžā§ āĻāϏāĻžāϰ āĻāĻž āĻā§āϰā§āĻ āϏāϰāĻāĻžāϰāĨ¤ āϤāĻžāϰāĻž āϏāĻŦāĻžāĻ āύā§āĻāĻžāϰ āĻŽā§āύāĻžāύā§āύ āĻžāĻĨāĨ¤ Evidence_Sentence: āϤāĻžāϰāĻž āϏāĻŦāĻžāĻ āύā§āĻāĻžāϰ āĻŽā§āύāĻžāύā§āύ āĻžāĻĨ Fig. 1. An illustration of MemeEvidenceDetect, a hybrid task where, given a meme, manually transcribed Bengali text, and a description of the memeâs context, the meme is supported by evidence sentences. These sentences are tagged with relevance scores, where a score of 2 denotes âRelevantâ. The recent surge in meme dissemination has led to a growing body of re- search on meme analysis, focusing on tasks such as multimodal hate speech detection from Bengali memes and texts [3], the development of datasets for Ben- gali abusive meme classification [4], multimodal cyberbullying meme detection from social media using deep learning approaches [5], and multimodal analysis of memes for sentiment extraction [6]. Other studies have worked on multimodal offensive meme classification using transformers [7], sarcasm detection in typo- graphic memes [8], and multi-task learning frameworks for multi-modal sarcasm, sentiment, and emotion analysis [9]. Despite the extensive work done in this field, there has been a noticeable lack of emphasis on multimodal evidence detection, especially concerning Bangla memes, which poses a significant challenge. To address this gap, we present innovative solutions aimed at automating the extraction of contextual evidence, thereby enhancing meme accessibility. Introducing a pioneering framework called BengaliMemeEvidenceNet, our approach focuses on identifying sentences within the provided context that can potentially elucidate the memeâs meaning. Figure 1 presents a meme categorized under Politics, featuring Bengali text that translates to âBut you are my opposing candidate.â This meme con- textualizes the 2024 elections, depicting a complex political scenario in which the Awami League leadership strategically nominates candidates from both its own ranks and rival parties. This strategic move seeks to influence world opinion and BanglaMemeEvidence: A Multimodal Benchmark Dataset3 strengthen domestic authority by creating the perception of democratic choice while concealing underlying political strategies. The goal of this research is to analyze this multimodal meme in order to uncover and investigate evidence of political interference during electoral processes. The embedded evidence sentence in the meme provides direct insight into the governmentâs strategy: intentional candidate nominations across the political spectrum to shape global perceptions and consolidate local power. This study seeks to explore methodologies for iden- tifying such evidence within related contextual documents. By doing so, it aims to elucidate the complex dynamics of communication in modern environments across diverse subjects and themes depicted in memes. To summarize, our main contributions are as follows: âĻ MemeEvidenceDetect is an hybrid task that predicts and analyzes the relevance of sentences to a given meme, aiming to clarify its meaning and humor. This task involves detecting relevance scores, with 0 indicating âNot relevant,â 1 representing âPartially relevant,â and 2 signifying âRelevant.â âĻ To bridge the gap in Bangla meme evidence research, we introduce a novel dataset called BanglaMemeEvidence, designed to facilitate the automatic detection of contextual evidence in such memes. The dataset comprises gold standard human annotations for 2,917 Bengali memes. The Bengali text from the meme image is manually transcribed to ensure accurate content representation. âĻ We introduce BengaliMemeEvidenceNet, a hybrid multimodal frame- work aimed at identifying evidence for memes from their related contexts. To achieve this, we employ various fusion techniques to integrate text and image features. 2 Related Work 2.1 Meme Analysis The study by Karim et al. [3] constructed a dataset integrating textual and vi- sual elements to tackle hate speech detection. Their study showcased the efficacy of Conv-LSTM for textual analysis and DenseNet-161 for visual interpretation, alongside highlighting the effectiveness of XLM-RoBERTa in textual analysis. In a parallel study [4], delved into abusive memes with the BanglaAbuseMeme dataset, favoring multimodal strategies, particularly CLIP. While XLMR showed promise in textual analysis and ViT in visual, it was CLIP that achieved the high- est macro score. Transitioning to cyberbullying detection, Ahmed et al. [5] pro- posed a VGG16-BiLSTM model, achieving notable accuracy in detecting harm- ful content within Bengali memes. Meanwhile, Alluri and Krishna [6] explored meme classification using the Memotion dataset, employing ViT for visuals and RoBERTa for textual analysis, alongside a transformer-based image captioning model for contextual understanding. 4F.T.J. Faria et al. 2.2 Visual Question Answering (VQA) Recent advancements in Bengali Visual Question Generation (VQG) and Vi- sual Question Answering (VQA) underscore the need for further exploration in these areas. Hasan et al. [10] have made significant strides in Bengali VQG by employing a transformer-based approach and introducing novel models, such as image-category and image-answer-category variants, which have demonstrated superior question generation capabilities. In contrast, Islam et al. [11] focused on Bengali VQA by adapting existing datasets and employing a hybrid model that combines CNN and bi-LSTM layers. Meanwhile, Rafi et al. [12] contributed to Bengali VQA by developing a human-annotated Bengali VQA dataset and proposing a Top-Down Attention-based model. Collectively, these studies high- light the need for continued research in VQG and VQA for underrepresented lan- guages, showcasing the potential of modern CNN architectures and pretrained models in advancing image-based question generation and answering. 3 BanglaMemeEvidence Dataset Annotation 3.1 Meme Collection Our meme collection process involved a comprehensive search across a diverse range of sources, ensuring a rich and varied dataset. We scoured popular so- cial media platforms such as Facebook, Instagram, and X (formerly known as Twitter) to capture the latest trends and viral content. Additionally, we delved into various online websites, forums, and communities known for their vibrant meme cultures. To ensure the datasetâs breadth and relevance, we cast our net wide, encompassing memes spanning multiple thematic categories. We explored content related to Entertainment, where memes often riff on celebrity culture, movies, and television shows, providing a lighthearted and humorous take on popular media. In the realm of Politics, we curated memes that reflect the dy- namic landscape of political discourse, capturing satirical commentary, politi- cal humor, and memes that encapsulate key events and figures in the political sphere. Our collection extended to Sports, where memes capture the passion, drama, and camaraderie of sporting events, as well as the lighter side of sports culture through memes celebrating iconic moments and poking fun at sporting rivalries. Education-themed memes offered insights into the world of academia, student life, and the challenges and quirks of learning environments, resonating with students and educators alike. In the realm of Technology, memes provided a humorous lens through which to explore the ever-evolving world of gadgets, software, and digital culture, offering witty observations on the quirks and frus- trations of technology users. 3.2 Context Document Curation In curating the contextual corpus corresponding to the memes collected, we em- ployed a multifaceted approach. For some topics, we relied on Bangla Wikipedia BanglaMemeEvidence: A Multimodal Benchmark Dataset5 Relevance_Score: 2 Meme_Category: Entertainment Meme_OCR: āϝāĻāύ āĻāĻ āĻĢāϏāĻŦā§ā§āĻ āĻāϰāĻŋāĻŦ āύāĻžāĻŽ āĻŦāĻšāĻžāϰ āĻā§āϰ: āĻāĻĒāĻŋāύ āϤāĻž āĻāϞā§āĻāϞāĻžāĻā§āĻž āĻŦā§āĻšā§ āϝāĻžāĻā§āĻŦāύ!! Meme_Context: āĻŽāĻžāύā§āώ āĻŋāύā§āĻā§āĻ āϧāĻžāĻŋāĻŽāĻ āĻĻāĻāĻžā§āύāĻžāϰ āĻāύ āϏāĻžāĻļāĻžāϞ āĻŋāĻŽāĻŋāĻĄā§āĻžā§āϤ āĻāϰāĻŋāĻŦ āύāĻžāĻŽ āĻŦāĻšāĻžāϰ āĻā§āϰāĨ¤ āĻāĻžāĻš āĻā§ āĻāύ āύ⧠āϧ⧠āϞāĻžāĻ āĻĻāĻāĻžā§āύāĻžāϰ āĻāύāĨ¤ āϤāĻžāĻ āĻāĻāĻāĻžā§āύ āĻāĻāĻž āĻā§āϰ āĻŦāϞāĻž āĻšā§ā§ā§āĻ āĻāĻĒāĻŋāύ āĻāϞā§āĻāϞāĻžāĻā§āĻž āĻŦā§āĻšā§ āϝāĻžāĻā§āĻŦāύāĨ¤ Evidence_Sentence: āĻāĻĒāĻŋāύ āĻāϞā§āĻāϞāĻžāĻā§āĻž āĻŦā§āĻšā§ āϝāĻžāĻā§āĻŦāύ Relevance_Score: 1 Meme_Category: Education Meme_OCR: āϝāĻāύ āϰāĻžāϤ ā§ŠāĻāĻž āĻāĻŦāĻ āϤāĻŋāĻŽ āϤāĻžāĻŽāĻžāϰ āĻ āĻžāϏāĻžāĻāύā§āĻŽ āĻāϰā§āϤā§āĻāĻž Meme_Context: āĻāĻŽāϰāĻž āĻ āϞāϏāϤāĻžāϰ āĻāĻžāϰā§āĻŖ āĻ āĻāĻŽā§āϤāĻž āĻĒā§āĻžā§āϞāĻāĻž āĻāĻŋāϰ āύāĻž āϏāĻŽā§ā§āϰ āĻĒā§āĻžā§āϞāĻāĻž āϏāĻŽā§ā§ āĻāĻŋāϰ āύāĻžāĨ¤ āĻāĻāύ āϏāĻŽā§āĻŽā§āϤāĻž āĻĒā§āĻžā§āϞāĻāĻž āύāĻž āĻāϰāĻžā§āϤ āϰāĻžāϤ āĻā§āĻ āϏāĻžāĻŦāĻŋāĻŽā§āĻāϰ āĻā§āĻāϰ āϰāĻžā§āϤ āĻāϏāĻžāĻāύā§āĻŽ āĻāĻŋāϰāĨ¤ āϰāĻžāϤ āĻŋāϤāύāĻāĻž āĻŦāĻžā§āĻ āĻ āĻžāϏāĻžāĻāύā§āĻŽ āĻāϰā§āϤā§āĻ, āϤāĻāύ āĻā§ā§āĻŽāϰ āĻāĻžāϰā§āύ āĻāĻžāύ āĻāĻžā§āĻ āĻŽā§āύāĻžā§āϝāĻžāĻ āĻāϰā§āϤ āĻĒāĻžā§āϰ āύāĻžāĨ¤ Evidence_Sentence: āϰāĻžāϤ āĻŋāϤāύāĻāĻž āĻŦāĻžā§āĻ āĻ āĻžāϏāĻžāĻāύā§āĻŽ āĻāϰā§āϤā§āĻ Relevance_Score: 1 Meme_Category: Others Meme_OCR: āĻŋāϰāĻāĻļāĻžāĻā§āĻžāϞāĻž āϝāĻāύ āĻ āϝāĻĨāĻž āĻā§āϏ āĻā§āϰ āĻāĻ āϝāĻžā§āĻŦāύ āϞ āĻāĻŋāĻŽ : āϏāĻžāĻāĻĨ āĻāĻžāĻŋāϰā§āĻž āϝāĻžā§āĻŦāĻž, āϝāĻžā§āĻŦāύ? Meme_Context: āĻŋāϰāĻžāĻā§āĻžāϞāĻžā§āĻĻāϰ āĻāĻžāĻ āϤāĻžāϰāĻž āĻŽāĻžāύā§āώā§āĻ āĻāĻžāĻāĻžāϰ āĻŋāĻŦāĻŋāύāĻŽā§ āĻāĻ āĻāĻžā§āĻāĻž āĻĨā§āĻ āĻ āύ āĻāĻžā§āĻāĻžā§ āĻŋāύā§ā§ āϝāĻžā§āĨ¤ āĻāĻ āϤāĻžā§āĻĻāϰ āĻā§ā§āϰ āĻā§āϏāĨ¤ āϝāĻžāϰ āĻāĻžāϰā§āĻŖ āϤāĻžāϰāĻž āϰāĻžā§ āĻāĻžāύ āĻŽāĻžāύā§āώ āĻĻāĻā§āϞ āĻā§āϏ āĻā§āϰ āĻĨāĻžā§āĻ āϏ āĻāĻžāĻĨāĻžāĻ āϝāĻžā§āĻŦ āĻŋāĻāύāĻžāĨ¤ āĻāĻžāϰāĻŖ āĻŋāϰāĻāĻļāĻžāĻā§āĻžāϞāĻž āϝāĻŋāĻĻ āϤāĻžāϰ āĻāĻŦā§āϞ āϤāĻžā§āĻ āĻĒā§ā§āĻ āĻŋāĻĻā§āϤ āĻĒāĻžā§āϰ āĻāϰ āĻŦāĻĻā§āϞ āϏ āĻŋāĻāĻ āĻā§ āϰāĻžāĻāĻāĻžāϰ āĻĒāĻžā§āĻŦāĨ¤ āĻāϤ āĻā§āϏāϰ āĻ ā§āύāĻ āĻŽāĻžāύā§āώ āĻŋāĻŦāϰ āĻšā§āĨ¤ Evidence_Sentence: āϝāĻžāϰ āĻāĻžāϰā§āĻŖ āϤāĻžāϰāĻž āϰāĻžā§ āĻāĻžāύ āĻŽāĻžāύā§āώ āĻĻāĻā§āϞ āĻā§āϏ āĻā§āϰ āĻĨāĻžā§āĻ āϏ āĻāĻžāĻĨāĻžāĻ āϝāĻžā§āĻŦ āĻŋāĻāύāĻž Relevance_Score: 2 Meme_Category: Education Meme_OCR: āĻŦāϤāĻŽāĻžāύ āϏāĻŽā§ā§āϰ āĻŋāĻļāĻžāϰ āĻ āĻŦāĻž āĻāĻĒāύāĻžā§āĻ āϏāĻŦāĻŋāĻāĻ ā§āϞ āĻĨā§āĻ āĻŋāύā§āϤ āĻšā§āĻŦ āϝāĻŽāύ- āĻŦāĻ, āĻāĻžāϤāĻž, āĻŦāĻžāĻ, āĻ ā§ āϤāĻž, āĻŽāĻžāĻāĻž āĻāϤāĻžāĻŋāĻĻ āĻāϰ āĻŋāĻļāĻž? āϤāĻžāϰ āĻāύ āĻŦāĻžāĻŋā§ā§āϤ āĻāĻāĻļāύ āĻāĻŋāϰā§ā§ āĻŋāύā§āĻŦāύ Meme_Context: āĻŋāĻļāĻž āĻŦāĻž āϤāĻĨ āĻāĻžāϰāĻžāĻĒ āĻĒāϝāĻžā§ā§ āĻā§āĻ āϝ āĻ ā§āύāĻ āĻāĻžāĻāĻž āĻŋāĻĻā§āϤ āĻšā§ āĻāĻŦāĻ āϏāĻžā§āĻĨ ā§ā§āϞāϰ āĻŦāĻžāĻ āĻŦāĻ āĻāĻžāϤāĻž āĻ ā§ āϤāĻž āĻāĻžāĻŽāĻž āϏāĻŦāĻŋāĻāĻ ā§āϞ āĻĨā§āĻ āĻŋāύā§āϞāĻ āϝāĻāύ āĻŋāĻļāĻžāϰ āĻŋāĻŦāώā§āĻāĻž āĻā§āϏ āϤāĻāύ āϏāĻāĻž āĻŦāĻžāϏāĻžā§ āĻāϞāĻžāĻĻāĻž āĻā§āϰ āĻāĻāĻžāϰ āϰā§āĻ āĻĒā§āĻžā§āύāĻžāϰ āĻāύ āĻŦā§āϞ Evidence_Sentence: āϝāĻāύ āĻŋāĻļāĻžāϰ āĻŋāĻŦāώā§āĻāĻž āĻā§āϏ āϤāĻāύ āϏāĻāĻž āĻŦāĻžāϏāĻžā§ āĻāϞāĻžāĻĻāĻž āĻā§āϰ āĻāĻāĻžāϰ āϰā§āĻ āĻĒā§āĻžā§āύāĻžāϰ āĻāύ āĻŦā§āϞ Relevance_Score: 1 Meme_Category: Technology Meme_OCR: āĻŦāĻžāĻŽ āĻā§āĻžāϞ āĻā§ āĻĄāĻžāύ āĻā§āĻžāϞ āĻā§ Meme_Context: āĻŋāĻā§āĻŦāĻžā§āĻĄāϰ āĻŽā§āϧ āĻŦāĻžāĻŽ āĻŋāĻĻā§āĻ āĻāϰ āĻā§āĻžāϞ āĻŋāĻ āĻāĻž āϏāĻŦā§āĻā§ā§ āĻŦāĻŋāĻļ āĻŦāϤ āĻšā§ āĻŋāĻ āϏāĻ āϤāϞāύāĻžā§ āĻĄāĻžāύ āĻŋāĻĻā§āĻāϰ āĻā§āĻžāϞ āϤāĻŽāύ āĻāĻžāύ āĻāĻžāĻŋāĻšāĻĻāĻž āύāĻāĨ¤ āĻāĻāĻž āĻŦāĻžāĻāĻžā§āϤāĻ āĻāĻāĻžā§āύ āĻāĻāĻ āĻŋāĻŽ āĻŦāĻšāĻžāϰ āĻāϰāĻž āĻšā§ā§ā§āĻ āϝāĻāĻžā§āύ āĻĻāĻāĻžā§āύāĻž āĻšā§ā§ā§āĻ āĻāĻāĻāύ āĻ āĻŋāĻā§āϰāĻ āϏāĻāĻžā§āϞāϰ āĻāĻš āĻŦāĻž āĻŋāĻ āĻā§āϰāĻāĻā§āύāϰ āĻŋāϤāĻāĻžāϰ āĻāĻžāύ āĻāĻš āύāĻāĨ¤ Evidence_Sentence: āĻāĻāĻāύ āĻ āĻŋāĻā§āϰāĻ āϏāĻāĻžā§āϞāϰ āĻāĻš āĻŦāĻž āĻŋāĻ āĻā§āϰāĻāĻā§āύāϰ āĻŋāϤāĻāĻžāϰ āĻāĻžāύ āĻāĻš āύāĻāĨ¤ Relevance_Score: 2 Meme_Category: Sports Meme_OCR: āĻšāĻžāĻ āĻāĻžāĻ āĻāĻžāϤā§ā§ āĻĻāϞ + āĻžāĻŦ āĻāĻā§ āĻāĻžā§āĻāĻžā§ āϏāĻŽāĻžāύ āĻāĻžā§āĻŦ āĻĒāĻžāϰāĻĢāĻŽ āĻā§āϰ āϝāĻž Meme_Context: āĻāĻŋāĻŽāĻŋāϞā§āĻžā§āύāĻž āĻŽāĻžāĻā§āύāĻ āĻāĻāĻāύ āĻā§āĻāĻŋāύāĻžāϰ āĻāĻžāϞāĻŋāĻāĻĒāĻžāϰāĨ¤āϏ āĻāĻžāϤā§ā§ āĻĻāϞ āĻ āĻžāĻŦ āĻāĻā§ āĻāĻžā§āĻāĻžā§ āĻāĻāĻ āϰāĻāĻŽ āĻāĻžā§āĻŦ āĻāϞā§āĻāĨ¤ āĻŽāĻžāĻŋāϤā§āύā§āĻāϰ āϏāĻŦā§āĻļāώ āĻā§āĻŋāϤ āĻŋāύā§āĻ āĻāĻžāύāĻžāĨ¤ āĻŦāϤāĻŽāĻžāύ āĻŋāĻŦā§āϰ āϏāϰāĻž āĻāĻžāϞāĻŋāĻāĻĒāĻžāϰ āĻāĻŋāĻŽāĻŋāϞā§āĻžā§āύāĻž āĻŽāĻžāĻŋāϤā§āύāĻāĨ¤ Evidence_Sentence: āϏ āĻāĻžāϤā§ā§ āĻĻāϞ āĻ āĻžāĻŦ āĻāĻā§ āĻāĻžā§āĻāĻžā§ āĻāĻāĻ āϰāĻāĻŽ āĻāĻžā§āĻŦ āĻāϞā§āĻ Relevance_Score: 1 Meme_Category: Politics Meme_OCR: āϏā§āϰ āϰāĻžāĻā§āύāĻŋāϤāĻ āĻŦāĻž āĻšā§āϰāĻ āϰāĻžāĻāĻžāϰ āĻĻā§āĻļ Meme_Context: āĻĻā§āĻļāϰ āϧāĻžāύ āĻĻ ā§ āĻ āϰāĻžāĻā§āύāĻŋāϤāĻ āĻĻāϞ āĻāĻā§āĻžāĻŽā§ āϞā§āĻ āĻāĻŦāĻ āĻŋāĻŦāĻāύāĻŋāĻĒ āĻĒāϰ āĻŋāĻŦā§āϰāĻžāϧ⧠āĻŽā§āĻŦāϰ āĻŽāĻžāϧā§āĻŽ āĻĻāĻļā§āĻāϰ āĻāĻĒāϰ āĻĻāĻļāĻ āϧā§āϰ āĻŦāĻžāĻāϞāĻžā§āĻĻā§āĻļāϰ āϰāĻžāĻā§āύāĻŋāϤāĻ āĻŽāĻšāϞ āĻ āĻ āĻŋāϰ āĻā§āϰ āϰā§āĻā§āĻ āĻŋāĻ āĻāϰ āĻŽāĻžā§āĻ āĻšāĻžāĻŋāϰā§ā§ āĻā§āĻ āĻŦāĻžāĻāϞāĻžā§āĻĻā§āĻļāϰ āĻāĻāĻ āϏāĻšāĻ āĻāĻŦāĻ āϏā§āϰ āϰāĻžāĻā§āύāĻŋāϤāĻ āĻŦāĻž āϝāĻž āĻāĻāύ āϧā§āĻ āĻāύāĻžāĻŋāϤāĨ¤ Evidence_Sentence: āĻŦāĻžāĻāϞāĻžā§āĻĻā§āĻļāϰ āĻāĻāĻ āϏāĻšāĻ āĻāĻŦāĻ āϏā§āϰ āϰāĻžāĻā§āύāĻŋāϤāĻ āĻŦāĻž āϝāĻž āĻāĻāύ āϧā§āĻ āĻāύāĻžāĻŋāϤāĨ¤ Relevance_Score: 0 Meme_Category: Sports Meme_OCR: āϏāĻĢāϞāĻāĻžā§āĻŦ āĻāĻžāύāĻžāϞāĻžāϰ āĻāϏāύ āĻā§āϤā§āĻ Meme_Context: ā§āύāϰ āĻāĻžāύāĻžāϞāĻž āĻŋāϏāĻā§āϞāĻž āϏāĻžāϧāĻžāϰāĻŖāϤ āĻā§āĻŦ āĻāύāĻŋā§ āĻšā§ āĻāĻŦāĻ āϝāĻžā§āϰāĻž āĻāĻ āĻŋāϏāĻāĻŋāϞā§āϤ āĻŦāϏā§āϤ āĻŦāĻŋāĻļ āĻāĻšā§ āĻĨāĻžā§āĻāĨ¤ āĻāĻžāύāĻžāϞāĻž āĻŋāϏāĻ āĻĨā§āĻ āĻŦāĻžāĻā§āϰ āϤāĻžāĻŋāĻā§ā§ āĻŽāĻ, āĻā§āĻŋāĻŽ, āĻļāĻšāϰ āĻāĻŦāĻ āĻāĻāĻžā§āĻļāϰ āϏā§āϝ āĻāĻĒā§āĻāĻžāĻ āĻāϰāĻž āϝāĻžā§, āϝāĻž āĻ ā§āύāĻ āϝāĻžā§āϰ āĻāύ āĻāĻāώāĻŖā§ā§āĨ¤ āĻŦāĻžāĻāϞāĻžā§āĻĻāĻļ āĻĻā§āϞāϰ ā§āĻāĻāĻžāϰāĻžāĻ āĻāĻžāύāĻžāϞāĻžāϰ āĻĒāĻžā§āĻļ āĻŋāϏāĻ āĻĒā§ā§ āϤāĻž āϏāĻžāĻŽāĻžāĻāĻ āϝāĻžāĻāĻžā§āϝāĻžāĻ āĻŽāĻžāϧā§āĻŽ āĻĒāĻž āĻā§āϰāύāĨ¤ Evidence_Sentence: āĻŦāĻžāĻāϞāĻžā§āĻĻāĻļ āĻĻā§āϞāϰ ā§āĻāĻāĻžāϰāĻžāĻ āĻāĻžāύāĻžāϞāĻžāϰ āĻĒāĻžā§āĻļ āĻŋāϏāĻ āĻĒā§ā§ āϤāĻž āϏāĻžāĻŽāĻžāĻāĻ āϝāĻžāĻāĻžā§āϝāĻžāĻ āĻŽāĻžāϧā§āĻŽ āĻĒāĻž āĻā§āϰāύāĨ¤ Relevance_Score: 1 Meme_Category: Politics Meme_OCR: āĻāĻāĻŦāĻžā§āϰ āĻŋāύāĻŦāĻžāĻāύ āĻāĻāĻĻāĻŽ āĻŋā§āĻžāϰ Meme_Context: ⧍ā§Ļā§§ā§Ē āĻāĻŦāĻ ā§¨ā§Ļā§§ā§Ž āĻāϰ āĻĒāϰ ⧍ā§Ļā§¨ā§Š-āĻ āĻŋāύāĻŦāĻžāĻāύāĻ āĻŋāĻāϞ āĻŦāϤāĻŽāĻžāύ āĻŽāϤāĻžāϏā§āύ āĻāĻā§āĻžāĻŽā§ āϞā§āĻ āϏāϰāĻāĻžā§āϰ āĻāύ āϤāĻžāϰ āĻ āĻŦāύāϤāĻžāĻŋāĻ ā§āĻžā§ āĻŋāύāĻŦāĻžāĻāύā§āĻ āĻŦāĻžāĻšāϤ āĻāϰāĻžāϰ āϝ āĻ āĻŋāĻā§āϝāĻžāĻ āĻŋāĻāϞ āϤāĻž āĻĨā§āĻ āĻŽā§ āĻšāĻā§āĻžāϰāĨ¤ āĻŋāĻ āĻĻāĻāĻž āϝāĻžā§ āϝ āĻāĻāĻ āĻŽā§ āĻŋāύāĻŦāĻžāĻāύ āĻŦāĻžāϰ āĻŋāϤāĻŋāϤ āĻŋāĻĻāϞ āϤāĻž āĻ āĻ āĻŽāϤāύ āĻĒāĻžāϞāύ āĻāϰā§āϤ āĻĒāĻžāĻŋāϰāĻŋāύ āϤāĻžāĻ āĻāύāĻāĻŖ āĻāĻ āĻŋāύāĻŦāĻžāĻāύ āĻĨā§āĻ āĻŽā§āĻ āĻŋāĻĢāĻŋāϰā§ā§ āĻŋāύā§ā§ā§āĻāĨ¤ āĻāĻžāĻ āĻŋāĻĻā§āϤ āϝāĻžā§ āύāĻž āĻŽāĻžāύā§āώ āϤāĻžāĻšā§āϞ āĻŋāĻĻā§ āĻāĻžāϰāĻžāĨ¤ āĻāĻŦāĻžāϰ āϏāϰāĻāĻžāϰ āĻĒ āĻŦāϞā§āϤā§āĻ āĻŋāύāĻŦāĻžāĻāύ āϏ⧠āĻšā§ā§ā§āĻāĨ¤ Evidence_Sentence: āĻāĻžāĻ āĻŋāĻĻā§āϤ āϝāĻžā§ āύāĻž āĻŽāĻžāύā§āώ āϤāĻžāĻšā§āϞ āĻŋāĻĻā§ āĻāĻžāϰāĻžāĨ¤ āĻāĻŦāĻžāϰ āϏāϰāĻāĻžāϰ āĻĒ āĻŦāϞā§āϤā§āĻ āĻŋāύāĻŦāĻžāĻāύ āϏ⧠āĻšā§ā§ā§āĻāĨ¤ . Fig. 2. Overview of BanglaMemeEvidence dataset across different meme categories 1 as a primary source, leveraging its wealth of information to provide evidence for the background of the memes. Additionally, we explored community-based discussion forums and question-answering websites such as Quora 2 , as well as other general-purpose websites. Our thorough search spanned traditional me- dia and digital publications, including newspapers 3 , online article portals 4 , and specialized media outlets 5 , where we gathered in-depth analysis and expert com- mentary on meme-related events and phenomena. This comprehensive process ensured that each context document captured the essence of meme culture, of- fering valuable insights into its social, cultural, and political underpinnings. 3.3 Dataset Description âĻ Meme_ID: Each meme is assigned a unique image ID for identification purposes. âĻ Meme_Category: We have taken Politics, Sports, Entertainment, Educa- tion, Technology, and Others as meme categories. âĻ Meme_OCR: The Bengali text from the meme image is manually tran- scribed, ensuring accurate representation of the content. âĻ Meme_Context: A concise description of the main Bengali context or theme depicted in the meme is provided. 1 https://bn.wikipedia.org/ 2 https://bn.quora.com/ 3 https://w.prothomalo.com/ 4 https://blog.muktomona.com/ 5 https://ekattor.tv/ 6F.T.J. Faria et al. PoliticsSportsTechnologyEntertainmentEducationOthers Meme_Category 0 200 400 600 800 1000 Values 1,002.0 498.0 317.0 357.0 309.0 434.0 BanglaMemeEvidence Dataset Fig. 3. An illustration showcasing the diversity of meme categories, including Politics, Sports, Entertainment, and more from the BanglaMemeEvidence dataset. âĻ Evidence_Sentence: This section includes a sentence or short text ex- tracted from contextual documents, serving as evidence to support the in- terpretation or understanding of the meme. âĻ Relevance_Score: Each evidence sentence in the dataset is tagged with a relevance score, indicating its relationship to the meme. This score operates on a scale where (0) denotes âNot relevant,â (1) signifies âPartially relevant,â and (2) represents âRelevant.â 3.4 Annotation Guideline We provided the following annotation guideline to the annotators: a) Comprehensive Comprehension: We ensure understanding of both the meme and its associated context before annotation. This comprehensive comprehension is crucial for accurate and insightful annotation. b) Semantics Alignment: We let the semantics of the meme guide the annotation process. By aligning with the intended meaning of the meme, our annotations maintain relevance and fidelity. c) Unit of Information: We recognize that self-contained, minimal units of information can serve as evidence. Each piece of evidence contributes to our understanding of the memeâs background and significance. d) Non-Contiguous Evidence: We acknowledge that valid evidence may not always occur contiguously. This understanding allows us to identify and annotate relevant information regardless of its spatial arrangement. BanglaMemeEvidence: A Multimodal Benchmark Dataset7 e) Completeness Assurance: In cases where the context document does not support a meme, we diligently search for corroborating evidence from other established sources. This ensures the completeness and accuracy of our annota- tions. f) Caution with Ambiguity: We exercise caution with ambiguous cases, opting to skip them to maintain the integrity of the annotation process. This approach prevents potential misinterpretations and ensures the reliability of our dataset. 3.5 Annotation Process In our annotation process, we engaged the expertise of six male annotators, all hailing from Bangladesh and currently pursuing undergraduate studies. With ages ranging between 21 to 25 years, these individuals possessed not only a knack for meme creation but also an intricate understanding of meme culture, coupled with proficiency in navigating various social media platforms. Acknowledging the invaluable contribution of our annotators, we ensured fair compensation in accordance with Bangladeshi standards, recognizing the time and expertise they dedicated to the task. Equipped with a comprehensive set of guidelines, our annotators embarked on their mission: to scour context documents in search of succinct sentences that provided essential background information for each meme. These identified sentences, termed âevidence sentences,â served as the cornerstone of our dataset, offering profound insights into the genesis, signifi- cance, and cultural context of each meme. To ensure the reliability and con- sistency of the annotations, we implemented a rigorous process for addressing disagreements between annotators. Whenever annotators encountered differing interpretations of a meme or evidence sentences, they engaged in discussions to reach a consensus, guided by the annotation guidelines. In cases where dis- agreements persisted, a senior annotator, with deeper expertise in meme culture and contextual analysis, reviewed the conflicting annotations and made the final decision. Table 1. Dataset Distribution Across Train, Test, and Validation Subsets Based on Relevance Score. Relevance Score Train Test Validation 06788585 1819102102 2836105105 3.6 Annotation Quality Maintenance In our annotation process, we utilized Fleiss Kappa [13] to ensure annotation quality. Fleiss Kappa assesses agreement among annotators, considering chance 8F.T.J. Faria et al. occurrences. It provides a score indicating agreement beyond chance. By moni- toring Fleiss Kappa regularly, we maintained consistency and reliability in anno- tations. This practice ensured the datasetâs accuracy and reliability for research and analysis. We achieved a Fleiss Kappa score of 0.87, indicating strong agree- ment among annotators. 3.7 Dataset Statistics We partitioned the dataset into three subsets: training (80%), validation (10%), and testing (10%). This distribution ensures a balanced and representative split that supports effective model training, tuning, and evaluation. By reserving dis- tinct portions of the data for each stage, we aim to promote generalization and minimize overfitting. Table 1 presents the data split by relevance score across train, test, and validation sets. Figure 2 illustrates representative examples from the dataset. The first meme critiques the superficial display of religious identity on social media through the adoption of Arabic names. The second highlights issues in the education system, where material provisions are present, but ef- fective learning often relies on external tutoring. The third meme provides a humorous observation on technology usage, contrasting the frequent use of the left Ctrl key with the neglect of its counterpart on the right. The fourth meme celebrates the consistent athletic performance of footballer Emiliano Martinez at both the national and club levels. The fifth meme offers a pointed critique of electoral integrity, emphasizing low voter participation and the questionable fairness of the process. The sixth brings a lighthearted moment from sports, de- picting Bangladeshi cricketers happily claiming window seats during air travel, capturing a relatable and simple joy. Figure 3 showcases its category-wise diver- sity. 4 Implementation Details In our approach, BengaliMemeEvidenceNet, for detecting explanatory evi- dence in memes, we explore multiple fusion techniques, including Early Fusion, Late Fusion, and Intermediate Fusion. By investigating and comparing the per- formance of these individual approaches, we develop a hybrid framework that combines the strengths of each fusion method. This exploration allows us to propose the best-performing fusion strategy, leveraging the advantages of each technique to achieve superior results in detecting explanatory evidence in memes. Step 1) Text Preprocessing: Before delving into analysis, we meticulously preprocess textual data extracted from various sources within memes, includ- ing Meme_OCR, Meme_Context, and Evidence_Sentence. This comprehensive preprocessing involves several steps aimed at ensuring the cleanliness and stan- dardization of the text: âĻ Punctuation Removal: Removing punctuation marks such as exclama- tion points, question marks, underscores, and quotation marks. This can be BanglaMemeEvidence: A Multimodal Benchmark Dataset9 BanglaMemeEvidence Dataset Meme_Context : āĻŦāĻŽā§ā§āϞāϰ āĻāĻāĻŋāϤāϰ āĻāĻĨāĻžāĻ āĻāĻāĻžā§āύ āĻŦā§āĻāĻžā§āύāĻž āĻšā§ā§ā§āĻāĨ¤ āĻŽāĻžāύā§ā§āώāϰ ā§ā§āĻžāĻāύā§ā§ āĻ āĻĒā§āĻŖ ā§āĻŦāϰ āĻĻāĻžāĻŽ āϝāύ āĻāĻžāύāĻāĻžā§āĻŦāĻ āĻāĻŽā§āĻ āύāĻžāĨ¤ āϏāϰāĻāĻžāϰ āĻŦāĻžāϰāĻŦāĻžāϰ āĻŦāĻŽā§āϞ āĻŋāύā§ā§āĻŖāϰ āĻŦāĻĨāϤāĻžāϰ āĻĒāĻŋāϰāĻā§ āĻŋāĻĻā§āĨ¤ āĻāĻ āĻŽā§āϞ āĻŦā§āϰ āĻĢā§āϞ āϏāĻžāϧāĻžāϰāĻŖ āĻāύāĻā§āĻŖāϰ āĻā§āĻŦāύ āĻŦāĻž āĻ ā§āύāĻ āĻāĻ āύ āĻšā§ā§ āĻĒā§ā§āĻāĨ¤ (English Translation: The price rise is explained here. The price of the imperfect products needed by people is not decreasing in any way. The government has repeatedly demonstrated its failure to control commodity prices. As a result of this increase in prices, the standard of living of the common people is becoming very dificult.) Meme_OCR : āĻāĻĢ ! āĻšāĻ āĻŦ āĻ āĻŽā§āϞ (English Translation: Oops! Hot products and prices) Relevance_Score (Ground Truth): 2 Evidence_Sentence : āĻŦāĻŽā§ā§āϞāϰ āĻāĻāĻŋāϤāϰ āĻāĻĨāĻžāĻ āĻāĻāĻžā§āύ āĻŦā§āĻāĻžā§āύāĻž āĻšā§ā§ā§āĻāĨ¤ (English Translation:The price rise is explained here.) Image Features Vector Text Features Vector Text Features Image Features Early F usion Late F usion Visual Feature Extactor Textual Feature Extactor Relevance Score Fig. 4. Illustration of BengaliMemeEvidenceNet, a hybrid multimodal framework de- signed for identifying evidence in Bengali memes by leveraging related contexts and predicting relevance scores. mathematically represented as: T cleaned = remove_punctuation(T)(1) where T is the original text, and T cleaned is the text with punctuation re- moved. âĻ Whitespace Elimination: Systematically eliminating any extraneous white spaces. This step is represented as: T cleaned = remove_whitespace(T cleaned )(2) where we remove any unnecessary whitespace from the text. âĻ Emoji Removal: Meticulously scanning and removing emojis to prevent noise and ambiguity. This is expressed as: T cleaned = remove_emoji(T cleaned )(3) where emojis are identified and removed from the text. âĻ Removing Non-Textual Content: Remove any URLs, HTML tags, and special characters that are not relevant to the text analysis. This can be represented as: T cleaned = remove_non_textual(T cleaned )(4) where any non-relevant content is stripped from the text. âĻ Spelling Correction: Correct misspelled words to standardize the text. This is typically done by applying a spelling correction algorithm such as: T corrected = spell_check(T cleaned )(5) where T corrected is the text with corrected spelling. Step 2) Image Preprocessing: To ensure uniformity and high quality across our dataset, we standardized the size of images extracted from memes to 224à 224 pixels, facilitating reliable and comparable analysis. Our image preprocessing pipeline includes several critical steps aimed at enhancing image quality and preparing the visual data for robust feature extraction and analysis: 10F.T.J. Faria et al. âĻ Normalizing the images to maintain consistency in pixel values. The normal- ization process is expressed as: I normalized = Iâ Îŧ Ī (6) where I is the original image, Îŧ is the mean pixel value, and Ī is the standard deviation of the pixel values. âĻ Applying edge detection algorithms (such as the Sobel operator) to highlight the boundaries within the images. This is represented as: I edges = Sobel(I)(7) where I edges contains the detected edges of the image. âĻ Applying noise reduction techniques, such as Gaussian smoothing, to reduce unwanted artifacts. This can be expressed as: I denoised = Gaussian_filter(I)(8) where I denoised is the filtered image with reduced noise. âĻ Adjusting the contrast of the images to highlight significant features. This adjustment is expressed as: I contrast_adjusted = adjust_contrast(I denoised )(9) where I contrast_adjusted has improved contrast to better emphasize visual de- tails. Step 3) Feature Extraction: To capture the rich and intricate nuances of explanatory information within memes, we employ advanced pre-trained lan- guage models and cutting-edge image analysis techniques. Our feature extraction process is two-fold, focusing on both textual and visual data: âĻ Textual Feature Extraction: We utilize state-of-the-art (SOTA) pre-trained language models to extract semantic insights from various textual com- ponents of memes. The models we employ include mBERT [14], XLM- RoBERTa [15], and distilBERT [16]. These models excel in understanding context and semantics, allowing us to accurately capture the subtleties and nuances inherent in the textual data of memes. We preferred mBERT, XLM- RoBERTa, and distilBERT due to their excellent multilingual capabilities, which are helpful when performing cross-lingual tasks and data collection from related languages. âĻ Visual Feature Extraction: For the visual aspect, we utilize advanced models specifically tailored for meme analysis to extract relevant features from images. These models include Vision Transformers (ViTs) [17], Swin Transformer [18], SwiftFormer [19], PoolFormer [20]. Each model is trained to detect and emphasize significant visual elements that contribute to the explanatory content of memes. BanglaMemeEvidence: A Multimodal Benchmark Dataset11 Step 4) Fusion Techniques: After extracting text and image features, we employ fusion techniques tailored to each approach: Early Fusion [21] and Late Fusion [22]. a) Early Fusion for Evidence Detection: In early fusion, we integrate features from both text and image modalities at a raw level and then feed them into a joint representation function. This joint representation is then used for further analysis. The equation for early fusion can be mathematically represented as: Z joint = f early (X text ,X image )(10) Where: âĻZ joint represents the joint representation combining features from text and image. âĻ f early is the early fusion function. âĻX text denotes the feature representation extracted from text sources such as Meme_OCR, Meme_Context, and Evidence_Sentence. âĻX image represents the feature representation extracted from the meme image. b) Late Fusion for Evidence Detection: In late fusion, predictions from text and image classification models are aggregated at a later stage to make the final decision. Each modality produces its prediction, and these predictions are combined using a fusion function. The equation for late fusion can be mathe- matically represented as: Ëy final = f late (Ëy text , Ëy image )(11) Where: âĻ Ëy final represents the final prediction. âĻ f late is the late fusion function. âĻ Ëy text denotes the prediction obtained from the text classification model based on sources like Meme_OCR, Meme_Context, and Evidence_Sentence. âĻ Ëy image represents the prediction obtained from the image classification model. c) Intermediate Fusion for Evidence Detection: In intermediate fu- sion, features from both text and image modalities are fused at an earlier stage, before the final classification decision is made. Rather than aggregating the pre- dictions directly, the features from both modalities are combined and passed through a shared model to produce the final prediction. The fusion function in this case operates on the feature vectors of both modalities. Let the feature vectors from the text and image models be denoted asf text and f image , respectively. The intermediate fusion can be mathematically represented as: Ëy final = f intermediate (f text ,f image ) =f text +f image (12) Where: 12F.T.J. Faria et al. âĻ Ëy final represents the final prediction after intermediate fusion. âĻ f intermediate is the intermediate fusion function, which in this case is the element-wise sum of the feature vectors from both modalities. âĻf text denotes the feature vector obtained from the text classification model, derived from sources such as Meme_OCR, Meme_Context, and Evidence_Sentence. âĻf image represents the feature vector obtained from the image classification model. In this implementation, the intermediate fusion function f intermediate com- bines the feature vectors from the text and image modalities by performing an element-wise sum of the two feature vectors. Step 5) Hyperparameter Tuning: Hyperparameter tuning plays a vital role in optimizing the performance of fusion strategies such as Early, Late, and Intermediate Fusion. For Early Fusion, the key hyperparameters include the learning rate (Ρ), fusion weight (w fusion ), and dropout rate (d). The optimal learning rate Ρ opt is determined by minimizing the loss function: Ρ opt = arg min Ρ L(θ;Ρ)(13) The fusion weight w fusion is tuned to balance the contribution of each modality in the fused feature vector, with the optimal value found as: w fusion = arg min w fusion L fusion (X text ,X image ,w fusion )(14) The dropout rate d opt is selected to prevent overfitting: d opt = arg min d L(θ;d)(15) For Late Fusion, the model combines predictions from different modalities, and key hyperparameters include the learning rate (Ρ), fusion weight (w fusion ), and batch size (b). The learning rate Ρ opt is tuned similarly to Early Fusion, while the optimal fusion weight w fusion is selected by minimizing: w fusion = arg min w fusion L fusion (y text ,y image ,w fusion )(16) The batch size b opt is optimized to ensure efficient training: b opt = arg min b L(θ;b)(17) For Intermediate Fusion, which combines features after extraction but before final decision-making, the learning rate (Ρ), number of layers (L), fusion weight (w fusion ), and dropout rate (d) are crucial. The learning rate Ρ opt and dropout rate d opt are optimized in the same way as in Early Fusion. The number of layers L opt is determined by minimizing the loss function: L opt = arg min L L(θ;L)(18) BanglaMemeEvidence: A Multimodal Benchmark Dataset13 The fusion weight w fusion is optimized similarly to Early Fusion: w fusion = arg min w fusion L fusion (X text ,X image ,w fusion )(19) Bayesian Optimization is employed to efficiently search for the optimal hyperpa- rameters by modeling the loss function and predicting the best combination of hyperparameters. The overall optimization problem for all fusion methods can be formalized as: θ opt ,Ρ opt ,w fusion , d opt ,b opt ,L opt = arg min θ,Ρ,w,d,b,L L(θ;Ρ,w,d,b,L) (20) where L represents the loss function. This approach ensures that the optimal hyperparameters are selected to minimize the overall loss across different fusion strategies. Step 6) Evaluation of Meme Relevant Score Detection: Evaluation involves assessing model performance in detecting explanatory evidence within memes using metrics such as accuracy, precision, recall, and F1-score. Precision measures the proportion of true positives among predicted positives, while recall evaluates the proportion of true positives among actual positives. The F1-score is crucial as it balances precision and recall into a single metric, which is par- ticularly useful for handling imbalanced datasets where certain relevance scores are less frequent. This task includes classifying evidence into relevance scores: 0 for âNot relevant,â 1 for âPartially relevant,â and 2 for âRelevant,â making it challenging due to the imbalanced distribution of these scores. 5 Result Analysis Table 2 present a comparative analysis of early and late fusion approaches for relevance score detection in Bengali memes, using F1-score as the primary eval- uation metric. Among the early fusion models, Swin-m achieves the highest F1- score (0.71), while other models exhibit moderate performance. However, the late fusion method outperforms both strategies, with our proposed model, Ben- galiMemeEvidenceNet , achieving the highest F1-score of 0.74, surpassing all other architectures. This result underscores the effectiveness of BengaliMemeEv- idenceNet in capturing multimodal interactions, demonstrating its superiority in Bengali meme relevance detection over traditional fusion techniques. 6 Limitations Although our approach, BengaliMemeEvidenceNet, empirically outperforms sev- eral competitive baselines, we observe certain limitations in the modeling capac- ity towards BanglaMemeEvidence. As depicted in Table 6, there are three pos- sible scenarios of ineffective detection: (a) no predictions, (b) partial match, and (c) incorrect predictions. The key challenges stem from the limitations in model- ing the complex level of abstractions that a meme exhibits. These are primarily encountered in the following scenarios: 14F.T.J. Faria et al. Table 2. Performance Metrics of Various Fusion Approaches for Relevance Score De- tection in Bengali Memes. ApproachModelsAccuracy Precision Recall F1-Score ViT+mBERT0.670.650.660.66 Swin Transformer+mBERT0.750.680.730.71 SwiftFormer+mBERT0.740.650.730.69 PoolFormer+mBERT0.720.750.700.72 EarlyViT+XLM0.680.680.660.67 FusionSwin+XLM0.740.720.670.69 Swift+XLM0.740.650.720.69 PoolFormer+XLM0.720.740.660.70 ViT+DistilBERT0.660.710.660.68 Swin+DistilBERT0.650.650.740.69 Swift+DistilBERT0.750.720.670.70 Pool+DistilBERT0.650.680.710.70 ViT+mBERT0.680.720.650.68 Swin Transformer+mBERT0.690.680.700.69 SwiftFormer+mBERT0.720.730.710.72 PoolFormer+mBERT0.700.710.690.70 LateViT+XLM0.670.720.650.67 FusionSwin Transformer+XLM0.710.700.720.71 SwiftFormer+XLM0.680.690.670.68 BengaliMemeEvidenceNet0.740.750.730.74 ViT+DistilBERT0.730.740.720.73 Swin Transformer+DistilBERT0.660.670.640.65 SwiftFormer+DistilBERT0.710.690.730.71 PoolFormer+DistilBERT0.650.640.660.65 âĻ Integration of Visual and Factual Knowledge: A critical yet cryptic piece of information within memes often comes from the visuals, which typ- ically require systematic integration of factual knowledge. BengaliMemeEv- idenceNet currently lacks this capability, making it difficult to accurately interpret and explain the visual elements in memes. âĻ Lexical Bias and Spurious Evidence: The model is prone to picking up potentially spurious pieces of evidence due to lexical biasing within the related context. This can lead to incorrect predictions, as the model may focus on irrelevant or misleading textual features rather than the intended meaning of the meme. âĻ Complex Abstractions: Memes often exhibit a complex level of abstrac- tion that is challenging for the model to capture. This includes subtle humor, cultural references, and context-specific nuances that require advanced rea- soning capabilities beyond the current scope of BengaliMemeEvidenceNet. âĻ Inadequate Contextual Understanding: The modelâs current contex- tual understanding is often inadequate for fully grasping the meaning con- veyed by memes. This limitation is particularly evident when memes rely on BanglaMemeEvidence: A Multimodal Benchmark Dataset15 intricate social or cultural contexts that the model has not been trained to recognize. 7 Future Works There are several areas where further research could significantly enhance the robustness and applicability of our models. Below, we outline future work that will build on our current findings: âĻ Dialect-Based Meme Annotation: We will explore dialect-based meme annotation and subsequent performance analysis to account for variations in the Bengali language. This approach will improve model accuracy and effectiveness in diverse linguistic contexts, ensuring that regional dialects and linguistic variations are accurately represented and understood. âĻ Banglish Meme Detection: We aim to develop capabilities for detecting Banglish memes, which involve a mixture of Bengali and English, particularly prevalent in spoken communication by Bangladeshis. We will create methods to handle language switching within sentences or conversations, enhancing the modelâs ability to accurately interpret and explain Banglish memes. âĻ Fine-Grained Semantic Role Analysis: We will conduct a fine-grained analysis of the semantic roles in memes to explore nuances in connotation and interpretation. We will investigate how factors such as image compo- sition, text placement, and meme format influence the perceived roles of entities like heroes, villains, and victims. This detailed analysis will provide deeper insights into how different elements of a meme contribute to its overall message. âĻ Applications Beyond Memes: We will investigate how the concept of visual semantic role labeling and natural language explanation generation can be applied to other forms of visual communication, such as advertise- ments, political cartoons, and social media posts. Expanding the application of these techniques will provide valuable tools for analyzing a wide range of visual media, enhancing our understanding of multimodal communication across different contexts. 8 Conclusion We have introduced and explored the hybrid task of meme evidence detection, emphasizing the need for an in-depth understanding of the visual and textual semantics embedded in memes. By presenting BanglaMemeEvidence, a unique dataset of 2,917 annotated Bengali memes, we address a significant gap in meme analysis for low-resource languages. Our dataset includes rich annotations and relevance scores, providing a comprehensive resource for further studies in this domain. To effectively detect explanatory evidence within memes, we developed BengaliMemeEvidenceNet, a hybrid approach that integrates textual and visual features. This pioneering work not only sets a benchmark in the field of meme 16F.T.J. Faria et al. analysis for Bengali but also opens avenues for future research in other low- resource languages. References 1. Shivam Sharma, Ramaneswaran S, Udit Arora, Md. Shad Akhtar, and Tanmoy Chakraborty. Memex: Detecting explanatory evidence for memes via knowledge- enriched contextualization, 2023. 2. Shivam Sharma, Siddhant Agarwal, Tharun Suresh, Preslav Nakov, Md. Shad Akhtar, and Tanmoy Chakraborty. What do you meme? generating explanations for visual semantic role labelling in memes, 2022. 3. Md. Rezaul Karim, Sumon Kanti Dey, Tanhim Islam, Md. Shajalal, and Bharathi Raja Chakravarthi. Multimodal hate speech detection from bengali memes and texts, 2022. 4. Mithun Das and Animesh Mukherjee. Banglaabusememe: A dataset for bengali abusive meme classification, 2023. 5. Md.Tofael Ahmed, Nahida Akter, Maqsudur Rahman, Abu Islam, Dipankar Das, and Md. Golam Rashed. Multimodal cyberbullying meme detection from social media using deep learning approach. International Journal of Computer Science and Information Technology, 15:27â37, 08 2023. 6. Nayan Alluri and Neeli Krishna. Multi modal analysis of memes for sentiment extraction. pages 213â217, 11 2021. 7. Roshan Nayak, Ullas Kannantha, Kruthi S, and C. Gururaj. Multimodal offen- sive meme classification using transformers and bilstm. International Journal of Engineering and Advanced Technology, 11:96â102, 02 2022. 8. Akshi Kumar and Geetanjali Garg. Sarc-m: Sarcasm detection in typo-graphic memes. SSRN Electronic Journal, 01 2019. 9. Dushyant Chauhan, Dhanush R, Asif Ekbal, and Pushpak Bhattacharyya. Senti- ment and emotion help sarcasm? a multi-task learning framework for multi-modal sarcasm, sentiment and emotion analysis. pages 4351â4360, 01 2020. 10. Mahmud Hasan, Labiba Islam, Jannatul Ferdous Ruma, Tasmiah Tahsin May- eesha, and Rashedur M Rahman. Visual question generation in bengali. arXiv preprint arXiv:2310.08187, 2023. 11. SM Shahriar Islam, Riyad Ahsan Auntor, Minhajul Islam, Mohammad Yousuf Hos- sain Anik, ABM Alim Al Islam, and Jannatun Noor. Note: Towards devising an effi- cient vqa in the bengali language. In Proceedings of the 5th ACM SIGCAS/SIGCHI Conference on Computing and Sustainable Societies, pages 632â637, 2022. 12. Mahamudul Hasan Rafi, Shifat Islam, SM Hasan Imtiaz Labib, SM Sajid Hasan, Faisal Muhammad Shah, and Sifat Ahmed. A deep learning-based bengali visual question answering system. In 2022 25th International Conference on Computer and Information Technology (ICCIT), pages 114â119. IEEE, 2022. 13. Iman Albakkosh. Using fleissâ kappa coefficient to measure the intra and inter- rater reliability of three ai software programs in the assessment of efl learnersâ story writing. International Journal of Educational Sciences and Arts, 3(1):69â96, January 2024. 14. Jacob Devlin, Ming-Wei Chang, Kenton Lee, and Kristina Toutanova. BERT: pre-training of deep bidirectional transformers for language understanding. CoRR, abs/1810.04805, 2018. BanglaMemeEvidence: A Multimodal Benchmark Dataset17 15. Alexis Conneau, Kartikay Khandelwal, Naman Goyal, Vishrav Chaudhary, Guil- laume Wenzek, Francisco GuzmÃĄn, Edouard Grave, Myle Ott, Luke Zettlemoyer, and Veselin Stoyanov. Unsupervised cross-lingual representation learning at scale, 2020. 16. Victor Sanh, Lysandre Debut, Julien Chaumond, and Thomas Wolf. Distilbert, a distilled version of bert: smaller, faster, cheaper and lighter, 2020. 17. Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xi- aohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale, 2021. 18. Ze Liu, Yutong Lin, Yue Cao, Han Hu, Yixuan Wei, Zheng Zhang, Stephen Lin, and Baining Guo. Swin transformer: Hierarchical vision transformer using shifted win- dows. In 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pages 9992â10002, 2021. 19. Abdelrahman Shaker, Muhammad Maaz, Hanoona Rasheed, Salman Khan, Ming- Hsuan Yang, and Fahad Shahbaz Khan. Swiftformer: Efficient additive attention for transformer-based real-time mobile vision applications, 2023. 20. Weihao Yu, Mi Luo, Pan Zhou, Chenyang Si, Yichen Zhou, Xinchao Wang, Jiashi Feng, and Shuicheng Yan. Metaformer is actually what you need for vision, 2022. 21. George Barnum, Sabera Talukder, and Yisong Yue. On the benefits of early fusion in multimodal representation learning, 2020. 22. Yagya Raj Pandeya and Joonwhoan Lee. Deep learning-based late fusion of multi- modal information for emotion classification of music video. Multimedia Tools and Applications, 80(2):2887â2905, September 2020.