Paper deep dive
Is AI Ready for Multimodal Hate Speech Detection? A Comprehensive Dataset and Benchmark Evaluation
Rui Xing, Qi Chai, Jie Ma, Jing Tao, Pinghui Wang, Shuming Zhang, Xinping Wang, Hao Wang
Intelligence
Status: succeeded | Model: anthropic/claude-sonnet-4.6 | Prompt: intel-v1 | Confidence: 97%
Last extracted: 3/24/2026, 1:30:13 AM
Summary
This paper presents M³ (Multi-platform, Multi-lingual, and Multimodal Meme), a dataset of 2,455 memes collected from X, 4chan, and Weibo for multimodal hate speech detection. The authors propose an agentic annotation framework with seven specialized agents (collector, extractor, cleaner, annotator, arbiter, explicator, validator) to generate hierarchical labels and rationales. The dataset features fine-grained hate categories (religion, politics, race, gender, health status, violence, public health, international relations) and human-verified rationales. Benchmarking state-of-the-art MLLMs (Gemini-3, GPT-4o, Qwen-VL, LLaVA-v1.6, GLM-4.1V-9B-Thinking) reveals that incorporating surrounding post context often degrades detection performance, highlighting the need for context-aware multimodal architectures.
Entities (41)
Relation Signals (32)
Qi Chai → affiliatedwith → Hong Kong University of Science and Technology (Guangzhou)
confidence 99% · The Hong Kong University of Science and Technology (Guangzhou)
Jie Ma → affiliatedwith → Xi'an Jiaotong University
confidence 99% · MOE KLINNS Lab, Xi'an Jiaotong University; School of Cyber Science and Engineering
Shuming Zhang → affiliatedwith → Northwest University
confidence 99% · Shuming Zhang - Northwest University
Hao Wang → affiliatedwith → Hong Kong University of Science and Technology (Guangzhou)
confidence 99% · The Hong Kong University of Science and Technology (Guangzhou)
Rui Xing → affiliatedwith → Xi'an Jiaotong University
confidence 99% · MOE KLINNS Lab, Xi'an Jiaotong University
M³ Dataset → benchmarks → Multimodal Large Language Models
confidence 99% · Benchmarking state-of-the-art Multimodal Large Language Models reveals that these models struggle to effectively utilize surrounding post context
M³ Dataset → collectedfrom → 4chan /pol/
confidence 99% · 2,455 memes collected from X, 4chan, and Weibo
M³ Dataset → collectedfrom → X (Twitter)
confidence 99% · 2,455 memes collected from X, 4chan, and Weibo
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Hate speech online targets individuals or groups based on identity attributes and spreads rapidly, posing serious social risks. Memes, which combine images and text, have emerged as a nuanced vehicle for disseminating hate speech, often relying on cultural knowledge for interpretation. However, existing multimodal hate speech datasets suffer from coarse-grained labeling and a lack of integration with surrounding discourse, leading to imprecise and incomplete assessments. To bridge this gap, we propose an agentic annotation framework that coordinates seven specialized agents to generate hierarchical labels and rationales. Based on this framework, we construct M^3 (Multi-platform, Multi-lingual, and Multimodal Meme), a dataset of 2,455 memes collected from X, 4chan, and Weibo, featuring fine-grained hate labels and human-verified rationales. Benchmarking state-of-the-art Multimodal Large Language Models reveals that these models struggle to effectively utilize surrounding post context, which often fails to improve or even degrades detection performance. Our finding highlights the challenges these models face in reasoning over memes embedded in real-world discourse and underscores the need for a context-aware multimodal architecture. Our dataset and code are available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2603.21686v1
- Canonical: https://arxiv.org/abs/2603.21686v1
Trouble viewing inline? Open PDF directly →
Full Text
41,554 characters extracted from source content.
Expand or collapse full text
Is AI Ready for Multimodal Hate Speech Detection? A Comprehensive Dataset and Benchmark Evaluation Rui Xing 1∗ , Qi Chai 2∗ , Jie Ma 1,3† , Jing Tao 1 , Pinghui Wang 1 , Shuming Zhang 4 , Xinping Wang 3 , Hao Wang 2 1 MOE KLINNS Lab, Xi’an Jiaotong University 2 The Hong Kong University of Science and Technology (Guangzhou) 3 School of Cyber Science and Engineering, Xi’an Jiaotong University 4 Northwest University * Equal contribution † Corresponding author jiema@xjtu.edu.cn Abstract Hate speech online targets individuals or groups based on identity attributes and spreads rapidly, posing serious social risks. Memes, which com- bine images and text, have emerged as a nu- anced vehicle for disseminating hate speech, of- ten relying on cultural knowledge for interpreta- tion. However, existing multimodal hate speech datasets suffer from coarse-grained labeling and a lack of integration with surrounding discourse, leading to imprecise and incomplete assessments. To bridge this gap, we propose an agentic anno- tation framework that coordinates seven special- ized agents to generate hierarchical labels and ratio- nales. Based on this framework, we construct M 3 (Multi-platform, Multi-lingual, and Multimodal Meme), a dataset of 2,455 memes collected from X, 4chan, and Weibo, featuring fine-grained hate la- bels and human-verified rationales. Benchmarking state-of-the-art Multimodal Large Language Mod- els reveals that these models struggle to effec- tively utilize surrounding post context, which often fails to improve or even degrades detection perfor- mance. Our finding highlights the challenges these models face in reasoning over memes embedded in real-world discourse and underscores the need for a context-aware multimodal architecture. Our dataset and code are available at https://github.com/ mira-ai-lab/M3. Disclaimer: This paper includes content that may be considered offensive or disturbing to some read- ers. 1 Introduction Hate speech [ Guterres, 2019 ] refers to “any kind of commu- nication in speech, writing, or behavior that attacks or uses pejorative or discriminatory language with reference to a per- son or a group on the basis of who they are; in other words, Existing Dataset Rationale: No hateful content. Label: Our Dataset (M ퟑ ) Post:Absolute mongoloids. Objectively ugly. Abominations. Rationale: Discriminate against individuals with Down‘s syndrome. Label: HateNormal HateNormal Health Status Other categories Figure 1: Comparison between existing datasets and ours (M 3 ). Ex- isting datasets typically label the meme (right) as normal. However, our dataset labels it as hate with a refined classification of healthy state-because of its accompanying post. based on their religion, ethnicity, nationality, race, color, de- scent, gender, or other identity factor.” Its rapid dissemi- nation across online platforms poses a serious threat to so- cial stability [ Velasquez et al., 2021 ] . For instance, the 2019 Christchurch mosque shootings in New Zealand and the ter- rorist attack in El Paso, Texas, that same year. Related reports and studies [ Barnes, 2019; Ware, 2022 ] indicate that the per- petrators incorporated Internet Memes into their manifestos or posts to spread hate speech and resonate with specific on- line subcultures, thus inciting extremist ideologies. A meme [ Shifman, 2013 ] is a multimodal composite con- sisting of an image and short text. As one of the vehicles for disseminating hate speech [ Pandiani et al., 2025 ] , the meme is easily produced. It may combine a humorous image with a slogan that incites violence [ Zhou et al., 2021 ] , or convey discriminatory meanings through visual elements, with the embedded textual content appearing neutral in isolation [ Hee et al., 2024 ] . Within online communities, memes are often in- conspicuously embedded in otherwise ordinary posts and are interpretable only by users who share specific cultural or sub- cultural knowledge. This subtle embedding—combined with their multimodal ambiguity—facilitates rapid dissemination through densely connected social networks [ Brown, 2018; Schmid, 2025 ] . These characteristics not only make hateful arXiv:2603.21686v1 [cs.MA] 23 Mar 2026 DatasetDomain Label Img textPostRationaleMethodSource Hateful?Categories The Hateful Memes Challenge Set Multiple fields Hateful, Not-hateful - --Human-onlySynthetical HatReDMultiple fields Hateful, Not-hateful --Human-onlySynthetical HarMemeCOVID-19 Very harmful, Partially harmful, Harmless - --Human-only Google Image Search, Reddit, Facebook, Instagram MMHS150KMultiple fields- Not hate, Racist, Sexist, Homophobic, Religion, Other hate --Human-onlyTwitter MAMIMisogynyMisogynous, Not-misogynous Shaming, Stereotype, Objectification, Violence --Human-onlyTwitter, Reddit ExMute-Hateful, Non-hateful Religious, Celebrity, Political, Male, Female, Others --Human-onlyFacebook, Reddit, Instagram M 3 (Ours)Multiple fieldsHate, Normal Religion, Politics, Race, Gender, Health Status, Violence, Public Health, International Relations Human-validate agentic X, 4chan, Weibo Table 1: Comparison of representative multimodal hate speech datasets. While existing datasets often provide coarse labels or limited context, OURS introduces fine-grained multi-dimensional annotations and additionally includes the surrounding post content, enabling richer and more context-aware hate speech analysis. memes particularly insidious but also expose critical gaps in current multimodal hate speech detection datasets. Existing datasets [ Chhabra and Vishwakarma, 2023; Jiang and Zubiaga, 2024; Nayak and Agrawal, 2022 ] predomi- nantly support superficial evaluation, which fails to accu- rately assess the true performance of hateful meme detection methods, thereby offering limited constructive feedback for improving hate speech detection systems. On the one hand, they still adopt coarse-grained or flat labeling schemes, creat- ing a significant discrepancy with the multi-dimensional com- plexity of hate speech as defined by UN protocols. As il- lustrated in Figure 1, the simplistic binary labels (i.e., Hate vs. Normal) preclude the representation of intersectional offenses, thus failing to reflect model efficacy in nuanced, multi-class scenarios. On the other hand, they typically iso- late memes from their surrounding discourse, focusing exclu- sively on their immediate content. Since memes in real-world social media are intrinsically tied to accompanying posts, the absence of such contextual information may lead to the erro- neous label or even incomplete and misleading interpretations (see Figure 1). To address the challenges above, we propose an agen- tic annotation framework that coordinates seven specialized agents including a collector, extractor, cleaner, annotator, ar- biter, explicator, and validator. The collector gathers memes and associated posts from X (formerly Twitter), 4chan, and Weibo, enabling broad cross-cultural coverage. After pro- cessing by the extractor and cleaner, the annotator and ar- biter conduct multi-round hierarchical annotation with expert adjudication. Subsequently, the explicator generates hate ra- tionales, while the validator performs quality control. This framework yields M 3 , the Multi-platform, Multi-lingual, and Multimodal Meme dataset, which contains 2,455 high- quality multimodal instances with two top-level labels (hate and normal), and eight fine-grained hate categories. We fur- ther evaluate M 3 on state-of-the-art Multimodal Large Lan- guage Models (MLLMs), including Gemini-3, GPT-4o mod- els and representative open-source models such as Qwen-VL, LLaVA-v1.6, and GLM-4.1V-9B-Thinking. Among them, Qwen3-VL-8B-Instruct achieves the highest overall score across the evaluation metrics. The results show that incorpo- rating surrounding post context leads to degraded accuracy in hate detection and fine-grained classification. These findings suggest that current AI is not yet ready to fully replace human moderation. Moreover, hate speech detection should consider the linguistic context in which memes are embedded. The main contributions are summarized as follows: • We develop an agentic annotation framework, which co- ordinates multiple agents for scalable labeling and ra- tionale generation, while ensuring annotation reliability through systematic human verification. • We introduce M 3 , a multi-platform, multi-lingual, and multimodal meme dataset spanning X, 4chan, and Weibo, featuring fine-grained hate annotations and human-verified rationales to support explainable multi- modal hate analysis. • We provide extensive benchmarking results on state-of- the-art MLLMs, demonstrating the effectiveness of M 3 in evaluating multimodal hate meme detection capabili- ties and revealing current limitations in rationale gener- ation. 2 Related Work 2.1 Unimodal Hate Speech Datasets Early research primarily focused on plain-text content via manual annotation (see Table 6 in the supplementary material section 4). Initial English-centric efforts [ Waseem and Hovy, 2016 ] were later scaled through crowdsourcing [ Founta et al., 2018 ] and extended to non-English contexts, including Ara- bic [ Mulki et al., 2019 ] , Chinese [ Rao et al., 2023 ] , and cross- lingual datasets [ Tonneau et al., 2025 ] . Parallel to this lingual diversification, the annota- tion paradigms shifted toward multifaceted schemes. OLID [ Zampieri et al., 2019 ] employs a three-tier hierar- chy, while HateXplain [ Mathew et al., 2021 ] incorporated community-specific labels and rationales to enhance transparency.This explainability-oriented paradigm was later extended to non-English settings through HateBRX- plain [ Salles et al., 2025 ] . Despite these advances, unimodal datasets remain limited in dynamic real-world scenarios. They strip away the visual and semiotic cues prevalent in social media platforms, creat- ing a critical gap in capturing the full spectrum of hate speech. 2.2 Multimodal Hate Speech Datasets Existing datasets vary in data sources, annotation granular- ity, and target applications, as summarized in Table 1. The Hateful Memes Challenge Set [ Kiela et al., 2020 ] is among the earliest multimodal hate speech datasets, offering synthet- ically constructed samples with binary labels to assess the vision-language fusion capability of MLLMs. HatReD [ Hee et al., 2023 ] further enriches annotations by capturing meme entities and additional socio-cultural context. Efforts to capture real-world data, such as HarMeme [ Pra- manick et al., 2021 ] and MAMI [ Fersini et al., 2022 ] , pro- vide authentic samples beyond binary labels but are largely restricted to single-domain issues (e.g., COVID-19, misog- yny). While MMHS150K [ Gomez et al., 2020 ] incorporates post text, its labels remain relatively flat and lack the struc- tured explanations. Unlike most datasets that treat memes as isolated images, M 3 preserves the multimodal context (memes + post) and pro- vides hierarchical annotations that align more closely with complex real-world hate speech dynamics. Furthermore, un- like ExMute [ Debnath et al., 2025 ] , which mainly focuses on Bengali and language-level phenomena, M 3 provides broader multilingual coverage such as English and Chinese, bridging Western and Eastern social media ecosystems. 3 Agentic Annotation Framework The Agentic Annotation Framework (illustrated in Figure 2) is a collaborative, multi-agent system designed to construct the M 3 with hierarchical labels and rationales from raw social media data. It orchestrates seven specialized agents across four phases, forming a systematic pipeline for large-scale data acquisition and preprocessing, hierarchical annotation, and expert-driven quality assurance. 3.1 Data Acquisition Collector. The workflow begins with the Collector, the primary gateway for data acquisition. Its input consists of platform-specific data streams, and it outputs images and raw metadata including post id, posttime, imgid, imgurl, and posttext (❶ in Figure 2). To capture the evolving nature of meme-based hate expression, the Collector continuously harvests data via APIs from X 1 , Weibo 2 , and the /pol/ (politically incorrect) board of 4chan 3 , spanning content from January to March 2024. In total, the Collector acquires 3,811,443 image–metadata pairs, including 43,567 from X, 2,090,793 from Weibo, and 1,677,083 from 4chan. Through- out the process, the Collector enforces privacy-preserving constraints by collecting only essential content without per- sonal identifiers and securely storing all data in a local envi- ronment. 1 https://docs.x.com/x-api/introduction 2 https://open.weibo.com/ 3 https://github.com/4chan/4chan-API 3.2 Preprocessing Extractor. Following data acquisition, the Extractor ini- tiates the processing phase by bridging visual and textual modalities in memes. Given the raw images gathered by the Collector, the Extractor invokes an OCR tool (PaddleOCR 4 in our implementation) to extract text embedded within images (img text), which is then combined with the original post context (post text) to form a complete textual represen- tation. Beyond recovering textual cues, the Extractor makes textual content observable to support strict downstream fil- tering, enabling subsequent agents to systematically identify images with excessively long text or no text at all. During this stage, some instances are inevitably filtered out due to unsuccessful tool execution. Cleaner. The Cleaner subsequently refines the combined textual outputs (post text and imgtext) from the Col- lector and Extractor. It removes images whose embedding text is overly verbose or entirely absent, as well as samples containing platform-specific noise such as URLs, hashtags (#), user mentions (@), and quoted content (<<). While im- ages without textual content may still convey hateful intent through visual symbolism alone, enforcing such constraints is necessary to maintain high precision in large-scale meme fil- tering. Through text normalization and length-based filtering, the Cleaner produces a high-density multimodal candidate set of 535,471 samples (❷ in Figure 2). 3.3 Hierarchical Annotation Annotator. As an MLLM-driven expert, the three Anno- tators independently execute hate speech detection and cat- egory classification. Given the candidate memes, the An- notator performs coarse-to-fine annotation. It first conducts general hate speech detection to assign a binary Normal/Hate label. Then, conditioned on a hate decision, it carries out fine- grained classification across eight predefined domains (❸ in Figure 2). The Annotator leverages the zero-shot capabili- ties of MLLMs to capture subtle, multilingual, and cultur- ally contextualized hate expressions (implementation details in the supplementary material section 1.3). Arbiter. To ensure the reliability of these automated labels, the Arbiter acts as a consensus-monitoring unit. It aggre- gates multiple independent outputs produced by Annotators and evaluates their agreement at both the label and category levels. High-consistency predictions like unanimous hate la- bels or largely aligned category assignments (❸ in Figure 2) are automatically accepted, while low-consistency samples are routed to a GUI (see Figure 1 and Figure 2 in the supple- mentary material) for manual review. Samples with persistent ambiguities are excluded to maintain data integrity, leaving 3,179 high-quality verified annotations after this arbitration process. Explicator. For samples identified as hate speech, the Ex- plicator introduces an additional semantic layer by generat- ing structured natural language rationales. Taking memes to- gether with their accepted labels and categories as input, the 4 https://github.com/PaddlePaddle/PaddleOCR Images Metadata •post_id•post_time •img_id•img_url •post_text 12a 2b Memes Metadata MetadataHate memes? Normal Normal Normal Hate Normal Hate 3a3b Label Hate Normal Race Religion Race Gender Gender, Religion Religion Category of hate memes? International Relation Violence Politics Religion Gender Public Health Race Health Status Politics Violence Race M ퟑ Agentic Annotation Framework X Weibo 4chan Collector 12a2b3b3a ExtractorCleaner Annotator ArbiterExplicator Validator 4 4 •post_id•post_time •img_id•img_url •post_text+img_text post_text& img_text: no URLno quote no tagno mention img: img_id img_text: EXT... post_text: Kill all the kikes ... label: hate category: race, violence rationale: Insult Jews; Spread threats of violence Figure 2: The agentic annotation framework for M 3 .❶ Acquisition: Collector harvests multi-platform images and metadata.❷ Preprocessing: Extractor and Cleaner perform OCR and metadata refinement.❸ Annotation: Annotators, Arbiter, and Explicator collaborate on classification and rationale generation.❹ Validation: Validator conducts quality assurance to finalize the M 3 dataset (sample entry on the right). Explicator synthesizes visual cues and textual context to artic- ulate the rationale behind each classification. Concretely, its output follows a controlled verb–object phrase format, such as “mock a religious group” or “incite violence against immi- grants”, which explicitly encodes the action and the targeted entity. 3.4 Quality Assurance Validator. Finally, the workflow concludes with the Val- idator (4), which audits all candidate labels, categories, and rationales. This stage adopts voting without modification, in order to strictly assess the reliability of Arbiter’s deci- sions.Three graduates with complementary disciplinary backgrounds (one in sociology and two in computer science) independently voted on 3,179 samples. Instances receiving only one vote are discarded, resulting in a final set of 2,455 samples. For a random audit of 200 high-consistency sam- ples, inter-rater agreement remains high (93.7%), further val- idating the effectiveness of the Arbiter in generating high- consistency annotations and stability of the proposed agentic annotation framework. 4 M 3 As the final outcome of our agentic annotation frame- work, Figure 2 presents an example from the M 3 dataset. M 3 consists of several structured fields (img, img text, Figure 3: Visualizing linguistic patterns in M 3 . The top panel dis- plays the word cloud of posts in hate samples from X, while the bottom-left and bottom-right panels illustrate the word cloud of Weibo and 4chan, respectively. post text, label, category, and rationale), en- abling a comprehensive assessment of MLLMs with respect to hateful meme detection, categorization, and explanation. We provide a detailed analysis of M 3 below. 4.1 Overview Following rigorous annotation and filtering, M 3 comprises 2,455 multimodal samples with 1,400 from 4chan, 526 from X, and 529 from Weibo, covering multiple languages. Each sample consists of an image paired with an accompanying Single-label samples 1018 Multi-label samples 300 public health 85 race 334 violence 297 health status 106 politics 348 gender 172 religion 108 international relations 180 hate 1318 Figure 4: Hierarchical categories in M 3 . Hate samples are catego- rized into eight themes, with 1,018 single-labeled and 300 multi- labeled samples. textual post. On average, the post length is 125.96 charac- ters, with a maximum of 789 and a minimum of 20. M 3 is balanced across the top-level labels, comprising 1,318 hate samples and 1,137 normal samples. 4.2 Multi-lingual and Multi-Platform Diversity Motivation for Platform Selection. We select platforms that differ substantially in moderation intensity, linguistic coverage, and cultural style to increase data diversity. • X: As a global social media platform with rapid informa- tion diffusion [ Ferrara et al., 2016 ] , X provides multilin- gual content spanning English, Arabic, and other Latin scripts, contributing to critical linguistic diversity. • Weibo: As one of the largest social media platforms in China [ Li et al., 2023 ] , Weibo serves as a primary source of large-scale Chinese-language multimodal con- tent, thereby enriching the cultural diversity. • 4chan (/pol/): The /pol/ board on 4chan is character- ized by anonymity and minimal moderation, resulting in a high density of extreme and explicit hate content [ Col- ley and Moore, 2022 ] . It allows M 3 to capture the upper bound of hateful visual–textual expressions rarely ob- served on mainstream platforms. M 3 encompasses a wide variety of languages, such as En- glish, Chinese, and Arabic (details in the supplementary ma- terial section 2), reflecting the globalized nature of online hate speech. As shown in Figure 3, the word cloud re- veals pronounced multi-platform heterogeneity in hate ex- pressions, indicating that M 3 spans a wide spectrum of hate explicitness. On 4chan, high-frequency terms such as “nig- ger”, “kike”, and “faggot” co-occur with explicit profan- ity (e.g., “fuck”, “shit”), indicating direct group targeting and overt dehumanization. In contrast, hate expressions on X are largely embedded within political and conflict-related discourse, with frequent references to entities such as “Is- rael”, “Gaza”, “Trump”, and “Biden”. On Weibo, high- frequency terms center on national identities (“Chia”, “Amer- ica”, “Japan”), indicating that hate is predomaintly articulated through event-driven discourse. 4.3 Hierarchical Multi-label Categorization Category Definition. We further operationalize the defini- tion of hate speech by UN [ Guterres, 2019 ] into eight the- matic categories to facilitate fine-grained analysis. Each cate- gory represents a distinct and socially significant form of hate commonly observed online: • Religion: Memes that promote harmful content related to religious conflict, such as disputes or hostility be- tween religious sects or groups. • Politics: Memes about political conflict, including hate stemming from government policy failures or partisan disputes, especially during elections. • Race: Memes containing racially discriminatory con- tent, targeting individuals or groups based on ethnicity or race. • Gender: Memes that involve gender-based discrimina- tion, including harmful content targeting women or the LGBTQ+ community. • Health Status: Memes that mock, insult, or discrimi- nate against individuals based on their health conditions, such as disabilities or chronic illnesses. • Violence: Memes that incite or glorify acts of violence, encompassing content that explicitly or implicitly advo- cates targeted shootings, promotes or incites online ha- rassment or cyber violence. • Public Health: Memes spreading misinformation or fear-mongering about public health crises, like famine panic or pandemic-related hate speech. • International Relations: Memes targeting international relations, for example, provoking inter-country hostility during conflicts or fueling anti-refugee sentiment. The per-category sample distribution is shown in Figure 4. Notably, due to data collection from the /pol/ board on 4chan, the politics category accounts for 26.4% of hate sam- ples. Among hate samples in M 3 , 22.76% contain multiple category labels. Statistical analysis shows that the most fre- quent multi-label combinations include (“race”, “violence”) and (“international relations”, “politics”) (58 samples each), followed by (“politics”, race”) and (“politics”, violence”) (40 samples each), which is consistent with intuitive discourse patterns on these platforms. 4.4 Rationales for hate memes For the hate samples, we annotate 1,557 rationales describ- ing why the content is hateful. Each rationale follows a <verb><object> structure (e.g., insult black people), with an average length of 32.35 characters, the longest being 101, and the shortest 9. As shown in Figure 5, hate samples are an- notated with a single rationale, while samples associated with multiple rationales constitute a smaller portion across all plat- forms. The three most common rationales include “express political hatred” (44 times in the politics category), “depre- ciate transgender individuals” (33 times in the gender cate- gory), and “discriminate against people with intellectual dis- abilities” (29 times in the health status category). ModelMemePostOverall↑ BinaryMulti-classRationale Acc↑P↑R↑F1↑Macro-P↑Macro-R↑Macro-F1↑HL↓Subset Acc↑BLEU↑ROUGE↑BERTScore↑ LLaVA-v1.6-Vicuna-7B-hf 19.8353.6953.69100.0069.8647.9229.4832.8916.189.480.434.3397.39 21.8253.6953.69100.0069.8626.4069.4635.6135.0713.200.061.8997.41 LLaVA-v1.6-Vicuna-13B-hf 19.9766.5261.8997.9575.8547.9229.4832.8916.189.480.405.0396.38 20.0359.4768.8444.7654.2540.7561.3944.7521.4814.260.294.5195.90 GLM-4.1V-9B-Thinking 77.4276.7498.1957.7472.7258.7086.5069.0211.6640.521.048.6797.76 80.6177.1192.0962.7574.6454.6881.9864.4913.9633.761.199.1797.79 Qwen2.5-VL-3B-Instruct 51.6376.2178.6376.4877.5456.5175.5662.0014.9828.910.333.7197.19 58.4174.7588.6160.7772.1051.8474.4758.3316.2724.130.255.2497.07 Qwen2.5-VL-7B-Instruct 82.1990.2691.1590.6790.9959.0879.5065.5412.4239.530.857.1997.68 77.4386.4886.6388.4787.5456.4873.6459.5914.7430.800.646.6397.34 Qwen3-VL-8B-Instruct 87.6386.8091.7682.8587.0861.7179.8668.1711.8642.111.489.1897.51 84.8385.9587.0586.7286.8955.3676.6462.5314.6333.161.508.6497.42 GPT-4o 60.3786.2794.1780.0886.5650.679.8760.9315.1626.260.107.9597.47 62.9685.4783.6589.4686.4549.0462.5753.4016.0419.890.147.8797.43 Gemini-3 59.4266.9792.7541.7357.5653.9785.6364.6015.2433.231.7411.6497.28 75.1673.4486.4359.9470.7948.1884.4259.7418.2025.172.1713.0797.51 Table 2: Comparison of state-of-the-art MLLMs on the M 3 dataset. The models are listed in order from open-source to proprietary, following their chronological release or version evolution: (1) the early LLaVA-v1.6 series (7B, 13B), (2) the reasoning-enhanced GLM-4.1V-9B- Thinking, (3) the latest Qwen series (Qwen2.5-VL-3B/7B and Qwen3-VL-8B), and (4) the proprietary GPT-4o and Gemini-3. Results cover binary and multi-class classification, and rationale quality under meme-only and meme + post settings. Boldface denotes the extremal values under the meme-only setting, while boldface with underline denotes the extremal values under the meme + post setting. 8.6% 7.4% 66.2% 14.3% 0.8% 2.6% WeiboX4chanSingle rationaleMultiple rationales Figure 5: The distribution of single rationale and multiple rationales across X, Weibo, and 4chan. In summary, M 3 is a thematically diverse multimodal benchmark with broad coverage. Its inclusion of real-world social media contexts, hierarchical multi-label annotations, and comprehensive rationales makes it a valuable resource for evaluating the nuanced hate recognition and interpreta- tion capabilities of MLLMs, especially within the real-world dynamic environments. 5 Experiments 5.1 Experiment Setups Dataset and Baselines. Experiments are conducted on M 3 , containing 2,455 memes paired with corresponding posts. Each instance is annotated with binary labels, and instances labeled as hate are further annotated with fine-grained cate- gories and phrase-level rationales. We compare several rep- resentative MLLMs, categorized into open-source and pro- prietary models, and ordered by their release dates or ver- sions: (1) Open-source models include LLaVA-v1.6 series (7B and 13B, Feb 2024), GLM-4.1V-9B-Thinking (Jul 2025), Qwen2.5-VL series (3B and 7B, Jan 2025), Qwen3-VL-8B- Instruct (Oct 2025); (2) Proprietary models include GPT-4o (May 2024) and Gemini-3 (Nov 2025). GPT-4o and Gemini- 3 are accessed via the official API, whereas the other models are deployed locally. Tasks and Evaluation Metrics. We investigate three meme understanding tasks: (1) Binary hate detection, which pre- dicts whether a meme contains hateful content (hate vs. nor- mal); (2) Multi-label fine-grained classification, identifying specific categories present in the hateful meme; and (3) Ra- tionale generation, which produces a concise verb–object phrase explaining why the meme is hateful. For binary clas- sification, we report Accuracy, Precision, Recall, and F1- score. For multi-label fine-grained classification, we compute macro-Precision, macro-Recall, macro-F1, Hamming Loss, and Subset Accuracy [ Zhang and Zhou, 2013 ] . For rationale generation, we assess the quality of generated rationales us- ing BLEU [ Papineni et al., 2002 ] , ROUGE [ Lin, 2004 ] and BERTScore [ Zhang et al., 2019 ] . We define an Overall score by first normalizing all metrics to [0, 1] (inverting where nec- essary) and taking average across the three tasks. Implementation Details. For each model, we explore two input settings: (i) Meme-only, where only the meme image is provided; and (i) Meme + Post, where the meme is paired with its associated post text. All experiments are conducted on a single NVIDIA A800 GPU (80GB). To ensure fair com- parison, no task-specific fine-tuning is performed; instead, models are directly evaluated on downstream tasks in a zero- shot setting. 5.2 Results and Discussion Different Tasks. According to Table 2, a clear task-level performance hierarchy emerges across all evaluated mod- els. models achieve robust results in binary classification (typically> 85% accuracy) but struggle with multi-class tasks (Macro-F1: 32.89%–69.02%). Notably, scaling model Figure 6: Performance comparison of MLLMs across three platforms. From left to right: 4chan (English-centric), X (Multi-lingual, e.g., Latin and Arabic), and Weibo (Chinese-dominant). capacity yields non-uniform gains. For example, increas- ing model capacity from LLaVA-7B to LLaVA-13B yields only marginal improvements in binary accuracy but leads to a noticeable increase in multi-class macro-F1. Similarly, within the Qwen family, scaling from 3B to 7B results in clear gains in binary classification (+13.7 accuracy points), whereas further scaling to Qwen3-VL-8B produces diminish- ing returns for binary accuracy but more pronounced benefits for multi-class macro-F1 and rationale metrics. A unique pat- tern emerges in rationale generation where high BERTScore values (above 95.0) coexist with low lexical-overlap metrics, including BLEU scores below 2.2 and a peak ROUGE of 11.64 (Gemini-3, meme-only). This discrepancy arises be- cause BLEU and ROUGE operate at the word level, failing to capture the semantic alignment of our generated rationales, which primarily consist of verb-object phrases with an aver- age length of 32.35 characters. Multimodal Inputs. As shown in Table 2, incorporating posts alongside memes produces heterogeneous effects across models and tasks, yielding inconsistent performance gains across the evaluated dimensions. Adding post information yields a negligible impact on binary classification (F1 ±3 points) but triggers a consistent performance decline in multi- class tasks. Specifically, macro-F1 scores drop in 6 of 9 evaluated models, with significant in high-performing mod- els like Qwen3-VL and GPT-4o. Although LLaVA showed an 11.86-point improvement, this behavior represents an ex- ception rather than the dominant pattern. Across newly re- leased models, BLEU and ROUGE improve when post text is included, suggesting that posts provide complementary se- mantic cues that facilitate rationale construction. Taken to- gether, these results indicate that current MLLMs struggle to integrate meme and post information for fine-grained in- tent understanding robustly. While additional context may help rationale generation, it often introduces ambiguity that degrades classification performance, highlighting a key limi- tation in real-world hate speech moderation scenarios where user intent is distributed across modalities. Multi-platform and Multi-lingual Evaluation. We se- lect four multimodal models (LLaVA-v1.6-Vicuna-13B- hf, GLM-4.1V-9B-Thinking, Qwen3-VL-8B-Instruct, and Gemini-3) from different model families that demonstrate strong overall performance in prior experiments and evaluate them across 4chan, X, and Weibo. A pronounced precision- recall trade-off emerges on X and Weibo, particularly for Gemini-3, which achieves perfect precision but a meager 16.94 recall on Weibo. These results suggest that certain models adopt overly conservative prediction strategies, pri- oritizing precision while sacrificing recall. Although this re- duces false positives, it leads to a large number of false nega- tives and limits practical usefulness. The effect is particularly pronounced under domain and language shifts, indicating that alignment or thresholding mechanisms may bias mod- els toward excessively “safe” predictions in multi-platform settings. From a multi-lingual perspective, models gener- ally perform better on English-dominated platforms (4chan) compared to non-English or mixed-language platforms like X and Weibo. Performance degradation is particularly no- ticeable for Chinese-language content, where high-precision models like Gemini-3 still suffer from dramatic recall drops, highlighting challenges in cross-linguistic generalization. For clarity, we visualize only the results under the meme + post setting. Results are reported in full in the supplementary ma- terial section 3.1. 6 Conclusion In this work, we introduce M 3 , a multi-platform, multi- lingual, and multimodal meme dataset constructed through an agentic annotation framework with systematic human ver- ification. M 3 contains 2,455 multimodal instances from X, 4chan, and Weibo, annotated with binary hate labels, fine- grained categories, and human-verified rationales. We bench- mark M 3 on state-of-the-art MLLMs. Results show that in- corporating surrounding post context does not consistently improve hate detection and often degrades fine-grained clas- sification, although it can benefit rationale generation. These findings highlight current limitations of MLLMs in integrat- ing multimodal contextual information. Overall, M 3 provides a challenging benchmark for multimodal hate speech analy- sis and offers insights into the gap between existing model capabilities and real-world moderation needs. Ethical Statement This work involves the analysis of hateful memes that may contain offensive content. All data in M 3 are collected from publicly available platforms and are used solely for research purposes. Personally identifiable information is removed dur- ing data processing. Annotations are conducted with human verification, and annotators are informed of the sensitive nature of the content. This study does not endorse hateful expressions. The dataset is intended to support research on multimodal hate detection and should be used responsibly in accordance with ethical guidelines. Acknowledgments and Disclosure of Funding This work was supported in part by the National Natural Sci- ence Foundation of China (62306229), the Youth Talent Sup- port Program of Shaanxi Science and Technology Associa- tion (20240113), the China Postdoctoral Science Foundation (2025T180425). References [ Barnes, 2019 ] Luke Barnes. With each new attack, far-right extremists’ manifestos are being ‘memeticized,’, 2019. [ Brown, 2018 ] Alexander Brown. What is so special about online (as compared to offline) hate speech? Ethnicities, 2018. [ Chhabra and Vishwakarma, 2023 ] Anusha Chhabra and Di- nesh Kumar Vishwakarma. A literature survey on multi- modal and multilingual automatic hate speech identifica- tion. Multimedia Systems, pages 1203–1230, 2023. [ Colley and Moore, 2022 ] ThomasColleyandMartin Moore. The challenges of studying 4chan and the alt- right:‘come on in the water’s fine’. New Media & Society, 24(1):5–30, 2022. [ Debnath et al., 2025 ] RiddhimanSwananDebnath, Nahian Beente Firuj, Abdul Wadud Shakib, Sadia Sultana, and Md Saiful Islam.ExMute: A context- enriched multimodal dataset for hateful memes.In NLPAIDL, pages 83–89, 2025. [ Ferrara et al., 2016 ] Emilio Ferrara, Onur Varol, Clayton Davis, Filippo Menczer, and Alessandro Flammini. The rise of social bots. Communications of the ACM, 59(7):96– 104, 2016. [ Fersini et al., 2022 ] Elisabetta Fersini, Francesca Gasparini, Giulia Rizzi, Aurora Saibene, Berta Chulvi, Paolo Rosso, Alyssa Lees, and Jeffrey Sorensen. Semeval-2022 task 5: Multimedia automatic misogyny identification. In Se- mEval, pages 533–549, 2022. [ Founta et al., 2018 ] Antigoni Founta, Constantinos Djou- vas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Black- burn, Gianluca Stringhini, Athena Vakali, Michael Siriv- ianos, and Nicolas Kourtellis. Large scale crowdsourc- ing and characterization of twitter abusive behavior. In ICWSM, 2018. [ Gomez et al., 2020 ] Raul Gomez, Jaume Gibert, Lluis Gomez, and Dimosthenis Karatzas. Exploring hate speech detection in multimodal publications. In WACV, pages 1470–1478, 2020. [ Guterres, 2019 ] Ant ́ onio Guterres. United nations strategy and plan of action on hate speech. Technical report, United Nations, 2019. [ Hee et al., 2023 ] Ming Shan Hee, Wen-Haw Chong, and Roy Ka-Wei Lee.Decoding the underlying mean- ing of multimodal hateful memes.arXiv preprint arXiv:2305.17678, 2023. [ Hee et al., 2024 ] Ming Shan Hee, Shivam Sharma, Rui Cao, Palash Nandi, Tanmoy Chakraborty, and Roy Ka-Wei Lee. Recent advances in hate speech moderation: Multimodal- ity and the role of large models. In EMNLP, 2024. [ Jiang and Zubiaga, 2024 ] Aiqi Jiang and Arkaitz Zubiaga. Cross-lingual offensive language detection: A systematic review of datasets, transfer approaches and challenges. arXiv preprint arXiv:2401.09244, 2024. [ Kiela et al., 2020 ] Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine.The hateful memes challenge: Detecting hate speech in multimodal memes. NeurIPS, pages 2611–2624, 2020. [ Li et al., 2023 ] Lifang Li, Hong Wen, and Qingpeng Zhang. Characterizing the role of weibo and wechat in sharing original information in a crisis. Journal of Contingencies and Crisis Management, 31(2):236–248, 2023. [ Lin, 2004 ] Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In TSB, pages 74–81, 2004. [ Mathew et al., 2021 ] BinnyMathew,PunyajoySaha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. Hatexplain: A benchmark dataset for explainable hate speech detection. In AAAI, pages 14867–14875, 2021. [ Mulki et al., 2019 ] HalaMulki,HatemHaddad, Chedi Bechikh Ali, and Halima Alshabani.L-hsab: A levantine twitter dataset for hate speech and abusive language. In ALW, pages 111–118, 2019. [ Nayak and Agrawal, 2022 ] AjayNayakandAnupam Agrawal.Detection of hate speech in social media memes: A comparative analysis.In ICICICT, pages 1179–1185, 2022. [ Pandiani et al., 2025 ] Delfina S Martinez Pandiani, Erik Tjong Kim Sang, and Davide Ceolin. ‘toxic’memes: A survey of computational perspectives on the detection and explanation of meme toxicities. Online Social Networks and Media, page 100317, 2025. [ Papineni et al., 2002 ] Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for au- tomatic evaluation of machine translation. In ACL, pages 311–318, 2002. [ Pramanick et al., 2021 ] Shraman Pramanick, Dimitar Dim- itrov, Rituparna Mukherjee, Shivam Sharma, Md Shad Akhtar, Preslav Nakov, and Tanmoy Chakraborty. De- tecting harmful memes and their targets. In ACL-IJCNLP, pages 2783–2796, 2021. [ Rao et al., 2023 ] Xiaojun Rao, Yangsen Zhang, Qilong Jia, Xueyang Liu, Shuang Peng, et al. Chinese hate speech detection method based on roberta-wwm. In CCL, pages 501–511, 2023. [ Salles et al., 2025 ] Isadora Salles, Francielle Vargas, and Fabr ́ ıcio Benevenuto. Hatebrxplain: A benchmark dataset with human-annotated rationales for explainable hate speech detection in brazilian portuguese. In COLING, pages 6659–6669, 2025. [ Schmid, 2025 ] Ursula Kristin Schmid.Humorous hate speech on social media: A mixed-methods investigation of users’ perceptions and processing of hateful memes. New Media & Society, pages 1588–1606, 2025. [ Shifman, 2013 ] Limor Shifman. Memes in digital culture. MIT press, 2013. [ Tonneau et al., 2025 ] Manuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A Hale, Samuel Fraiberger, Victor Orozco-Olvera, and Paul R ̈ ottger. Hateday: Insights from a global hate speech dataset representative of a day on twit- ter. In ACL, pages 2297–2321, 2025. [ Velasquez et al., 2021 ] Nicolas Velasquez, Rhys Leahy, N Johnson Restrepo, Yonatan Lupu, Richard Sear, Nicholas Gabriel, OK Jha, Beth Goldberg, and NF John- son. Online hate network spreads malicious covid-19 con- tent outside the control of individual social media plat- forms. Scientific reports, page 11549, 2021. [ Ware, 2022 ] Jacob Ware. Testament to Murder: The Violent Far-Right’s Increasing Use of Terrorist Manifestos. JS- TOR, 2022. [ Waseem and Hovy, 2016 ] Zeerak Waseem and Dirk Hovy. Hateful symbols or hateful people? predictive features for hate speech detection on twitter. In NAACL-HLT, pages 88–93, 2016. [ Zampieri et al., 2019 ] Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. Predicting the type and target of offensive posts in social media. arXiv preprint arXiv:1902.09666, 2019. [ Zhang and Zhou, 2013 ] Min-Ling Zhang and Zhi-Hua Zhou.A review on multi-label learning algorithms. IEEE transactions on knowledge and data engineering, 26(8):1819–1837, 2013. [ Zhang et al., 2019 ] Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert.arXiv preprint arXiv:1904.09675, 2019. [ Zhou et al., 2021 ] Yi Zhou, Zhenhao Chen, and Huiyuan Yang. Multimodal learning for hateful memes detection. In ICMEW, pages 1–6, 2021.