Paper deep dive
Is AI Ready for Multimodal Hate Speech Detection? A Comprehensive Dataset and Benchmark Evaluation
Rui Xing, Qi Chai, Jie Ma, Jing Tao, Pinghui Wang, Shuming Zhang, Xinping Wang, Hao Wang
Intelligence
Status: succeeded | Model: google/gemini-3.1-flash-lite-preview | Prompt: intel-v1 | Confidence: 96%
Last extracted: 3/26/2026, 1:27:10 AM
Summary
The paper introduces M^3, a multi-platform, multi-lingual, and multimodal meme dataset containing 2,455 instances collected from X, 4chan, and Weibo. It proposes an agentic annotation framework utilizing seven specialized agents to generate hierarchical hate speech labels and rationales. Benchmarking against state-of-the-art Multimodal Large Language Models (MLLMs) reveals that current models struggle to effectively integrate surrounding post context, often leading to performance degradation, highlighting the need for context-aware architectures.
Entities (6)
Relation Signals (3)
M^3 ā collectedfrom ā X
confidence 100% Ā· a dataset of 2,455 memes collected from X, 4chan, and Weibo
Agentic Annotation Framework ā produced ā M^3
confidence 100% Ā· Based on this framework, we construct M^3
MLLMs ā benchmarkedon ā M^3
confidence 95% Ā· We further evaluate M^3 on state-of-the-art Multimodal Large Language Models
Cypher Suggestions (0)
No Cypher suggestions yet.
Abstract
Abstract:Hate speech online targets individuals or groups based on identity attributes and spreads rapidly, posing serious social risks. Memes, which combine images and text, have emerged as a nuanced vehicle for disseminating hate speech, often relying on cultural knowledge for interpretation. However, existing multimodal hate speech datasets suffer from coarse-grained labeling and a lack of integration with surrounding discourse, leading to imprecise and incomplete assessments. To bridge this gap, we propose an agentic annotation framework that coordinates seven specialized agents to generate hierarchical labels and rationales. Based on this framework, we construct M^3 (Multi-platform, Multi-lingual, and Multimodal Meme), a dataset of 2,455 memes collected from X, 4chan, and Weibo, featuring fine-grained hate labels and human-verified rationales. Benchmarking state-of-the-art Multimodal Large Language Models reveals that these models struggle to effectively utilize surrounding post context, which often fails to improve or even degrades detection performance. Our finding highlights the challenges these models face in reasoning over memes embedded in real-world discourse and underscores the need for a context-aware multimodal architecture. Our dataset and code are available at this https URL.
Tags
Links
- Source: https://arxiv.org/abs/2603.21686v2
- Canonical: https://arxiv.org/abs/2603.21686v2
Trouble viewing inline? Open PDF directly ā
Full Text
65,119 characters extracted from source content.
Expand or collapse full text
Is AI Ready for Multimodal Hate Speech Detection? A Comprehensive Dataset and Benchmark Evaluation Rui Xing 1ā , Qi Chai 2ā , Jie Ma 1,3ā , Jing Tao 1 , Pinghui Wang 1 , Shuming Zhang 4 , Xinping Wang 3 , Hao Wang 2 1 MOE KLINNS Lab, Xiāan Jiaotong University 2 The Hong Kong University of Science and Technology (Guangzhou) 3 School of Cyber Science and Engineering, Xiāan Jiaotong University 4 Northwest University * Equal contribution ā Corresponding author jiema@xjtu.edu.cn Abstract Hate speech online targets individuals or groups based on identity attributes and spreads rapidly, posing serious social risks. Memes, which com- bine images and text, have emerged as a nu- anced vehicle for disseminating hate speech, of- ten relying on cultural knowledge for interpreta- tion. However, existing multimodal hate speech datasets suffer from coarse-grained labeling and a lack of integration with surrounding discourse, leading to imprecise and incomplete assessments. To bridge this gap, we propose an agentic anno- tation framework that coordinates seven special- ized agents to generate hierarchical labels and ratio- nales. Based on this framework, we construct M 3 (Multi-platform, Multi-lingual, and Multimodal Meme), a dataset of 2,455 memes collected from X, 4chan, and Weibo, featuring fine-grained hate la- bels and human-verified rationales. Benchmarking state-of-the-art Multimodal Large Language Mod- els reveals that these models struggle to effec- tively utilize surrounding post context, which often fails to improve or even degrades detection perfor- mance. Our finding highlights the challenges these models face in reasoning over memes embedded in real-world discourse and underscores the need for a context-aware multimodal architecture. Our dataset and code are available at https://github.com/ mira-ai-lab/M3. Disclaimer: This paper includes content that may be considered offensive or disturbing to some read- ers. 1 Introduction Hate speech [ Guterres, 2019 ] refers to āany kind of commu- nication in speech, writing, or behavior that attacks or uses pejorative or discriminatory language with reference to a per- son or a group on the basis of who they are; in other words, Existing Dataset Rationale: No hateful content. Label: Our Dataset (M ķ ) Post:Absolute mongoloids. Objectively ugly. Abominations. Rationale: Discriminate against individuals with Downās syndrome. Label: HateNormal HateNormal Health Status Other categories Figure 1: Comparison between existing datasets and ours (M 3 ). Ex- isting datasets typically label the meme (right) as normal. However, our dataset labels it as hate with a refined classification of healthy state-because of its accompanying post. based on their religion, ethnicity, nationality, race, color, de- scent, gender, or other identity factor.ā Its rapid dissemi- nation across online platforms poses a serious threat to so- cial stability [ Velasquez et al., 2021 ] . For instance, the 2019 Christchurch mosque shootings in New Zealand and the ter- rorist attack in El Paso, Texas, that same year. Related reports and studies [ Barnes, 2019; Ware, 2022 ] indicate that the per- petrators incorporated Internet Memes into their manifestos or posts to spread hate speech and resonate with specific on- line subcultures, thus inciting extremist ideologies. A meme [ Shifman, 2013 ] is a multimodal composite con- sisting of an image and short text. As one of the vehicles for disseminating hate speech [ Pandiani et al., 2025 ] , the meme is easily produced. It may combine a humorous image with a slogan that incites violence [ Zhou et al., 2021 ] , or convey discriminatory meanings through visual elements, with the embedded textual content appearing neutral in isolation [ Hee et al., 2024 ] . Within online communities, memes are often in- conspicuously embedded in otherwise ordinary posts and are interpretable only by users who share specific cultural or sub- cultural knowledge. This subtle embeddingācombined with their multimodal ambiguityāfacilitates rapid dissemination through densely connected social networks [ Brown, 2018; Schmid, 2025 ] . These characteristics not only make hateful arXiv:2603.21686v2 [cs.MA] 25 Mar 2026 DatasetDomain Label Img textPostRationaleMethodSource Hateful?Categories The Hateful Memes Challenge Set Multiple fields Hateful, Not-hateful - --Human-onlySynthetical HatReDMultiple fields Hateful, Not-hateful --Human-onlySynthetical HarMemeCOVID-19 Very harmful, Partially harmful, Harmless - --Human-only Google Image Search, Reddit, Facebook, Instagram MMHS150KMultiple fields- Not hate, Racist, Sexist, Homophobic, Religion, Other hate --Human-onlyTwitter MAMIMisogynyMisogynous, Not-misogynous Shaming, Stereotype, Objectification, Violence --Human-onlyTwitter, Reddit ExMute-Hateful, Non-hateful Religious, Celebrity, Political, Male, Female, Others --Human-onlyFacebook, Reddit, Instagram M 3 (Ours)Multiple fieldsHate, Normal Religion, Politics, Race, Gender, Health Status, Violence, Public Health, International Relations Human-validate agentic X, 4chan, Weibo Table 1: Comparison of representative multimodal hate speech datasets. While existing datasets often provide coarse labels or limited context, OURS introduces fine-grained multi-dimensional annotations and additionally includes the surrounding post content, enabling richer and more context-aware hate speech analysis. memes particularly insidious but also expose critical gaps in current multimodal hate speech detection datasets. Existing datasets [ Chhabra and Vishwakarma, 2023; Jiang and Zubiaga, 2024; Nayak and Agrawal, 2022 ] predomi- nantly support superficial evaluation, which fails to accu- rately assess the true performance of hateful meme detection methods, thereby offering limited constructive feedback for improving hate speech detection systems. On the one hand, they still adopt coarse-grained or flat labeling schemes, creat- ing a significant discrepancy with the multi-dimensional com- plexity of hate speech as defined by UN protocols. As il- lustrated in Figure 1, the simplistic binary labels (i.e., Hate vs. Normal) preclude the representation of intersectional offenses, thus failing to reflect model efficacy in nuanced, multi-class scenarios. On the other hand, they typically iso- late memes from their surrounding discourse, focusing exclu- sively on their immediate content. Since memes in real-world social media are intrinsically tied to accompanying posts, the absence of such contextual information may lead to the erro- neous label or even incomplete and misleading interpretations (see Figure 1). To address the challenges above, we propose an agen- tic annotation framework that coordinates seven specialized agents including a collector, extractor, cleaner, annotator, ar- biter, explicator, and validator. The collector gathers memes and associated posts from X (formerly Twitter), 4chan, and Weibo, enabling broad cross-cultural coverage. After pro- cessing by the extractor and cleaner, the annotator and ar- biter conduct multi-round hierarchical annotation with expert adjudication. Subsequently, the explicator generates hate ra- tionales, while the validator performs quality control. This framework yields M 3 , the Multi-platform, Multi-lingual, and Multimodal Meme dataset, which contains 2,455 high- quality multimodal instances with two top-level labels (hate and normal), and eight fine-grained hate categories. We fur- ther evaluate M 3 on state-of-the-art Multimodal Large Lan- guage Models (MLLMs), including Gemini-3, GPT-4o mod- els and representative open-source models such as Qwen-VL, LLaVA-v1.6, and GLM-4.1V-9B-Thinking. Among them, Qwen3-VL-8B-Instruct achieves the highest overall score across the evaluation metrics. The results show that incorpo- rating surrounding post context leads to degraded accuracy in hate detection and fine-grained classification. These findings suggest that current AI is not yet ready to fully replace human moderation. Moreover, hate speech detection should consider the linguistic context in which memes are embedded. The main contributions are summarized as follows: ⢠We develop an agentic annotation framework, which co- ordinates multiple agents for scalable labeling and ra- tionale generation, while ensuring annotation reliability through systematic human verification. ⢠We introduce M 3 , a multi-platform, multi-lingual, and multimodal meme dataset spanning X, 4chan, and Weibo, featuring fine-grained hate annotations and human-verified rationales to support explainable multi- modal hate analysis. ⢠We provide extensive benchmarking results on state-of- the-art MLLMs, demonstrating the effectiveness of M 3 in evaluating multimodal hate meme detection capabili- ties and revealing current limitations in rationale gener- ation. 2 Related Work 2.1 Unimodal Hate Speech Datasets Early research primarily focused on plain-text content via manual annotation (see Table 8 in Section D of the appendix. Initial English-centric efforts [ Waseem and Hovy, 2016 ] were later scaled through crowdsourcing [ Founta et al., 2018 ] and extended to non-English contexts, including Arabic [ Mulki et al., 2019 ] , Chinese [ Rao et al., 2023 ] , and cross-lingual datasets [ Tonneau et al., 2025 ] . Parallel to this lingual diversification, the annota- tion paradigms shifted toward multifaceted schemes. OLID [ Zampieri et al., 2019 ] employs a three-tier hierar- chy, while HateXplain [ Mathew et al., 2021 ] incorporated community-specific labels and rationales to enhance transparency.This explainability-oriented paradigm was later extended to non-English settings through HateBRX- plain [ Salles et al., 2025 ] . Despite these advances, unimodal datasets remain limited in dynamic real-world scenarios. They strip away the visual and semiotic cues prevalent in social media platforms, creat- ing a critical gap in capturing the full spectrum of hate speech. 2.2 Multimodal Hate Speech Datasets Existing datasets vary in data sources, annotation granular- ity, and target applications, as summarized in Table 1. The Hateful Memes Challenge Set [ Kiela et al., 2020 ] is among the earliest multimodal hate speech datasets, offering synthet- ically constructed samples with binary labels to assess the vision-language fusion capability of MLLMs. HatReD [ Hee et al., 2023 ] further enriches annotations by capturing meme entities and additional socio-cultural context. Efforts to capture real-world data, such as HarMeme [ Pra- manick et al., 2021 ] and MAMI [ Fersini et al., 2022 ] , pro- vide authentic samples beyond binary labels but are largely restricted to single-domain issues (e.g., COVID-19, misog- yny). While MMHS150K [ Gomez et al., 2020 ] incorporates post text, its labels remain relatively flat and lack the struc- tured explanations. Unlike most datasets that treat memes as isolated images, M 3 preserves the multimodal context (memes + post) and pro- vides hierarchical annotations that align more closely with complex real-world hate speech dynamics. Furthermore, un- like ExMute [ Debnath et al., 2025 ] , which mainly focuses on Bengali and language-level phenomena, M 3 provides broader multilingual coverage such as English and Chinese, bridging Western and Eastern social media ecosystems. 3 Agentic Annotation Framework The Agentic Annotation Framework (illustrated in Figure 2) is a collaborative, multi-agent system designed to construct the M 3 with hierarchical labels and rationales from raw social media data. It orchestrates seven specialized agents across four phases, forming a systematic pipeline for large-scale data acquisition and preprocessing, hierarchical annotation, and expert-driven quality assurance. 3.1 Data Acquisition Collector. The workflow begins with the Collector, the primary gateway for data acquisition. Its input consists of platform-specific data streams, and it outputs images and raw metadata including post id, posttime, imgid, imgurl, and posttext (ā¶ in Figure 2). To capture the evolving nature of meme-based hate expression, the Collector continuously harvests data via APIs from X 1 , Weibo 2 , and the /pol/ (politically incorrect) board of 4chan 3 , spanning content from January to March 2024. In total, the Collector acquires 3,811,443 imageāmetadata pairs, including 43,567 from X, 2,090,793 from Weibo, and 1,677,083 from 4chan. Through- out the process, the Collector enforces privacy-preserving constraints by collecting only essential content without per- sonal identifiers and securely storing all data in a local envi- ronment. 1 https://docs.x.com/x-api/introduction 2 https://open.weibo.com/ 3 https://github.com/4chan/4chan-API 3.2 Preprocessing Extractor. Following data acquisition, the Extractor ini- tiates the processing phase by bridging visual and textual modalities in memes. Given the raw images gathered by the Collector, the Extractor invokes an OCR tool (PaddleOCR 4 in our implementation) to extract text embedded within images (img text), which is then combined with the original post context (post text) to form a complete textual represen- tation. Beyond recovering textual cues, the Extractor makes textual content observable to support strict downstream fil- tering, enabling subsequent agents to systematically identify images with excessively long text or no text at all. During this stage, some instances are inevitably filtered out due to unsuccessful tool execution. Cleaner. The Cleaner subsequently refines the combined textual outputs (post text and imgtext) from the Col- lector and Extractor. It removes images whose embedding text is overly verbose or entirely absent, as well as samples containing platform-specific noise such as URLs, hashtags (#), user mentions (@), and quoted content (<<). While im- ages without textual content may still convey hateful intent through visual symbolism alone, enforcing such constraints is necessary to maintain high precision in large-scale meme fil- tering. Through text normalization and length-based filtering, the Cleaner produces a high-density multimodal candidate set of 535,471 samples (ā· in Figure 2). 3.3 Hierarchical Annotation Annotator. As an MLLM-driven expert, the three Anno- tators independently execute hate speech detection and cat- egory classification. Given the candidate memes, the An- notator performs coarse-to-fine annotation. It first conducts general hate speech detection to assign a binary Normal/Hate label. Then, conditioned on a hate decision, it carries out fine- grained classification across eight predefined domains (āø in Figure 2). The Annotator leverages the zero-shot capabilities of MLLMs to capture subtle, multilingual, and culturally con- textualized hate expressions (implementation details in Sec- tion A.3 of the appendix). Arbiter. To ensure the reliability of these automated labels, the Arbiter acts as a consensus-monitoring unit. It aggre- gates multiple independent outputs produced by Annotators and evaluates their agreement at both the label and category levels. High-consistency predictions like unanimous hate la- bels or largely aligned category assignments (āø in Figure 2) are automatically accepted, while low-consistency samples are routed to a GUI (see Figure 7 and Figure 8 in Section A.3 of the appendix) for manual review. Samples with persistent ambiguities are excluded to maintain data integrity, leaving 3,179 high-quality verified annotations after this arbitration process. Explicator. For samples identified as hate speech, the Ex- plicator introduces an additional semantic layer by generat- ing structured natural language rationales. Taking memes to- gether with their accepted labels and categories as input, the 4 https://github.com/PaddlePaddle/PaddleOCR Images Metadata ā¢post_idā¢post_time ā¢img_idā¢img_url ā¢post_text 12a 2b Memes Metadata MetadataHate memes? Normal Normal Normal Hate Normal Hate 3a3b Label Hate Normal Race Religion Race Gender Gender, Religion Religion Category of hate memes? International Relation Violence Politics Religion Gender Public Health Race Health Status Politics Violence Race M ķ Agentic Annotation Framework X Weibo 4chan Collector 12a2b3b3a ExtractorCleaner Annotator ArbiterExplicator Validator 4 4 ā¢post_idā¢post_time ā¢img_idā¢img_url ā¢post_text+img_text post_text& img_text: no URLno quote no tagno mention img: img_id img_text: EXT... post_text: Kill all the kikes ... label: hate category: race, violence rationale: Insult Jews; Spread threats of violence Figure 2: The agentic annotation framework for M 3 .ā¶ Acquisition: Collector harvests multi-platform images and metadata.ā· Preprocessing: Extractor and Cleaner perform OCR and metadata refinement.āø Annotation: Annotators, Arbiter, and Explicator collaborate on classification and rationale generation.ā¹ Validation: Validator conducts quality assurance to finalize the M 3 dataset (sample entry on the right). Explicator synthesizes visual cues and textual context to artic- ulate the rationale behind each classification. Concretely, its output follows a controlled verbāobject phrase format, such as āmock a religious groupā or āincite violence against immi- grantsā, which explicitly encodes the action and the targeted entity. 3.4 Quality Assurance Validator. Finally, the workflow concludes with the Valida- tor (ā¹ in Figure 2), which audits all candidate labels, cate- gories, and rationales. This stage adopts voting without mod- ification, in order to strictly assess the reliability of Arbiterās decisions. Three graduates with complementary disciplinary backgrounds (one in sociology and two in computer science) independently voted on 3,179 samples. Instances receiving only one vote are discarded, resulting in a final set of 2,455 samples. For a random audit of 200 high-consistency sam- ples, inter-rater agreement remains high (93.7%), further val- idating the effectiveness of the Arbiter in generating high- consistency annotations and stability of the proposed agentic annotation framework. 4 M 3 As the final outcome of our agentic annotation frame- work, Figure 2 presents an example from the M 3 dataset. M 3 consists of several structured fields (img, img text, Figure 3: Visualizing linguistic patterns in M 3 . The top panel dis- plays the word cloud of posts in hate samples from X, while the bottom-left and bottom-right panels illustrate the word cloud of Weibo and 4chan, respectively. post text, label, category, and rationale), en- abling a comprehensive assessment of MLLMs with respect to hateful meme detection, categorization, and explanation. We provide a detailed analysis of M 3 below. 4.1 Overview Following rigorous annotation and filtering, M 3 comprises 2,455 multimodal samples with 1,400 from 4chan, 526 from X, and 529 from Weibo, covering multiple languages. Each sample consists of an image paired with an accompanying Single-label samples 1018 Multi-label samples 300 public health 85 race 334 violence 297 health status 106 politics 348 gender 172 religion 108 international relations 180 hate 1318 Figure 4: Hierarchical categories in M 3 . Hate samples are catego- rized into eight themes, with 1,018 single-labeled and 300 multi- labeled samples. textual post. On average, the post length is 125.96 charac- ters, with a maximum of 789 and a minimum of 20. M 3 is balanced across the top-level labels, comprising 1,318 hate samples and 1,137 normal samples. 4.2 Multi-lingual and Multi-Platform Diversity Motivation for Platform Selection. We select platforms that differ substantially in moderation intensity, linguistic coverage, and cultural style to increase data diversity. ⢠X: As a global social media platform with rapid informa- tion diffusion [ Ferrara et al., 2016 ] , X provides multilin- gual content spanning English, Arabic, and other Latin scripts, contributing to critical linguistic diversity. ⢠Weibo: As one of the largest social media platforms in China [ Li et al., 2023 ] , Weibo serves as a primary source of large-scale Chinese-language multimodal con- tent, thereby enriching the cultural diversity. ⢠4chan (/pol/): The /pol/ board on 4chan is character- ized by anonymity and minimal moderation, resulting in a high density of extreme and explicit hate content [ Col- ley and Moore, 2022 ] . It allows M 3 to capture the upper bound of hateful visualātextual expressions rarely ob- served on mainstream platforms. M 3 encompasses a wide variety of languages, such as English, Chinese, and Arabic (details in Section B of the appendix), reflecting the globalized nature of online hate speech. As shown in Figure 3, the word cloud reveals pro- nounced multi-platform heterogeneity in hate expressions, in- dicating that M 3 spans a wide spectrum of hate explicitness. On 4chan, high-frequency terms such as āniggerā, ākikeā, and āfaggotā co-occur with explicit profanity (e.g., āfuckā, āshitā), indicating direct group targeting and overt dehuman- ization. In contrast, hate expressions on X are largely em- bedded within political and conflict-related discourse, with frequent references to entities such as āIsraelā, āGazaā, āTrumpā, and āBidenā. On Weibo, high-frequency terms cen- ter on national identities (āChiaā, āAmericaā, āJapanā), in- dicating that hate is predomaintly articulated through event- driven discourse. 4.3 Hierarchical Multi-label Categorization Category Definition. We further operationalize the defini- tion of hate speech by UN [ Guterres, 2019 ] into eight the- matic categories to facilitate fine-grained analysis. Each cate- gory represents a distinct and socially significant form of hate commonly observed online: ⢠Religion: Memes that promote harmful content related to religious conflict, such as disputes or hostility be- tween religious sects or groups. ⢠Politics: Memes about political conflict, including hate stemming from government policy failures or partisan disputes, especially during elections. ⢠Race: Memes containing racially discriminatory con- tent, targeting individuals or groups based on ethnicity or race. ⢠Gender: Memes that involve gender-based discrimina- tion, including harmful content targeting women or the LGBTQ+ community. ⢠Health Status: Memes that mock, insult, or discrimi- nate against individuals based on their health conditions, such as disabilities or chronic illnesses. ⢠Violence: Memes that incite or glorify acts of violence, encompassing content that explicitly or implicitly advo- cates targeted shootings, promotes or incites online ha- rassment or cyber violence. ⢠Public Health: Memes spreading misinformation or fear-mongering about public health crises, like famine panic or pandemic-related hate speech. ⢠International Relations: Memes targeting international relations, for example, provoking inter-country hostility during conflicts or fueling anti-refugee sentiment. The per-category sample distribution is shown in Figure 4. Notably, due to data collection from the /pol/ board on 4chan, the politics category accounts for 26.4% of hate sam- ples. Among hate samples in M 3 , 22.76% contain multiple category labels. Statistical analysis shows that the most fre- quent multi-label combinations include (āraceā, āviolenceā) and (āinternational relationsā, āpoliticsā) (58 samples each), followed by (āpoliticsā, raceā) and (āpoliticsā, violenceā) (40 samples each), which is consistent with intuitive discourse patterns on these platforms. 4.4 Rationales for hate memes For the hate samples, we annotate 1,557 rationales describ- ing why the content is hateful. Each rationale follows a <verb><object> structure (e.g., insult black people), with an average length of 32.35 characters, the longest being 101, and the shortest 9. As shown in Figure 5, hate samples are an- notated with a single rationale, while samples associated with multiple rationales constitute a smaller portion across all plat- forms. The three most common rationales include āexpress political hatredā (44 times in the politics category), ādepre- ciate transgender individualsā (33 times in the gender cate- gory), and ādiscriminate against people with intellectual dis- abilitiesā (29 times in the health status category). ModelMemePostOverallā BinaryMulti-classRationale AccāPāRāF1āMacro-PāMacro-RāMacro-F1āHLāSubset AccāBLEUāROUGEāBERTScoreā LLaVA-v1.6-Vicuna-7B-hf 19.8353.6953.69100.0069.8647.9229.4832.8916.189.480.434.3397.39 21.8253.6953.69100.0069.8626.4069.4635.6135.0713.200.061.8997.41 LLaVA-v1.6-Vicuna-13B-hf 19.9766.5261.8997.9575.8547.9229.4832.8916.189.480.405.0396.38 20.0359.4768.8444.7654.2540.7561.3944.7521.4814.260.294.5195.90 GLM-4.1V-9B-Thinking 77.4276.7498.1957.7472.7258.7086.5069.0211.6640.521.048.6797.76 80.6177.1192.0962.7574.6454.6881.9864.4913.9633.761.199.1797.79 Qwen2.5-VL-3B-Instruct 51.6376.2178.6376.4877.5456.5175.5662.0014.9828.910.333.7197.19 58.4174.7588.6160.7772.1051.8474.4758.3316.2724.130.255.2497.07 Qwen2.5-VL-7B-Instruct 82.1990.2691.1590.6790.9959.0879.5065.5412.4239.530.857.1997.68 77.4386.4886.6388.4787.5456.4873.6459.5914.7430.800.646.6397.34 Qwen3-VL-8B-Instruct 87.6386.8091.7682.8587.0861.7179.8668.1711.8642.111.489.1897.51 84.8385.9587.0586.7286.8955.3676.6462.5314.6333.161.508.6497.42 GPT-4o 60.3786.2794.1780.0886.5650.679.8760.9315.1626.260.107.9597.47 62.9685.4783.6589.4686.4549.0462.5753.4016.0419.890.147.8797.43 Gemini-3 59.4266.9792.7541.7357.5653.9785.6364.6015.2433.231.7411.6497.28 75.1673.4486.4359.9470.7948.1884.4259.7418.2025.172.1713.0797.51 Table 2: Comparison of state-of-the-art MLLMs on the M 3 dataset. The models are listed in order from open-source to proprietary, following their chronological release or version evolution: (1) the early LLaVA-v1.6 series (7B, 13B), (2) the reasoning-enhanced GLM-4.1V-9B- Thinking, (3) the latest Qwen series (Qwen2.5-VL-3B/7B and Qwen3-VL-8B), and (4) the proprietary GPT-4o and Gemini-3. Results cover binary and multi-class classification, and rationale quality under meme-only and meme + post settings. Boldface denotes the extremal values under the meme-only setting, while boldface with underline denotes the extremal values under the meme + post setting. 8.6% 7.4% 66.2% 14.3% 0.8% 2.6% WeiboX4chanSingle rationaleMultiple rationales Figure 5: The distribution of single rationale and multiple rationales across X, Weibo, and 4chan. In summary, M 3 is a thematically diverse multimodal benchmark with broad coverage. Its inclusion of real-world social media contexts, hierarchical multi-label annotations, and comprehensive rationales makes it a valuable resource for evaluating the nuanced hate recognition and interpreta- tion capabilities of MLLMs, especially within the real-world dynamic environments. 5 Experiments 5.1 Experiment Setups Dataset and Baselines. Experiments are conducted on M 3 , containing 2,455 memes paired with corresponding posts. Each instance is annotated with binary labels, and instances labeled as hate are further annotated with fine-grained cate- gories and phrase-level rationales. We compare several rep- resentative MLLMs, categorized into open-source and pro- prietary models, and ordered by their release dates or ver- sions: (1) Open-source models include LLaVA-v1.6 series (7B and 13B, Feb 2024), GLM-4.1V-9B-Thinking (Jul 2025), Qwen2.5-VL series (3B and 7B, Jan 2025), Qwen3-VL-8B- Instruct (Oct 2025); (2) Proprietary models include GPT-4o (May 2024) and Gemini-3 (Nov 2025). GPT-4o and Gemini- 3 are accessed via the official API, whereas the other models are deployed locally. Tasks and Evaluation Metrics. We investigate three meme understanding tasks: (1) Binary hate detection, which pre- dicts whether a meme contains hateful content (hate vs. nor- mal); (2) Multi-label fine-grained classification, identifying specific categories present in the hateful meme; and (3) Ra- tionale generation, which produces a concise verbāobject phrase explaining why the meme is hateful. For binary clas- sification, we report Accuracy, Precision, Recall, and F1- score. For multi-label fine-grained classification, we compute macro-Precision, macro-Recall, macro-F1, Hamming Loss, and Subset Accuracy [ Zhang and Zhou, 2013 ] . For rationale generation, we assess the quality of generated rationales us- ing BLEU [ Papineni et al., 2002 ] , ROUGE [ Lin, 2004 ] and BERTScore [ Zhang et al., 2019 ] . We define an Overall score by first normalizing all metrics to [0, 1] (inverting where nec- essary) and taking average across the three tasks. Implementation Details. For each model, we explore two input settings: (i) Meme-only, where only the meme image is provided; and (i) Meme + Post, where the meme is paired with its associated post text. All experiments are conducted on a single NVIDIA A800 GPU (80GB). To ensure fair com- parison, no task-specific fine-tuning is performed; instead, models are directly evaluated on downstream tasks in a zero- shot setting. 5.2 Results and Discussion Different Tasks. According to Table 2, a clear task-level performance hierarchy emerges across all evaluated mod- els. models achieve robust results in binary classification (typically> 85% accuracy) but struggle with multi-class tasks (Macro-F1: 32.89%ā69.02%). Notably, scaling model Figure 6: Performance comparison of MLLMs across three platforms. From left to right: 4chan (English-centric), X (Multi-lingual, e.g., Latin and Arabic), and Weibo (Chinese-dominant). capacity yields non-uniform gains. For example, increas- ing model capacity from LLaVA-7B to LLaVA-13B yields only marginal improvements in binary accuracy but leads to a noticeable increase in multi-class macro-F1. Similarly, within the Qwen family, scaling from 3B to 7B results in clear gains in binary classification (+13.7 accuracy points), whereas further scaling to Qwen3-VL-8B produces diminish- ing returns for binary accuracy but more pronounced benefits for multi-class macro-F1 and rationale metrics. A unique pat- tern emerges in rationale generation where high BERTScore values (above 95.0) coexist with low lexical-overlap metrics, including BLEU scores below 2.2 and a peak ROUGE of 11.64 (Gemini-3, meme-only). This discrepancy arises be- cause BLEU and ROUGE operate at the word level, failing to capture the semantic alignment of our generated rationales, which primarily consist of verb-object phrases with an aver- age length of 32.35 characters. Multimodal Inputs. As shown in Table 2, incorporating posts alongside memes produces heterogeneous effects across models and tasks, yielding inconsistent performance gains across the evaluated dimensions. Adding post information yields a negligible impact on binary classification (F1 ±3 points) but triggers a consistent performance decline in multi- class tasks. Specifically, macro-F1 scores drop in 6 of 9 evaluated models, with significant in high-performing mod- els like Qwen3-VL and GPT-4o. Although LLaVA showed an 11.86-point improvement, this behavior represents an ex- ception rather than the dominant pattern. Across newly re- leased models, BLEU and ROUGE improve when post text is included, suggesting that posts provide complementary se- mantic cues that facilitate rationale construction. Taken to- gether, these results indicate that current MLLMs struggle to integrate meme and post information for fine-grained in- tent understanding robustly. While additional context may help rationale generation, it often introduces ambiguity that degrades classification performance, highlighting a key limi- tation in real-world hate speech moderation scenarios where user intent is distributed across modalities. Multi-platform and Multi-lingual Evaluation. We se- lect four multimodal models (LLaVA-v1.6-Vicuna-13B- hf, GLM-4.1V-9B-Thinking, Qwen3-VL-8B-Instruct, and Gemini-3) from different model families that demonstrate strong overall performance in prior experiments and evaluate them across 4chan, X, and Weibo. A pronounced precision- recall trade-off emerges on X and Weibo, particularly for Gemini-3, which achieves perfect precision but a meager 16.94 recall on Weibo. These results suggest that certain models adopt overly conservative prediction strategies, pri- oritizing precision while sacrificing recall. Although this re- duces false positives, it leads to a large number of false nega- tives and limits practical usefulness. The effect is particularly pronounced under domain and language shifts, indicating that alignment or thresholding mechanisms may bias mod- els toward excessively āsafeā predictions in multi-platform settings. From a multi-lingual perspective, models gener- ally perform better on English-dominated platforms (4chan) compared to non-English or mixed-language platforms like X and Weibo. Performance degradation is particularly no- ticeable for Chinese-language content, where high-precision models like Gemini-3 still suffer from dramatic recall drops, highlighting challenges in cross-linguistic generalization. For clarity, we visualize only the results under the meme + post setting. Results are reported in full in Section C.1 of the ap- pendix. 6 Conclusion In this work, we introduce M 3 , a multi-platform, multi- lingual, and multimodal meme dataset constructed through an agentic annotation framework with systematic human ver- ification. M 3 contains 2,455 multimodal instances from X, 4chan, and Weibo, annotated with binary hate labels, fine- grained categories, and human-verified rationales. We bench- mark M 3 on state-of-the-art MLLMs. Results show that in- corporating surrounding post context does not consistently improve hate detection and often degrades fine-grained clas- sification, although it can benefit rationale generation. These findings highlight current limitations of MLLMs in integrat- ing multimodal contextual information. Overall, M 3 provides a challenging benchmark for multimodal hate speech analy- sis and offers insights into the gap between existing model capabilities and real-world moderation needs. Ethical Statement This work involves the analysis of hateful memes that may contain offensive content. All data in M 3 are collected from publicly available platforms and are used solely for research purposes. Personally identifiable information is removed dur- ing data processing. Annotations are conducted with human verification, and annotators are informed of the sensitive nature of the content. This study does not endorse hateful expressions. The dataset is intended to support research on multimodal hate detection and should be used responsibly in accordance with ethical guidelines. Acknowledgments and Disclosure of Funding This work was supported in part by the National Natural Sci- ence Foundation of China (62306229), the Youth Talent Sup- port Program of Shaanxi Science and Technology Associa- tion (20240113), the China Postdoctoral Science Foundation (2025T180425). References [ Barnes, 2019 ] Luke Barnes. With each new attack, far-right extremistsā manifestos are being āmemeticized,ā, 2019. [ Brown, 2018 ] Alexander Brown. What is so special about online (as compared to offline) hate speech? Ethnicities, 2018. [ Chhabra and Vishwakarma, 2023 ] Anusha Chhabra and Di- nesh Kumar Vishwakarma. A literature survey on multi- modal and multilingual automatic hate speech identifica- tion. Multimedia Systems, pages 1203ā1230, 2023. [ Colley and Moore, 2022 ] ThomasColleyandMartin Moore. The challenges of studying 4chan and the alt- right:ācome on in the waterās fineā. New Media & Society, 24(1):5ā30, 2022. [ Debnath et al., 2025 ] RiddhimanSwananDebnath, Nahian Beente Firuj, Abdul Wadud Shakib, Sadia Sultana, and Md Saiful Islam.ExMute: A context- enriched multimodal dataset for hateful memes.In NLPAIDL, pages 83ā89, 2025. [ Ferrara et al., 2016 ] Emilio Ferrara, Onur Varol, Clayton Davis, Filippo Menczer, and Alessandro Flammini. The rise of social bots. Communications of the ACM, 59(7):96ā 104, 2016. [ Fersini et al., 2022 ] Elisabetta Fersini, Francesca Gasparini, Giulia Rizzi, Aurora Saibene, Berta Chulvi, Paolo Rosso, Alyssa Lees, and Jeffrey Sorensen. Semeval-2022 task 5: Multimedia automatic misogyny identification. In Se- mEval, pages 533ā549, 2022. [ Founta et al., 2018 ] Antigoni Founta, Constantinos Djou- vas, Despoina Chatzakou, Ilias Leontiadis, Jeremy Black- burn, Gianluca Stringhini, Athena Vakali, Michael Siriv- ianos, and Nicolas Kourtellis. Large scale crowdsourc- ing and characterization of twitter abusive behavior. In ICWSM, 2018. [ Gomez et al., 2020 ] Raul Gomez, Jaume Gibert, Lluis Gomez, and Dimosthenis Karatzas. Exploring hate speech detection in multimodal publications. In WACV, pages 1470ā1478, 2020. [ Guterres, 2019 ] Ant Ģ onio Guterres. United nations strategy and plan of action on hate speech. Technical report, United Nations, 2019. [ Hee et al., 2023 ] Ming Shan Hee, Wen-Haw Chong, and Roy Ka-Wei Lee.Decoding the underlying mean- ing of multimodal hateful memes.arXiv preprint arXiv:2305.17678, 2023. [ Hee et al., 2024 ] Ming Shan Hee, Shivam Sharma, Rui Cao, Palash Nandi, Tanmoy Chakraborty, and Roy Ka-Wei Lee. Recent advances in hate speech moderation: Multimodal- ity and the role of large models. In EMNLP, 2024. [ Jiang and Zubiaga, 2024 ] Aiqi Jiang and Arkaitz Zubiaga. Cross-lingual offensive language detection: A systematic review of datasets, transfer approaches and challenges. arXiv preprint arXiv:2401.09244, 2024. [ Kiela et al., 2020 ] Douwe Kiela, Hamed Firooz, Aravind Mohan, Vedanuj Goswami, Amanpreet Singh, Pratik Ringshia, and Davide Testuggine.The hateful memes challenge: Detecting hate speech in multimodal memes. NeurIPS, pages 2611ā2624, 2020. [ Li et al., 2023 ] Lifang Li, Hong Wen, and Qingpeng Zhang. Characterizing the role of weibo and wechat in sharing original information in a crisis. Journal of Contingencies and Crisis Management, 31(2):236ā248, 2023. [ Lin, 2004 ] Chin-Yew Lin. Rouge: A package for automatic evaluation of summaries. In TSB, pages 74ā81, 2004. [ Mathew et al., 2021 ] BinnyMathew,PunyajoySaha, Seid Muhie Yimam, Chris Biemann, Pawan Goyal, and Animesh Mukherjee. Hatexplain: A benchmark dataset for explainable hate speech detection. In AAAI, pages 14867ā14875, 2021. [ Mulki et al., 2019 ] HalaMulki,HatemHaddad, Chedi Bechikh Ali, and Halima Alshabani.L-hsab: A levantine twitter dataset for hate speech and abusive language. In ALW, pages 111ā118, 2019. [ Nayak and Agrawal, 2022 ] AjayNayakandAnupam Agrawal.Detection of hate speech in social media memes: A comparative analysis.In ICICICT, pages 1179ā1185, 2022. [ Pandiani et al., 2025 ] Delfina S Martinez Pandiani, Erik Tjong Kim Sang, and Davide Ceolin. ātoxicāmemes: A survey of computational perspectives on the detection and explanation of meme toxicities. Online Social Networks and Media, page 100317, 2025. [ Papineni et al., 2002 ] Kishore Papineni, Salim Roukos, Todd Ward, and Wei-Jing Zhu. Bleu: a method for au- tomatic evaluation of machine translation. In ACL, pages 311ā318, 2002. [ Pramanick et al., 2021 ] Shraman Pramanick, Dimitar Dim- itrov, Rituparna Mukherjee, Shivam Sharma, Md Shad Akhtar, Preslav Nakov, and Tanmoy Chakraborty. De- tecting harmful memes and their targets. In ACL-IJCNLP, pages 2783ā2796, 2021. [ Rao et al., 2023 ] Xiaojun Rao, Yangsen Zhang, Qilong Jia, Xueyang Liu, Shuang Peng, et al. Chinese hate speech detection method based on roberta-wwm. In CCL, pages 501ā511, 2023. [ Salles et al., 2025 ] Isadora Salles, Francielle Vargas, and Fabr Ģ Ä±cio Benevenuto. Hatebrxplain: A benchmark dataset with human-annotated rationales for explainable hate speech detection in brazilian portuguese. In COLING, pages 6659ā6669, 2025. [ Schmid, 2025 ] Ursula Kristin Schmid.Humorous hate speech on social media: A mixed-methods investigation of usersā perceptions and processing of hateful memes. New Media & Society, pages 1588ā1606, 2025. [ Shifman, 2013 ] Limor Shifman. Memes in digital culture. MIT press, 2013. [ Tonneau et al., 2025 ] Manuel Tonneau, Diyi Liu, Niyati Malhotra, Scott A Hale, Samuel Fraiberger, Victor Orozco-Olvera, and Paul R Ģ ottger. Hateday: Insights from a global hate speech dataset representative of a day on twit- ter. In ACL, pages 2297ā2321, 2025. [ Velasquez et al., 2021 ] Nicolas Velasquez, Rhys Leahy, N Johnson Restrepo, Yonatan Lupu, Richard Sear, Nicholas Gabriel, OK Jha, Beth Goldberg, and NF John- son. Online hate network spreads malicious covid-19 con- tent outside the control of individual social media plat- forms. Scientific reports, page 11549, 2021. [ Ware, 2022 ] Jacob Ware. Testament to Murder: The Violent Far-Rightās Increasing Use of Terrorist Manifestos. JS- TOR, 2022. [ Waseem and Hovy, 2016 ] Zeerak Waseem and Dirk Hovy. Hateful symbols or hateful people? predictive features for hate speech detection on twitter. In NAACL-HLT, pages 88ā93, 2016. [ Zampieri et al., 2019 ] Marcos Zampieri, Shervin Malmasi, Preslav Nakov, Sara Rosenthal, Noura Farra, and Ritesh Kumar. Predicting the type and target of offensive posts in social media. arXiv preprint arXiv:1902.09666, 2019. [ Zhang and Zhou, 2013 ] Min-Ling Zhang and Zhi-Hua Zhou.A review on multi-label learning algorithms. IEEE transactions on knowledge and data engineering, 26(8):1819ā1837, 2013. [ Zhang et al., 2019 ] Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q Weinberger, and Yoav Artzi. Bertscore: Evaluating text generation with bert.arXiv preprint arXiv:1904.09675, 2019. [ Zhou et al., 2021 ] Yi Zhou, Zhenhao Chen, and Huiyuan Yang. Multimodal learning for hateful memes detection. In ICMEW, pages 1ā6, 2021. A The Details of the Agentic Annotation Framework This section provides detailed implementation and repro- ducibility information for the Agentic Annotation Framework (Section 3), organized according to the four main processes used to construct the M 3 dataset: Data Acquisition, Prepro- cessing, Multimodal Annotation, and Expert-driven Quality Assurance. A.1 Process 1: Data Acquisition Collector. The Collector harvests raw images and associ- ated JSON metadata (Table 3) from three diverse and comple- mentary social platforms: X (formerly Twitter), Weibo, and 4chanās /pol/ board. This selection ensures a broad linguistic and cultural scope, capturing a wide spectrum of image-based hate speech. We collected 43,567 image-post pairs includ- ing English, Arabic, Japanese, and other Latin and non-Latin script languages from X. The smaller volume is due to lim- ited API stability and access constraints during our specific collection window in February 2024. FieldDescription postidUnique identifier of the post posttimeTimestamp the post was published imgidUnique identifier for the meme imgurlURL linking to the meme posttextCleaned textual content of the post imgtextOCR-extracted textual content within the meme Table 3: Fields in JSON Metadata Files. posttext contains the textual content of the post after Extractor and Cleaner. A.2 Process 2: Preprocessing Extractor. The Extractor uses PaddleOCR to extract em- bedded image text (img text) and combines it with post text (posttext) to create unified textual representations. Processing is parallelized, skips images with pre-existing OCR results, and aggregates line-level outputs into single strings per image: ⢠Batch Processing and Parallelism: Images are pro- cessed in parallel using a thread pool executor, which en- ables concurrent OCR inference to accelerate through- put on multi-core systems. ⢠OCR Model Configuration: PaddleOCR is initialized with angle classification enabled and set to recognize English text, allowing detection of rotated or stylized text commonly found in memes. ⢠Incremental Processing: For each image, the corre- sponding JSON metadata file is checked for existing OCR results. If recognized text is already present, the image is skipped to avoid redundant computation. ⢠Text Extraction and Aggregation: The OCR outputs are parsed to extract line-level recognized text, which are concatenated with newline separators to form a single textual string representing the image content. ⢠Result Persistence and Error Logging: The extracted text is appended to the imageās JSON metadata under the recognized text field. Any OCR processing failures are logged into a dedicated failure record file for later inspection and potential reprocessing. Cleaner. To reduce noise and enhance the semantic quality of textual data, we apply platform-specific cleaning proce- dures to both the post content and the OCR-extracted embed- ded text. These include: ⢠Removal of URLs and web links, which are common but irrelevant for hate speech semantics. ⢠Elimination of user mentions (e.g., @username), with distinct regular expressions adapted to each platformās syntax to accurately remove references without harming context. ⢠Filtering out hashtags, including specially format- ted tags on platforms like Weibo (e.g., #...# or [#...#]), to focus on natural language content. ⢠Replacement of newlines and excessive whitespace with standardized delimiters (commas or spaces) to normal- ize textual structure and facilitate downstream process- ing. ⢠For 4chan data, additional removal of quoting symbols and post references (e.g., >>12345) that do not con- tribute to meme semantics. These tailored cleaning steps ensure that the textual modal- ities reflect core semantic information relevant for hate speech detection, while minimizing platform-specific noise artifacts. Given the heterogeneous nature of user-generated content, we impose length-based thresholds on the combined textual representation (post content plus embedded image text) to re- tain only samples likely to carry meaningful semantic con- tent. Specifically: ⢠Samples with overly short text (e.g., fewer than 20 char- acters for 4chan posts, or fewer than 10 characters for embedded image text) are discarded as they lack suffi- cient context for interpretation. ⢠Samples with excessively long text (e.g., beyond 500 characters for 4chan post content, or 100 characters for embedded text) are also excluded to avoid noise from spam, off-topic content, or multi-topic posts. ⢠Corresponding image files are removed in tandem with filtered JSON metadata to maintain dataset integrity. This length-based filtering step balances data quality and scale, ensuring the final dataset consists of content-rich mul- timodal memes suitable for detailed annotation and model evaluation. Following this pipeline, we retain 16,179 meme samples from X, 298,043 from Weibo, and 221,249 from 4chan. A.3 Process 3: Multimodal Annotation Annotator. Three MLLMs (GLM-4v-flash, llava-v1.6- mistral-7b-hf, Qwen2.5-VL-7B-Instruct) as the Annotator predict for each meme, including a binary hate label (Nor- mal, Hate) and Fine-grained category label(s) (Religion, Pol- itics, Race, Gender, Health Status, Violence, Public Health, International Relations). Specifically, the three models come from different model families and training paradigms, pro- viding complementary strengths in multimodal reasoning, in- struction following, and cross-lingual understanding: GLM- 4v-flash is optimized for fast and low-cost inference, LLaVA- v1.6-Mistral-7B represents a strong open-source baseline with competitive multimodal understanding, and Qwen2.5- VL-7B-Instruct offers robust visionālanguage alignment and multilingual performance. Restricting the ensemble to three models significantly reduces inference cost and latency, which is crucial for large-scale meme annotation We adopt a prompt-based annotation strategy to guide the annotators. The specific prompts used for annotation are as follows: Prompt P1 You are a hateful meme detection assistant. Based on the input image and post: posttext, please select the most appropriate option from [āhateā, ānormalā]. And give your confidence score for this option as a decimal between 0 and 1. Please output in this format:ālabelā: āhateā or ānormalā, āconfidencescoreā: 0.0-1.0 P2 You are a harmful meme classification assistant. Based on the the content of the image, please se- lect the most appropriate one or more of the fol- lowing categories: Category list:[āreligionā, āpoliticsā, āraceā, āgenderā, āhealth statusā, āviolenceā, āpublic healthā, āinternational relationsā]. Please output in json format:ācategoryā: [ācate- gory1ā, ācategory2ā], āconfidence scoreā: [0.78, 0.56] Arbiter. The web-based interface is used only for man- ual review of low-consistency samples identified by the Ar- biter. It displays the meme image together with predictions from multiple LLM annotators for two tasks: binary hate labeling (see Figure 7) and hate category annotation (see Figure 8). As a representative example, the category ar- bitration interface (Figure 7) presents eight predefined hate categories.For each meme, the predictions from three LLMs are displayed as checkbox options, allowing human arbiters to select one or more suggestions, combine them, or reject all and leave the label blank. Upon submission, the results are written incrementally to a local JSON file (finalannotations.json), and the interface automat- ically proceeds to the next sample. This interface is built on the Flask web framework and uti- lizes dynamic HTML templates for rendering. Static image assets are served locally to ensure low latency and data secu- rity during annotation. Hate label selection and category cor- rection modules follow the same annotation flow and backend architecture, differing only in label type and instruction text. A.4 Process 4: Expert-driven Quality Assurance Explicator. The Explicator generates structured verb- object rationales for hate memes by combining visual and textual cues. First, a prompt is given to GPT-4o to produce an initial rationale or output. The prompt instructs the model to strictly output rationales for a memeās harmfulness using only verb-object phrases, separated by commas, with a max- imum of 30 words. Prompt You are a hateful meme explicator. Strictly output rationale for hate in English, using only verb-object phrases. If the meme have many rationales, separate verb- object phrases with commas. Do not add anything else. No more than 30 words. Given: - post:post content - category:category The output of the model serves as a draft, which may con- tain imprecise wording, incomplete logic, or formatting in- consistencies. Humans then review and edit this draft to cor- rect verb choices, clarify objects, ensure word limits, and en- force the specified format. During manual refinement stage, explicators are are provided with explicit guidance, includ- ing the attack target (individual, group, society, and nation) and the attack type (abuse, discriminate, ridicule, dehuman- ize, and incite violence), to ensure accurate editing. Validator. To further assess the reliability of high- consistency samples, we conducted a random audit of 200 instances that had unanimous agreement in the arbitration. Table 4 reports the detailed voting outcomes of three grad- uate validators. Each validator independently marked agree- ment (1) or opposition (0) for every sample. At the individual level, each validator maintained high agreement rates, rang- ing from 94.5% to 97.5%, with opposition limited to 5ā11 cases per validator. These results confirm the robustness of the Arbiterās initial decisions. B Dataset Details and Examples In this section, we provide a supplementary overview and analysis of our dataset M 3 . We describe its overall com- position, representative examples and detailed multi-lingual statistics. B.1 Representative Examples per Category Table 5 presents representative examples across the eight hate-related categories. Each entry includes a meme iden- tifier, the post text, the annotated category, and a rationale explaining its hateful content. These examples highlight the spectrum of hateful expres- sion in social and political contexts, demonstrating that effec- tive detection requires identifying not only toxicity but also Image 1 / 29114 Current status: Unlabeled There was a site that made magic cards from prompt. Is that what yours are? MLLM 1: normal (Conf: 1.0) MLLM 2: hate (Conf: 0.95) MLLM 3: Normal (Confidence: 0.99) Please make your final determination based on the image, post content, and reference tags: HateNormal next Image 1 / 662 Current status: Unlabeled of course you can't. MLLM 1: normal (Conf: 1.0) MLLM 2: hate (Conf: 0.95) MLLM 3: Hate (Confidence: 0.85) Please make your final determination based on the image, post content, and reference tags: HateNormal next Image 2 / 3252 Current status: Unlabeled é山愼å¤ę„¼ļ¼å¤§A让人ęļ¼ę„é£äøåŗ¦ēéØå ³ļ¼ē£Øå¾äŗŗé½č¦åē«ćå©ęä¹ę„ļ¼ēē»å®å们诓ć MLLM 1: normal (Conf: 1) MLLM 2: hate (Conf: 0.95) MLLM 3: Normal (Confidence: 0.82) Please make your final determination based on the image, post content, and reference tags: Hate Normal next Figure 7: Web-based arbitration interface for binary hate annotation. From left to right, examples are sourced from 4chan, X, and Weibo. The interface displays the meme alongside predictions from annotators to support human review of low-consistency samples. ValidatorAgreement (1)Opposition (0) Graduate 11955 Graduate 21946 Graduate 318911 Vote Statistics Three votes in favor181 Two votes in favor16 One vote in favor3 No votes in favor0 Table 4: Validator voting results for a random audit of 200 high- consistency samples. The table reports individual agreement/oppo- sition counts and aggregated vote statistics. the specific thematic target. M 3 captures a range of rhetorical strategies, from overt slurs and explicit hostility (ID 714, ID 653) to satire, indirect blame, and ideological dog whis- tles (ID 405, ID 1104), necessitating nuanced cultural and contextual understanding. Consequently, robust detec- tion must integrate multimodal and sociolinguistic inference. Category-level distinctions further underscore varied man- ifestations of hate: attacks on marginalized identities (e.g., race, gender, health status) often involve dehumanization or invalidation, whereas political and international content fre- quently conveys ideological vilification or geopolitical mock- ery. B.2 Multi-lingual Diversity To examine the linguistic and cultural diversity of the dataset, we analyze post and image text using a Unicode scriptābased approach. Latin-script languages are grouped together, and posts containing multiple scripts are labeled as multilingual (Table 6). Most post texts are monolingual, dominated by Latin-script content, followed by Chinese, while multilingual posts of- ten mix ChineseāLatin or ArabicāLatin, with rarer combi- nations involving Japanese, Korean, Hebrew, and Cyrillic scripts. OCR-extracted image text shows similar patterns: 2,237 monolingual and 194 multilingual samples, primar- ily Latin and Chinese, with multilingual cases largely Chi- neseāLatin. Cross-modal analysis reveals that image and post text are mostly linguistically aligned, though some instances ex- hibit explicit mismatches, reflecting cross-lingual divergence. Overall, the dataset spans Western, East Asian, Middle East- ern, and Slavic scripts, capturing both monolingual and mul- tilingual communication across modalities. C Experiments C.1 Multi-platform, Multi-lingual, and Multimodal Evaluation. We evaluate four representative multimodal models (LLaVA- v1.6-Vicuna-13B-hf, GLM-4.1V-9B-Thinking, Qwen3-VL- 8B-Instruct, and Gemini-3) across three platforms (4chan, X, and Weibo) under two input settings: Meme + Post and Meme Only. Overall, incorporating post text consistently improves per- formance across platforms (Table 6), particularly in recall and F1, indicating the importance of complementary linguis- tic context. Among the evaluated models, Qwen3-VL-8B- Instruct achieves the strongest and most stable performance under both input settings, yielding the highest accuracy and F1 on all three platforms. GLM-4.1V-9B-Thinking also per- forms competitively, especially on X and Weibo, though it exhibits reduced recall in the meme-only setting. In contrast, LLaVA-v1.6-Vicuna-13B-hf consistently un- derperforms, suggesting limited robustness to cross-platform Image Annotation 91 / 24553 Post: SUNNI ISLAM 2050..TOTAL KAFFIR DEATH LLM Output Categories & Confidence: ['religion', 'politics', 'violence'] ā [0.9, 0.8, 0.85] ['violence', 'international relations'] ā [0.78, 0.56] ['religion', 'politics'] ā [0.85, 0.72] Select Categories (multiple or none): religion politics race gender health status violence public health international relations Next Image Image Annotation 2 / 250 Post: Definitely perfect description of the dictator. LLM Output Categories & Confidence: ['politics', 'international relations'] ā [0.9, 0.85] ['religion', 'violence'] ā [0.78, 0.56] ['politics'] ā [0.92] Select Categories (multiple or none): religion politics race gender health status violence public health international relations Next Image Image Annotation 6 / 199 Select Categories (multiple or none): religion politics race gender health status violence public health international relations Next Image ę„ę¬7.4ēŗ§å°éļ¼å¤å°ååøęµ·åøč¦ę„ Figure 8: Web-based arbitration interface for hate category annotation. From left to right, examples are sourced from 4chan, X, and Weibo. The interface presents hate categories with suggestions of annotators, allowing human arbiters to select, combine, or override model predic- tions. IDCategoryPostRationale 405ReligionEveryone but Christians refuse to participate in usury.Blame Christians while using apocalyptic language 502PoliticsMana. I had such high hopes for her... Then she came out as a commie. Incite anti-communist hatred 653RaceIf thereās one speck of... is it still just good olāe milk?Insult Black people 714GenderTrannies be like āitās not a mental illnessā...Depreciate transgender individuals 832Health StatusA big thing I used to do... but just for me these days.Discriminate against people with depres- sion 930ViolenceCanāt come soon enoughEncourage suicide 1011Public HealthWe care... watch you cope in your 55K a year shill cubi- cle. Spread vaccine death conspiracy 1104International RelationsNo freedom bucks for youSatirize Europeās dependence on the United States Table 5: Representative single-label examples for each hate category. and cross-lingual variation. Gemini-3 demonstrates high pre- cision but notably low recall on X and Weibo, indicating a conservative prediction tendency that limits its effectiveness in multi-label hate detection. These results highlight substantial performance variability across platforms and input modalities, underscoring the chal- lenges posed by linguistic diversity, cultural context, and mul- timodal interactions in real-world hate speech detection. C.2 Case Study To qualitatively illustrate model behavior on M 3 , we examine representative examples shown in Figure 9, highlighting the impact of textual context and differences in rationale genera- tion across models. The case studies in Figure 9 demonstrate the critical role of accompanying textual context in meme understanding. In the first example, the meme image alone appears ambiguous and is classified as normal. However, when the post text is intro- duced, implicit hostility becomes explicit, leading the model to correctly revise its prediction to hate. This highlights how textual cues can surface latent intent that is not visually ap- parent. In the second example, the inclusion of post text reduces model uncertainty by narrowing overgeneralized predictions. Without textual context, the model assigns multiple hate- related categories, reflecting ambiguity in visual interpreta- tion. The post text provides additional constraints, enabling the model to refine its output to a smaller, more precise set of categories. The bottom part of Figure 9 compares rationales generated by different models. GLM-4.1V-9B-Thinking produces ratio- nales that are more semantically faithful to the ground truth, often using concise and well-aligned verbāobject phrases. In contrast, GPT-4o and Qwen2.5-VL tend to generate broader StatisticsValue Post Text Language Distribution Total posts2455 Monolingual posts2117 * Monolingual ā Latin1794 * Monolingual ā Chinese323 Multilingual posts338 * Multilingual ā Chinese + Latin216 * Multilingual ā Arabic + Latin101 * Multilingual ā Chinese + Japanese + Latin15 * Multilingual ā Korean + Latin3 * Multilingual ā Arabic + Hebrew + Latin1 * Multilingual ā Chinese + Korean + Latin1 * Multilingual ā Chinese + Cyrillic + Latin1 Image Text Language Distribution Total images2455 Monolingual images2261 * Monolingual ā Latin1993 * Monolingual ā Chinese268 Multilingual images194 * Multilingual ā Chinese + Latin194 Post-Image Language Comparison Language-aligned pairs2326 Language-misaligned pairs129 Alignment ratio94.7% Misalignment ratio5.3% Table 6: Language statistics of post and image text. or more generic rationales that capture the overall tone but miss specific discriminatory intent. These qualitative differ- ences suggest that effective rationale generation favors se- mantic alignment over surface-level lexical overlap, partic- ularly in multimodal hate analysis tasks. D Related Work Details regarding the size, annotation status, and language of unimodal hate speech datasets are presented in Table 8, as supplemented in Section 2.1 of the paper. Figure 9: Case study showing model predictions under meme-only and meme+post settings. Adding a post changes the prediction from normal to hate (top), and refines category predictions (middle). The bottom part compares model-generated rationales, where GLM- 4.1V aligns better with ground truth than GPT-4o and Qwen2.5-VL. Meme + Post PlatformModelAccPRF1Macro-PMacro-RMacro-F1Subset AccBERTScore 4chan LLaVA-v1.6-Vicuna-13B-hf5581.5852.5463.9243.0363.1347.5616.195.89 GLM-4.1V-9B-Thinking72.9391.870.6279.8353.9282.7363.7929.6697.76 Qwen3-VL-8B-Instruct82.4387.7189.3688.5354.5678.6662.0629.2897.41 Gemini-369.3686.0971.0977.8846.986.0158.7420.8397.46 X LLaVA-v1.6-Vicuna-13B-hf64.2623.0818.1820.3437.4253.5539.879.8596.04 GLM-4.1V-9B-Thinking80.6189.4725.764048.4776.5657.7945.4597.94 Qwen3-VL-8B-Instruct89.7377.0884.0980.4350.9866.4556.0946.9797.54 Gemini-377.1987.510.6118.9255.9773.1661.7645.897.8 Weibo LLaVA-v1.6-Vicuna-13B-hf66.5411.596.458.2922.952.120.33.2395.87 GLM-4.1V-9B-Thinking84.6910034.6851.558.3678.3865.5456.4597.89 Qwen3-VL-8B-Instruct91.4995.466.9478.6756.0971.5657.7251.6197.35 Gemini-380.5310016.9428.9746.2991.2957.740.3297.61 Meme - Only PlatformModelAccPRF1Macro-PMacro-RMacro-F1Subset AccBERTScore 4chan LLaVA-v1.6-Vicuna-13B-hf78.2978.0799.2587.448.2530.9934.1910.5596.36 GLM-4.1V-9B-Thinking70.598.2262.2476.258.487.6268.8338.0497.73 Qwen3-VL-8B-Instruct82.9392.6484.1888.2161.3381.1867.938.9897.51 Gemini-358.7193.0649.2564.4152.987.0863.892997.22 X LLaVA-v1.6-Vicuna-13B-hf51.934.1298.4850.6846.420.3126.46.8196.46 GLM-4.1V-9B-Thinking82.3297.5630.346.2449.7281.4959.640.1597.9 Qwen3-VL-8B-Instruct90.4983.6177.2780.3154.5670.0259.350.7697.6 Gemini-376.8191.678.3315.2854.3870.1160.1846.9797.53 Weibo LLaVA-v1.6-Vicuna-13B-hf49.9130.1486.2944.6838.7216.2616.153.2396.5 GLM-4.1V-9B-Thinking87.7198.3648.3964.8657.4291.7667.4662.197.91 Qwen3-VL-8B-Instruct93.3893.277.4284.5854.8777.9761.4559.6897.47 Gemini-379.0284.2112.922.3850.3592.1462.2554.8497.57 Table 7: Multi-platform and multi-lingual evaluation results under two input settings: Meme + Post and Meme Only. DatasetSizeLabelClassificationLanguagemethods Waseemās16,914Sexist, Racist, NeitherMulti-classEnglishManual Fountaās80,000 Offensive, Abusive, Hate speech, Aggressive, Cyberbullying, Spam, Normal Multi-classEnglishManual L-HSAB5846hate, abusive, normalMulti-classArabicManual CHSD17,430hate, normalbinaryChineseManual OLID14,100 offensive, not offensive Hierarchical multi-label Englishmanual targeted insult, untargeted individual, group, other HateXplain20,148 hate, offensive, normal Hierarchical multi-label Englishmanual African, Islam, Jewish, Heterosexual, Women, Refugee, Arab, Caucasian, Hispanic, Asian HATEDAY240,000 Hateful, Offensive, Neutral Hierarchical multi-label Arabic, English, French, German, Indonesian, Portuguese, Spanish, Turkish manual Politics, National Origin, Gender, Religion, Sexual Orientation HateBRXplain7,000 Offensive, Non-offensive Hierarchical multi-labelPortuguesemanual highly, moderately, slightly xenophobia, racism, homophobia, sexism, religious intolerance, partyism, apology for the dictatorship, antisemitism, fatphobia Table 8: Summary of unimodal hate speech datasets. For OLID, targeted insult and untargeted are secondary labels under the primary label offensive; further, individual, group, and other are tertiary labels under targeted insult. For HateXplain, hate, offensive, and normal are primary labels, while secondary labels denote specific targeted communities, such as African, Islam, and others. For HATEDAY, hateful, offensive, and neutral are primary labels; politics, national origin, gender, religion, sexual orientation are secondary labels under hateful. For HateBRXplain, highly, moderately, slightly are secondary labels under the primary label offensive; Tertiary labels such as xenophobia are hate speech groups.